WO2016124116A1 - 信息传播方法和装置 - Google Patents

信息传播方法和装置 Download PDF

Info

Publication number
WO2016124116A1
WO2016124116A1 PCT/CN2016/072783 CN2016072783W WO2016124116A1 WO 2016124116 A1 WO2016124116 A1 WO 2016124116A1 CN 2016072783 W CN2016072783 W CN 2016072783W WO 2016124116 A1 WO2016124116 A1 WO 2016124116A1
Authority
WO
WIPO (PCT)
Prior art keywords
user
propagation
information
probability
relationship network
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/072783
Other languages
English (en)
French (fr)
Inventor
李朝
王志荣
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba Group Holding Ltd
Original Assignee
Alibaba Group Holding Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba Group Holding Ltd filed Critical Alibaba Group Holding Ltd
Priority to JP2017541018A priority Critical patent/JP2018511851A/ja
Publication of WO2016124116A1 publication Critical patent/WO2016124116A1/zh
Priority to US15/662,188 priority patent/US20170323313A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q30/00Commerce
    • G06Q30/02Marketing; Price estimation or determination; Fundraising
    • G06Q30/0201Market modelling; Market analysis; Collecting market data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N7/00Computing arrangements based on specific mathematical models
    • G06N7/01Probabilistic graphical models, e.g. probabilistic networks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L51/00User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
    • H04L51/52User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail for supporting social networking services
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/01Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/40Business processes related to social networking or social networking services
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/40Business processes related to social networking or social networking services
    • G06Q10/42Determination of affinities or common interests between users
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/40Business processes related to social networking or social networking services
    • G06Q10/46Determination of level of influence of users within social networking services
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/40Business processes related to social networking or social networking services
    • G06Q10/48Business processes related to social networking or social networking services using social graphs

Definitions

  • the present invention relates to the field of Internet technologies, and in particular, to an information dissemination method and apparatus.
  • the propagation of information can be controlled by establishing a probability model to learn the probability of information propagation between users.
  • the maximum expectation model EM (Expectation-maximuzation) can be used to learn the propagation probability between users.
  • the EM model method can easily calculate the extreme probability of probability 0 or probability 1. The resulting propagation probability tends to be large, and the propagation efficiency obtained after actual application is still not high.
  • the present invention aims to solve at least one of the technical problems in the related art to some extent.
  • an object of the present invention is to provide an information dissemination method which can improve the efficiency and credibility of information dissemination.
  • Another object of the present invention is to provide an information dissemination apparatus.
  • the information dissemination method includes: determining a first user corresponding to the information to be propagated, where the first user is the influence type of the first user belonging to the interest type network is greater than the preset a user of the value; acquiring a user relationship network starting from the first user, and propagating the information in the user relationship network with the first user as a starting point.
  • the first user corresponding to the information to be propagated is determined by the first user, and the first user is the user whose influence is greater than the preset value, and the first user is used as the starting point for information dissemination, which may be High-impact users disseminate information, improve the credibility of information dissemination, and improve the efficiency of information dissemination.
  • an information dissemination apparatus includes: a determining module, configured to determine a first user corresponding to the propagated information, where the first user is a user whose influence type is greater than a preset value in the network of the interest type that the first user belongs to; the propagation module is configured to obtain the first user as a starting point a user relationship network in which the information is propagated starting from the first user in the user relationship network.
  • the information dissemination device by determining the first user corresponding to the information to be propagated, the first user is a user whose influence is greater than a preset value, and the first user is used as a starting point for information dissemination, which may be High-impact users disseminate information, improve the credibility of information dissemination, and improve the efficiency of information dissemination.
  • FIG. 1 is a schematic flow chart of an information dissemination method according to an embodiment of the present invention.
  • FIG. 2 is a schematic diagram of an interest type network according to an embodiment of the present invention.
  • FIG. 3 is a schematic diagram of a process for establishing a preset number of interest type networks and determining a corresponding first user in each interest type network according to an embodiment of the present invention
  • FIG. 4 is a schematic diagram of a first user corresponding to determining information to be propagated according to an embodiment of the present invention
  • FIG. 5 is a schematic diagram of a probability of propagation of a user relationship network according to an embodiment of the present invention.
  • FIG. 6 is a schematic flow chart of obtaining propagation probabilities between users according to an embodiment of the present invention.
  • FIG. 7 is a schematic structural diagram of an information dissemination apparatus according to another embodiment of the present invention.
  • FIG. 8 is a schematic structural diagram of an information disseminating apparatus according to another embodiment of the present invention.
  • FIG. 1 is a schematic flowchart of an information dissemination method according to an embodiment of the present invention, where the method includes:
  • S101 Determine a first user corresponding to the information to be propagated, where the first user is a user whose influence type is greater than a preset value in the interest type network to which the first user belongs.
  • the information to be transmitted may be commodity promotion information or other information, and the present invention does not do this. limited.
  • the first user corresponding to the information to be propagated may be one or more.
  • the interest type network is a name of a category obtained by dividing a user based on the user's interest.
  • the user's interest may be determined according to the label possessed by the user, and the label possessed by the user may be predetermined according to the user's purchase or browsing history commodity information.
  • a preset number of interest type networks may be established in advance, and in each interest type network, a corresponding first user is determined.
  • multiple tags can be preset to divide users into different interest type networks.
  • the interest type network includes tags, such as fashion, outdoor, business, sports, travel, electronics, etc., each user can correspond to one or more tags.
  • the first user is a user whose influence type is greater than a preset value in the interest type network.
  • the influence is a kind of attribute of the user.
  • the influence of a user is used to measure the difficulty level of the information transmitted by the user, and the information transmitted by the user with greater influence is more easily Others accept.
  • the first user can also be called a person.
  • the number of people in each interest type network can be one or more.
  • the information to be transmitted is the information of the commodity.
  • a predetermined number of interest type networks are established, and in each interest type network, the corresponding first user is determined, specifically include:
  • the user-tag matrix is obtained according to the label propagation learning algorithm, and may include:
  • the similarity matrix of goods and goods can be used to indicate the similarity between goods in terms of user behavior, product titles, and product attributes.
  • the commodity used for calculating the similarity matrix may be a commodity processed by a quality buyer, and the processing may refer to one or more of purchasing, browsing, clicking, and collecting, and the quality buyer may determine according to the quality buyer model. For example, a buyer with a high credit rating or a large number of purchases is identified as a quality buyer. Specifically, it is possible to obtain information on all buyers, and then determine the quality buyers from all the buyers according to the quality buyer model, and then obtain the products processed by the quality buyers, and then according to the two products processed by the quality buyers. The two products calculate the similarity, and the similarity matrix W is obtained.
  • the commodity may be hash-mapped by a minimum hash algorithm to obtain a similarity matrix of the commodity and the commodity, where pid is the ID (Identity) of the commodity, and vid is the value of the commodity attribute.
  • ID, pid, and vid are usually available from the underlying data table.
  • the product in the product-tag information matrix F may also specifically refer to the product processed by the quality buyer, and the label refers to the updated product label, and after obtaining the product processed by the quality buyer, the product may be
  • the initial label is calculated through an iterative process to obtain a commodity-label information matrix F, wherein the initial label of each commodity is It is pre-recorded in the database as an attribute of the item, so that the initial label of the item can be obtained from the database.
  • the commodity-tag information matrix F can be obtained according to an iterative formula of the label propagation learning algorithm, and the iteration formula is as follows:
  • the user-product information matrix can be obtained according to the purchase corresponding to the quality buyer and the quality buyer, clicking or collecting the goods.
  • V the initial value of the item-tag information matrix F described above can be obtained.
  • the label possessed by the commodity can be obtained according to a statistical or HITS (Hyperlink-Induced Topic Search) sorting algorithm.
  • HITS Hydrolink-Induced Topic Search
  • the final commodity-tag information matrix F can be obtained when the iterative convergence condition is satisfied according to the above iterative formula.
  • the iterative convergence condition may include: setting a maximum number of iterations, satisfying the iterative convergence condition when the number of iterations reaches the maximum number of iterations; or satisfying the difference between the value after the iteration and the value before the iteration, when the difference is greater than a preset threshold
  • indicate that the iterative convergence condition is satisfied
  • the user in the user-tag matrix L may also specifically refer to a quality buyer, and the tag refers to a tag that the user has, and the tag that the user has may be determined according to the updated tag of the product processed by the user.
  • the quality buyers can be determined in the manner shown above, as well as the products processed by the quality buyers, and the initial labels of the products processed by the quality buyers can be obtained from the database, and then processed according to the quality buyers.
  • the product and the above-mentioned method (1) calculate the similarity matrix W of the product and the product, and then according to the similarity matrix W of the product and the product, and the initial label of the product processed by the quality buyer and the above-mentioned manner (2)
  • the product-label information matrix F is calculated, and the user-commodity information matrix V can be established according to the products processed by the quality buyer and the quality buyer, and then the user-label is obtained in the following manner according to the above V and F.
  • Matrix L Matrix L.
  • S32 Perform clustering on the user-tag matrix to obtain a preset number of interest type networks, and obtain a first user in each interest type network.
  • the matrix L can be clustered. For example, if the preset number is k, the matrix L can be double-clustered to obtain k categories, and each category corresponds to an interest type network. .
  • the center point of each category may be determined as the first user of the interest type network.
  • a list of the first users composed of the first users in the different interest type networks may be obtained, and the list of the people is obtained.
  • the first user in the network of different interest types is included, and when the information needs to be propagated, the first user corresponding to the information to be propagated may be first determined.
  • the first user corresponding to the information to be propagated is determined, including:
  • the first label is a label included in the information to be propagated
  • the first user including the first tag is determined as the first user corresponding to the information to be propagated.
  • the list of people includes: clothing up, 3C up, and home up, if the information to be spread includes the label 3C, then The first user corresponding to the disseminated information is a 3C person.
  • S102 Acquire a user relationship network starting from the first user, and use the first user as a starting point to propagate the information in the user relationship network.
  • the user relationship network is a network for describing an association relationship between the user and the user, and can directly obtain a user relationship network from an existing social network type application.
  • the user relationship network can be pre-established by adding friends or increasing attention.
  • the first user's friend may be obtained from the first user's application, and the second user's friend may be acquired in the second user's application.
  • the user relationship network includes: a first user -> a second user -> a third user.
  • the user relationship network starting from the first user can be imported from the existing data of the application, for example, from The user network of the social network is imported to determine the first user as the starting point.
  • the first user corresponding to the information to be transmitted is a 3C person
  • the user relationship network starting from the existing data and starting from 3C is the user relationship network 41, as shown in FIG.
  • the information to be propagated can be propagated in the user relationship network 41 starting from the 3C person.
  • the transmitting by using the first user as the starting point in the user relationship network, the information includes:
  • the information is propagated in the user relationship network by using the first user as a starting point according to a preset policy, where the preset policy includes a propagation range policy or a propagation speed policy.
  • the propagation range strategy refers to prioritizing the scope of propagation
  • the propagation speed strategy refers to prioritizing the speed of propagation
  • the probability of propagation between the user and the user in the user relationship network can be obtained.
  • the propagation range policy is adopted, information propagation can be performed regardless of the propagation probability.
  • the propagation speed strategy is adopted, the propagation probability can be greater than the preset. Information is propagated on the path of the value.
  • the user relationship network includes a first path 51, a second path 52, a third path 53, a fourth path 54, and a fifth path 55, assuming the first path 51,
  • the propagation probability between the users included in the second path 52 and the third path 53 is greater than a preset value, and the propagation probability less than the preset value exists between the users included in the fourth path 54 and the fifth path 55, then the information It is possible to propagate on the first path 51, the second path 52 and the third path 53, without propagating on the fourth path 54 and the fifth path 55.
  • the first user acts as a seed node for information propagation at the initial moment, and the seed node is responsible for propagating information to the neighbor node.
  • the first user is a 3C person, and the 3C person is associated with the 3C person.
  • the neighboring neighbor node includes the first node and the second node, and then the 3C person is set as the seed node at the initial time t, and the information is propagated by the 3C person to the first node and the second node, and the seed node transmits the information to the seed node. After the neighbor node, the neighbor node becomes the new seed node at the next time.
  • the seed node is the first node, and is no longer the 3C leader, and so on, from the initial first user in turn according to the user relationship.
  • User neighbor relationships in the network propagate information until there are no new seed nodes.
  • the propagation probability between neighboring node users in the user relationship network is independent and is not affected by the relationship between other neighboring nodes.
  • each seed node has only one opportunity to propagate information to the non-seed neighbor node. For example, the user becomes a seed node at time t, and there is only one chance to try to propagate information to the non-seed neighbor node at time t.
  • the neighbor is The node becomes the seed node at time t+1, regardless of whether the user successfully propagates at time t, and the user can no longer attempt to propagate information to its neighbor nodes at other times. If at the same time, there are multiple seed nodes trying to propagate information to the same node, the order of propagation can be arbitrary.
  • the acquiring the probability of propagation between the user and the user in the user relationship network includes:
  • the propagation probability learning model introducing the propagation probability variance control factor, the propagation probability between the user and the user in the user relationship network is obtained.
  • the propagation probability learning model may be an EM (Expectation-maximuzation) model. Due to the sparsity of the data, in the process of propagation probability learning, the propagation probability learned according to the EM model tends to be large. This is mainly because the calculation method of the EM model over-fitting in the case of sparse data results in uneven data distribution, and it is easy to estimate the extreme probability case where the probability of acquisition is 0 or the probability is 1.
  • EM Emergectation-maximuzation
  • a propagation probability variance control factor is introduced in the EM model to prevent the EM model from fluctuating violently during the iterative process.
  • the propagation probability learning model is introduced according to the propagation probability variance control factor, and the probability of propagation between the user and the user in the user relationship network is obtained, including:
  • the propagation probability variance control factor is introduced into the propagation probability learning model, and the propagation probability learning model with the propagation probability variance control factor is introduced.
  • the information propagation model is learned according to the propagation probability learning model that introduces the propagation probability variance control factor.
  • the propagation probability between the updated users is determined as the probability of propagation between the user and the user in the user relationship network.
  • the process of obtaining the propagation probability between users may include:
  • the independent cascading model is a basic propagation model that can be established in the existing way based on the user relationship network.
  • nodes and edges may be included, wherein each node may correspond to a user in a user relationship network, and each edge is a line segment composed of two adjacent users in the user relationship network.
  • the EM (Expectation-maximuzation) model is an optimization algorithm.
  • the EM model can be used to learn the independent cascade model, thereby obtaining the propagation probability of each edge included in the independent cascade model. That is, the probability of propagation between users and users in a user relationship network.
  • the traditional EM model can be expressed as:
  • the EM model of the propagation probability variance control factor can be:
  • is the control factor and k v,w is the propagation probability of the edge (v, w).
  • S64 Acquire a first update rule and a second update rule according to an EM model that introduces a propagation probability variance control factor.
  • the optimization equation can be determined according to the EM model introducing ⁇ , and then the optimization equation is solved to obtain the first update rule.
  • the first update rule is:
  • the first update rule is:
  • S65 It is judged whether or not the time segment data is finished. If not, the process proceeds to S66, and if so, S68 is executed.
  • the time segment data is preset to indicate the information dissemination time.
  • the seed node may be selected in the user relationship network, and then the preset information is propagated according to the user relationship network starting from the seed node, and the propagation time is preset time segment data.
  • the difference between the current time and the time when the information starts to propagate may be obtained. If the difference is smaller than the preset time segment data, it is determined that the time segment data does not end, otherwise it ends.
  • S66 It is judged whether the edge to be calculated is activated in the time segment data, and if so, S67 is executed; otherwise, S65 and its subsequent steps are repeatedly executed.
  • the edge to be calculated is the edge composed of user A and user B.
  • the information propagated through user A and user B then it can be determined that the edge composed of user A and user B is activated within the time. Otherwise it is not activated.
  • S67 The first update rule is used to update the propagation probability of the edge to be calculated, and then S69 is executed.
  • each side can set the initial propagation probability.
  • S68 The second update rule is used to update the propagation probability of the unactivated edge in the entire time segment data, and then execute S69.
  • the edges composed of user A and user C are not activated, that is, the information is not propagated between user A and user C, and the second update rule as shown above may be employed.
  • the propagation probability of the edge composed of user A and user C is updated.
  • propagation probability learning model is an example of an EM model, and the propagation probability learning model may also be other models, for example, a Markov model.
  • the first user by determining the first user corresponding to the information to be propagated, the first user is a user whose influence is greater than a preset value, and the first user is used as a starting point for information dissemination, and may be a user with greater influence. Disseminate information, improve the credibility of information dissemination, and improve the efficiency of information dissemination.
  • the first user can be determined by the label propagation learning algorithm to improve the effectiveness.
  • the accuracy of the propagation probability can be improved by introducing a control factor into the propagation probability learning model.
  • information diversity communication can be achieved by setting different propagation strategies.
  • the present invention also proposes an information dissemination apparatus.
  • FIG. 7 is a schematic structural diagram of an information disseminating apparatus according to another embodiment of the present invention. As shown in FIG. 7, the information dissemination apparatus includes a determination module 100 and a propagation module 200.
  • the determining module 100 is configured to determine a first user corresponding to the information to be propagated, where the first user is a user whose influence type is greater than a preset value in the network of interest types to which the first user belongs.
  • the information to be transmitted may be the product promotion information or other information, which is not limited by the present invention.
  • the first user corresponding to the information to be propagated may be one or more.
  • the interest type network may be a network that classifies users or information according to the type of interest, and may also be referred to as an interest network.
  • a preset number of interest type networks may be established in advance, and in each interest type network, a corresponding first user is determined.
  • multiple tags can be preset to divide users into different interest type networks.
  • the interest type network includes tags, such as fashion, outdoor, business, sports, travel, electronics, etc., each user can correspond to one or more tags. The process of specifically establishing an interest type network will be introduced in the following embodiments.
  • the first user is a user whose influence is greater than the preset value in the interest type network, and the influence is an attribute of the user.
  • the influence of one user is used to measure that the information transmitted by the user is accepted by others.
  • the degree of difficulty, in which the information transmitted by users with great influence is more easily accepted by others.
  • the first user can also be called a person.
  • the number of people in each interest type network can be one or more.
  • the list of people includes: clothing up, 3C up, and home up, if the information to be spread includes the label 3C, then The first user corresponding to the disseminated information is a 3C person.
  • the propagation module 200 is configured to acquire a user relationship network starting from the first user, where the information is propagated by using the first user as a starting point.
  • the user relationship network is a network for describing an association relationship between the user and the user, and can directly obtain a user relationship network from an existing social network type application. In the social network type application, between users The user relationship network can be pre-established by adding friends or increasing attention. For example, the first user's friend may be obtained from the first user's application, and the second user's friend may be acquired in the second user's application.
  • the user relationship network includes: a first user -> a second user -> a third user.
  • the user relationship network starting from the first user may be imported from the existing data of the application, for example, a user relationship network starting from the application of the social network to determine the first user as a starting point.
  • the first user corresponding to the information to be transmitted is a 3C person
  • the user relationship network starting from the existing data and starting from 3C is the user relationship network 41, as shown in FIG.
  • the information to be propagated can be propagated in the user relationship network 41 starting from the 3C person.
  • the first user is a user whose influence is greater than a preset value, and the first user is used as a starting point for information dissemination, and may be a user with greater influence. Disseminate information, improve the credibility of information dissemination, and improve the efficiency of information dissemination.
  • FIG. 8 is a schematic structural diagram of an information disseminating apparatus according to another embodiment of the present invention.
  • the information dissemination apparatus includes: a determination module 100, a second acquisition submodule 110, a first determination submodule 120, a propagation module 200, a third acquisition submodule 210, an acquisition unit 211, a modeling unit 212, The updating unit 213, the determining unit 214, the second determining submodule 220, the establishing module 300, the first obtaining submodule 310, and the clustering submodule 320.
  • the establishing module 300 includes a first obtaining submodule 310 and a clustering submodule 320.
  • the determining module 100 includes a second obtaining submodule 110 and The first determining sub-module 120; the propagating module 200 includes a third obtaining sub-module 210 and a second determining sub-module 220; the third obtaining sub-module 210 includes an obtaining unit 211, a modeling unit 212, an updating unit 213, and a determining unit 214.
  • the establishing module 300 is configured to establish a preset number of interest type networks, and determine a corresponding first user in each interest type network.
  • the information about the information to be transmitted is the information of the commodity, and the establishing module 300 may specifically include:
  • the first obtaining submodule 310 is configured to obtain a user-tag matrix according to the label propagation learning algorithm. Specifically, it may include:
  • the similarity matrix of goods and goods can be used to indicate the similarity between goods in terms of user behavior, product titles, and product attributes.
  • the commodity used for calculating the similarity matrix may be a commodity processed by a quality buyer, and the processing may refer to one or more of purchasing, browsing, clicking, and collecting, and the quality buyer may determine according to the quality buyer model. For example, a buyer with a high credit rating or a large number of purchases is identified as a quality buyer. Specifically, it is possible to obtain information on all buyers, and then determine the quality buyers from all the buyers according to the quality buyer model, and then obtain the products processed by the quality buyers, and then according to the two products processed by the quality buyers. The two products calculate the similarity, and the similarity matrix W is obtained.
  • the first obtaining sub-module 310 may hash the commodity (pid, vid) by a minimum hash algorithm to obtain a similarity matrix of the commodity and the commodity, where pid is the ID of the commodity (identity). , vid is the ID of the item attribute value, and pid and vid can usually be obtained from the underlying data table.
  • the product in the product-tag information matrix F may also specifically refer to the product processed by the quality buyer, and the label refers to the updated product label, and after obtaining the product processed by the quality buyer, the product may be
  • the initial label is calculated through an iterative process to obtain a commodity-label information matrix F, wherein the initial label of each commodity may be pre-recorded in the database as an attribute of the commodity, so that the initial label of the commodity can be obtained from the database.
  • the commodity-tag information matrix F can be obtained according to an iterative formula of the label propagation learning algorithm, and the iteration formula is as follows:
  • the user-product information matrix can be obtained according to the purchase corresponding to the quality buyer and the quality buyer, clicking or collecting the goods.
  • V the initial value of the item-tag information matrix F described above can be obtained.
  • the label possessed by the commodity can be obtained according to a statistical or HITS (Hyperlink-Induced Topic Search) sorting algorithm.
  • HITS Hydrolink-Induced Topic Search
  • the final commodity-tag information matrix F can be obtained when the iterative convergence condition is satisfied according to the above iterative formula.
  • the iterative convergence condition may include: setting a maximum number of iterations, satisfying the iterative convergence condition when the number of iterations reaches the maximum number of iterations; or satisfying the difference between the value after the iteration and the value before the iteration, when the difference is greater than a preset threshold
  • indicate that the iterative convergence condition is satisfied
  • the user in the user-tag matrix L may also specifically refer to a quality buyer, and the tag refers to a tag that the user has, and the tag that the user has may be determined according to the updated tag of the product processed by the user.
  • the quality buyers can be determined in the manner shown above, as well as the products processed by the quality buyers, and the initial labels of the products processed by the quality buyers can be obtained from the database, and then processed according to the quality buyers.
  • the product and the above-mentioned method (1) calculate the similarity matrix W of the product and the product, and then according to the similarity matrix W of the product and the product, and the initial label of the product processed by the quality buyer and the above-mentioned manner (2)
  • the product-tag information matrix F is calculated, and the user-commodity information matrix V can be established based on the products processed by the quality buyer and the quality buyer, and then the user-tag matrix L is obtained in the following manner according to the above V and F.
  • the clustering sub-module 320 is configured to cluster the user-tag matrix to obtain a preset number of interest type networks, and obtain a first user in each interest type network. After the user-tag matrix L is obtained, the matrix L can be clustered. For example, if the preset number is k, the matrix L can be double-clustered to obtain k categories, and each category corresponds to an interest type network. .
  • the center point of each category may be determined as the first user of the interest type network.
  • a list of the first users composed of the first users in the different interest type networks may be obtained, and the list of the people is obtained.
  • the first user in the network of different interest types is included, and when the information needs to be propagated, the first user corresponding to the information to be propagated may be first determined.
  • the determining module 100 specifically includes:
  • the second obtaining sub-module 110 is configured to acquire a first label, where the first label is a label included in the information to be propagated;
  • the first determining sub-module 120 is configured to determine, by the first user that includes the first label, a first user corresponding to the information to be propagated.
  • the person list includes: a clothing person, a 3C person, and a home person
  • the second obtaining sub-module 110 obtains the information to be transmitted
  • the label is 3C
  • the first determining sub-module 120 determines that the first user corresponding to the information to be propagated is a 3C person.
  • the propagation module 200 is further configured to: in the user relationship network, propagate the information by using the first user as a starting point according to a preset policy, where the preset policy includes a propagation range policy or a propagation speed policy.
  • the propagation range strategy refers to prioritizing the scope of propagation
  • the propagation speed strategy refers to prioritizing the speed of propagation.
  • the third obtaining sub-module 210 can obtain the propagation probability between the user and the user in the user relationship network.
  • the propagation range policy is adopted, the information can be propagated regardless of the propagation probability.
  • the propagation speed policy is adopted, Information is propagated only on paths where the probability of propagation is greater than the preset value. For example, taking the propagation speed policy as an example, referring to FIG.
  • the user relationship network includes a first path 51, a second path 52, a third path 53, a fourth path 54, and a fifth path 55, assuming the first path 51,
  • the propagation probability between the users included in the second path 52 and the third path 53 is greater than a preset value, and the existence probability of the user included in the fourth path 54 and the fifth path 55 is less than a preset value.
  • Information may propagate on the first path 51, the second path 52, and the third path 53 without propagating on the fourth path 54 and the fifth path 55.
  • the first user acts as a seed node for information propagation at the initial moment, and the seed node is responsible for propagating information to its neighbor nodes, for example, the first user is a 3C person, and the 3C person
  • the neighboring neighbor node includes the first node and the second node, then the 3C person is set as the seed node at the initial time t, and the information is propagated by the 3C person to the first node and the second node, when the seed node propagates the information
  • the neighbor node becomes the new seed node at the next time. For example, at time t+1, the seed node is the first node, and is no longer the 3C leader.
  • the information is propagated from the initial first user according to the user neighbor relationship in the user relationship network, until there is no new seed node.
  • the propagation probability between neighboring node users in the user relationship network is independent and is not affected by the relationship between other neighboring nodes.
  • each seed node has only one opportunity to propagate information to the non-seed neighbor node. For example, the user becomes a seed node at time t, and there is only one chance to try to propagate information to the non-seed neighbor node at time t.
  • the neighbor is The node becomes the seed node at time t+1, regardless of whether the user successfully propagates at time t, and the user can no longer attempt to propagate information to its neighbor nodes at other times. If at the same time, there are multiple seed nodes trying to propagate information to the same node, the order of propagation can be arbitrary.
  • the third obtaining sub-module 210 is further configured to obtain a propagation probability between the user and the user in the user relationship network according to a propagation probability learning model that introduces a propagation probability variance control factor.
  • the propagation probability learning model may be an EM (Expectation-maximuzation) model. Due to the sparsity of the data, in the process of propagation probability learning, the propagation probability learned according to the EM model tends to be large. This is mainly because the calculation method of the EM model over-fitting in the case of sparse data results in uneven data distribution, and it is easy to estimate the extreme probability case where the probability of acquisition is 0 or the probability is 1.
  • a propagation probability variance control factor is introduced in the EM model to prevent the EM model from fluctuating violently during the iterative process.
  • the third obtaining submodule 210 includes:
  • the obtaining unit 211 is configured to obtain the user relationship network, for example, importing a user relationship network from an application of an existing social network, and establishing an information propagation model according to the user relationship network and time segment data, for example, establishing an independent Cascading model.
  • the independent cascading model is a basic propagation model that can be established in the existing way based on the user relationship network.
  • the time segment data is preset to indicate the information dissemination time.
  • nodes and edges may be included, wherein each node may correspond to a user in a user relationship network, and each edge is a line segment composed of two adjacent users in the user relationship network.
  • the modeling unit 212 is configured to introduce a propagation probability variance control factor into the propagation probability learning model, obtain a propagation probability learning model that introduces a propagation probability variance control factor, and according to the propagation probability learning model that introduces a propagation probability variance control factor,
  • the information dissemination model performs learning to obtain a propagation probability update rule, and the update rule includes a first update rule and a second update rule.
  • the EM (Expectation-maximuzation) model is an optimization algorithm.
  • the EM model can be used to learn the independent cascade model, thereby obtaining the propagation probability of each edge included in the independent cascade model. That is, the probability of propagation between users and users in a user relationship network.
  • the traditional EM model can be expressed as:
  • the EM model of the propagation probability variance control factor can be:
  • is the control factor and k v,w is the propagation probability of the edge (v, w).
  • the optimization equation can be determined according to the EM model introducing ⁇ , and then the optimization equation is solved to obtain the first update rule.
  • the first update rule is:
  • the first update rule is:
  • the updating unit 213 is configured to update a propagation probability between the first group of users by using the first update rule, and update a propagation probability between the second group of users by using the second update rule, between the first group of users Edge at the time piece Within the data is activated and the edges between the second set of users are not activated within the time segment data.
  • the seed node may be selected in the user relationship network, and then the preset information is propagated according to the user relationship network starting from the seed node, and the propagation time is preset time segment data.
  • time segment data may be determined whether the time segment data ends, for example, a difference between a current time and a time when the information starts to propagate may be acquired, and if the difference is less than the preset time segment data, it is determined that the time segment data is not ended. Otherwise it ends.
  • the edge to be calculated is activated in the time segment data.
  • the edge to be calculated is the edge composed of user A and user B, and the information transmitted during the information propagation time passes. User A and User B can then determine that the edge composed of User A and User B is activated during this time, otherwise it is not activated. If activated, the first update rule is used to update the propagation probability of the edge to be calculated, and then the propagation probability of each edge is written into the propagation probability update library. If not activated, return to continue to determine if the time segment data is over.
  • the second update rule is used to update the propagation probability of the inactive side of the entire time segment data, and then the propagation probability of each edge update is written into the propagation probability update library.
  • propagation probability learning model is an example of an EM model, and the propagation probability learning model may also be other models, for example, a Markov model.
  • the determining unit 214 is configured to determine a propagation probability between the updated users as a propagation probability between the user and the user in the user relationship network.
  • the second determining sub-module 220 is configured to determine the path whose propagation probability is greater than a preset value as a propagation path, and propagate the information according to the propagation path to achieve a maximum propagation speed.
  • the first user by determining the first user corresponding to the information to be propagated, the first user is a user whose influence is greater than a preset value, and the first user is used as a starting point for information dissemination, and may be a user with greater influence. Disseminate information, improve the credibility of information dissemination, and improve the efficiency of information dissemination.
  • the first user can be determined by the label propagation learning algorithm to improve the effectiveness.
  • the accuracy of the propagation probability can be improved by introducing a control factor into the propagation probability learning model.
  • information diversity communication can be achieved by setting different propagation strategies.
  • portions of the invention may be implemented in hardware, software, firmware or a combination thereof.
  • multiple steps or methods may be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system.
  • a suitable instruction execution system For example, if implemented in hardware, as in another embodiment, it can be implemented by any one or combination of the following techniques well known in the art: having logic gates for implementing logic functions on data signals. Discrete logic circuits, application specific integrated circuits with suitable combinational logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
  • each functional unit in each embodiment of the present invention may be integrated into one processing module, or each unit may exist physically separately, or two or more units may be integrated into one module.
  • the above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
  • the integrated modules, if implemented in the form of software functional modules and sold or used as stand-alone products, may also be stored in a computer readable storage medium.
  • the above mentioned storage medium may be a read only memory, a magnetic disk or an optical disk or the like.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Business, Economics & Management (AREA)
  • General Physics & Mathematics (AREA)
  • Accounting & Taxation (AREA)
  • Development Economics (AREA)
  • Finance (AREA)
  • Strategic Management (AREA)
  • Data Mining & Analysis (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Entrepreneurship & Innovation (AREA)
  • Artificial Intelligence (AREA)
  • General Engineering & Computer Science (AREA)
  • Evolutionary Computation (AREA)
  • Mathematical Physics (AREA)
  • General Business, Economics & Management (AREA)
  • Marketing (AREA)
  • Economics (AREA)
  • Game Theory and Decision Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Probability & Statistics with Applications (AREA)
  • Algebra (AREA)
  • Computational Mathematics (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Pure & Applied Mathematics (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Information Transfer Between Computers (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

本发明提出一种信息传播方法和装置,该信息传播方法包括:确定要传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户;获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。该方法能够提高信息传播的效率和可信性。

Description

信息传播方法和装置 技术领域
本发明涉及互联网技术领域,尤其涉及一种信息传播方法和装置。
背景技术
随着社会信息化的发展,许多信息都需要得到有效的传播。几年来,社交网络已经成为人们获取、分享信息的主要渠道。通过社交网络传播信息,例如通过用户之间的信息分享来传播信息等,变得更容易被用户所接受。由于社交网络中的信息传播还处于初步阶段,许多信息传播的因素,例如信息传播速度、信息传播范围等参数,还处于难以预测的状态。目前,信息传播时可以采用专门的传播方式,例如,广告、营销推广等,但是,这种传播方式不容易被用户接受,效率不高。
现有技术中,可通过建立概率模型学习用户之间的信息传播概率来控制信息的传播。在传播概率学习的过程中,可以利用最大期望模型EM(Expectation-maximuzation)来学习用户间的传播概率。但由于数据的稀疏性导致数据分布不均匀,EM模型方法很容易计算得到概率为0或概率为1的极端概率情况,导致得到的传播概率往往方差比较大,实际应用后得到的传播效率仍然不高。
发明内容
本发明旨在至少在一定程度上解决相关技术中的技术问题之一。
为此,本发明的一个目的在于提出一种信息传播方法,该方法可以提高信息传播的效率和可信性。
本发明的另一个目的在于提出一种信息传播装置。
为达到上述目的,本发明实施例提出的信息传播方法,包括:确定要传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户;获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。
本发明实施例提出的信息传播方法,通过确定要传播的信息对应的第一用户,第一用户是影响力大于预设值的用户,并由第一用户为起点进行信息传播,可以由具有较大影响力的用户传播信息,提高信息传播的可信性,提高信息传播效率。
为达到上述目的,本发明实施例提出的信息传播装置,包括:确定模块,用于确定要 传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户;传播模块,用于获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。
本发明实施例提出的信息传播装置,通过确定要传播的信息对应的第一用户,第一用户是影响力大于预设值的用户,并由第一用户为起点进行信息传播,可以由具有较大影响力的用户传播信息,提高信息传播的可信性,提高信息传播效率。
本发明附加的方面和优点将在下面的描述中部分给出,部分将从下面的描述中变得明显,或通过本发明的实践了解到。
附图说明
本发明上述的和/或附加的方面和优点从下面结合附图对实施例的描述中将变得明显和容易理解,其中:
图1是本发明一实施例提出的信息传播方法的流程示意图;
图2是本发明一实施例的兴趣类型网络的示意图;
图3是本发明一实施例的建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户的流程示意图;
图4是本发明一实施例的确定要传播的信息对应的第一用户的示意图;
图5是本发明一实施例的用户关系网络传播概率的示意图;
图6是本发明一实施例的获取用户之间的传播概率的流程示意图;
图7是本发明另一实施例的信息传播装置的结构示意图;
图8是本发明另一实施例的信息传播装置的结构示意图。
具体实施方式
下面详细描述本发明的实施例,所述实施例的示例在附图中示出,其中自始至终相同或类似的标号表示相同或类似的元件或具有相同或类似功能的元件。下面通过参考附图描述的实施例是示例性的,仅用于解释本发明,而不能理解为对本发明的限制。相反,本发明的实施例包括落入所附加权利要求书的精神和内涵范围内的所有变化、修改和等同物。
下面参考附图描述根据本发明实施例的信息传播方法和装置。
图1是本发明一实施例提出的信息传播方法的流程示意图,该方法包括:
S 101:确定要传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户。
其中,要传播的信息可以是商品推广信息,也可以是其他信息,本发明对此不做 限定。要传播的信息对应的第一用户可以是一个或多个。
兴趣类型网络是基于用户的兴趣对用户进行划分后得到的类别的名称,用户的兴趣可以根据用户具有的标签确定的,用户具有的标签可以根据用户的购买或浏览历史商品信息等预先确定。
具体地,可以预先建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户。例如,可以预先设置多个标签,按照标签将用户划分到不同的兴趣类型网络中。如图2所示,兴趣类型网络包括标签,例如时尚、户外、商务、运动、旅行、电子等,每个用户都可以对应一个或多个标签。
第一用户是兴趣类型网络中影响力大于预设值的用户。影响力是用户的一种属性,在本实施例中,一个用户的影响力用于衡量该用户传播的信息被其他人接受的难易程度,其中,影响力大的用户传播的信息更容易被其他人接受。第一用户也可以称为达人。每个兴趣类型网络中达人可以是一个或者多个。
可选的,以要传播的信息是商品的信息为例,如图3所示,建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户,具体可以包括:
S31:根据标签传播学习算法,获取用户-标签矩阵;
具体的,根据标签传播学习算法,获取用户-标签矩阵,可以包括:
(1)计算得到商品与商品的相似度矩阵W。
商品与商品的相似度矩阵可以用于表示商品之间在用户行为、商品标题和商品属性上的相似度。
其中,用于计算相似度矩阵的商品可以是优质买家处理过的商品,处理具体可以是指购买,浏览,点击,收藏中的一项或者多项,优质买家可以根据优质买家模型确定,例如,将信用等级高或者购买次数多的买家确定为优质买家。具体的,可以获取所有买家的信息,再根据优质买家模型从所有买家中确定出优质买家,再获取优质买家处理过的商品,再根据优质买家处理过的商品中的两两商品计算相似度,得到相似度矩阵W。
具体的,可以通过最小哈希算法对商品(pid,vid)进行哈希映射,得到商品与商品的相似度矩阵,其中,pid是商品的ID(Identity,身份标识),vid是商品属性值的ID,pid和vid通常可以从基础数据表中获取。
(2)计算得到商品-标签信息矩阵F。
其中,商品-标签信息矩阵F中的商品也可以具体是指优质买家处理过的商品,标签是指商品更新后的标签,在获取优质买家处理过的商品后,可以根据每个商品的初始标签,经过迭代过程计算得到商品-标签信息矩阵F,其中,每个商品的初始标签可 以是作为商品的一个属性预先记录在数据库中的,从而可以从数据库中获取商品的初始标签。
具体的,商品-标签信息矩阵F可以根据标签传播学习算法的迭代公式得到,迭代公式如下:
While(F收敛)
F(t+1)=αSF(t)+(1-α)Y
end
其中,要计算的商品-标签信息矩阵F是上述公式中收敛时得到F(t+1),0≤α≤1为预设的加权参数,S是根据上述的商品与商品的相似度矩阵W计算得到的,S=D-1W∈Rn×n
Figure PCTCN2016072783-appb-000001
Figure PCTCN2016072783-appb-000002
Y是初始标签值,F(t)的初始值可以是根据已有的买家信息得到的商品-标签信息矩阵F的初始值,已有的买家信息可以是根据优质买家模型得到的。例如,根据买家的信用等级从多个买家中确定出预设个数的优质买家,再根据优质买家与优质买家对应的购买,点击或收藏的商品可以得到用户-商品信息矩阵V,以及,根据优质买家购买,点击或收藏的商品与商品具有的标签可以得到上述的商品-标签信息矩阵F的初始值。
其中,商品具有的标签可以根据统计或者HITS(Hyperlink-Induced Topic Search,超链诱导主题搜索)排序算法得到。
在获取F的初始值后,可以根据上述迭代公式,在满足迭代收敛条件时得到最终的商品-标签信息矩阵F。
迭代收敛条件可以包括:设置最大迭代次数,当迭代次数达到最大迭代次数时满足迭代收敛条件;或者,根据迭代后的值与迭代前的值的差值,在该差值大于预设阈值时满足迭代收敛条件,例如,||F(t+1)-F(t)||<β时表明满足迭代收敛条件,||F(t+1)-F(t)||表示F(t+1)与F(t)的欧氏距离,β表示预设阈值。
(3)计算得到用户-标签矩阵L。
其中,用户-标签矩阵L中的用户也可以具体是指优质买家,标签是指用户具有的标签,用户具有的标签可以根据用户处理过的商品的更新后的标签确定。
具体的,可以采用如上所示的方式确定出优质买家,以及获取优质买家处理过的商品,以及,从数据库中获取优质买家处理过的商品的初始标签,之后可以根据优质买家处理过的商品以及上述的方式(1)计算得到商品与商品的相似度矩阵W,再根据商品与商品的相似度矩阵W和优质买家处理过的商品具有的初始标签以及上述的方式(2)计算得到商品-标签信息矩阵F,再根据优质买家以及优质买家处理过的商品可以建立用户-商品信息矩阵V,之后在根据上述的V和F采用如下的方式得到用户-标签 矩阵L。
具体的,计算公式可以是:L=V*F,其中,V是上述得到的用户-商品信息矩阵,F是上述得到的收敛时的最终的商品-标签信息矩阵。
S32:对所述用户-标签矩阵进行聚类,得到预设个数的兴趣类型网络,并获取每个兴趣类型网络中的第一用户。
在得到用户-标签矩阵L后,可以对该矩阵L进行聚类,例如,预设个数是k个,则可以对矩阵L进行双聚类得到k个类别,每个类别对应一个兴趣类型网络。
在对矩阵L进行聚类得到k个类别后,以每个兴趣类型网络包括一个第一用户为例,每个类别的中心点可以确定为该兴趣类型网络的第一用户。不同兴趣类型网络的第一用户可以组成列表,该列表可以称为达人列表,达人列表例如表示为:P={p1,p2,…,pk},其中,pi(i=1,2,…,k)是第i个兴趣类型网络中的第一用户,也可以称为达人,pi可以由用户ID和该用户具有的标签组成。
在上述预先建立多个兴趣类型网络,并确定每个兴趣类型网络中的第一用户后,如上所述,可以得到由不同的兴趣类型网络中的第一用户组成的达人列表,达人列表中包括不同兴趣类型网络中的第一用户,在当前需要传播信息时,可以首先确定要传播的信息对应的第一用户。
可选的,确定要传播的信息对应的第一用户,包括:
获取第一标签,所述第一标签是所述要传播的信息包括的标签;
将包括所述第一标签的第一用户,确定为所述要传播的信息对应的第一用户。
例如,假设第一用户称为达人,如图4所示,达人列表中包括:服装达人,3C达人和家居达人,则如果要传播的信息包括的标签是3C,则该要传播的信息对应的第一用户是3C达人。
S102:获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。
其中,用户关系网络是用于描述用户与用户之间的关联关系的网络,可以直接从已有的社交网络类型的应用程序中获取用户关系网络,在社交网络类型的应用程序中,用户之间可以通过增加好友或者增加关注等方式预先建立用户关系网络。例如,可以先从第一用户的应用程序中获取到第一用户的好友包括第二用户,再在第二用户的应用程序中获取到第二用户的好友包括第三用户,则可以获取到的用户关系网络包括:第一用户—>第二用户—>第三用户。
以所述第一用户为起点的用户关系网络可以从应用程序的已有数据中导入,例如,从 社交网络的应用程序中导入以确定出的第一用户为起点的用户关系网络。
例如,如图4所示,假设要传播的信息对应的第一用户是3C达人,从已有数据中获取的以3C达人为起点的用户关系网络是用户关系网络41,则如图4所示,则可以将要传播的信息以3C达人为起点在用户关系网络41中传播。
可选的,所述在所述用户关系网络中以所述第一用户为起点传播所述信息,包括:
根据预设策略,在所述用户关系网络中以所述第一用户为起点传播所述信息,所述预设策略包括传播范围策略,或者,传播速度策略。
其中,传播范围策略是指优先考虑传播范围,传播速度策略是指优先考虑传播速度。
具体的,可以获取用户关系网络中用户与用户之间的传播概率,当采用传播范围策略时,可以不论传播概率高低都进行信息传播,当采用传播速度策略时,可以只在传播概率大于预设值的路径上进行信息传播。
例如,以传播速度策略为例,参见图5,假设用户关系网络包括第一路径51,第二路径52,第三路径53,第四路径54和第五路径55,假设第一路径51,第二路径52和第三路径53中包括的用户之间的传播概率都大于预设值,而第四路径54和第五路径55上包括的用户之间存在小于预设值的传播概率,则信息可以在第一路径51,第二路径52和第三路径53上传播,而不在第四路径54和第五路径55上传播。
具体的,信息在用户关系网络中传播时,第一用户在初始时刻作为信息传播的种子节点,种子节点负责向其邻居节点传播信息,例如,第一用户是3C达人,与3C达人相邻的邻居节点包括第一节点和第二节点,则在初始时刻t设置3C达人是种子节点,并且由3C达人将信息传播给第一节点和第二节点,当种子节点将信息传播给邻居节点后,邻居节点在下一时间成为新的种子节点,例如,在t+1时刻种子节点是第一节点,而不再是3C达人,依次类推,从初始的第一用户依次根据用户关系网络中的用户相邻关系进行信息传播,直至没有新的种子节点。另外,用户关系网络中邻居节点用户之间的传播概率是独立的,不受其他邻居节点之间的关系影响。并且,每个种子节点只有一次机会向非种子邻居节点传播信息,例如,用户在t时刻成为种子节点,仅在t时刻有一次机会尝试对非种子邻居节点传播信息,如果传播成功,则该邻居节点成为t+1时刻的种子节点,而不管该用户在t时刻是否传播成功,该用户再也不能在其他时刻试图传播信息给它的邻居节点。如果在同一时刻,有多个种子节点试图传播信息给同一节点,其传播的顺序可以是任意的。
可选的,所述获取所述用户关系网络中用户与用户之间的传播概率,包括:
根据引入传播概率方差控制因子的传播概率学习模型,获取所述用户关系网络中用户与用户之间的传播概率。
例如,传播概率学习模型可以是EM(Expectation-maximuzation,最大期望)模型。 由于数据的稀疏性,在传播概率学习的过程中,根据EM模型学习到的传播概率往往方差比较大。这主要是因为EM模型的计算方法在稀疏数据情况下过拟合导致数据分布不均匀,很容易估计获得概率为0或概率为1的极端概率情况。
本申请实施例中,为了解决传统EM模型存在的上述问题,在EM模型中引入了传播概率方差控制因子,防止EM模型在迭代过程中发生剧烈的波动。
可选的,所述根据引入传播概率方差控制因子的传播概率学习模型,获取所述用户关系网络中用户与用户之间的传播概率,包括:
获取所述用户关系网络,并根据所述用户关系网络和时间片断数据建立信息传播模型,所述时间片段数据是预设的信息传播扩散时间;
将传播概率方差控制因子引入传播概率学习模型中,得到引入传播概率方差控制因子的传播概率学习模型,并根据所述引入传播概率方差控制因子的传播概率学习模型,对所述信息传播模型进行学习,获取传播概率更新规则,所述更新规则包括第一更新规则和第二更新规则;
采用所述第一更新规则更新第一组用户之间的传播概率,采用所述第二更新规则更新第二组用户之间的传播概率,所述第一组用户之间的边在所述时间片断数据内被激活,所述第二组用户之间的边在所述时间片断数据内没有被激活;
将更新后的用户之间的传播概率确定为所述用户关系网络中用户与用户之间的传播概率。
具体的,如图6所示,获取用户之间的传播概率的流程可以包括:
S61:导入用户关系网络。
例如,从已有的社交网络的应用程序中导入用户关系网络。
S62:建立独立级联模型。
独立级联模型是一种基本的传播模型,可以采用现有的方式根据用户关系网络建立。
在传播模型中,可以包括节点和边,其中,每个节点可以对应用户关系网络中的一个用户,每个边是由用户关系网络中两个相邻用户组成的线段。
S63:将传播概率方差控制因子引入EM模型中。
EM(Expectation-maximuzation,最大化期望)模型是一种优化算法,在本实施例中,可以采用EM模型对独立级联模型进行学习,从而得到独立级联模型中包括的每个边的传播概率,也就是用户关系网络中用户与用户之间的传播概率。
传统的EM模型可以表示为:
Figure PCTCN2016072783-appb-000003
在引入传播概率方差控制因子后,可以根据求解过程是否收敛得到不同的引入传播概率方差控制因子的EM模型,采用哪种引入传播概率方差控制因子的EM模型可以根据实际需要确定,具体的,引入传播概率方差控制因子的EM模型可以是:
Figure PCTCN2016072783-appb-000004
或者,
Figure PCTCN2016072783-appb-000005
其中,λ是控制因子,kv,w是边(v,w)的传播概率。
S64:根据引入传播概率方差控制因子的EM模型获取第一更新规则和第二更新规则。
其中,可以先根据引入λ的EM模型确定优化方程,再对优化方程进行求解,得到第一更新规则。
具体的,如果引入λ的EM模型是:
Figure PCTCN2016072783-appb-000006
其对应的优化方程是:
Figure PCTCN2016072783-appb-000007
对该优化方程进行求解后,得到第一更新规则是:
Figure PCTCN2016072783-appb-000008
其中,
Figure PCTCN2016072783-appb-000009
表示v∈Ds(t),w∈Ds(t+1),
Figure PCTCN2016072783-appb-000010
表示v∈Ds(t),
Figure PCTCN2016072783-appb-000011
Ds(t)表示在t时刻激活的点的集合,Pw(s)表示w被激活的概率。
对该优化方程进行求解后,得到第二更新规则是:
Figure PCTCN2016072783-appb-000012
得到,
Figure PCTCN2016072783-appb-000013
如果引入λ的EM模型是:
Figure PCTCN2016072783-appb-000014
其对应的优化方程是:
Figure PCTCN2016072783-appb-000015
对该优化方程进行求解后,得到第一更新规则是:
Figure PCTCN2016072783-appb-000016
对该优化方程进行求解后,得到第二更新规则是:
Figure PCTCN2016072783-appb-000017
S65:判断时间片断数据是否结束,若否,执行S66,若是,执行S68。
其中,时间片断数据是预设的,用于表明信息传播扩散时间。
在得到第一更新规则和第二更新规则后,可以在用户关系网络中选取种子节点,然后以种子节点为起点根据用户关系网络传播预设信息,传播时间是预设的时间片断数据。
具体的,可以得到当前时间与信息开始传播的时间之间的差值,如果该差值小于预设的时间片断数据,则确定时间片断数据没有结束,否则结束。
S66:判断要计算的边在该时间片断数据内是否被激活,若是,执行S67,否则,重复执行S65及其后续步骤。
例如,要计算的边是用户A与用户B组成的边,在信息传播时间内,传播的信息经过用户A和用户B,则可以确定用户A和用户B组成的边在该时间内被激活了,否则未激活。
S67:采用第一更新规则,对该要计算的边的传播概率进行更新,之后,执行S69。
其中,第一更新规则的具体公式可以参见上述描述。
另外,每个边可以设置初始传播概率。
S68:采用第二更新规则,对整个时间片断数据内未被激活的边的传播概率进行更新,之后,执行S69。
例如,在整个预设的时间片断数据内,用户A和用户C组成的边都没有被激活,也就是信息没有在用户A和用户C之间传播,则可以采用如上所示的第二更新规则对用户A和用户C组成的边的传播概率进行更新。
S69:将每个边更新后的传播概率写入传播概率更新库。
可以理解的是,上述以传播概率学习模型是EM模型为例,传播概率学习模型也可以是其他模型,例如,马尔科夫模型。
本实施例中,通过确定要传播的信息对应的第一用户,第一用户是影响力大于预设值的用户,并由第一用户为起点进行信息传播,可以由具有较大影响力的用户传播信息,提高信息传播的可信性,提高信息传播效率。本实施例通过标签传播学习算法可以确定出第一用户,提高有效性。本实施例通过在传播概率学习模型中引入控制因子,可以提高传播概率的准确性。本实施例通过设置不同的传播策略,可以实现信息多样性传播。
为了实现上述实施例,本发明还提出一种信息传播装置。
图7是本发明另一实施例的信息传播装置的结构示意图。如图7所示,该信息传播装置包括:确定模块100和传播模块200。
具体地,确定模块100用于确定要传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户。其中,要传播的信息可以是商品推广信息,也可以是其他信息,本发明对此不做限定。要传播的信息对应的第一用户可以是一个或多个。
兴趣类型网络可以是根据兴趣类型对用户或信息进行分类标记的网络,也可以称为兴趣网络。
具体地,可以预先建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户。例如,可以预先设置多个标签,按照标签将用户划分到不同的兴趣类型网络中。如图2所示,兴趣类型网络包括标签,例如时尚、户外、商务、运动、旅行、电子等,每个用户都可以对应一个或多个标签。具体建立兴趣类型网络的过程将在后续实施例中介绍。
第一用户是兴趣类型网络中影响力大于预设值的用户,影响力是用户的一种属性,在本实施例中,一个用户的影响力用于衡量该用户传播的信息被其他人接受的难易程度,其中,影响力大的用户传播的信息更容易被其他人接受。
第一用户也可以称为达人。每个兴趣类型网络中达人可以是一个或者多个。
例如,假设第一用户称为达人,如图4所示,达人列表中包括:服装达人,3C达人和家居达人,则如果要传播的信息包括的标签是3C,则该要传播的信息对应的第一用户是3C达人。
传播模块200用于获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。其中,用户关系网络是用于描述用户与用户之间的关联关系的网络,可以直接从已有的社交网络类型的应用程序中获取用户关系网络,在社交网络类型的应用程序中,用户之间可以通过增加好友或者增加关注等方式预先建立用户关系网络。例如,可以先从第一用户的应用程序中获取到第一用户的好友包括第二用户,再在第二用户的应用程序中获取到第二用户的好友包括第三用户,则可以获取到的用户关系网络包括:第一用户—>第二用户—>第三用户。
以所述第一用户为起点的用户关系网络可以从应用程序的已有数据中导入,例如,从社交网络的应用程序中导入以确定出的第一用户为起点的用户关系网络。
例如,如图4所示,假设要传播的信息对应的第一用户是3C达人,从已有数据中获取的以3C达人为起点的用户关系网络是用户关系网络41,则如图4所示,则可以将要传播的信息以3C达人为起点在用户关系网络41中传播。
本实施例中,通过确定要传播的信息对应的第一用户,第一用户是影响力大于预设值的用户,并由第一用户为起点进行信息传播,可以由具有较大影响力的用户传播信息,提高信息传播的可信性,提高信息传播效率。
图8是本发明另一实施例的信息传播装置的结构示意图。如图8所示,该信息传播装置包括:确定模块100、第二获取子模块110、第一确定子模块120、传播模块200、第三获取子模块210、获取单元211、建模单元212、更新单元213、确定单元214、第二确定子模块220、建立模块300、第一获取子模块310和聚类子模块320。其中,建立模块300包括第一获取子模块310和聚类子模块320;确定模块100包括第二获取子模块110和 第一确定子模块120;传播模块200包括第三获取子模块210和第二确定子模块220;第三获取子模块210包括获取单元211、建模单元212、更新单元213和确定单元214。
具体地,建立模块300用于建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户。以要传播的信息是商品的信息为例,建立模块300具体可以包括:
第一获取子模块310,用于根据标签传播学习算法,获取用户-标签矩阵。具体可以包括:
(1)计算得到商品与商品的相似度矩阵W。
商品与商品的相似度矩阵可以用于表示商品之间在用户行为、商品标题和商品属性上的相似度。
其中,用于计算相似度矩阵的商品可以是优质买家处理过的商品,处理具体可以是指购买,浏览,点击,收藏中的一项或者多项,优质买家可以根据优质买家模型确定,例如,将信用等级高或者购买次数多的买家确定为优质买家。具体的,可以获取所有买家的信息,再根据优质买家模型从所有买家中确定出优质买家,再获取优质买家处理过的商品,再根据优质买家处理过的商品中的两两商品计算相似度,得到相似度矩阵W。
更具体的,第一获取子模块310可以通过最小哈希算法对商品(pid,vid)进行哈希映射,得到商品与商品的相似度矩阵,其中,pid是商品的ID(Identity,身份标识),vid是商品属性值的ID,pid和vid通常可以从基础数据表中获取。
(2)计算得到商品-标签信息矩阵F。
其中,商品-标签信息矩阵F中的商品也可以具体是指优质买家处理过的商品,标签是指商品更新后的标签,在获取优质买家处理过的商品后,可以根据每个商品的初始标签,经过迭代过程计算得到商品-标签信息矩阵F,其中,每个商品的初始标签可以是作为商品的一个属性预先记录在数据库中的,从而可以从数据库中获取商品的初始标签。
更具体的,商品-标签信息矩阵F可以根据标签传播学习算法的迭代公式得到,迭代公式如下:
While(F收敛)
F(t+1)=αSF(t)+(1-α)Y
end
其中,要计算的商品-标签信息矩阵F是上述公式中收敛时得到F(t+1),0≤α≤1为预设的加权参数,S是根据上述的商品与商品的相似度矩阵W计算得到的,S=D-1W∈Rn×n
Figure PCTCN2016072783-appb-000018
Figure PCTCN2016072783-appb-000019
Y是初始标签值, F(t)的初始值可以是根据已有的买家信息得到的商品-标签信息矩阵F的初始值,已有的买家信息可以是根据优质买家模型得到的。例如,根据买家的信用等级从多个买家中确定出预设个数的优质买家,再根据优质买家与优质买家对应的购买,点击或收藏的商品可以得到用户-商品信息矩阵V,以及,根据优质买家购买,点击或收藏的商品与商品具有的标签可以得到上述的商品-标签信息矩阵F的初始值。
其中,商品具有的标签可以根据统计或者HITS(Hyperlink-Induced Topic Search,超链诱导主题搜索)排序算法得到。
在获取F的初始值后,可以根据上述迭代公式,在满足迭代收敛条件时得到最终的商品-标签信息矩阵F。
迭代收敛条件可以包括:设置最大迭代次数,当迭代次数达到最大迭代次数时满足迭代收敛条件;或者,根据迭代后的值与迭代前的值的差值,在该差值大于预设阈值时满足迭代收敛条件,例如,||F(t+1)-F(t)||<β时表明满足迭代收敛条件,||F(t+1)-F(t)||表示F(t+1)与F(t)的欧氏距离,β表示预设阈值。
(3)计算得到用户-标签矩阵L。
其中,用户-标签矩阵L中的用户也可以具体是指优质买家,标签是指用户具有的标签,用户具有的标签可以根据用户处理过的商品的更新后的标签确定。
具体的,可以采用如上所示的方式确定出优质买家,以及获取优质买家处理过的商品,以及,从数据库中获取优质买家处理过的商品的初始标签,之后可以根据优质买家处理过的商品以及上述的方式(1)计算得到商品与商品的相似度矩阵W,再根据商品与商品的相似度矩阵W和优质买家处理过的商品具有的初始标签以及上述的方式(2)计算得到商品-标签信息矩阵F,再根据优质买家以及优质买家处理过的商品可以建立用户-商品信息矩阵V,之后在根据上述的V和F采用如下的方式得到用户-标签矩阵L。
更具体的,计算公式可以是:L=V*F,其中V是上述得到的用户-商品信息矩阵,F是上述得到的收敛时的最终的商品-标签信息矩阵。
聚类子模块320,用于对所述用户-标签矩阵进行聚类,得到预设个数的兴趣类型网络,并获取每个兴趣类型网络中的第一用户。在得到用户-标签矩阵L后,可以对该矩阵L进行聚类,例如,预设个数是k个,则可以对矩阵L进行双聚类得到k个类别,每个类别对应一个兴趣类型网络。
在对矩阵L进行聚类得到k个类别后,以每个兴趣类型网络包括一个第一用户为例,每个类别的中心点可以确定为该兴趣类型网络的第一用户。不同兴趣类型网络的 第一用户可以组成列表,该列表可以称为达人列表,达人列表例如表示为:P={p1,p2,…,pk},其中,pi(i=1,2,…,k)是第i个兴趣类型网络中的第一用户,也可以称为达人,pi可以由用户ID和该用户具有的标签组成。
在上述预先建立多个兴趣类型网络,并确定每个兴趣类型网络中的第一用户后,如上所述,可以得到由不同的兴趣类型网络中的第一用户组成的达人列表,达人列表中包括不同兴趣类型网络中的第一用户,在当前需要传播信息时,可以首先确定要传播的信息对应的第一用户。
所述确定模块100具体包括:
第二获取子模块110用于获取第一标签,所述第一标签是所述要传播的信息包括的标签;
第一确定子模块120用于将包括所述第一标签的第一用户,确定为所述要传播的信息对应的第一用户。
例如,假设第一用户称为达人,如图4所示,达人列表中包括:服装达人,3C达人和家居达人,则如果第二获取子模块110获取到要传播的信息包括的标签是3C,则第一确定子模块120确定该要传播的信息对应的第一用户是3C达人。
传播模块200还用于根据预设策略,在所述用户关系网络中以所述第一用户为起点传播所述信息,所述预设策略包括传播范围策略,或者,传播速度策略。其中,传播范围策略是指优先考虑传播范围,传播速度策略是指优先考虑传播速度。
更具体的,第三获取子模块210可以获取用户关系网络中用户与用户之间的传播概率,当采用传播范围策略时,可以不论传播概率高低都进行信息传播,当采用传播速度策略时,可以只在传播概率大于预设值的路径上进行信息传播。例如,以传播速度策略为例,参见图5,假设用户关系网络包括第一路径51,第二路径52,第三路径53,第四路径54和第五路径55,假设第一路径51,第二路径52和第三路径53中包括的用户之间的传播概率都大于预设值,而第四路径54和第五路径55上包括的用户之间的存在小于预设值的传播概率,则信息可以在第一路径51,第二路径52和第三路径53上传播,而不在第四路径54和第五路径55上传播。
更具体的,信息在用户关系网络中传播时,第一用户在初始时刻作为信息传播的种子节点,种子节点负责向其邻居节点传播信息,例如,第一用户是3C达人,与3C达人相邻的邻居节点包括第一节点和第二节点,则在初始时刻t设置3C达人是种子节点,并且由3C达人将信息传播给第一节点和第二节点,当种子节点将信息传播给邻居节点后,邻居节点在下一时间成为新的种子节点,例如,在t+1时刻种子节点是第一节点,而不再是3C达人, 依次类推,从初始的第一用户依次根据用户关系网络中的用户相邻关系进行信息传播,直至没有新的种子节点。另外,用户关系网络中邻居节点用户之间的传播概率是独立的,不受其他邻居节点之间的关系影响。并且,每个种子节点只有一次机会向非种子邻居节点传播信息,例如,用户在t时刻成为种子节点,仅在t时刻有一次机会尝试对非种子邻居节点传播信息,如果传播成功,则该邻居节点成为t+1时刻的种子节点,而不管该用户在t时刻是否传播成功,该用户再也不能在其他时刻试图传播信息给它的邻居节点。如果在同一时刻,有多个种子节点试图传播信息给同一节点,其传播的顺序可以是任意的。
可选地,第三获取子模块210还用于根据引入传播概率方差控制因子的传播概率学习模型,获取所述用户关系网络中用户与用户之间的传播概率。例如,传播概率学习模型可以是EM(Expectation-maximuzation,最大化期望)模型。由于数据的稀疏性,在传播概率学习的过程中,根据EM模型学习到的传播概率往往方差比较大。这主要是因为EM模型的计算方法在稀疏数据情况下过拟合导致数据分布不均匀,很容易估计获得概率为0或概率为1的极端概率情况。
本申请实施例中,为了解决传统EM模型存在的上述问题,在EM模型中引入了传播概率方差控制因子,防止EM模型在迭代过程中发生剧烈的波动。
可选的,所述第三获取子模块210,包括:
获取单元211用于获取所述用户关系网络,例如,从已有的社交网络的应用程序中导入用户关系网络,并根据所述用户关系网络和时间片断数据建立信息传播模型,例如,可以建立独立级联模型。独立级联模型是一种基本的传播模型,可以采用现有的方式根据用户关系网络建立。
其中,时间片断数据是预设的,用于表明信息传播扩散时间。
在传播模型中,可以包括节点和边,其中,每个节点可以对应用户关系网络中的一个用户,每个边是由用户关系网络中两个相邻用户组成的线段。
建模单元212用于将传播概率方差控制因子引入传播概率学习模型中,得到引入传播概率方差控制因子的传播概率学习模型,并根据所述引入传播概率方差控制因子的传播概率学习模型,对所述信息传播模型进行学习,获取传播概率更新规则,所述更新规则包括第一更新规则和第二更新规则。
EM(Expectation-maximuzation,最大化期望)模型是一种优化算法,在本实施例中,可以采用EM模型对独立级联模型进行学习,从而得到独立级联模型中包括的每个边的传播概率,也就是用户关系网络中用户与用户之间的传播概率。
传统的EM模型可以表示为:
Figure PCTCN2016072783-appb-000020
在引入传播概率方差控制因子后,可以根据求解过程是否收敛得到不同的引入传播概率方差控制因子的EM模型,采用哪种引入传播概率方差控制因子的EM模型可以根据实际需要确定,具体的,引入传播概率方差控制因子的EM模型可以是:
Figure PCTCN2016072783-appb-000021
或者,
Figure PCTCN2016072783-appb-000022
其中,λ是控制因子,kv,w是边(v,w)的传播概率。
其中,可以先根据引入λ的EM模型确定优化方程,再对优化方程进行求解,得到第一更新规则。
具体的,如果引入λ的EM模型是:
Figure PCTCN2016072783-appb-000023
其对应的优化方程是:
Figure PCTCN2016072783-appb-000024
对该优化方程进行求解后,得到第一更新规则是:
Figure PCTCN2016072783-appb-000025
其中,
Figure PCTCN2016072783-appb-000026
表示v∈Ds(t),w∈Ds(t+1),
Figure PCTCN2016072783-appb-000027
表示v∈Ds(t),
Figure PCTCN2016072783-appb-000028
Ds(t)表示在t时刻激活的点的集合,Pw(s)表示w被激活的概率。
对该优化方程进行求解后,得到第二更新规则是:
Figure PCTCN2016072783-appb-000029
得到,
Figure PCTCN2016072783-appb-000030
如果引入λ的EM模型是:
Figure PCTCN2016072783-appb-000031
其对应的优化方程是:
Figure PCTCN2016072783-appb-000032
对该优化方程进行求解后,得到第一更新规则是:
Figure PCTCN2016072783-appb-000033
对该优化方程进行求解后,得到第二更新规则是:
Figure PCTCN2016072783-appb-000034
更新单元213用于采用所述第一更新规则更新第一组用户之间的传播概率,采用所述第二更新规则更新第二组用户之间的传播概率,所述第一组用户之间的边在所述时间片断 数据内被激活,所述第二组用户之间的边在所述时间片断数据内没有被激活。在得到第一更新规则和第二更新规则后,可以在用户关系网络中选取种子节点,然后以种子节点为起点根据用户关系网络传播预设信息,传播时间是预设的时间片断数据。更具体地,可以判断时间片断数据是否结束,例如,可以获取当前时间与信息开始传播的时间之间的差值,如果该差值小于预设的时间片断数据,则确定时间片断数据没有结束,否则结束。
若时间片断数据没有结束,则可以判断要计算的边在该时间片断数据内是否被激活,例如,要计算的边是用户A与用户B组成的边,在信息传播时间内,传播的信息经过用户A和用户B,则可以确定用户A和用户B组成的边在该时间内被激活了,否则未激活。如果被激活,则采用第一更新规则,对该要计算的边的传播概率进行更新,之后,将每个边更新后的传播概率写入传播概率更新库。如果未被激活,则返回继续判断时间片断数据是否结束。
若时间片断数据已结束,则采用第二更新规则,对整个时间片断数据内未被激活的边的传播概率进行更新,之后,将每个边更新后的传播概率写入传播概率更新库。
可以理解的是,上述以传播概率学习模型是EM模型为例,传播概率学习模型也可以是其他模型,例如,马尔科夫模型。
确定单元214用于将更新后的用户之间的传播概率确定为所述用户关系网络中用户与用户之间的传播概率。
第二确定子模块220用于将所述传播概率大于预设值的路径确定为传播路径,并根据所述传播路径传播所述信息,以实现最大的传播速度。
本实施例中,通过确定要传播的信息对应的第一用户,第一用户是影响力大于预设值的用户,并由第一用户为起点进行信息传播,可以由具有较大影响力的用户传播信息,提高信息传播的可信性,提高信息传播效率。本实施例通过标签传播学习算法可以确定出第一用户,提高有效性。本实施例通过在传播概率学习模型中引入控制因子,可以提高传播概率的准确性。本实施例通过设置不同的传播策略,可以实现信息多样性传播。
需要说明的是,在本发明的描述中,术语“第一”、“第二”等仅用于描述目的,而不能理解为指示或暗示相对重要性。此外,在本发明的描述中,除非另有说明,“多个”的含义是两个或两个以上。
流程图中或在此以其他方式描述的任何过程或方法描述可以被理解为,表示包括一个或更多个用于实现特定逻辑功能或过程的步骤的可执行指令的代码的模块、片段或部分,并且本发明的优选实施方式的范围包括另外的实现,其中可以不按所示出或讨论的顺序,包括根据所涉及的功能按基本同时的方式或按相反的顺序,来执行功能,这应被本发明的 实施例所属技术领域的技术人员所理解。
应当理解,本发明的各部分可以用硬件、软件、固件或它们的组合来实现。在上述实施方式中,多个步骤或方法可以用存储在存储器中且由合适的指令执行系统执行的软件或固件来实现。例如,如果用硬件来实现,和在另一实施方式中一样,可用本领域公知的下列技术中的任一项或他们的组合来实现:具有用于对数据信号实现逻辑功能的逻辑门电路的离散逻辑电路,具有合适的组合逻辑门电路的专用集成电路,可编程门阵列(PGA),现场可编程门阵列(FPGA)等。
本技术领域的普通技术人员可以理解实现上述实施例方法携带的全部或部分步骤是可以通过程序来指令相关的硬件完成,所述的程序可以存储于一种计算机可读存储介质中,该程序在执行时,包括方法实施例的步骤之一或其组合。
此外,在本发明各个实施例中的各功能单元可以集成在一个处理模块中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。所述集成的模块如果以软件功能模块的形式实现并作为独立的产品销售或使用时,也可以存储在一个计算机可读取存储介质中。
上述提到的存储介质可以是只读存储器,磁盘或光盘等。
在本说明书的描述中,参考术语“一个实施例”、“一些实施例”、“示例”、“具体示例”、或“一些示例”等的描述意指结合该实施例或示例描述的具体特征、结构、材料或者特点包含于本发明的至少一个实施例或示例中。在本说明书中,对上述术语的示意性表述不一定指的是相同的实施例或示例。而且,描述的具体特征、结构、材料或者特点可以在任何的一个或多个实施例或示例中以合适的方式结合。
尽管上面已经示出和描述了本发明的实施例,可以理解的是,上述实施例是示例性的,不能理解为对本发明的限制,本领域的普通技术人员在本发明的范围内可以对上述实施例进行变化、修改、替换和变型。

Claims (14)

  1. 一种信息传播方法,其特征在于,包括:
    确定要传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户;
    获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。
  2. 根据权利要求1所述的方法,其特征在于,还包括:
    建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户;
    所述建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户,包括:
    根据标签传播学习算法,获取用户-标签矩阵;
    对所述用户-标签矩阵进行聚类,得到预设个数的兴趣类型网络,并获取每个兴趣类型网络中的第一用户。
  3. 根据权利要求2所述的方法,其特征在于,所述第一用户的标识信息包括:用户标识和标签,兴趣类型网络包括标签,所述确定要传播的信息对应的第一用户,包括:
    获取第一标签,所述第一标签是所述要传播的信息包括的标签;
    将包括所述第一标签的第一用户,确定为所述要传播的信息对应的第一用户。
  4. 根据权利要求1所述的方法,其特征在于,所述在所述用户关系网络中以所述第一用户为起点传播所述信息,包括:
    根据预设策略,在所述用户关系网络中以所述第一用户为起点传播所述信息,所述预设策略包括传播范围策略,或者,传播速度策略。
  5. 根据权利要求4所述的方法,其特征在于,当所述预设策略是传播速度策略时,所述根据预设策略,在所述用户关系网络中以所述第一用户为起点传播所述信息,包括:
    获取所述用户关系网络中用户与用户之间的传播概率;
    将所述传播概率大于预设值的路径确定为传播路径,并根据所述传播路径传播所述信息。
  6. 根据权利要求5所述的方法,其特征在于,所述获取所述用户关系网络中用户与用户之间的传播概率,包括:
    根据引入传播概率方差控制因子的传播概率学习模型,获取所述用户关系网络中用户与用户之间的传播概率。
  7. 根据权利要求6所述的方法,其特征在于,所述根据引入传播概率方差控制因子的 传播概率学习模型,获取所述用户关系网络中用户与用户之间的传播概率,包括:
    获取所述用户关系网络,并根据所述用户关系网络和时间片断数据建立信息传播模型,所述时间片段数据是预设的信息传播扩散时间;
    将传播概率方差控制因子引入传播概率学习模型中,得到引入传播概率方差控制因子的传播概率学习模型,并根据所述引入传播概率方差控制因子的传播概率学习模型,对所述信息传播模型进行学习,获取传播概率更新规则,所述更新规则包括第一更新规则和第二更新规则;
    采用所述第一更新规则更新第一组用户之间的传播概率,采用所述第二更新规则更新第二组用户之间的传播概率,所述第一组用户之间的边在所述时间片断数据内被激活,所述第二组用户之间的边在所述时间片断数据内没有被激活;
    将更新后的用户之间的传播概率确定为所述用户关系网络中用户与用户之间的传播概率。
  8. 一种信息传播装置,其特征在于,包括:
    确定模块,用于确定要传播的信息对应的第一用户,所述第一用户是所述第一用户属于的兴趣类型网络中影响力大于预设值的用户;
    传播模块,用于获取以所述第一用户为起点的用户关系网络,在所述用户关系网络中以所述第一用户为起点传播所述信息。
  9. 根据权利要求8所述的装置,其特征在于,还包括:
    建立模块,用于建立预设个数的兴趣类型网络,并在每个兴趣类型网络中,确定对应的第一用户;
    所述建立模块,包括:
    第一获取子模块,用于根据标签传播学习算法,获取用户-标签矩阵;
    聚类子模块,用于对所述用户-标签矩阵进行聚类,得到预设个数的兴趣类型网络,并获取每个兴趣类型网络中的第一用户。
  10. 根据权利要求9所述的装置,其特征在于,所述第一用户的标识信息包括:用户标识和标签,兴趣类型网络包括标签,所述确定模块,包括:
    第二获取子模块,用于获取第一标签,所述第一标签是所述要传播的信息包括的标签;
    第一确定子模块,用于将包括所述第一标签的第一用户,确定为所述要传播的信息对应的第一用户。
  11. 根据权利要求8所述的装置,其特征在于,所述传播模块还用于根据预设策略,在所述用户关系网络中以所述第一用户为起点传播所述信息,所述预设策略包括传播范围策略,或者,传播速度策略。
  12. 根据权利要求11所述的装置,其特征在于,当所述预设策略是传播速度策略时,所述传播模块,包括:
    第三获取子模块,用于获取所述用户关系网络中用户与用户之间的传播概率;
    第二确定子模块,用于将所述传播概率大于预设值的路径确定为传播路径,并根据所述传播路径传播所述信息。
  13. 根据权利要求12所述的装置,其特征在于,所述第三获取子模块还用于根据引入传播概率方差控制因子的传播概率学习模型,获取所述用户关系网络中用户与用户之间的传播概率。
  14. 根据权利要求13所述的装置,其特征在于,所述第三获取子模块,包括:
    获取单元,用于获取所述用户关系网络,并根据所述用户关系网络和时间片断数据建立信息传播模型,所述时间片段数据是预设的信息传播扩散时间;
    建模单元,用于将传播概率方差控制因子引入传播概率学习模型中,得到引入传播概率方差控制因子的传播概率学习模型,并根据所述引入传播概率方差控制因子的传播概率学习模型,对所述信息传播模型进行学习,获取传播概率更新规则,所述更新规则包括第一更新规则和第二更新规则;
    更新单元,用于采用所述第一更新规则更新第一组用户之间的传播概率,采用所述第二更新规则更新第二组用户之间的传播概率,所述第一组用户之间的边在所述时间片断数据内被激活,所述第二组用户之间的边在所述时间片断数据内没有被激活;
    确定单元,用于将更新后的用户之间的传播概率确定为所述用户关系网络中用户与用户之间的传播概率。
PCT/CN2016/072783 2015-02-04 2016-01-29 信息传播方法和装置 Ceased WO2016124116A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2017541018A JP2018511851A (ja) 2015-02-04 2016-01-29 情報伝達方法及び装置
US15/662,188 US20170323313A1 (en) 2015-02-04 2017-07-27 Information propagation method and apparatus

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201510058167.8 2015-02-04
CN201510058167.8A CN105991397B (zh) 2015-02-04 2015-02-04 信息传播方法和装置

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US15/662,188 Continuation US20170323313A1 (en) 2015-02-04 2017-07-27 Information propagation method and apparatus

Publications (1)

Publication Number Publication Date
WO2016124116A1 true WO2016124116A1 (zh) 2016-08-11

Family

ID=56563459

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/072783 Ceased WO2016124116A1 (zh) 2015-02-04 2016-01-29 信息传播方法和装置

Country Status (4)

Country Link
US (1) US20170323313A1 (zh)
JP (1) JP2018511851A (zh)
CN (1) CN105991397B (zh)
WO (1) WO2016124116A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113902578A (zh) * 2021-10-21 2022-01-07 南京邮电大学 一种动态社交网络中商品传播最大化方法、系统、装置及存储介质
CN117611374A (zh) * 2024-01-23 2024-02-27 深圳博十强志科技有限公司 一种基于多元化大数据分析的信息传播分析方法及系统
CN120804739A (zh) * 2025-09-15 2025-10-17 中国科学技术大学 一种基于圈图随机游走的高影响力信息传播者识别方法

Families Citing this family (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20180018709A1 (en) * 2016-05-31 2018-01-18 Ramot At Tel-Aviv University Ltd. Information spread in social networks through scheduling seeding methods
CN108322316B (zh) * 2017-01-17 2021-10-19 阿里巴巴(中国)有限公司 确定信息传播热度的方法、装置及计算设备
CN107566244B (zh) * 2017-07-24 2019-05-24 平安科技(深圳)有限公司 一种网络账户的选取方法及其设备
CN108734380B (zh) * 2018-04-08 2022-02-01 创新先进技术有限公司 风险账户判定方法、装置及计算设备
CN109583620B (zh) * 2018-10-11 2024-03-01 平安科技(深圳)有限公司 企业潜在风险预警方法、装置、计算机设备和存储介质
CN110932909B (zh) * 2019-12-05 2022-02-18 中国传媒大学 信息传播预测方法、系统及存储介质
CN111159437B (zh) * 2019-12-26 2023-08-22 中国传媒大学 影视作品的传播结果及类型的预测方法及系统
CN111882343A (zh) * 2020-06-12 2020-11-03 智云众(北京)信息技术有限公司 基于达人价值指数的广告投放方法、装置及设备
CN111814065B (zh) * 2020-06-24 2022-05-06 平安科技(深圳)有限公司 信息传播路径分析方法、装置、计算机设备及存储介质
CN112511411A (zh) * 2020-12-07 2021-03-16 郁剑 一种5g背景下新媒体影像的视觉传播方法
CN113222774B (zh) * 2021-04-19 2023-05-23 浙江大学 社交网络种子用户选择方法和装置、电子设备、存储介质
CN114254251B (zh) * 2021-05-31 2025-04-25 大连交通大学 直接免疫scir的舆情传播模型构建方法
CN115660089B (zh) * 2022-09-16 2025-12-12 北海淇昂信息科技有限公司 基于概率分布的层级增量式标签传播方法及装置
CN117151914B (zh) * 2023-11-01 2024-01-30 中国人民解放军国防科技大学 基于综合影响力评估的群智感知用户选择方法和装置

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103177382A (zh) * 2013-03-19 2013-06-26 武汉大学 微博平台上的关键传播路径和中心节点的探测方法
CN103279512A (zh) * 2013-05-17 2013-09-04 湖州师范学院 利用社会网络上最有影响力节点实现高效病毒营销的方法
CN103412872A (zh) * 2013-07-08 2013-11-27 西安交通大学 一种基于有限节点驱动的微博社会网络信息推荐方法

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120001919A1 (en) * 2008-10-20 2012-01-05 Erik Lumer Social Graph Based Recommender
US8312056B1 (en) * 2011-09-13 2012-11-13 Xerox Corporation Method and system for identifying a key influencer in social media utilizing topic modeling and social diffusion analysis
WO2013126144A2 (en) * 2012-02-20 2013-08-29 Aptima, Inc. Systems and methods for network pattern matching
WO2013175410A1 (en) * 2012-05-22 2013-11-28 Thakker Mitesh L Systems and methods for authenticating, tracking, and rewarding word of mouth propagation
US9247020B2 (en) * 2012-08-07 2016-01-26 Google Inc. Media content receiving device and distribution of media content utilizing social networks and social circles
US20140115010A1 (en) * 2012-10-18 2014-04-24 Google Inc. Propagating information through networks
CN103064917B (zh) * 2012-12-20 2016-08-17 中国科学院深圳先进技术研究院 一种面向微博的特定倾向的高影响力用户群发现方法
CN103106616B (zh) * 2013-02-27 2016-01-20 中国科学院自动化研究所 基于资源整合与信息传播特征的社区发现及演化方法
CN103678669B (zh) * 2013-12-25 2017-02-08 福州大学 一种社交网络中的社区影响力评估系统及方法

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103177382A (zh) * 2013-03-19 2013-06-26 武汉大学 微博平台上的关键传播路径和中心节点的探测方法
CN103279512A (zh) * 2013-05-17 2013-09-04 湖州师范学院 利用社会网络上最有影响力节点实现高效病毒营销的方法
CN103412872A (zh) * 2013-07-08 2013-11-27 西安交通大学 一种基于有限节点驱动的微博社会网络信息推荐方法

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113902578A (zh) * 2021-10-21 2022-01-07 南京邮电大学 一种动态社交网络中商品传播最大化方法、系统、装置及存储介质
CN117611374A (zh) * 2024-01-23 2024-02-27 深圳博十强志科技有限公司 一种基于多元化大数据分析的信息传播分析方法及系统
CN117611374B (zh) * 2024-01-23 2024-05-07 深圳博十强志科技有限公司 一种基于多元化大数据分析的信息传播分析方法及系统
CN120804739A (zh) * 2025-09-15 2025-10-17 中国科学技术大学 一种基于圈图随机游走的高影响力信息传播者识别方法
CN120804739B (zh) * 2025-09-15 2025-12-30 中国科学技术大学 一种基于圈图随机游走的高影响力信息传播者识别方法

Also Published As

Publication number Publication date
CN105991397B (zh) 2020-03-03
JP2018511851A (ja) 2018-04-26
CN105991397A (zh) 2016-10-05
US20170323313A1 (en) 2017-11-09

Similar Documents

Publication Publication Date Title
CN105991397B (zh) 信息传播方法和装置
US11748379B1 (en) Systems and methods for generating and implementing knowledge graphs for knowledge representation and analysis
CN111382283B (zh) 资源类别标签标注方法、装置、计算机设备和存储介质
WO2021164390A1 (zh) 冷链配送的路线确定方法、装置、服务器及存储介质
US10459996B2 (en) Big data based cross-domain recommendation method and apparatus
US20150242447A1 (en) Identifying effective crowdsource contributors and high quality contributions
WO2017080176A1 (zh) 个体用户画像方法和系统
Xiao et al. A truth discovery approach with theoretical guarantee
CN107451894A (zh) 数据处理方法、装置和计算机可读存储介质
CN106126669A (zh) 基于标签的用户协同过滤内容推荐方法及装置
US12456031B2 (en) Neural architecture search via similarity-based operator ranking
CN115293919A (zh) 面向社交网络分布外泛化的图神经网络预测方法及系统
CN112989169A (zh) 目标对象识别方法、信息推荐方法、装置、设备及介质
CN112380433A (zh) 面向冷启动用户的推荐元学习方法
Chan et al. Continuous model selection for large-scale recommender systems
CN110348947A (zh) 对象推荐方法及装置
Salehi et al. KATZ centrality with biogeography-based optimization for influence maximization problem
US9477757B1 (en) Latent user models for personalized ranking
CN115880024A (zh) 一种基于预训练的轻量化图卷积神经网络的推荐方法
Mokhtari et al. Community detection by influential nodes based on random walk distance
CN115345635A (zh) 推荐内容的处理方法、装置、计算机设备和存储介质
KR102144122B1 (ko) 적합도 기반의 온라인 광고 효과 계산 방법 및 장치
Chen et al. From tie strength to function: Home location estimation in social network
HK1229578B (zh) 信息传播方法和装置
Goyal Social influence and its applications: An algorithmic and data mining study

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16746111

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2017541018

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16746111

Country of ref document: EP

Kind code of ref document: A1