WO2022013909A1 - グラフ計算装置、グラフ計算方法及びグラフ計算プログラム - Google Patents
グラフ計算装置、グラフ計算方法及びグラフ計算プログラム Download PDFInfo
- Publication number
- WO2022013909A1 WO2022013909A1 PCT/JP2020/027218 JP2020027218W WO2022013909A1 WO 2022013909 A1 WO2022013909 A1 WO 2022013909A1 JP 2020027218 W JP2020027218 W JP 2020027218W WO 2022013909 A1 WO2022013909 A1 WO 2022013909A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- node
- pair
- evaluation value
- nodes
- main
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
Definitions
- This disclosure relates to a graph calculation device, a graph calculation method, and a graph calculation program.
- the b-matching graph is known as a graph calculated from multidimensional data.
- the b-matching graph is a graph in which b edges are connected to each of a plurality of nodes which are multidimensional data.
- Non-Patent Document 1 discloses a method for obtaining a b-matching graph based on a belief propagation method.
- Non-Patent Document 1 updates the message related to the belief propagation method for all node pairs selected from N nodes represented by M-dimensional vectors until they converge. Therefore, if the number of repetitions until convergence is set to T, it can be seen that the calculation cost is as high as O (N 2 M 2 T).
- An object of the present disclosure is to provide a graph calculation device, a graph calculation method, and a graph calculation program for calculating a b-matching graph while suppressing a calculation cost.
- the graph computing device is selected based on the distance from the main node among the N nodes, with the node as the main node (b + 1).
- the initialization unit that generates a node pair set consisting of a plurality of node pairs having more than (N-1) nodes as slave nodes, and for each node pair related to the node pair set, the evaluation value related to the node pair is set to the node pair.
- the reference evaluation value which is the b-smallest evaluation value among the evaluation values of other node pairs having the slave node of the node pair as the main node, and the distance of the node pair related to the reference evaluation value. It is provided with an update unit that updates until the convergence condition of is satisfied, and an edge determination unit that determines b slave nodes connected by an edge for each main node based on the updated evaluation value.
- the graph calculation device can calculate the b-matching graph while suppressing the calculation cost.
- FIG. 1 is a schematic block diagram showing the configuration of the graph calculation device 10 according to the first embodiment.
- the graph calculation device 10 calculates a b-matching graph in which b edges having weights are connected to each node for N nodes which are M-dimensional vectors.
- the graph calculation device 10 includes an input unit 11, a dimension reduction unit 12, a vector storage unit 13, an initialization unit 14, a pair storage unit 15, an update unit 16, a continuation determination unit 17, an additional unit 18, an edge determination unit 19, and a weight calculation.
- a unit 20 and an output unit 21 are provided.
- the input unit 11 receives the input of N nodes to be calculated in the graph.
- N and M are natural numbers.
- the vector input to the input unit 11 is also referred to as a raw vector.
- the dimension reduction unit 12 reduces the dimensions of each low vector (x 1 to x N ) input to the input unit 11.
- the dimension reduction unit 12 reduces the dimension of each node by, for example, Singular Value Decomposition.
- the low-dimensional vector which is the vector after the dimension reduction, is an m-dimensional vector. m is a natural number less than M.
- the dimension reduction unit 12 generates a low-dimensional vector by the following procedure. First, the dimension reduction unit 12 obtains a unitary matrix, a singular value matrix, and a conjugate matrix by singular value decomposition of an N ⁇ M node matrix X in which N low vectors (x 1 to x N) are put together.
- the dimension reduction unit 12 extracts an N ⁇ m submatrix from the unitary matrix and an m ⁇ m submatrix from the singular value matrix, and multiplies them to generate an N ⁇ m low-dimensional matrix.
- the dimension reduction unit 12 extracts each row vector of the low-dimensional matrix as a low-dimensional vector (x to 1 to x to N : where x to is a tilde on top of x). Further, the dimension reduction unit 12 calculates the residual r representing the root mean square error of the low vector and the low dimension vector by the following equation (1).
- the vector storage unit 13 stores the low vector, the low-dimensional vector, and the residual of the node in association with each of the N nodes.
- the initialization unit 14 extracts a node pair to be calculated by the belief propagation method from N nodes, and generates an initial value of a message for the extracted node pair.
- the message of the belief propagation method is an example of the evaluation value.
- the pair storage unit 15 stores the message of the node pair in association with each of the node pairs to be calculated.
- the update unit 16 updates the message related to each node pair stored in the pair storage unit 15.
- the continuation determination unit 17 determines whether or not the update of the message by the update unit 16 has converged.
- the addition unit 18 determines whether or not to add the exclusion pair, which is a node pair in which the pair storage unit 15 is not stored, to the pair storage unit 15.
- the edge determination unit 19 determines a node pair to be connected at the edge based on the message stored in the pair storage unit 15.
- the weight calculation unit 20 gradually calculates the edge weight based on the message stored in the pair storage unit 15.
- the output unit 21 outputs a b-matching graph based on the calculation results of the edge determination unit 19 and the weight calculation unit 20.
- FIG. 2 is a flowchart showing a calculation method of the b-matching graph according to the first embodiment.
- the input unit 11 receives inputs of N nodes (x 1 to x N ) to be calculated and a low vector related to each node (step S1).
- the dimension reduction unit 12 calculates N low-dimensional vectors (x to 1 to x to N) and a residual r from the N low vectors acquired in step S1 (step S2).
- the dimension reduction unit 12 records the low vector, the low dimension vector, and the residual in the vector storage unit 13 in association with the node (step S3).
- FIG. 3 is a flowchart showing a method of initializing a node pair and a message according to the first embodiment.
- the initialization unit 14 initializes a set of node pairs for which the message is calculated. Specifically, the initialization unit 14 secures a storage area of a node pair set having the node as the main node for each of the N nodes in the pair storage unit 15 (step S101).
- the main node for convenience.
- all node pair sets are empty sets.
- the initialization unit 14 selects N nodes one by one (step S102), and executes the processes of steps S103 to S110 below.
- the initialization unit 14 extracts a node set consisting of (N-1) nodes (candidate nodes) other than the selected node as a candidate node set (step S103).
- the initialization unit 14 calculates the distance between the selected node and each candidate node based on the low vector stored in the vector storage unit 13 by the following equation (2) (step S104).
- the initialization unit 14 identifies the candidate node having the shortest distance from the main node from the candidate node set (step S105).
- the initialization unit 14 refers to the pair storage unit 15 and determines whether or not the number of node pairs having the candidate node as the main node is less than (b + 1) (step S106).
- step S106 determines whether or not the number of node pairs having the candidate node as the main node is less than (b + 1) (step S106: YES)
- the initialization unit 14 sets the node selected in step S102 as the main node and sets the candidate node specified in step S105 as the main node.
- the node pair to be a slave node is recorded in the pair storage unit 15 (step S107).
- the initialization unit 14 records in the pair storage unit 15 a node pair in which the candidate node specified in step S105 is the main node and the node selected in step S102 is the slave node (step S108). At this time, the initialization unit 14 records a message having an initial value of 1 in the pair storage unit 15 in association with each node pair.
- step S106 When the number of node pairs having the candidate node as the main node is (b + 1) or more (step S106: NO), or when the node pair is recorded in the pair storage unit 15 (steps S107 and S108), the initialization unit 14 is a candidate.
- the candidate node specified in step S105 is deleted from the node set (step S109).
- the initialization unit 14 determines whether or not the number of node pairs having the node selected in step S102 as the main node is less than (b + 1) (step S110).
- step S110 When the number of node pairs having the node selected in step S102 as the main node is less than (b + 1) (step S110: YES), the initialization unit 14 returns the process to step S105.
- step S110 On the other hand, when the number of node pairs having the node selected in step S102 as the main node is (b + 1) or more (step S110: NO), the initialization unit 14 returns to step S102 and selects the next node.
- the initialization unit 14 can initialize the node pair set and the message. As described above, in the initialization unit 14, N (b + 1) node pairs are recorded in the pair storage unit 15. Since the number of all node pairs that can be specified from N nodes is N (N-1), the initialization unit 14 excludes (N-b-2) node pairs from the calculation target. You can see that. As a result, the graph calculation device 10 can reduce the calculation cost of the b-matching graph.
- the update unit 16 stores each message stored in the pair storage unit 15. Update (step S5).
- FIG. 4 is a flowchart showing a method of updating a message according to the first embodiment.
- the update unit 16 selects N nodes stored by the pair storage unit 15 one by one (step S201), and performs the processes from step S202 to step S204 below.
- the update unit 16 has a node pair (x i , x j ) whose main node is the node x i selected in step S201, and has a distance D of the node pair whose base is e, as shown in the following equation (2).
- the weighted message m w [x i , x j ] is obtained (step S202).
- the update unit 16 has the b-th smallest ⁇ [x i ] and the (b + 1) th-smallest ⁇ . identifying a [x i] (step S203).
- the b-th smallest weighted message related to the main node x i is referred to as a lower reference message ⁇ [x i ] related to the node x i.
- the (b + 1) smallest weighted message related to the main node x i is called an upper reference message ⁇ [x i ] related to the node x i.
- the update unit 16 specifies the slave node v [x i ] related to the lower reference message ⁇ [x i ] and the slave node w [x i ] related to the upper reference message ⁇ [x i ] (step S204). ..
- a node x i the lower reference message mu [x i] slave node according to according to referred to as a lower reference node v [i] according to the node x i.
- the update unit 16 selects the node pairs stored in the pair storage unit 15 one by one (step S205), and processes the following steps S206 to S208.
- the update unit 16 determines whether or not the main node x i related to the node pair selected in step S8 is the lower reference node v [j] in the slave node x j related to the node pair (step S206). When the main node x i is not the lower reference node v [j] (step S206: NO), the update unit 16 sends a message related to the node pair selected in step S205 to the bottom e as shown in the equation (4).
- the exponential function of the distance D [i, j] of the node pair is rewritten to the value divided by the lower reference message ⁇ [x j ] related to the slave node x j (step S207). If the node x i is not the lower reference node v [j], the lower reference message ⁇ [x j ] is the b-th smallest evaluation value among other node pairs having the node x j as the main node. It can be said that it is the product of the value and the exponential function of the distance of the node pair with e as the base.
- step S206 when the main node x i is the lower reference node v [j] (step S206: YES), the update unit 16 sends a message related to the node pair selected in step S8 as shown in the equation (5).
- the exponential function of the distance D [i, j] of the node pair having e as the base is rewritten to the value divided by the upper reference message ⁇ [x j ] related to the slave node x j (step S208).
- the upper reference message ⁇ [x j ] is the b-th smallest evaluation value among other node pairs having the node x j as the main node. It can be said that it is the product of the evaluation value and the exponential function of the distance of the node pair with e as the base.
- the update unit 16 can initialize the candidate node set and the message. Before updating the message in steps S205 to S208, the update unit 16 specifies the lower reference message, the upper reference message, the lower reference node, and the upper reference node for each node in advance in steps S201 to S204. .. As a result, the update unit 16 does not need to calculate the b-th smallest evaluation value among the evaluation values related to other node pairs having the slave node x j of the node pair related to the message to be updated as the main node, so that the calculation is performed. The cost can be reduced.
- step S6 determines whether or not the message has converged.
- the continuation determination unit 17 determines that the message has converged, for example, when the root mean square of the amount of change in the message is less than a predetermined value.
- the message converges after repeating O (N) times. If the message has not converged (step S6: NO), the graph calculation device 10 returns the process to step S5 and repeats updating the message.
- step S6 when the message has converged (step S6: YES), the additional unit 18 is a node pair that is not stored by the pair storage unit 15 among all the node pairs that can be specified from the N nodes, that is, a node pair that is not calculated.
- an exclusion pair one by one (step S7), and the processing of the following steps S8 to S12 is executed.
- the additional unit 18 is a low-dimensional vector x to i of the main node and a low dimension of the residual r [x i ] and the subordinate node related to the exclusion pair (x i , x j ) selected in step S7 from the vector storage unit 13.
- the vectors x to j and the residual r [x j ] are read out, and the approximate distances D to [i, j] are calculated by the equation (6) (step S8).
- the additional unit 18 includes an upper reference message ⁇ [x i ] related to the main node x i , an upper reference message ⁇ [x j ] related to the slave node x j , and an approximate distance D to [i, j] calculated in step S8. Based on the above, it is determined whether or not the excluded pair satisfies the first additional condition shown in the equation (7) (step S9).
- the first additional condition is a condition for determining whether or not the excluded pair may be connected at an edge in the b-matching graph.
- the additional unit 18 calculates the distance D [i, j] related to the excluded pair (step S10).
- the additional unit 18 is based on the upper reference message ⁇ [x i ] related to the main node x i , the upper reference message ⁇ [x j ] related to the slave node x j , and the distance D [i, j] calculated in step S10. Then, it is determined whether or not the excluded pair satisfies the second additional condition shown in the equation (8) (step S11).
- the second additional condition is a condition for strictly determining whether or not the excluded pair may be connected at an edge in the b-matching graph. That is, when all the excluded pairs do not satisfy the second addition condition, the addition unit 18 can determine that the optimum solution of the b-matching graph has been obtained. The reason why the optimum solution can be determined by the first additional condition and the second additional condition will be described later.
- the calculation cost of the distance D [i, j] is O (M), which is higher than the calculation cost O (m) of the approximate distances D to [i, j]. That is, by determining the first additional condition, the additional unit 18 can reduce the number of calculations of the distance D [i, j], which requires a large amount of calculation for the determination of the second additional condition.
- step S11 When the selected exclusion pair satisfies the first addition condition (step S11: YES), the addition unit 18 generates a message relating to the exclusion pair based on the equation (4), and displays the exclusion pair and the message. Recording is performed in the pair storage unit 15 (step S12). That is, the addition unit 18 adds the exclusion pair to the node pair set.
- the continuation determination unit 17 determines whether or not there is an excluded pair added to the node pair set (step S13). If there is an excluded pair added to the node pair set (step S13: YES), the graph computing device 10 returns the process to step S5 and repeats updating the message. On the other hand, when there is no exclusion pair added to the node pair set (step S13: NO), the edge determination unit 19 has, for each of the N nodes, among the node pairs whose main node is the node stored in the pair storage unit 15. The node pair from the one with the smallest message m [x i, x j ] to the bth node pair is determined as the node pair to be connected at the edge (step S14).
- FIG. 5 is a flowchart showing an edge weight calculation method according to the first embodiment.
- the weight calculation unit 20 selects N nodes one by one (step S301), and executes the processes of steps S302 to S309 below.
- the weight calculation unit 20 determines whether or not the weight calculation is completed for at least one slave node with respect to the node selected in step S301 (step S302). If there is no slave node weight calculation is completed (step S302: NO), the weight calculator 20, the primary node according to the node pair to the selected node in the step S301 and the main node of the vector x i and the slave node and b dimensions of the inner product vector q i of the inner product of vectors x j as elements, calculates the number b of the inverse matrix P i -1 of gram matrix P i of b ⁇ b in accordance with the vector of the slave nodes directly (step S303 ). Elements q i of the inner product vector q i [j] is expressed by the following equation (9), elements P i of the Gram matrix P i [j, k] is represented by the following formula (10).
- step S302 when there is a slave node for which the weight calculation has been completed (step S302: YES), the weight calculation unit 20 sets the node pair whose main node is the node among the nodes for which the weight calculation has been completed.
- the node having the largest number of slave nodes in common with the node pair having the node selected in step S301 as the main node is specified (step S304). That is, the node y i represented by the following equation (11) is specified.
- the node for which the weight has been calculated having the largest number of common slave nodes is referred to as a similar node x i'.
- the set C [x] is a set of subordinate nodes related to a node pair having the node x as the main node.
- the set C [x] is referred to as a slave node set related to the node x.
- the set D is a set of slave nodes for which the weight calculation has been completed.
- the node x i is the node selected in step S301.
- the weight calculation unit 20 identifies one set of non-common subordinate nodes between the node pair whose main node is the similar node x j and the node pair whose main node is the node selected in step S301 (step S305). Specifically, the slave node according to the node pair to the similar node x j main node, x 1, x 2, x 3, is x 4, x 5, the selected node in the step S301 and the main node node pair If the slave node according to is x 3, x 4, x 5 , x 6, x 7, a slave node which is not common and the (x 1, x 2) and (x 6, x 7). In this case, the weight calculation unit 20, for example, specifying a combination of x 1 and x 6.
- the weight calculation unit 20 uses one slave node x j'that is not common to the slave node set C [x i' ] related to the similar node x i', and the slave node set C [] related to the node x i selected in step S301.
- the set replaced with one node x j that is not common to x i ] is set as the update similar set C [x i ′′ ]
- the update similar set C [x i ′′ ] is shown by the following equation (12).
- the vector q j ′′ to be obtained and the matrix P j ′′ -1 represented by the equation (13) are calculated (step S306).
- the weight calculation unit 20 determines whether or not the update similarity set C [x i ′′ ] and the slave node set related to the node x i selected in step S301 match (step S307). When C [x i ′′ ] and the slave node set C [x i ] do not match (step S307: NO), the weight calculation unit 20 updates the set C [x i ′′ ] and the similar set C [x i ′′ ]. The process is returned to step S305, and the vector q j " and the matrix P j" are gradually calculated.
- the weight calculation unit 20 sets the vector q j ′′ and the matrix P j ′′ -1 .
- each inner product vector q i considered as the inverse matrix P i -1 of the gram matrix P i equation (11), by the equation (12), progressively inner product vector q i and the inverse matrix of the gram matrix P i P i. - It is clear from the Sherman-Morrison formula that one is required.
- the inverse matrix P i -1 of the Gram matrix P i by the expression (13) is the inner product by the equation (9) and (10) It is less than the calculation amount of the calculation to directly obtain the inverse matrix Pi -1 of the vector q i and the gram matrix P i.
- progressive inner product vector q i whereas the calculation amount of calculation for obtaining the inverse matrix P i -1 of the Gram matrix P i is O (b 2 + bM), Equation (9) and ( The amount of calculation in 10) is O (b 3 + b 2 M).
- the weight calculation unit 20 by obtaining the product of the inverse matrix P i -1 and the inner product vector q i of the Gram matrix P i, and calculates the weight vectors w i (step S308). That is, the weight calculator 20 uses the following equation (14), calculates the weight vector w i.
- Weight calculation unit 20 all elements of the weight vector w i determines whether a non-negative (step S309). If all the elements of weight vectors w i is positive (step S309: YES), the weight calculation unit 20, a weight vector w i, the minimum value is 0, it is normalized so that the sum of the elements becomes 1 Then, the weight of each edge connected to the node selected in step S301 is calculated (step S310).
- Step S309 if at least one of the elements of the weight vector w i is negative (Step S309: NO), the weight calculation unit 20, among the edges connected to the node selected in step S301, what weight is positive , Extracted as a target edge set for weight calculation by the quadratic programming method (step S311).
- the weight calculation unit 20 rewrites the value of which is not extracted as a calculation target of the elements of the weight vector w i to 0, the value of which is not extracted as a calculation target, the minimum value is 0, so that the sum of the elements becomes 1 Normalize to (step S312).
- the weight calculation unit 20 calculates the weight of the target edge set extracted in step S311 by the quadratic programming method (step S313).
- the weight calculation unit 20 determines whether or not there is an edge connected to the node selected in step S301 whose slave node x j satisfies the following equation (15) for the edge not included in the target edge set. (Step S314).
- step S314 determines whether there is an edge that is not included in the target edge set and satisfies the equation (15) (step S314: YES). If there is an edge that is not included in the target edge set and satisfies the equation (15) (step S314: YES), the edge is added to the target edge set (step S315), the process returns to step S312, and the quadratic Recalculate the weight by the planning method. On the other hand, a edge which is not included in the target edge set if there is nothing to satisfy the equation (15) (step S314: YES), determines the weights of the edges of the value of the weight vector w i calculated in step S313.
- the output unit 21 outputs a b-matching graph based on the edges and weights determined in steps S14 and S15, as shown in FIG. (Step S16).
- the graph calculation device 10 evaluates the node pair set having (b + 1) node pairs selected based on the distance as the initial set for each of the N nodes. By calculating the message, b edges are determined. That is, where the number of node pairs to be calculated for the message is (N-1) in the case of brute force calculation, it can be reduced to (b + 1) according to the first embodiment. This makes it possible to calculate the b-matching graph at a low calculation cost for determining the edge. In the first embodiment, the graph calculation device 10 extracts (b + 1) node pairs as an initial set, but the present invention is not limited to this.
- the graph computing device 10 may extract (b + 2) or more and less than (N-1) node pairs as an initial set. However, the smaller the number of node pairs included in the initial set, the smaller the calculation cost. Further, the graph calculation device 10 according to another embodiment may initialize the node pair set based on the approximate distance calculated from the low-dimensional vector and the residual, instead of the distance based on the low vector. Approximate distance is an example of distance.
- the graph computing device 10 adds to the node pair set the excluded pairs that are not included in the node pair set and satisfy the equation (8).
- the graph calculation device 10 can extract all the candidates for the node pair to be connected by the edge.
- the exclusion pair satisfying the first additional condition shown in the equation (7) using the approximate distance is further represented by the second equation (8) using the distance. If the additional condition is satisfied, the excluded pair is added to the node pair set.
- the graph calculation device 10 can reduce the amount of calculation related to the additional determination of the node pair. The reason is as follows.
- the graph calculation device 10 can select the exclusion pair that can satisfy the second additional condition by determining the first additional condition.
- the present invention is not limited to this, and the second additional condition may be determined for all the excluded pairs without determining the first additional condition. In this case, for example, by storing the distance calculated by the initialization unit 14 at the time of initialization and reading it out, the calculation may be omitted each time.
- the graph calculator 10 calculates the edge weight based on the b ⁇ b Gram matrix related to the vector of the slave node, and gradually calculates the inverse matrix of the Gram matrix. Therefore, the calculation cost can be reduced. Further, the graph calculation device 10 calculates the weight by the quadratic programming method when the weight cannot be calculated based on the gram matrix, but excludes the edge having a negative weight based on the gram matrix from the calculation target of the quadratic programming method. By doing so, the number of edges subject to the quadratic programming can be reduced. This is based on the inventor's finding that when the weight calculated based on the Gram matrix is negative, the weight related to the edge is often zero.
- the graph calculation device 10 may be configured by a single computer, or the configuration of the graph calculation device 10 may be divided into a plurality of computers so that the plurality of computers cooperate with each other. By doing so, it may function as the graph calculation device 10.
- FIG. 6 is a schematic block diagram showing a computer configuration of the graph calculation device 10.
- the graph calculation device 10 includes a processor 51, a memory 53, an auxiliary storage device 55, an interface 57, and the like connected by a bus, and by executing a graph calculation program, an input unit 11, a dimension reduction unit 12, and a vector storage unit 13 It functions as a device including an initialization unit 14, a pair storage unit 15, an update unit 16, a continuation determination unit 17, an addition unit 18, an edge determination unit 19, a weight calculation unit 20, and an output unit 21.
- Examples of the processor 51 include a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), a microprocessor, and the like.
- the graph calculation program may be recorded on a computer-readable recording medium such as the auxiliary storage device 55.
- the computer-readable recording medium is, for example, a storage device such as a magnetic disk, a magneto-optical disk, an optical disk, or a semiconductor memory.
- the graph calculation program may be transmitted via a telecommunication line. All or part of each function of graph calculation may be realized by using a custom LSI (Large Scale Integrated Circuit) such as ASIC (Application Specific Integrated Circuit) or PLD (Programmable Logic Device). Examples of PLDs include PAL (Programmable Array Logic), GAL (Generic Array Logic), CPLD (Complex Programmable Logic Device), and FPGA (Field Programmable Gate Array). Such an integrated circuit is also included in an example of the processor 51.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Mathematical Physics (AREA)
- Data Mining & Analysis (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Computational Mathematics (AREA)
- Pure & Applied Mathematics (AREA)
- Databases & Information Systems (AREA)
- Software Systems (AREA)
- General Engineering & Computer Science (AREA)
- Algebra (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
初期化部は、計算対象のN個のノードそれぞれについて、当該ノードを主ノードとし、N個のノードのうち前記主ノードからの距離に基づいて選択される(b+1)個以上(N-1)個未満のノードを従ノードとする複数のノードペアからなるノードペア集合を生成する。更新部は、ノードペア集合に係る各ノードペアについて、ノードペアに係る評価値を、ノードペアの距離、基準評価値、及び基準評価値に係るノードペアの距離に基づく計算により、所定の収束条件を満たすまで更新する。基準評価値は、ノードペアの従ノードを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値である。エッジ決定部は、更新された評価値に基づいて、各主ノードについてエッジで接続するb個の従ノードを決定する。
Description
本開示は、グラフ計算装置、グラフ計算方法及びグラフ計算プログラムに関する。
多次元データから計算されるグラフとして、b-matchingグラフが知られている。b-matchingグラフは、多次元データである複数のノードそれぞれにb個のエッジが接続されるグラフである。非特許文献1には、確率伝搬法(belief propagation)に基づいてb-matchingグラフを求める手法が開示されている。
Tony Jebara, Jun Wang, Shih-Fu Chang, Graph construction and b-matching for semi-supervised learning, ICML, 2009.
非特許文献1に記載の手法は、M次元のベクトルで表されるN個のノードから選択されるすべてのノードペアについて、収束するまで確率伝搬法に係るメッセージを更新するものである。そのため、収束までの繰り返し回数をTとおくと、計算コストがO(N2M2T)と高いことが分かる。
本開示の目的は、計算コストを抑えてb-matchingグラフを計算するグラフ計算装置、グラフ計算方法及びグラフ計算プログラムを提供することにある。
本開示の目的は、計算コストを抑えてb-matchingグラフを計算するグラフ計算装置、グラフ計算方法及びグラフ計算プログラムを提供することにある。
一態様によれば、グラフ計算装置は、計算対象のN個のノードそれぞれについて、当該ノードを主ノードとし、前記N個のノードのうち前記主ノードからの距離に基づいて選択される(b+1)個以上(N-1)個未満のノードを従ノードとする複数のノードペアからなるノードペア集合を生成する初期化部と、前記ノードペア集合に係る各ノードペアについて、前記ノードペアに係る評価値を、前記ノードペアの距離、前記ノードペアの従ノードを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値である基準評価値、及び前記基準評価値に係るノードペアの距離に基づく計算により、所定の収束条件を満たすまで更新する更新部と、更新された前記評価値に基づいて、各主ノードについてエッジで接続するb個の従ノードを決定するエッジ決定部とを備える。
上記態様によれば、グラフ計算装置は、計算コストを抑えてb-matchingグラフを計算することができる。
〈第1の実施形態〉
《グラフ計算装置10の構成》
以下、図面を参照しながら実施形態について詳しく説明する。
図1は、第1の実施形態に係るグラフ計算装置10の構成を示す概略ブロック図である。
グラフ計算装置10は、M次元のベクトルであるN個のノードについて、各ノードにウエイトを有するb本のエッジが接続されるb-matchingグラフを計算する。
グラフ計算装置10は、入力部11、次元削減部12、ベクトル記憶部13、初期化部14、ペア記憶部15、更新部16、継続判定部17、追加部18、エッジ決定部19、ウエイト計算部20、出力部21を備える。
《グラフ計算装置10の構成》
以下、図面を参照しながら実施形態について詳しく説明する。
図1は、第1の実施形態に係るグラフ計算装置10の構成を示す概略ブロック図である。
グラフ計算装置10は、M次元のベクトルであるN個のノードについて、各ノードにウエイトを有するb本のエッジが接続されるb-matchingグラフを計算する。
グラフ計算装置10は、入力部11、次元削減部12、ベクトル記憶部13、初期化部14、ペア記憶部15、更新部16、継続判定部17、追加部18、エッジ決定部19、ウエイト計算部20、出力部21を備える。
入力部11は、グラフの計算対象となるN個のノードの入力を受け付ける。各ノードは、M次元のベクトル(xi=xi[1]、…xi[M])で表される多次元データである。N及びMは自然数である。以下、入力部11に入力されたベクトルをロー(Raw)ベクトルともいう。
次元削減部12は、入力部11に入力された各ローベクトル(x1~xN)の次元を削減する。次元削減部12は、例えば特異値分解(Singular Value Decomposition)により各ノードの次元を削減する。次元削減後のベクトルである低次元ベクトルは、m次元のベクトルである。mはM未満の自然数である。具体的には、次元削減部12は、以下の手順で低次元ベクトルを生成する。まず、次元削減部12は、N個のローベクトル(x1~xN)をまとめたN×Mのノード行列Xを特異値分解することで、ユニタリ行列、特異値行列及び随伴行列を得る。次元削減部12は、ユニタリ行列からN×mの部分行列を、特異値行列からm×mの部分行列を抽出し、これらを乗算することで、N×mの低次元行列を生成する。次元削減部12は、当該低次元行列の各行ベクトルを、低次元ベクトル(x~
1~x~
N:ただし、x~は、xの上にチルダが乗ったもの)として抽出する。また、次元削減部12は、ローベクトルと低次元ベクトルの二乗平均誤差を表す残差rを以下の式(1)により算出する。
ベクトル記憶部13は、N個のノードそれぞれに関連付けて、当該ノードのローベクトルと低次元ベクトルと残差とを記憶する。
初期化部14は、N個のノードから、確率伝搬法の計算対象となるノードペアを抽出し、抽出したノードペアについてメッセージの初期値を生成する。確率伝搬法のメッセージは、評価値の一例である。
ペア記憶部15は、計算対象となるノードペアそれぞれに関連付けて、当該ノードペアのメッセージを記憶する。
ペア記憶部15は、計算対象となるノードペアそれぞれに関連付けて、当該ノードペアのメッセージを記憶する。
更新部16は、ペア記憶部15が記憶する各ノードペアに係るメッセージを更新する。
継続判定部17は、更新部16によるメッセージの更新が収束したか否かを判定する。
追加部18は、更新部16が更新したメッセージに基づいて、ペア記憶部15が記憶されないノードペアである除外ペアを、ペア記憶部15に追加するか否かを判定する。
継続判定部17は、更新部16によるメッセージの更新が収束したか否かを判定する。
追加部18は、更新部16が更新したメッセージに基づいて、ペア記憶部15が記憶されないノードペアである除外ペアを、ペア記憶部15に追加するか否かを判定する。
エッジ決定部19は、ペア記憶部15が記憶するメッセージに基づいて、エッジで接続するノードペアを決定する。
ウエイト計算部20は、ペア記憶部15が記憶するメッセージに基づいて、漸進的にエッジのウエイトを計算する。
出力部21は、エッジ決定部19及びウエイト計算部20の計算結果に基づいて、b-matchingグラフを出力する。
ウエイト計算部20は、ペア記憶部15が記憶するメッセージに基づいて、漸進的にエッジのウエイトを計算する。
出力部21は、エッジ決定部19及びウエイト計算部20の計算結果に基づいて、b-matchingグラフを出力する。
《グラフ計算装置10の処理》
図2は、第1の実施形態に係るb-matchingグラフの計算方法を示すフローチャートである。
入力部11は、計算対象となるN個のノード(x1~xN)と各ノードに係るローベクトルの入力を受け付ける(ステップS1)。次元削減部12は、ステップS1で取得したN個のローベクトルから、N個の低次元ベクトル(x~ 1~x~ N)と残差rとを算出する(ステップS2)。次元削減部12は、ノードに関連付けて、ローベクトル、低次元ベクトル、及び残差をベクトル記憶部13に記録する(ステップS3)。
図2は、第1の実施形態に係るb-matchingグラフの計算方法を示すフローチャートである。
入力部11は、計算対象となるN個のノード(x1~xN)と各ノードに係るローベクトルの入力を受け付ける(ステップS1)。次元削減部12は、ステップS1で取得したN個のローベクトルから、N個の低次元ベクトル(x~ 1~x~ N)と残差rとを算出する(ステップS2)。次元削減部12は、ノードに関連付けて、ローベクトル、低次元ベクトル、及び残差をベクトル記憶部13に記録する(ステップS3)。
次に、初期化部14は、N個のノードから計算対象となるノードペア及びメッセージの初期値を生成する(ステップS4)。初期化部14は、以下の手順で初期化処理を行う。
図3は、第1の実施形態に係るノードペア及びメッセージの初期化方法を示すフローチャートである。
初期化部14は、メッセージの計算対象とするノードペアの集合を初期化する。具体的には、初期化部14は、ペア記憶部15に、N個のノードそれぞれについて、当該ノードを主ノードとするノードペア集合の記憶領域を確保する(ステップS101)。ここでは、ノードペアに係る2つのノードのうちメッセージの計算において基準となるものを便宜的に主ノードと呼ぶ。のなお、この時点においてすべてのノードペア集合は空集合である。
図3は、第1の実施形態に係るノードペア及びメッセージの初期化方法を示すフローチャートである。
初期化部14は、メッセージの計算対象とするノードペアの集合を初期化する。具体的には、初期化部14は、ペア記憶部15に、N個のノードそれぞれについて、当該ノードを主ノードとするノードペア集合の記憶領域を確保する(ステップS101)。ここでは、ノードペアに係る2つのノードのうちメッセージの計算において基準となるものを便宜的に主ノードと呼ぶ。のなお、この時点においてすべてのノードペア集合は空集合である。
初期化部14は、N個のノードを1つずつ選択し(ステップS102)、以下のステップS103からステップS110の処理を実行する。
初期化部14は、選択したノード以外の(N-1)個のノード(候補ノード)からなるノード集合を、候補ノード集合として抽出する(ステップS103)。初期化部14は、以下の式(2)により、ベクトル記憶部13が記憶するローベクトルに基づいて選択したノードと各候補ノードとの距離を算出する(ステップS104)。
初期化部14は、選択したノード以外の(N-1)個のノード(候補ノード)からなるノード集合を、候補ノード集合として抽出する(ステップS103)。初期化部14は、以下の式(2)により、ベクトル記憶部13が記憶するローベクトルに基づいて選択したノードと各候補ノードとの距離を算出する(ステップS104)。
初期化部14は、候補ノード集合のうち、主ノードとの距離が最も短い候補ノードを特定する(ステップS105)。初期化部14は、ペア記憶部15を参照し、当該候補ノードを主ノードとするノードペアの数が(b+1)未満であるか否かを判定する(ステップS106)。候補ノードを主ノードとするノードペアの数が(b+1)未満である場合(ステップS106:YES)、初期化部14は、ステップS102で選択したノードを主ノードとし、ステップS105で特定した候補ノードを従ノードするノードペアを、ペア記憶部15に記録する(ステップS107)。また、初期化部14は、ステップS105で特定した候補ノードを主ノードとし、ステップS102で選択したノードを従ノードするノードペアを、ペア記憶部15に記録する(ステップS108)。このとき、初期化部14は、各ノードペアに関連付けて初期値1のメッセージをペア記憶部15に記録する。
候補ノードを主ノードとするノードペアの数が(b+1)以上である場合(ステップS106:NO)、またはペア記憶部15にノードペアを記録した場合(ステップS107及びS108)、初期化部14は、候補ノード集合からステップS105で特定した候補ノードを削除する(ステップS109)。
初期化部14は、ステップS102で選択したノードを主ノードとするノードペアの数が(b+1)未満であるか否かを判定する(ステップS110)。ステップS102で選択したノードを主ノードとするノードペアの数が(b+1)未満である場合(ステップS110:YES)、初期化部14は、ステップS105に処理を戻す。他方、ステップS102で選択したノードを主ノードとするノードペアの数が(b+1)以上である場合(ステップS110:NO)、初期化部14は、ステップS102に戻り、次のノードを選択する。
初期化部14は、ステップS102で選択したノードを主ノードとするノードペアの数が(b+1)未満であるか否かを判定する(ステップS110)。ステップS102で選択したノードを主ノードとするノードペアの数が(b+1)未満である場合(ステップS110:YES)、初期化部14は、ステップS105に処理を戻す。他方、ステップS102で選択したノードを主ノードとするノードペアの数が(b+1)以上である場合(ステップS110:NO)、初期化部14は、ステップS102に戻り、次のノードを選択する。
上記手順により、初期化部14は、ノードペア集合及びメッセージを初期化することができる。なお、上述の通り、初期化部14は、ペア記憶部15には、N(b+1)個のノードペアが記録される。N個のノードから特定可能なすべてのノードペアの数は、N(N-1)個であるため、初期化部14によって、(N-b-2)個のノードペアが、計算対象外となったことが分かる。これにより、グラフ計算装置10は、b-matchingグラフの計算コストを低減することができる。
図2に示すように、初期化部14がステップS4で、N個のノードから計算対象となるノードペア及びメッセージの初期値を生成すると、更新部16は、ペア記憶部15が記憶する各メッセージを更新する(ステップS5)。
更新部16は、以下の手順で更新処理を行う。図4は、第1の実施形態に係るメッセージの更新方法を示すフローチャートである。
更新部16は、ペア記憶部15が記憶するN個のノードを1つずつ選択し(ステップS201)、以下のステップS202からステップS204の処理を行う。
更新部16は、ペア記憶部15が記憶するN個のノードを1つずつ選択し(ステップS201)、以下のステップS202からステップS204の処理を行う。
まず更新部16は、ステップS201で選択したノードxiを主ノードとするノードペア(xi,xj)について、以下の式(2)に示すように、eを底とする当該ノードペアの距離D[i,j]の指数関数と、メッセージm[xi,xj]とを乗算することで、重み付きメッセージmw[xi,xj]を求める(ステップS202)。
更新部16は、ステップS201で選択したノードxiに係る重み付きメッセージmw[xi,xj]のうち、b番目に小さいものμ[xi]と、(b+1)番目に小さいものν[xi]を特定する(ステップS203)。以下、主ノードxiに係るb番目に小さい重み付きメッセージを、ノードxiに係る下側基準メッセージμ[xi]と呼ぶ。また、主ノードxiに係る(b+1)番目に小さい重み付きメッセージを、ノードxiに係る上側基準メッセージν[xi]と呼ぶ。
また、更新部16は、下側基準メッセージμ[xi]に係る従ノードv[xi]及び上側基準メッセージν[xi]に係る従ノードw[xi]を特定する(ステップS204)。以下、ノードxiに係る下側基準メッセージμ[xi]に係る従ノードを、ノードxiに係る下側基準ノードv[i]とよぶ。また、ノードxiに係る上側基準メッセージν[xi]に係る従ノードを、ノードxiに係る上側基準ノードw[i]とよぶ。つまり、下側基準メッセージμ[xi]及び上側基準メッセージν[xi]は、以下の式(3)のように表せる。
また、更新部16は、下側基準メッセージμ[xi]に係る従ノードv[xi]及び上側基準メッセージν[xi]に係る従ノードw[xi]を特定する(ステップS204)。以下、ノードxiに係る下側基準メッセージμ[xi]に係る従ノードを、ノードxiに係る下側基準ノードv[i]とよぶ。また、ノードxiに係る上側基準メッセージν[xi]に係る従ノードを、ノードxiに係る上側基準ノードw[i]とよぶ。つまり、下側基準メッセージμ[xi]及び上側基準メッセージν[xi]は、以下の式(3)のように表せる。
更新部16は、ペア記憶部15が記憶するノードペアを1つずつ選択し(ステップS205)、以下のステップS206からステップS208の処理を行う。
更新部16は、ステップS8で選択したノードペアに係る主ノードxiが、当該ノードペアに係る従ノードxjにおける下側基準ノードv[j]であるか否かを判定する(ステップS206)。主ノードxiが下側基準ノードv[j]でない場合(ステップS206:NO)、更新部16は、式(4)に示すように、ステップS205で選択したノードペアに係るメッセージを、eを底とする当該ノードペアの距離D[i,j]の指数関数を、従ノードxjに係る下側基準メッセージμ[xj]で除算した値に書き換える(ステップS207)。なお、ノードxiが下側基準ノードv[j]でない場合、下側基準メッセージμ[xj]は、ノードxjを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値と、eを底とする当該ノードペアの距離の指数関数との積であるといえる。
更新部16は、ステップS8で選択したノードペアに係る主ノードxiが、当該ノードペアに係る従ノードxjにおける下側基準ノードv[j]であるか否かを判定する(ステップS206)。主ノードxiが下側基準ノードv[j]でない場合(ステップS206:NO)、更新部16は、式(4)に示すように、ステップS205で選択したノードペアに係るメッセージを、eを底とする当該ノードペアの距離D[i,j]の指数関数を、従ノードxjに係る下側基準メッセージμ[xj]で除算した値に書き換える(ステップS207)。なお、ノードxiが下側基準ノードv[j]でない場合、下側基準メッセージμ[xj]は、ノードxjを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値と、eを底とする当該ノードペアの距離の指数関数との積であるといえる。
他方、主ノードxiが下側基準ノードv[j]である場合(ステップS206:YES)、更新部16は、式(5)に示すように、ステップS8で選択したノードペアに係るメッセージを、eを底とする当該ノードペアの距離D[i,j]の指数関数を、従ノードxjに係る上側基準メッセージν[xj]で除算した値に書き換える(ステップS208)。なお、主ノードxiが下側基準ノードv[j]である場合、上側基準メッセージν[xj]は、ノードxjを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値と、eを底とする当該ノードペアの距離の指数関数との積であるといえる。
上記手順により、更新部16は、候補ノード集合及びメッセージを初期化することができる。
なお、更新部16は、ステップS205~S208によるメッセージの更新前に、予めステップS201~S204で、ノードごとに下側基準メッセージ、上側基準メッセージ、下側基準ノード、上側基準ノードを特定しておく。これにより、更新部16は、更新対象のメッセージに係るノードペアの従ノードxjを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値を都度計算する必要がないため、計算コストを低減することができる。
なお、更新部16は、ステップS205~S208によるメッセージの更新前に、予めステップS201~S204で、ノードごとに下側基準メッセージ、上側基準メッセージ、下側基準ノード、上側基準ノードを特定しておく。これにより、更新部16は、更新対象のメッセージに係るノードペアの従ノードxjを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値を都度計算する必要がないため、計算コストを低減することができる。
図2に示すように、更新部16がステップS5で、メッセージを更新すると、継続判定部17は、メッセージが収束したか否かを判定する(ステップS6)。継続判定部17は、例えばメッセージの変化量の二乗平均が所定値未満になった場合に、メッセージが収束したと判定する。なお、メッセージは、O(N)回の繰り返しで収束する。
メッセージが収束していない場合(ステップS6:NO)、グラフ計算装置10はステップS5に処理を戻し、メッセージの更新を繰り返す。他方、メッセージが収束した場合(ステップS6:YES)、追加部18は、N個のノードから特定可能なすべてのノードペアのうち、ペア記憶部15が記憶しないノードペアすなわち計算対象外となったノードペア(以下、除外ペアと呼ぶ)を1つずつ選択し(ステップS7)、以下のステップS8からステップS12の処理を実行する。
メッセージが収束していない場合(ステップS6:NO)、グラフ計算装置10はステップS5に処理を戻し、メッセージの更新を繰り返す。他方、メッセージが収束した場合(ステップS6:YES)、追加部18は、N個のノードから特定可能なすべてのノードペアのうち、ペア記憶部15が記憶しないノードペアすなわち計算対象外となったノードペア(以下、除外ペアと呼ぶ)を1つずつ選択し(ステップS7)、以下のステップS8からステップS12の処理を実行する。
追加部18は、ベクトル記憶部13から、ステップS7で選択した除外ペア(xi,xj)に係る主ノードの低次元ベクトルx~
i及び残差r[xi]と従ノードの低次元ベクトルx~
j及び残差r[xj]を読み出し、式(6)により近似距離D~[i,j]を算出する(ステップS8)。
追加部18は、主ノードxiに係る上側基準メッセージν[xi]、従ノードxjに係る上側基準メッセージν[xj]、及びステップS8で算出した近似距離D~[i,j]に基づいて、除外ペアが式(7)に示す第1追加条件を満たすか否かを判定する(ステップS9)。第1追加条件は、除外ペアがb-matchingグラフにおいてエッジで接続される可能性があるか否かを判断するための条件である。
選択された除外ペアが第1追加条件を満たす場合(ステップS9:YES)、追加部18は、除外ペアに係る距離D[i,j]を算出する(ステップS10)。追加部18は、主ノードxiに係る上側基準メッセージν[xi]、従ノードxjに係る上側基準メッセージν[xj]、及びステップS10で算出した距離D[i,j]に基づいて、除外ペアが式(8)に示す第2追加条件を満たすか否かを判定する(ステップS11)。
第2追加条件は、除外ペアがb-matchingグラフにおいてエッジで接続される可能性があるか否かを厳格に判断するための条件である。つまり、すべての除外ペアが第2追加条件を満たさない場合、追加部18はb-matchingグラフの最適解が求められたと判断することができる。第1追加条件及び第2追加条件によって最適解の判定ができる理由については、後述する。
なお、距離D[i,j]の計算コストはO(M)であり、近似距離D~[i,j]の計算コストO(m)より高い。つまり、追加部18は、第1追加条件の判定を行うことで、第2追加条件の判定に必要な計算量の多い距離D[i,j]の計算回数を少なくすることができる。
なお、距離D[i,j]の計算コストはO(M)であり、近似距離D~[i,j]の計算コストO(m)より高い。つまり、追加部18は、第1追加条件の判定を行うことで、第2追加条件の判定に必要な計算量の多い距離D[i,j]の計算回数を少なくすることができる。
選択された除外ペアが第1追加条件を満たす場合(ステップS11:YES)、追加部18は、式(4)に基づいて当該除外ペアに係るメッセージを生成し、当該除外ペア及びそのメッセージを、ペア記憶部15に記録する(ステップS12)。つまり、追加部18は、除外ペアをノードペア集合に追加する。
継続判定部17は、ノードペア集合に追加された除外ペアがあるか否かを判定する(ステップS13)。ノードペア集合に追加された除外ペアがある場合(ステップS13:YES)、グラフ計算装置10はステップS5に処理を戻し、メッセージの更新を繰り返す。
他方、ノードペア集合に追加された除外ペアがない場合(ステップS13:NO)エッジ決定部19は、N個のノードそれぞれについて、ペア記憶部15が記憶する当該ノードを主ノードとするノードペアのうち、メッセージm[xi,xj]が小さいものからb個目までのノードペアを、エッジで接続するノードペアに決定する(ステップS14)。
他方、ノードペア集合に追加された除外ペアがない場合(ステップS13:NO)エッジ決定部19は、N個のノードそれぞれについて、ペア記憶部15が記憶する当該ノードを主ノードとするノードペアのうち、メッセージm[xi,xj]が小さいものからb個目までのノードペアを、エッジで接続するノードペアに決定する(ステップS14)。
次に、ウエイト計算部20は、決定された各エッジについて、ウエイトを計算する(ステップS15)。ウエイト計算部20は、以下の手順で初期化処理を行う。
図5は、第1の実施形態に係るエッジのウエイト計算方法を示すフローチャートである。
まず、ウエイト計算部20は、N個のノードを1つずつ選択し(ステップS301)、以下のステップS302からステップS309の処理を実行する。
図5は、第1の実施形態に係るエッジのウエイト計算方法を示すフローチャートである。
まず、ウエイト計算部20は、N個のノードを1つずつ選択し(ステップS301)、以下のステップS302からステップS309の処理を実行する。
ウエイト計算部20は、ステップS301で選択したノードに対する少なくとも1つの従ノードについて、ウエイトの計算が完了しているか否かを判定する(ステップS302)。ウエイトの計算が完了している従ノードがない場合(ステップS302:NO)、ウエイト計算部20は、ステップS301で選択したノードを主ノードとするノードペアに係る主ノードのベクトルxiと従ノードのベクトルxjの内積を要素とするb次元の内積ベクトルqiと、前記b個の従ノードのベクトルに係るb×bのグラム行列Piの逆行列Pi
-1を直接算出する(ステップS303)。内積ベクトルqiの要素qi[j]は、以下の式(9)で表され、グラム行列Piの要素Pi[j,k]は、以下の式(10)で表される。
他方、ウエイトの計算が完了している従ノードがある場合(ステップS302:YES)、ウエイト計算部20は、ウエイトの計算が完了しているノードのうち、当該ノードを主ノードとするノードペアと、ステップS301で選択したノードを主ノードとするノードペアとで共通する従ノードの数が最も多いものを特定する(ステップS304)。すなわち、以下の式(11)で示されるノードyiを特定する。以下、共通する従ノードの数が最も多いウエイト計算済みのノードを、類似ノードxi´とよぶ。
ウエイト計算部20は、類似ノードxjを主ノードとするノードペアと、ステップS301で選択したノードを主ノードとするノードペアとで、共通しない従ノードを1組特定する(ステップS305)。具体的には、類似ノードxjを主ノードとするノードペアに係る従ノードが、x1、x2、x3、x4、x5であり、ステップS301で選択したノードを主ノードとするノードペアに係る従ノードがx3、x4、x5、x6、x7である場合、(x1、x2)と(x6、x7)とが共通しない従ノードである。このとき、ウエイト計算部20は、例えばx1とx6の組み合わせを特定する。
ウエイト計算部20は、類似ノードxi´に係る従ノード集合C[xi´]の共通しない1つの従ノードxj´を、ステップS301で選択されたノードxiに係る従ノード集合C[xi]の共通しない1つのノードxjと入れ替えた集合を、更新類似集合C[xi″]とおいた場合に、更新類似集合C[xi″]について、以下の式(12)で示されるベクトルqj″と、式(13)で示される行列Pj″
-1を算出する(ステップS306)。
ウエイト計算部20は、更新類似集合C[xi″]と、ステップS301で選択されたノードxiに係る従ノード集合とが一致しているか否かを判定する(ステップS307)。更新類似集合C[xi″]と従ノード集合C[xi]とが一致しない場合(ステップS307:NO)、ウエイト計算部20は、集合C[xi″]を更新類似集合C[xi″]に置き換えて、ステップS305に処理を戻し、漸進的にベクトルqj″と行列Pj″を算出する。他方、更新類似集合C[xi″]と従ノード集合C[xi]とが一致する場合(ステップS307:YES)、ウエイト計算部20は、ベクトルqj″と行列Pj″
-1を、それぞれ内積ベクトルqi、グラム行列Piの逆行列Pi
-1とみなす。式(11)、式(12)により、漸進的に内積ベクトルqi及びグラム行列Piの逆行列Pi
-1を求められることは、Sherman-Morrisonの公式から明らかである。
なお、式(12)及び式(13)により漸進的に内積ベクトルqi、グラム行列Piの逆行列Pi
-1を求める計算の計算量は、式(9)及び式(10)により内積ベクトルqi及びグラム行列Piの逆行列Pi
-1を直接求める計算の計算量より少ない。具体的には、漸進的に内積ベクトルqi、グラム行列Piの逆行列Pi
-1を求める計算の計算量はO(b2+bM)であるのに対し、式(9)及び式(10)の計算量は、O(b3+b2M)である。
次に、ウエイト計算部20は、グラム行列Piの逆行列Pi
-1と内積ベクトルqiの積を求めることで、ウエイトベクトルwiを算出する(ステップS308)。すなわち、ウエイト計算部20は、以下の式(14)により、ウエイトベクトルwiを算出する。
ウエイト計算部20は、ウエイトベクトルwiのすべての要素が非負であるか否かを判定する(ステップS309)。ウェイトベクトルwiのすべての要素が正である場合(ステップS309:YES)、ウエイト計算部20は、ウエイトベクトルwiを、最小値が0、要素の総和が1となるように正規化することで、ステップS301で選択したノードに接続する各エッジのウエイトを算出する(ステップS310)。
他方、ウェイトベクトルwiの少なくとも1つの要素が負である場合(ステップS309:NO)、ウエイト計算部20は、ステップS301で選択したノードに接続するエッジのうち、ウエイトが正数であるものを、二次計画法によるウエイトの計算対象の対象エッジ集合として抽出する(ステップS311)。そして、ウエイト計算部20は、ウエイトベクトルwiの要素のうち計算対象として抽出されないものの値を0に書き換え、計算対象として抽出されないものの値を、最小値が0、要素の総和が1となるように正規化する(ステップS312)。
ウエイト計算部20は、ステップS311で抽出した対象エッジ集合について、二次計画法によりウエイトを計算する(ステップS313)。ウエイト計算部20は、ステップS301で選択したノードに接続するエッジのうち、対象エッジ集合に含まれないものについて、従ノードxjが以下の式(15)を満たすものがあるか否かを判定する(ステップS314)。
対象エッジ集合に含まれないエッジであって式(15)を満たすものが存在する場合(ステップS314:YES)、当該エッジを対象エッジ集合に追加し(ステップS315)、ステップS312に戻り、二次計画法によるウエイトの再計算を行う。他方、対象エッジ集合に含まれないエッジであって式(15)を満たすものが存在しない場合(ステップS314:YES)、エッジのウエイトをステップS313で算出したウエイトベクトルwiの値に決定する。
上述の手順でウエイト計算部20が各エッジのウエイトを算出すると、図2に示すように、出力部21は、ステップS14及びステップS15で決定したエッジ及びウエイトに基づいて、b-matchingグラフを出力する(ステップS16)。
《作用・効果》
このように、第1の実施形態によれば、グラフ計算装置10は、N個のノードそれぞれについて、距離に基づいて選択される(b+1)個のノードペアを初期集合とするノードペア集合について、評価値であるメッセージの計算を行うことで、b個のエッジを決定する。つまり、総当たりで計算する場合にメッセージの計算対象となるノードペアの数が(N-1)個であるところ、第1の実施形態によれば、(b+1)個まで低減することができる。これにより、エッジの決定のための計算コストを抑えてb-matchingグラフを計算することができる。なお、第1の実施形態では、グラフ計算装置10は、(b+1)個のノードペアを初期集合として抽出するが、これに限られない。例えば、他の実施形態においては、グラフ計算装置10は、(b+2)個以上(N-1)個未満のノードペアを初期集合として抽出してもよい。ただし、初期集合に含まれるノードペアの数が少ないほど計算コストが小さくなる。また、他の実施形態に係るグラフ計算装置10は、ローベクトルに基づく距離に代えて、低次元ベクトル及び残差から算出される近似距離に基づいてノードペア集合を初期化してもよい。近似距離は、距離の一例である。
このように、第1の実施形態によれば、グラフ計算装置10は、N個のノードそれぞれについて、距離に基づいて選択される(b+1)個のノードペアを初期集合とするノードペア集合について、評価値であるメッセージの計算を行うことで、b個のエッジを決定する。つまり、総当たりで計算する場合にメッセージの計算対象となるノードペアの数が(N-1)個であるところ、第1の実施形態によれば、(b+1)個まで低減することができる。これにより、エッジの決定のための計算コストを抑えてb-matchingグラフを計算することができる。なお、第1の実施形態では、グラフ計算装置10は、(b+1)個のノードペアを初期集合として抽出するが、これに限られない。例えば、他の実施形態においては、グラフ計算装置10は、(b+2)個以上(N-1)個未満のノードペアを初期集合として抽出してもよい。ただし、初期集合に含まれるノードペアの数が少ないほど計算コストが小さくなる。また、他の実施形態に係るグラフ計算装置10は、ローベクトルに基づく距離に代えて、低次元ベクトル及び残差から算出される近似距離に基づいてノードペア集合を初期化してもよい。近似距離は、距離の一例である。
第1の実施形態によれば、グラフ計算装置10は、ノードペア集合に含まれない除外ペアのうち式(8)を満たすものをノードペア集合に追加する。これにより、エッジによる接続対象となるノードペアの候補をグラフ計算装置10は漏れなく抽出することができる。
ここで、第1の実施形態に係るグラフ計算装置10は、近似距離を用いて式(7)に示す第1追加条件を満たす除外ペアについて、さらに距離を用いた式(8)に示す第2追加条件を満たす場合に、当該除外ペアをノードペア集合に追加する。この2段階の判定により、グラフ計算装置10は、ノードペアの追加の判定に係る計算量を低減することができる。理由は以下の通りである。近似距離は、要素数m個の低次元ベクトルに係る計算であるため、要素数M(M>m)個のベクトルに係る距離の計算より計算量が少ない。また、近似距離の値は、必ず距離の値以下となるため、グラフ計算装置10は、第1追加条件の判定をすることによって、第2追加条件を満たし得る除外ペアを選別することができる。なお、他の実施形態においては、これに限られず、第1追加条件の判定を行わずにすべての除外ペアについて第2追加条件の判定を行ってもよい。この場合に、例えば、初期化部14が初期化の際に計算した距離を記憶しておき、これを読み出すことで、都度の計算を省略してもよい。
ここで、第1の実施形態に係るグラフ計算装置10は、近似距離を用いて式(7)に示す第1追加条件を満たす除外ペアについて、さらに距離を用いた式(8)に示す第2追加条件を満たす場合に、当該除外ペアをノードペア集合に追加する。この2段階の判定により、グラフ計算装置10は、ノードペアの追加の判定に係る計算量を低減することができる。理由は以下の通りである。近似距離は、要素数m個の低次元ベクトルに係る計算であるため、要素数M(M>m)個のベクトルに係る距離の計算より計算量が少ない。また、近似距離の値は、必ず距離の値以下となるため、グラフ計算装置10は、第1追加条件の判定をすることによって、第2追加条件を満たし得る除外ペアを選別することができる。なお、他の実施形態においては、これに限られず、第1追加条件の判定を行わずにすべての除外ペアについて第2追加条件の判定を行ってもよい。この場合に、例えば、初期化部14が初期化の際に計算した距離を記憶しておき、これを読み出すことで、都度の計算を省略してもよい。
第1の実施形態によれば、グラフ計算装置10は、従ノードのベクトルに係るb×bのグラム行列に基づいてエッジのウエイトを算出し、当該グラム行列の逆行列を漸進的に計算することで、計算コストを低減することができる。また、グラフ計算装置10は、グラム行列に基づいてウエイトを計算できない場合に、二次計画法によってウエイトを計算するが、グラム行列に基づくウエイトが負数のエッジを二次計画法の計算対象から除外することで、二次計画法の対象となるエッジの数を低減することができる。これは、グラム行列に基づいて計算されるウエイトが負数になる場合、そのエッジに係るウエイトがゼロとなることが多いという発明者の知見に基づく。
〈他の実施形態〉
以上、図面を参照して一実施形態について詳しく説明してきたが、具体的な構成は上述のものに限られることはなく、様々な設計変更等をすることが可能である。すなわち、他の実施形態においては、上述の処理の順序が適宜変更されてもよい。また、一部の処理が並列に実行されてもよい。
上述した実施形態に係るグラフ計算装置10は、単独のコンピュータによって構成されるものであってもよいし、グラフ計算装置10の構成を複数のコンピュータに分けて配置し、複数のコンピュータが互いに協働することでグラフ計算装置10として機能するものであってもよい。
以上、図面を参照して一実施形態について詳しく説明してきたが、具体的な構成は上述のものに限られることはなく、様々な設計変更等をすることが可能である。すなわち、他の実施形態においては、上述の処理の順序が適宜変更されてもよい。また、一部の処理が並列に実行されてもよい。
上述した実施形態に係るグラフ計算装置10は、単独のコンピュータによって構成されるものであってもよいし、グラフ計算装置10の構成を複数のコンピュータに分けて配置し、複数のコンピュータが互いに協働することでグラフ計算装置10として機能するものであってもよい。
〈コンピュータ構成〉
図6は、グラフ計算装置10のコンピュータ構成を示す概略ブロック図である。
グラフ計算装置10は、バスで接続されたプロセッサ51、メモリ53、補助記憶装置55、インタフェース57などを備え、グラフ計算プログラムを実行することによって、入力部11、次元削減部12、ベクトル記憶部13、初期化部14、ペア記憶部15、更新部16、継続判定部17、追加部18、エッジ決定部19、ウエイト計算部20及び出力部21を備える装置として機能する。プロセッサ51の例としては、CPU(Central Processing Unit)、GPU(Graphic Processing Unit)、マイクロプロセッサなどが挙げられる。
グラフ計算プログラムは、補助記憶装置55などのコンピュータ読み取り可能な記録媒体に記録されてもよい。コンピュータ読み取り可能な記録媒体とは、例えば磁気ディスク、光磁気ディスク、光ディスク、半導体メモリ等の記憶装置である。グラフ計算プログラムは、電気通信回線を介して送信されてもよい。
なお、グラフ計算の各機能の全て又は一部は、ASIC(Application Specific Integrated Circuit)やPLD(Programmable Logic Device)等のカスタムLSI(Large Scale Integrated Circuit)を用いて実現されてもよい。PLDの例としては、PAL(Programmable Array Logic)、GAL(Generic Array Logic)、CPLD(Complex Programmable Logic Device)、FPGA(Field Programmable Gate Array)が挙げられる。このような集積回路も、プロセッサ51の一例に含まれる。
図6は、グラフ計算装置10のコンピュータ構成を示す概略ブロック図である。
グラフ計算装置10は、バスで接続されたプロセッサ51、メモリ53、補助記憶装置55、インタフェース57などを備え、グラフ計算プログラムを実行することによって、入力部11、次元削減部12、ベクトル記憶部13、初期化部14、ペア記憶部15、更新部16、継続判定部17、追加部18、エッジ決定部19、ウエイト計算部20及び出力部21を備える装置として機能する。プロセッサ51の例としては、CPU(Central Processing Unit)、GPU(Graphic Processing Unit)、マイクロプロセッサなどが挙げられる。
グラフ計算プログラムは、補助記憶装置55などのコンピュータ読み取り可能な記録媒体に記録されてもよい。コンピュータ読み取り可能な記録媒体とは、例えば磁気ディスク、光磁気ディスク、光ディスク、半導体メモリ等の記憶装置である。グラフ計算プログラムは、電気通信回線を介して送信されてもよい。
なお、グラフ計算の各機能の全て又は一部は、ASIC(Application Specific Integrated Circuit)やPLD(Programmable Logic Device)等のカスタムLSI(Large Scale Integrated Circuit)を用いて実現されてもよい。PLDの例としては、PAL(Programmable Array Logic)、GAL(Generic Array Logic)、CPLD(Complex Programmable Logic Device)、FPGA(Field Programmable Gate Array)が挙げられる。このような集積回路も、プロセッサ51の一例に含まれる。
10 グラフ計算装置
11 入力部
12 次元削減部
13 ベクトル記憶部
14 初期化部
15 ペア記憶部
16 更新部
17 継続判定部
18 追加部
19 エッジ決定部
20 ウエイト計算部
21 出力部
11 入力部
12 次元削減部
13 ベクトル記憶部
14 初期化部
15 ペア記憶部
16 更新部
17 継続判定部
18 追加部
19 エッジ決定部
20 ウエイト計算部
21 出力部
Claims (5)
- 計算対象のN個のノードそれぞれについて、当該ノードを主ノードとし、前記N個のノードのうち前記主ノードからの距離に基づいて選択される(b+1)個以上(N-1)個未満のノードを従ノードとする複数のノードペアからなるノードペア集合を生成する初期化部と、
前記ノードペア集合に係る各ノードペアについて、前記ノードペアに係る評価値を、前記ノードペアの距離、前記ノードペアの従ノードを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値である基準評価値、及び前記基準評価値に係るノードペアの距離に基づく計算により、所定の収束条件を満たすまで更新する更新部と、
更新された前記評価値に基づいて、各主ノードについてエッジで接続するb個の従ノードを決定するエッジ決定部と、
を備えるグラフ計算装置。 - 前記ノードペア集合に含まれないノードペアである除外ペアについて、前記除外ペアの距離、前記除外ペアと主ノードを同じくする他のノードペアに係る評価値のうちb+1番目に小さい評価値、及び前記除外ペアの従ノードを主ノードに持つ他のノードペアに係る評価値のうちb+1番目に小さい評価値が所定の関係を有する場合に、前記除外ペアを前記ノードペア集合に追加する追加部
を備える請求項1に記載のグラフ計算装置。 - 前記エッジで接続するb個のノードペアに係る主ノードのベクトルと従ノードのベクトルの内積を要素とするb次元の内積ベクトルと、前記b個の従ノードのベクトルに係るb×bのグラム行列の逆行列とに基づいて前記エッジのウエイトを算出するウエイト算出部
を備え、
前記ウエイト算出部は、前記内積ベクトル及び前記グラム行列の逆行列を、前記主ノードと共通の従ノードを有する他のノードである近傍ノードを用いた繰り返し計算により求める
請求項1または請求項2に記載のグラフ計算装置。 - 計算対象のN個のノードそれぞれについて、当該ノードを主ノードとし、前記N個のノードのうち前記主ノードからの距離に基づいて選択される(b+1)個以上(N-1)個未満のノードを従ノードとする複数のノードペアからなるノードペア集合を生成するステップと、
前記ノードペア集合に係る各ノードペアについて、前記ノードペアの関係を評価する評価値を、前記ノードペアの距離、前記ノードペアの従ノードを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値である基準評価値、及び前記基準評価値に係るノードペアの距離に基づく計算により、所定の収束条件を満たすまで更新するステップと、
更新された前記評価値に基づいて、各主ノードについてエッジで接続するb個の従ノードを決定するステップと、
を備えるグラフ計算方法。 - コンピュータに、
計算対象のN個のノードそれぞれについて、当該ノードを主ノードとし、前記N個のノードのうち前記主ノードからの距離に基づいて選択される(b+1)個以上(N-1)個未満のノードを従ノードとする複数のノードペアからなるノードペア集合を生成するステップと、
前記ノードペア集合に係る各ノードペアについて、前記ノードペアの関係を評価する評価値を、前記ノードペアの距離、前記ノードペアの従ノードを主ノードに持つ他のノードペアに係る評価値のうちb番目に小さい評価値である基準評価値、及び前記基準評価値に係るノードペアの距離に基づく計算により、所定の収束条件を満たすまで更新するステップと、
更新された前記評価値に基づいて、各主ノードについてエッジで接続するb個の従ノードを決定するステップと、
を実行させるためのグラフ計算プログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/027218 WO2022013909A1 (ja) | 2020-07-13 | 2020-07-13 | グラフ計算装置、グラフ計算方法及びグラフ計算プログラム |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/027218 WO2022013909A1 (ja) | 2020-07-13 | 2020-07-13 | グラフ計算装置、グラフ計算方法及びグラフ計算プログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022013909A1 true WO2022013909A1 (ja) | 2022-01-20 |
Family
ID=79555212
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/027218 Ceased WO2022013909A1 (ja) | 2020-07-13 | 2020-07-13 | グラフ計算装置、グラフ計算方法及びグラフ計算プログラム |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2022013909A1 (ja) |
-
2020
- 2020-07-13 WO PCT/JP2020/027218 patent/WO2022013909A1/ja not_active Ceased
Non-Patent Citations (2)
| Title |
|---|
| BIENKOWSKI MARCIN, FUCHSSTEINER DAVID, MARCINKOWSKI JAN, SCHMID STEFAN: "Online Dynamic B-Matching With Applications to Reconfigurable Datacenter Networks", 14 March 2021 (2021-03-14), XP055898089, Retrieved from the Internet <URL:https://arxiv.org/pdf/2006.10692.pdf> * |
| XIAOTIAN HAO; JUNQI JIN; JIANYE HAO; JIN LI; WEIXUN WANG; YI MA; ZHENZHE ZHENG; HAN LI; JIAN XU; KUN GAI: "Learning to Accelerate Heuristic Searching for Large-Scale Maximum Weighted b-Matching Problems in Online Advertising", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 12 May 2020 (2020-05-12), 201 Olin Library Cornell University Ithaca, NY 14853 , XP081671110 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Wang et al. | Membrane computing model for IIR filter design | |
| US8825573B2 (en) | Controlling quarantining and biasing in cataclysms for optimization simulations | |
| JP6623947B2 (ja) | 情報処理装置、イジング装置及び情報処理装置の制御方法 | |
| JP7512631B2 (ja) | イジングマシンデータ入力機器、及びイジングマシンにデータを入力する方法 | |
| JPWO2019064774A1 (ja) | 情報処理装置、および情報処理方法 | |
| CN112805768B (zh) | 秘密s型函数计算系统及其方法、秘密逻辑回归计算系统及其方法、秘密s型函数计算装置、秘密逻辑回归计算装置、程序 | |
| JP2012208924A (ja) | 適応的重み付けを用いた様々な文書間類似度計算方法に基づいた文書比較方法および文書比較システム | |
| US20210097397A1 (en) | Information processing apparatus and information processing method | |
| JP2017016384A (ja) | 混合係数パラメータ学習装置、混合生起確率算出装置、及び、これらのプログラム | |
| CN111160733B (zh) | 一种基于有偏样本的风险控制方法、装置及电子设备 | |
| WO2022270163A1 (ja) | 計算機システム及び介入効果予測方法 | |
| CN111373391A (zh) | 语言处理装置、语言处理系统和语言处理方法 | |
| CN103782290A (zh) | 建议值的生成 | |
| JP7294017B2 (ja) | 情報処理装置、情報処理方法および情報処理プログラム | |
| CN117253079A (zh) | 模型训练方法、装置、设备及存储介质 | |
| CN111046958A (zh) | 基于数据依赖的核学习和字典学习的图像分类及识别方法 | |
| JP2020027604A (ja) | 情報処理方法、及び情報処理システム | |
| WO2016002020A1 (ja) | 行列生成装置及び行列生成方法及び行列生成プログラム | |
| JP7472998B2 (ja) | パラメータ推定装置、秘密パラメータ推定システム、秘密計算装置、それらの方法、およびプログラム | |
| WO2020177863A1 (en) | Training of algorithms | |
| JP5244452B2 (ja) | 文書特徴表現計算装置、及びプログラム | |
| Skidmore et al. | The KaGE RLS algorithm | |
| JP2022158010A (ja) | 情報処理システム、情報処理方法、及び情報処理プログラム | |
| JP2021144659A (ja) | 計算機、計算方法及びプログラム | |
| WO2020054402A1 (ja) | ニューラルネットワーク処理装置、コンピュータプログラム、ニューラルネットワーク製造方法、ニューラルネットワークデータの製造方法、ニューラルネットワーク利用装置、及びニューラルネットワーク小規模化方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20945420 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20945420 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: JP |





