WO2021130867A1 - 均衡状態計算装置、均衡状態計算方法及びプログラム - Google Patents

均衡状態計算装置、均衡状態計算方法及びプログラム Download PDF

Info

Publication number
WO2021130867A1
WO2021130867A1 PCT/JP2019/050675 JP2019050675W WO2021130867A1 WO 2021130867 A1 WO2021130867 A1 WO 2021130867A1 JP 2019050675 W JP2019050675 W JP 2019050675W WO 2021130867 A1 WO2021130867 A1 WO 2021130867A1
Authority
WO
WIPO (PCT)
Prior art keywords
strategy
equilibrium state
cost
node
calculation device
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2019/050675
Other languages
English (en)
French (fr)
Inventor
中村 健吾
晋作 坂上
宜仁 安田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to US17/787,859 priority Critical patent/US20230052060A1/en
Priority to JP2021566610A priority patent/JP7279820B2/ja
Priority to PCT/JP2019/050675 priority patent/WO2021130867A1/ja
Publication of WO2021130867A1 publication Critical patent/WO2021130867A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • A—HUMAN NECESSITIES
    • A63—SPORTS; GAMES; AMUSEMENTS
    • A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
    • A63F13/80—Special adaptations for executing a specific game genre or game mode
    • A63F13/822—Strategy games; Role-playing games
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00—Administration; Management
    • G06Q10/04—Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
    • A—HUMAN NECESSITIES
    • A63—SPORTS; GAMES; AMUSEMENTS
    • A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
    • A63F13/70—Game security or game management aspects
    • A63F13/77—Game security or game management aspects involving data related to game devices or game servers, e.g. configuration data, software version or amount of memory
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N99/00—Subject matter not provided for in other groups of this subclass
    • A—HUMAN NECESSITIES
    • A63—SPORTS; GAMES; AMUSEMENTS
    • A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F2300/00—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game
    • A63F2300/50—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers
    • A63F2300/53—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers details of basic data processing
    • A63F2300/535—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers details of basic data processing for monitoring, e.g. of user parameters, terminal parameters, application parameters, network parameters
    • A—HUMAN NECESSITIES
    • A63—SPORTS; GAMES; AMUSEMENTS
    • A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F2300/00—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game
    • A63F2300/80—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game specially adapted for executing a specific type of game
    • A63F2300/807—Role playing or strategy games

Definitions

  • the present invention relates to an equilibrium state calculation device, an equilibrium state calculation method, and a program.
  • Congestion game is known as one of the non-cooperative games in game theory.
  • a crowded game is a model of a situation in which players who do not cooperate with each other compete for some resources or allocate resources to players who do not cooperate with each other.
  • Selfish Routing which is a type of congestion game, for example, in a communication network consisting of communication paths where the delay increases as the amount of communication increases, many people (players) try to communicate between two points with a small delay, or traffic. In a road network consisting of roads that take longer as the amount increases, it is possible to model a situation in which each player tries to move between two points in a short time.
  • Congested games are a more generalized version of Selfish Routing that can handle a wide range of strategic sets, for example, in situations where many people are trying to communicate between points with little delay, or even when communicating between two points.
  • the billing amount is set separately from the delay of the communication path, and it becomes possible to handle the situation where communication is performed by limiting the communication route to the communication route of the billing amount below the budget.
  • the set of such combinations is a strategy set, and each element of the contents of the strategy set (that is, the combination of items) is called a strategy.
  • Each item is set so that the higher the percentage of selected players, the higher the cost, and the cost received by the player is the sum of the costs of the items in the selected strategy. At this time, each player does not cooperate with each other and tries to select a strategy with the lowest cost as much as possible.
  • each side of the graph structure is represented by an item, a path from a certain vertex to a certain vertex on the graph structure. It is a crowded game with a set of combinations of items as a strategic set.
  • the above-mentioned situation of performing multipoint communication can be modeled as a congestion game in which a Steiner tree with a certain vertex set on the graph structure as a terminal is used as a strategic set, and the charge amount is limited to a certain amount or less.
  • the situation of communicating with each other can be modeled as a congestion game in which a set of combinations of items represented by a path that satisfies the charge amount in a path from a certain vertex to a certain vertex on the graph structure is used as a strategic set.
  • An important state in a crowded game is a state called an equilibrium state.
  • An equilibrium state is a state in which players are not dissatisfied, and is a state in which players who do not cooperate with each other end up aiming for a state of minimum cost on their own. If the equilibrium state in a congestion game can be calculated, for example, when designing a communication network or a road network, how much congestion will occur in each communication path or road due to the design, and how much the player's actual cost will be. It becomes possible to simulate whether or not it becomes.
  • Non-Patent Document 1 a technique for approximately obtaining the equilibrium state in Selfish Routing has been proposed.
  • a technique has been proposed in which the equilibrium state in Selfish Routing is theoretically obtained in polynomial time by repeatedly using a flow algorithm on a graph structure.
  • an optimization algorithm called the Frank-Wolfe algorithm is well known (Non-Patent Documents 2 and 3).
  • the number of elements of a set of combinations such as a strategic set is 2 n at the maximum with respect to the size n of the original set [n], and is often exponentially large in general. Therefore, holding all the elements of the strategy set requires a great deal of cost in terms of both calculation time and memory, and for example, it is often impossible to obtain an equilibrium state in practice even if n is about several tens.
  • One embodiment of the present invention has been made in view of the above points, and an object of the present invention is to calculate the equilibrium state of a crowded game.
  • the balance state calculation device is a balance state calculation device that calculates the balance state of the congestion game, and is a strategy used by the player of the congestion game, and is a combination of items.
  • the input means for inputting the graph information in which the set of strategies represented by is represented by the zero-suppress type dichotomy diagram and the graph information input by the input means, and the variation of the Frank-Wolfe algorithm the above-mentioned in the balanced state. It is characterized by a calculation means for calculating equilibrium state information including the percentage of players who have selected a strategy.
  • the equilibrium state of the crowded game can be calculated.
  • the equilibrium state calculation device 10 capable of calculating the equilibrium state of the congestion game will be described.
  • the equilibrium state calculation device 10 does not depend on the strategy set by incorporating the zero-suppressed Binary Decision Diagram (hereinafter referred to as “ZDD”) into the Frank-Wolfe algorithm. It enables the equilibrium state of a general congestion game to be calculated at high speed.
  • ZDD zero-suppressed Binary Decision Diagram
  • the fully-corrective Frank-Wolfe algorithm and the Away-step Frank-Wolfe algorithm, which are variants thereof, are used as the Frank-Wolfe algorithm.
  • the equilibrium state calculation device 10 according to the present embodiment can obtain an equilibrium state whose approximation accuracy is guaranteed.
  • ZDD is a structure that can compactly express a combination set such as a strategy set. For example, a set of paths from a certain vertex to a certain vertex on a graph structure, a set of Steiner trees, a set of paths satisfying a charge amount constraint, and the like can all be expressed by ZDD.
  • the ZDD representing the combination set can be constructed by, for example, the Frontier-based method. Further, Graphillion and the like are known as a library using the frontier method, and it is possible to efficiently construct ZDD by using such a library.
  • the frontier method for example, Reference 1 “Jun Kawahara, Takeru Inoue, Hiroaki Iwashita, and Shin-ichi Minato.
  • each player selects the item combination S ⁇ [n] for the item set [n].
  • what kind of item combination can be selected is decided in advance, and the set of such combinations is the strategic set.
  • each element of the contents of the strategy set (that is, the combination of items) is called a strategy.
  • a strategy In the crowded game dealt with in this embodiment, there are innumerable players, and the ratio z S of the player who chose the strategy to each strategy S is considered. In addition, it should be noted.
  • the usage rate y i of item i is the sum of the percentage of players who chose the strategy that includes item i, that is,
  • Wardrop equilibrium is the state where the cost is the lowest of all strategies for strategies with a player percentage greater than 0, that is,
  • the above-mentioned congestion game will be more generalized, and a case where different strategy sets are used depending on the player will be described.
  • the set of r player groups is expressed as [r]
  • the strategy set is not one, but a plurality of types of r strategy sets.
  • a player group is a set of 0 or more players. In addition, it should be noted.
  • the Wardrop equilibrium state at this time is a state in which the cost is the minimum in the strategy of the corresponding strategy set for the strategy in which the ratio of players is greater than 0, that is,
  • An ⁇ -approximate Wardrop equilibrium state is a state in which the cost for a strategy with a player ratio greater than 0 is not greater than the minimum cost in the strategy of the corresponding strategy set, that is, the tolerance ⁇ or more.
  • the i-th element is 1, otherwise the i-th element.
  • the i-th element is 1, otherwise the i-th element.
  • it be an n-dimensional vector in which the element of is 0.
  • n-dimensional vector x p for the i-th element of x i p above,
  • the n-dimensional vector y having the above usage rate y i as the i-th element can be obtained.
  • the ⁇ -approximate Wardrop equilibrium state is obtained by solving the minimization problem related to this potential function ⁇ by using the strategy set expressed by ZDD and the variant of the Frank-Wolfe algorithm. In addition, it should be noted.
  • ⁇ (y) i represents the i-th element of ⁇ (y) (that is, the partial differential of ⁇ with respect to y i).
  • FIG. 1 is a diagram showing an example of the overall configuration of the equilibrium state calculation device 10 according to the present embodiment.
  • the equilibrium state calculation device 10 includes an input unit 101, an optimization unit 102, an output unit 103, and a storage unit 104.
  • the storage unit 104 stores various information necessary for calculating the ⁇ -approximate Wardrop equilibrium state in the congestion game.
  • the information in the storage unit 104 are stored, for example, item set [n], the cost function c i of each item (y i), the set information, the group player representation of one or more of each of the strategy set by ZDD [R], the ratio of players using each strategy set m 1 , ..., mr , tolerance ⁇ , etc.
  • the strategy set expressed in ZDD will be used.
  • the storage unit 104 may store information such as the calculation process of the ⁇ -approximate Wardrop equilibrium state.
  • the ZDD representing the strategy set is a directed acyclic graph (DAG) composed of a node set and an edge set of directed edges connecting the nodes.
  • DAG directed acyclic graph
  • each node v is also referred to as a node r.
  • each node v is given an integer value l v ⁇ ⁇ 1, ..., N ⁇ called a label, and the item and the node are associated with each other by this label.
  • the label value may be set to n + 1 or the like.
  • ZDD ensures that the 0-branch and 1-branch of each node go from the node with the smaller label to the node with the larger label. That is, (label of node v) ⁇ (label of node v 0 ) and (label of node v) ⁇ (label of node v 1 ) are established for any node v.
  • the ZDD is layered by the label value. For example, the node included in the first layer (that is, the node r) corresponds to the item 1, and the node v included in the second layer corresponds to the item 2. To do. In this way, the node v included in the i-th layer of the ZDD corresponds to the item i.
  • each path from the root node r to the terminal node can represent a combination of items (that is, a strategy). That is, if the path includes an edge from node v to node v 1 , the item corresponding to the label of the node v is included in the strategy, and if the edge from node v to node v 0 is included in the path, the relevant item is included.
  • the combination can be expressed by the path.
  • Input unit 101 the item set [n], the cost function c i of each item (y i), 1 or more strategies set expressed in ZDD, a set of group players [r], the percentage of players using the strategy set Enter various information such as m 1 , ..., mr , and margin of error ⁇ .
  • the optimization unit 102 obtains various information in the ⁇ -approximate Wardrop equilibrium state by processing based on the Fully-corrective Frank-Wolfe algorithm. More specifically, the optimization unit 102 obtains various information in the ⁇ -approximate Wardrop equilibrium state by solving the minimization problem related to the potential function ⁇ using various information input by the input unit 101. As a result, the ratio z S p of player that chose each strategy S in ⁇ - approximate Wardrop equilibrium state is obtained.
  • the output unit 103 outputs various information obtained by the optimization unit 102 (e.g., .epsilon. approximate Wardrop percentage of players chose the strategy S at equilibrium z S p, etc.).
  • the output destination of the output unit 103 is not limited, and may be any output destination.
  • the output destination of the output unit 103 may be a storage unit 104, a display device such as a display, or a database server connected via a communication network.
  • the optimization unit 102 includes an initial setting unit 111, a shortest path calculation unit 112, and an update unit 113.
  • the initialization unit 111 initializes various variables (parameters) to be updated by a variant of the Frank-Wolfe algorithm.
  • Parameter is the ratio z S p of the active set, and the player chose the strategy S representing a set of strategies S of n-dimensional vector x p as described above, each player is currently selected.
  • the shortest path calculation unit 112 calculates the strategy that minimizes the cost in the strategy set by calculating the shortest path on the ZDD representing the strategy set by dynamic programming.
  • the update unit 113 updates various parameters by the correction process based on the Away-step Frank-Wolfe algorithm.
  • FIG. 2 is a flowchart showing an example of the equilibrium state calculation process according to the present embodiment.
  • step S1100 the input unit 101, various kinds of information (item set [n], the cost function c i of each item (y i), 1 or more strategies set expressed in ZDD, a set of group players [r], the Enter the percentage of players who use the strategy set m 1 , ..., mr , margin of error ⁇ , etc.).
  • item set [n] the cost function c i of each item (y i)
  • strategy set expressed in ZDD a set of group players [r]
  • step S1200 the optimization unit 102 uses the initialization unit 111 to perform a strategy for each p ⁇ [r].
  • step S1300 the optimization unit 102, the initial setting unit 111, various parameters initialized as follows respective (n-dimensional vector x 0 p, the Active Set, the ratio z S p players chose the strategy S) To be.
  • steps S1410 to S1450 the case where the number of repetitions is the kth time will be described, and the subscript at the lower right of each symbol except z indicates the number of repetitions.
  • the n-dimensional vector y k represents the n-dimensional vector y when the number of repetitions is k.
  • step S1410 the optimization unit 102 sets the n-dimensional vector y k having the utilization rate y i as the i-th element.
  • step S1420 the optimization unit 102 repeatedly executes steps S1421 to S1422 for each p ⁇ [r]. In the subsequent steps S1421 to S1422, steps S1421 to S1422 relating to a certain p will be described.
  • step S1421 the optimization unit 102 is ZDD by the shortest path calculation unit 112.
  • the strategy sk p that minimizes the cost in the strategy set corresponding to p is calculated. That is, the shortest path calculation unit 112 calculates the shortest path on the ZDD that represents the strategy set corresponding to p.
  • the shortest path calculation unit 112 calculates the strategy sk p as follows.
  • the shortest path calculation unit 112 sets the distance of 0-branch to 0 and the distance of 1-branch to each node v of ZDD.
  • the distance of the 1-branch of the node v is set as the cost of the item corresponding to the label of the node v.
  • the shortest path calculation unit 112 calculates a path (shortest path) that is a path from the root node r to the terminal node and that minimizes the total distance, by dynamic programming.
  • a path shortest path
  • the combination of items represented by the path that minimizes the distance on ZDD is obtained as the strategy sk p that minimizes the cost in the strategy set corresponding to p.
  • the method of calculating the shortest path on a directed acyclic graph such as ZDD is widely known. For example, reference 5 "Tetsuro Shibuya," Information Engineering Algorithm ", Maruzen Publishing, November 2016", etc. Please refer.
  • step S1422 the optimization unit 102 determines the average cost of the player using the strategy set corresponding to p in the current state and the strategy with the lowest cost in this strategy set (that is, the strategy sk p ). Calculate the difference g k p from the cost. That is, the optimization unit 102
  • ⁇ - ⁇ (y k) , d k p> ⁇ (y k), x k p> - ⁇ (y k), s k p> a, ⁇ (y k ), x k p> represents the average cost of the player you are using a strategy set corresponding to p, ⁇ (y k) , s k p> represents the cost of the strategy s k p ..
  • step S1430 the optimization unit 102 determines for all p ⁇ [r] whether or not the difference g k p between the average cost and the cost of the strategy with the lowest cost is less than or equal to the margin of error ⁇ . That is, the optimization unit 102
  • step S1440 is executed, otherwise (NO in step S1430) step S1440. Is not executed.
  • step S1440 the output unit 103 has the current parameters.
  • Is output are the n-dimensional vector x in the ⁇ -approximate Wardrop equilibrium, the active set, and the percentage of players who chose each strategy, respectively. The reason why these parameters satisfy the ⁇ -approximate Wardrop equilibrium state will be described later.
  • step S1450 the optimization unit 102 executes the correction process by the update unit 113 to update various parameters. That is, the optimization unit 102 is a subroutine.
  • FIG. 3 is a flowchart showing an example of the Correction process according to the present embodiment.
  • the index k in step S1400 of FIG. 2 is omitted and the subroutine is omitted.
  • step S2100 the update unit 113 repeatedly executes step S2110 for each p ⁇ [r].
  • step S2110 relating to a certain p will be described.
  • step S2110 the update unit 113
  • steps S2210 to S2280 the case where the number of repetitions is the lth time will be described, and the subscript at the lower right of each symbol except z indicates the number of repetitions.
  • the n-dimensional vector y l represents the n-dimensional vector y when the number of repetitions is the first.
  • step S2210 the update unit 113 sets an n-dimensional vector y l having the usage rate y i as the i-th element.
  • step S2220 the update unit 113 repeatedly executes step S2221 for each p ⁇ [r]. However, when step S2240 is executed, the correction process is terminated and the process returns to the caller of the subroutine. In the subsequent step S2221, step S2221 relating to a certain p will be described.
  • step S2221 the update unit 113, at the present time, strategy cost in new strategies set corresponding to p is minimized s l p and strategy cost is largest of the active set corresponding to the p v Calculate l p. Specifically, the update unit 113
  • d l p and FW represent the direction from x p to the strategy with the lowest cost s l p
  • d l p and A represent the direction opposite to the direction from x p to the strategy v l p with the highest cost.
  • the update unit 113 to calculate a cost for each strategy may be calculated and a strategy s l p and strategies v l p. This is because the size of the new strategic set or active set corresponding to p is very small (at most, about O (n)) as compared with the strategic set corresponding to p.
  • step S2230 the update unit 113 determines whether or not the difference between the cost of the strategy v l p and the cost of the strategy l p is equal to or less than the margin of error ⁇ for all p ⁇ [r]. That is, the update unit 113
  • step S2240 is executed. Otherwise (NO in step S2230), steps S2250 to S2280 are executed.
  • step S2240 the update unit 113 has a parameter.
  • the update unit 113 ends the correction process and returns to the caller of the subroutine.
  • step S2250 the update unit 113
  • step S2260 the update unit 113 calculates dl and ⁇ max according to the magnitude relationship between gl FW and gl A. That is, when the update unit 113 has gl FW ⁇ g l A ,
  • step S2270 the update unit 113
  • the update unit 113 finds the point ⁇ l at which the value of the function F becomes the minimum when traveling in the direction from x l to d l. This may be obtained by, for example, a line search.
  • step S2280 the update unit 113 updates the parameters.
  • the update unit 113 updates the parameters as follows.
  • Update by. For the method of updating the ratio z of the players who have selected each strategy, refer to, for example, Reference 4 above.
  • FIG. 4 is a diagram showing an example of the hardware configuration of the equilibrium state calculation device 10 according to the present embodiment.
  • the equilibrium state calculation device 10 is realized by a general computer or computer system, and includes an input device 201, a display device 202, an external I / F 203, and a communication I / F 204. , The processor 205 and the memory device 206. Each of these hardware is communicably connected via bus 207.
  • the input device 201 is, for example, a keyboard, a mouse, a touch panel, or the like.
  • the display device 202 is, for example, a display or the like.
  • the equilibrium state calculation device 10 does not have to have at least one of the input device 201 and the display device 202.
  • the external I / F 203 is an interface with an external device.
  • the external device includes, for example, a recording medium 203a and the like.
  • the equilibrium state calculation device 10 can read or write the recording medium 203a via the external I / F 203.
  • the recording medium 203a may store, for example, one or more programs that realize each functional unit (input unit 101, optimization unit 102, and output unit 103) of the equilibrium state calculation device 10.
  • the recording medium 203a includes, for example, a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), a USB (Universal Serial Bus) memory card, and the like.
  • a CD Compact Disc
  • DVD Digital Versatile Disk
  • SD memory card Secure Digital memory card
  • USB Universal Serial Bus
  • the communication I / F 204 is an interface for connecting the equilibrium state calculation device 10 to the communication network.
  • One or more programs that realize each functional unit of the equilibrium state calculation device 10 may be acquired (downloaded) from a predetermined server device or the like via the communication I / F 204.
  • the processor 205 is, for example, various arithmetic units such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). Each functional unit included in the equilibrium state calculation device 10 is realized by, for example, a process of causing the processor 205 to execute one or more programs stored in the memory device 206 or the like.
  • a CPU Central Processing Unit
  • GPU Graphics Processing Unit
  • the memory device 206 is, for example, various storage devices such as HDD (Hard Disk Drive), SSD (Solid State Drive), RAM (Random Access Memory), ROM (Read Only Memory), and flash memory.
  • the storage unit 104 included in the equilibrium state calculation device 10 can be realized by using, for example, the memory device 206.
  • the storage unit 104 may be realized by using a storage device or the like connected to the equilibrium state calculation device 10 via a communication network.
  • the equilibrium state calculation device 10 can realize the above-mentioned equilibrium state calculation process by having the hardware configuration shown in FIG.
  • the hardware configuration shown in FIG. 4 is an example, and the equilibrium state calculation device 10 may have another hardware configuration.
  • the equilibrium state calculation device 10 may have a plurality of processors 205 or a plurality of memory devices 206.
  • the equilibrium state calculation device 10 practically provides an equilibrium state ( ⁇ -approximate Wardrop equilibrium state) in which the accuracy of approximation is guaranteed even for a congestion game composed of a general strategy set. It can be obtained at high speed. This makes it possible to calculate the equilibrium state at a practically high speed even for a crowded game that models a complicated situation such as multi-point communication or two-point communication with budget constraints.
  • the present inventor may use the equilibrium state calculation device 10 according to the present embodiment to obtain a case where the calculation of the equilibrium state is about 1000 times faster than the case where all the contents of the strategy set are listed, or the contents of the strategy set. We have confirmed that there are cases where the calculation of the equilibrium state is completed in a few seconds even if it is not possible to list all of them due to memory and time constraints.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Business, Economics & Management (AREA)
  • Theoretical Computer Science (AREA)
  • General Business, Economics & Management (AREA)
  • Human Resources & Organizations (AREA)
  • Physics & Mathematics (AREA)
  • Economics (AREA)
  • General Physics & Mathematics (AREA)
  • Strategic Management (AREA)
  • Computer Security & Cryptography (AREA)
  • Quality & Reliability (AREA)
  • Tourism & Hospitality (AREA)
  • Game Theory and Decision Science (AREA)
  • Operations Research (AREA)
  • Marketing (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Software Systems (AREA)
  • Development Economics (AREA)
  • Entrepreneurship & Innovation (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

一実施形態に係る均衡状態計算装置は、混雑ゲームの均衡状態を計算する均衡状態計算装置であって、前記混雑ゲームのプレイヤーが使用する戦略であって、アイテムの組合せで表される戦略の集合をゼロサプレス型二分決定図で表現したグラフ情報を入力する入力手段と、前記入力手段により入力されたグラフ情報を用いて、Frank-Wolfeアルゴリズムの変種により前記均衡状態において前記戦略を選んだプレイヤーの割合を含む均衡状態情報を計算する計算手段と、ことを特徴とする。

Description

均衡状態計算装置、均衡状態計算方法及びプログラム
 本発明は、均衡状態計算装置、均衡状態計算方法及びプログラムに関する。
 ゲーム理論における非協力ゲームの1つとして混雑ゲーム(congestion game)が知られている。混雑ゲームとは、互いに協力しないプレイヤーたちがいくつかの資源を奪い合う状況又は互いに協力しないプレイヤーたちに資源を割り当てる状況をモデル化したものである。混雑ゲームの一種であるSelfish Routingでは、例えば通信量が多いほど遅延が増大する通信路からなる通信ネットワークにおいて多くの人(プレイヤー)が各々少ない遅延で2点間通信を行おうとする状況や、交通量が多いほど時間が掛かる道からなる道路ネットワークにおいてプレイヤーが各々短い時間で2点間を移動しようとする状況をモデル化できる。混雑ゲームはSelfish Routingをより一般化して、広範な戦略集合を扱えるようにしたもので、例えば、多くの人が各々少ない遅延で多点間通信を行おうとする状況や、2点間の通信でも通信路の遅延とは別に課金額が定められ、予算以下の課金額の通信ルートに制限して通信を行う状況も扱えるようになる。
 ここで、混雑ゲームでは、アイテム集合[n]:={1,・・・,n}に対して、各プレイヤーはアイテムの組合せS⊆[n]を選ぶことになる。ただし、どのようなアイテムの組合せを選べるかは予め決められていて、そのような組合せの集合が戦略集合であり、戦略集合の中身の各要素(つまり、アイテムの組合せ)は戦略と呼ばれる。各アイテムは選んだプレイヤーの割合が多いほどコストが高くなるように設定されており、プレイヤーが受けるコストは、選択した戦略中のアイテムのコストの和となる。このとき、各プレイヤーは互いに協力せず、各自勝手になるべくコストが小さい戦略を選ぼうとする。
 例えば、Selfish Routingは、通信ネットワークや道路ネットワーク等を抽象化したグラフ構造が与えられたときに、そのグラフ構造の各辺をアイテム、グラフ構造上の或る頂点から或る頂点までのパスが表すアイテムの組合せの集合を戦略集合とした混雑ゲームとなる。同様に、上述した多点間通信を各々行う状況は、グラフ構造上の或る頂点集合をターミナルとしたシュタイナー木を戦略集合とする混雑ゲームとしてモデル化できるし、一定以下の課金額に制限して通信を行う状況は、グラフ構造上の或る頂点から或る頂点までのパスの中で課金額を満たしたパスが表すアイテムの組合せの集合を戦略集合とする混雑ゲームとしてモデル化できる。
 混雑ゲームにおいて重要な状態として、均衡状態と呼ばれる状態がある。均衡状態とは、プレイヤーたちが不満に思わない状態のことであり、互いに協力しないプレイヤーたちが各自でコスト最小の状態に目指した結果として行き着く状態のことである。混雑ゲームにおける均衡状態を計算することができれば、例えば、通信ネットワークや道路ネットワークの設計時に、その設計により各通信路や道路にどの程度の混雑が発生するか、またプレイヤーの実際のコストがどの程度になるかをシミュレーションすることが可能となる。
 これまでに、Selfish Routingにおける均衡状態を近似的に求める手法が提案されている。例えば、グラフ構造上のフローのアルゴリズムを繰り返し用いることで、Selfish Routingにおける均衡状態を理論上多項式時間で求める技術が提案されている(非特許文献1)。また、実用的な均衡状態の計算手法としては、Frank-Wolfeアルゴリズムと呼ばれる最適化アルゴリズムがよく知られている(非特許文献2及び3)。
 また、戦略集合の各要素を全て保持した上でFrank-Wolfeアルゴリズムを用いることで、一般的な混雑ゲームにおける均衡状態を計算できることも知られている。
Alex Fabrikant, Christos Papadimitriou, and Kunal Talwar. The complexity of pure Nash equilibria. In Proceedings of the 36th Annual ACM Symposium on Theory of Computing, pp. 604-612, 2004. Marguerite Frank and Philip Wolfe. An algorithm for quadratic programming. Naval Research Logistics Quarterly, Vol. 3, pp. 95-110. Jose R. Correa and Nicolas Stier-Moses. Wardrop Equilibria. In Wiley Encyclopedia of Operations Research and Management Science.
 しかしながら、戦略集合のような組合せの集合の要素数は、もとのアイテムの集合[n]のサイズnに対して最大で2nとなり、一般にも指数的に大きくなることが多い。したがって、戦略集合の要素を全て保持するのは計算時間・メモリともに多大なコストを要し、例えば、nが数十程度でも実用上均衡状態を求めることが不可能になることが多い。
 本発明の一実施形態は、上記の点に鑑みてなされたもので、混雑ゲームの均衡状態を計算することを目的とする。
 上記目的を達成するため、一実施形態に係る均衡状態計算装置は、混雑ゲームの均衡状態を計算する均衡状態計算装置であって、前記混雑ゲームのプレイヤーが使用する戦略であって、アイテムの組合せで表される戦略の集合をゼロサプレス型二分決定図で表現したグラフ情報を入力する入力手段と、前記入力手段により入力されたグラフ情報を用いて、Frank-Wolfeアルゴリズムの変種により前記均衡状態において前記戦略を選んだプレイヤーの割合を含む均衡状態情報を計算する計算手段と、ことを特徴とする。
 混雑ゲームの均衡状態を計算することができる。
本実施形態に係る均衡状態計算装置の全体構成の一例を示す図である。 本実施形態に係る均衡状態計算処理の一例を示すフローチャートである。 本実施形態に係るCorrection処理の一例を示すフローチャートである。 本実施形態に係る均衡状態計算装置のハードウェア構成の一例を示す図である。
 以下、本発明の一実施形態について説明する。本実施形態では、混雑ゲームの均衡状態を計算することが可能な均衡状態計算装置10について説明する。
 本実施形態に係る均衡状態計算装置10は、ゼロサプレス型二分決定図(Zero-suppressed Binary Decision Diagram, 以降、「ZDD」と表記する。)をFrank-Wolfeアルゴリズムに組み入れることで、戦略集合に依存しない一般的な混雑ゲームの均衡状態を高速に計算可能にする。
 特に、本実施形態では、Frank-Wolfeアルゴリズムとして、その変種であるFully-corrective Frank-Wolfeアルゴリズム及びAway-step Frank-Wolfeアルゴリズムを用いる。これにより、本実施形態に係る均衡状態計算装置10は、その近似精度に保証がある均衡状態を求めることができる。
 なお、ZDDは、戦略集合のような組合せ集合をコンパクトに表現できる構造である。例えば、グラフ構造上の或る頂点から或る頂点までのパスの集合、シュタイナー木の集合、課金額の制約を満たすパスの集合等はいずれもZDDで表現することが可能である。組合せ集合を表現するZDDは、例えば、フロンティア法(Frontier-based method)等で構築することが可能である。また、フロンティア法を用いたライブラリとしてGraphillion等が知られており、このようなライブラリを利用することで効率的にZDDを構築することが可能である。フロンティア法については、例えば、参考文献1「Jun Kawahara, Takeru Inoue, Hiroaki Iwashita, and Shin-ichi Minato. Frontier-based search for enumerating all constrained subgraphs with compressed representation. IEICE TRANSACTIONS on Fundamentals of Electronics, Communications and Computer Sciences, Vol. E100-A, pp. 1773-1784, 2017.」等を参照されたい。Graphillionについては、例えば、参考文献2「GitHub - takemaru-graphillion Fast, lightweight graphset operation library, インターネット<URL:https://github.com/takemaru/graphillion/>」等を参照されたい。
 また、ZDDについては、例えば、参考文献3「Shin-ichi Minato. Zero-suppressed BDDs for set manipulation in combinatorial problems. In Proceedings of the 30th ACM/IEEE Design Automation Conference, pp. 272-277, 1993.」等を参照されたい。Fully-corrective Frank-Wolfeアルゴリズム及びAway-step Frank-Wolfeアルゴリズムについては、例えば、参考文献4「Simon Lacoste-Julien and Martin Jaggi. On the global linear convergence of Frank-Wolfe optimization variants. In Proceedings of the 28th International Conference on Neural Information Processing Systems, Vol. 1, pp. 496-504, 2015.」等を参照されたい。
 <混雑ゲーム>
 まず、混雑ゲームについて説明する。混雑ゲームでは、アイテム集合[n]:={1,・・・,n}の各アイテムに対して、その使用率yiに関して単調非減少なコスト関数ci(yi)が与えられる。また、アイテム集合[n]に対して、各プレイヤーはアイテムの組合せS⊆[n]を選ぶことになる。ただし、どのようなアイテムの組合せを選べるかは予め決められていて、そのような組合せの集合が戦略集合
Figure JPOXMLDOC01-appb-M000001
である。また、戦略集合の中身の各要素(つまり、アイテムの組合せ)は戦略と呼ばれる。本実施形態で扱う混雑ゲームでは、プレイヤーは無数に多くいて、各戦略Sに対してその戦略を選んだプレイヤーの割合zSを考えることとする。なお、
Figure JPOXMLDOC01-appb-M000002
である。
 各戦略に対してその戦略を選択したプレイヤーの割合が決まると、アイテムiの使用率yiは、アイテムiを含む戦略を選んだプレイヤーの割合の和、すなわち、
Figure JPOXMLDOC01-appb-M000003
で求まる。そして、戦略SのコストをcSと表すことにすると、これは戦略Sに含まれるアイテムのコストの和、すなわち、
Figure JPOXMLDOC01-appb-M000004
により求まる。このとき、各プレイヤーは互いに協力せず、各自勝手になるべくコストが小さい戦略を選ぼうとする。すると、もし或るプレイヤーについて、今選んでいる戦略よりもコストが小さい戦略があれば、その戦略を選び直すことでよりコストを小さくすることができるため、当該プレイヤーは戦略を変えることになる。このような戦略の変更が起こらない状態がWardrop均衡状態と呼ばれる均衡状態である。Wardrop均衡状態は、プレイヤーの割合が0より大きい戦略についてはコストが全戦略中で最小になっている状態、すなわち、
Figure JPOXMLDOC01-appb-M000005
が成立している状態と定義される。
 上述した混雑ゲームをより一般化して、プレイヤーによって違った戦略集合を使う場合について説明する。この場合、r個のプレイヤー群の集合を[r]と表記すれば、戦略集合は1つではなく、複数種類のr個の戦略集合
Figure JPOXMLDOC01-appb-M000006
が与えられ、各々の戦略集合を使うプレイヤーの割合m1,・・・,mrも同時に与えられるものとする。プレイヤー群とは0以上のプレイヤーの集合である。なお、
Figure JPOXMLDOC01-appb-M000007
である。
 この状況では、各プレイヤー群p∈[r]に対して、各戦略
Figure JPOXMLDOC01-appb-M000008
についてその戦略を選んだプレイヤーの割合zS pを考えることになる。なお、
Figure JPOXMLDOC01-appb-M000009
である。
 上記の割合zS pが決まれば、
Figure JPOXMLDOC01-appb-M000010
とおくことで、各アイテムの使用率は
Figure JPOXMLDOC01-appb-M000011
と計算することができる。したがって、戦略Sのコストは、
Figure JPOXMLDOC01-appb-M000012
と計算することができる。このときのWardrop均衡状態とは、プレイヤーの割合が0より大きい戦略についてはコストが該当の戦略集合の戦略中で最小になっている状態、すなわち、
Figure JPOXMLDOC01-appb-M000013
が成立している状態と定義される。
 本実施形態では、プレイヤーによって違った戦略集合を使う混雑ゲームを想定して、その均衡状態としてε-近似Wardrop均衡状態を求めるものとする。ε-近似Wardrop均衡状態とは、プレイヤーの割合が0より大きい戦略についてはコストが、該当の戦略集合の戦略中で最小のコストよりも許容誤差ε以上大きいことはない状態、すなわち、
Figure JPOXMLDOC01-appb-M000014
が成立している状態と定義される。これは、Wardrop均衡状態に対してその近似誤差がε以内であることを保証していることを意味する。
 また、1S∈{0,1}nをアイテムi∈[n]が戦略Sに含まれていれば(つまり、i∈Sであれば)i番目の要素が1、そうでなければi番目の要素が0であるn次元ベクトルとする。このn次元ベクトル1Sを用いれば、上記のxi pをi番目の要素するn次元ベクトルxpは、
Figure JPOXMLDOC01-appb-M000015
と表すことができる。また、このn次元ベクトルxpを用いれば、上記の使用率yiをi番目の要素とするn次元ベクトルyは、
Figure JPOXMLDOC01-appb-M000016
と表すことができる。したがって、戦略SのコストcSをベクトルyの関数と見做した場合、
Figure JPOXMLDOC01-appb-M000017
と表すことができる。また、このコスト関数コストcS(・)を用いれば、
Figure JPOXMLDOC01-appb-M000018
として、Rrn上のコスト関数CSを
Figure JPOXMLDOC01-appb-M000019
と定義することができる。
 また、ポテンシャル関数Φ:Rn→Rを
Figure JPOXMLDOC01-appb-M000020
と定義する。本実施形態では、ZDDで表現した戦略集合とFrank-Wolfeアルゴリズムの変種とを用いて、このポテンシャル関数Φに関する最小化問題を解くことで、ε-近似Wardrop均衡状態を求める。なお、
Figure JPOXMLDOC01-appb-M000021
である。ただし、∇Φ(y)iは∇Φ(y)のi番目の要素(つまり、Φのyiに関する偏微分)を表す。
 <全体構成>
 次に、本実施形態に係る均衡状態計算装置10の全体構成について、図1を参照しながら説明する。図1は、本実施形態に係る均衡状態計算装置10の全体構成の一例を示す図である。
 図1に示すように、本実施形態に係る均衡状態計算装置10は、入力部101と、最適化部102と、出力部103と、記憶部104とを有する。
 記憶部104には、混雑ゲームにおけるε-近似Wardrop均衡状態の計算に必要な各種情報が記憶されている。記憶部104に記憶されている情報としては、例えば、アイテム集合[n]、各アイテムのコスト関数ci(yi)、1以上の戦略集合の各々をZDDで表現した情報、プレイヤー群の集合[r]、各戦略集合を使うプレイヤーの割合m1,・・・,mr、許容誤差ε等がある。以降では、ZDDで表現された戦略集合を
Figure JPOXMLDOC01-appb-M000022
と表す。なお、上記の情報以外にも、記憶部104には、ε-近似Wardrop均衡状態の計算過程等の情報が記憶されてもよい。
 ここで、戦略集合を表すZDDは、ノード集合とノード間を接続する有向エッジのエッジ集合とで構成される有向非巡回グラフ(DAG:Directed Acyclic Graph))である。ノード集合の中には、アイテムを表すノードvの他に、終端ノード⊥と終端ノード
Figure JPOXMLDOC01-appb-M000023
とが含まれる。また、各ノードvからは「0-枝」及び「1-枝」と呼ばれる2つのエッジが出る。本実施形態では、ノードvから出る1-枝が指す先のノードを1-子ノードと呼び、v1で表す。同様に、ノードvから出る0-枝が指す先のノードを0-子ノードと呼び、v0で表す。また、各ノードvのうちの根ノードをノードrとも表す。
 更に、各ノードvには、ラベルと呼ばれる整数値lv∈{1,・・・,n}が付与されており、このラベルによってアイテムとノードとが対応付けられている。なお、終端ノードについては、例えば、ラベルの値をn+1等としておけばよい。
 このとき、ZDDは、各ノードの0-枝と1-枝が、必ずラベルが小さい方のノードから大きい方のノードに向かうようにする。すなわち、任意のノードvに対して、(ノードvのラベル)<(ノードv0のラベル)かつ(ノードvのラベル)<(ノードv1のラベル)が成立するようにする。これにより、ZDDはラベルの値で階層化され、例えば、第1層目に含まれるノード(つまり、ノードr)はアイテム1に対応し、第2層目に含まれるノードvはアイテム2に対応する。このように、ZDDの第i層目に含まれるノードvは、アイテムiに対応することになる。
 したがって、根ノードrから終端ノードへの各パス(経路)でアイテムの組合せ(つまり、戦略)を表現することができる。すなわち、ノードvからノードv1へのエッジがパスに含まれる場合は当該ノードvのラベルに対応するアイテムが戦略に含まれ、ノードvからノードv0へのエッジがパスに含まれる場合は当該ノードvのラベルに対応するアイテムは戦略に含まれないとすることで、パスで組合せ(戦略)を表現することができる。
 入力部101は、アイテム集合[n]、各アイテムのコスト関数ci(yi)、ZDDで表現された1以上の戦略集合、プレイヤー群の集合[r]、各戦略集合を使うプレイヤーの割合m1,・・・,mr、許容誤差ε等の各種情報を入力する。
 最適化部102は、Fully-corrective Frank-Wolfeアルゴリズムに基づく処理によりε-近似Wardrop均衡状態における各種情報を求める。より具体的には、最適化部102は、入力部101により入力された各種情報を用いてポテンシャル関数Φに関する最小化問題を解くことで、ε-近似Wardrop均衡状態における各種情報を求める。これにより、ε-近似Wardrop均衡状態において各戦略Sを選んだプレイヤーの割合zS pが得られる。
 出力部103は、最適化部102により求めた各種情報(例えば、ε-近似Wardrop均衡状態において各戦略Sを選んだプレイヤーの割合zS p等)を出力する。なお、出力部103の出力先は限定されず、任意の出力先としてよい。例えば、出力部103の出力先は、記憶部104であってもよいし、ディスプレイ等の表示装置であってもよいし、通信ネットワークを介して接続されるデータベースサーバ等であってもよい。
 ここで、最適化部102には、初期設定部111と、最短経路計算部112と、更新部113とが含まれる。
 初期設定部111は、Frank-Wolfeアルゴリズムの変種で更新対象となる各種変数(パラメータ)の初期化等を行う。パラメータは、上述したn次元ベクトルxp、各プレイヤーが現在選択している戦略Sの集合を表すアクティブ集合、及び各戦略Sを選んだプレイヤーの割合zS pである。
 最短経路計算部112は、戦略集合を表すZDD上の最短経路を動的計画法により計算することで、戦略集合の中でコストが最小になる戦略を計算する。
 更新部113は、Away-step Frank-Wolfeアルゴリズムに基づくCorrection処理により、各種パラメータの更新を行う。
 <均衡状態計算処理>
 次に、本実施形態に係る均衡状態計算装置10によって混雑ゲームのε-近似Wardrop均衡状態を計算する均衡状態処理について、図2を参照しながら説明する。図2は、本実施形態に係る均衡状態計算処理の一例を示すフローチャートである。
 ステップS1100では、入力部101は、各種情報(アイテム集合[n]、各アイテムのコスト関数ci(yi)、ZDDで表現された1以上の戦略集合、プレイヤー群の集合[r]、各戦略集合を使うプレイヤーの割合m1,・・・,mr、許容誤差ε等)を入力する。
 ステップS1200では、最適化部102は、初期設定部111により、各p∈[r]に対して、戦略
Figure JPOXMLDOC01-appb-M000024
を選択する。これらの戦略Spが、プレイヤー群pが最初に選択した戦略となる。
 ステップS1300では、最適化部102は、初期設定部111により、各種パラメータ(n次元ベクトルx0 p、アクティブ集合、各戦略Sを選んだプレイヤーの割合zS p)のそれぞれを以下のように初期化する。
Figure JPOXMLDOC01-appb-M000025
 また、簡単のため、
Figure JPOXMLDOC01-appb-M000026
と表記する。
 ステップS1400では、最適化部102は、繰り返し回数を表すインデックスをkとして、k=0,1,・・・,Kに対してステップS1410~ステップS1450を繰り返し実行する。Kは予め設定されたハイパーパラメータである。なお、以降のステップS1410~ステップS1450では、繰り返し回数がk回目である場合について説明し、zを除く各種記号の右下の添え字は繰り返し回数を表すものとする。例えば、n次元ベクトルykは、繰り返し回数がk回目のときのn次元ベクトルyを表す。
 ステップS1410では、最適化部102は、使用率yiをi番目の要素とするn次元ベクトルykを
Figure JPOXMLDOC01-appb-M000027
により計算する。
 ステップS1420では、最適化部102は、各p∈[r]に対して、ステップS1421~ステップS1422を繰り返し実行する。なお、以降のステップS1421~ステップS1422では、或るpに関するステップS1421~ステップS1422について説明する。
 ステップS1421では、最適化部102は、最短経路計算部112により、ZDD
Figure JPOXMLDOC01-appb-M000028
上の最短経路を動的計画法により計算することで、pに対応する戦略集合の中でコストが最小となる戦略sk pを計算する。すなわち、最短経路計算部112は、pに対応する戦略集合を表現するZDD上の最短経路を計算することで、
Figure JPOXMLDOC01-appb-M000029
を計算する。なお、<・,・>は内積を表す。
 具体的には、最短経路計算部112は、次のようにして戦略sk pを計算する。
 まず、最短経路計算部112は、ZDDの各ノードvに対して、0-枝の距離を0、1-枝の距離を
Figure JPOXMLDOC01-appb-M000030
とする。つまり、ノードvの1-枝の距離を、当該ノードvのラベルに対応するアイテムのコストとする。
 そして、最短経路計算部112は、根ノードrから終端ノードへのパスであって、その距離の合計が最小となるパス(最短経路)を動的計画法により計算する。これにより、ZDD上で距離が最小となるパスで表されるアイテムの組合せが、pに対応する戦略集合の中でコストが最小となる戦略sk pとして得られる。なお、ZDD等の有向非巡回グラフ上の最短経路を計算する方法は広く知られており、例えば、参考文献5「渋谷哲朗, "情報工学 アルゴリズム", 丸善出版, 2016年11月」等を参照されたい。
 ステップS1422では、最適化部102は、現在の状態においてpに対応する戦略集合を使用しているプレイヤーの平均コストと、この戦略集合の中でコスト最小の戦略(つまり、戦略sk p)のコストとの差gk pを計算する。すなわち、最適化部102は、
Figure JPOXMLDOC01-appb-M000031
として、
Figure JPOXMLDOC01-appb-M000032
を計算する。
 なお、<-∇Φ(yk),dk p>=<∇Φ(yk),xk p>-<∇Φ(yk),sk p>であり、<∇Φ(yk),xk p>はpに対応する戦略集合を使用しているプレイヤーの平均コストを表しており、<∇Φ(yk),sk p>は戦略sk pのコストを表している。
 ステップS1430では、最適化部102は、全てのp∈[r]に対して、平均コストとコスト最小の戦略のコストとの差gk pが許容誤差ε以下であるか否かを判定する。すなわち、最適化部102は、
Figure JPOXMLDOC01-appb-M000033
を満たすか否かを判定する。
 全てのp∈[r]に対してgk pが許容誤差ε以下であると判定された場合(ステップS1430でYES)はステップS1440が実行され、そうでなければ(ステップS1430でNO)ステップS1440は実行されない。
 ステップS1440では、出力部103は、現時点のパラメータ
Figure JPOXMLDOC01-appb-M000034
を出力する。これらのパラメータは、それぞれ、ε-近似Wardrop均衡状態におけるn次元ベクトルx、アクティブ集合及び各戦略を選んだプレイヤーの割合である。なお、これらのパラメータがε-近似Wardrop均衡状態を満たす理由については後述する。
 ステップS1450では、最適化部102は、更新部113によりCorrection処理を実行して各種パラメータの更新を行う。すなわち、最適化部102は、サブルーチン
Figure JPOXMLDOC01-appb-M000035
を呼び出して、更新後のパラメータ
Figure JPOXMLDOC01-appb-M000036
を得る。
 <Correction処理>
 ここで、上記のステップS1450のCorrection処理の詳細について、図3を参照しながら説明する。図3は、本実施形態に係るCorrection処理の一例を示すフローチャートである。なお、以降では、簡単のため、図2のステップS1400におけるインデックスkを省略して、サブルーチン
Figure JPOXMLDOC01-appb-M000037
が呼び出されたものとして説明をする。なお、x0=x=(x1,・・・,xr)とする。
 ステップS2100では、更新部113は、各p∈[r]に対して、ステップS2110を繰り返し実行する。なお、以降のステップS2110では、或るpに関するステップS2110について説明する。
 ステップS2110では、更新部113は、
Figure JPOXMLDOC01-appb-M000038
として、新たな戦略集合
Figure JPOXMLDOC01-appb-M000039
を作成する。
 ステップS2200では、更新部113は、繰り返し回数を表すインデックスをlとして、l=0,1,・・・,Lに対してステップS2210~ステップS2280を繰り返し実行する。Lは予め設定されたハイパーパラメータである。なお、以降のステップS2210~ステップS2280では、繰り返し回数がl回目である場合について説明し、zを除く各種記号の右下の添え字は繰り返し回数を表すものとする。例えば、n次元ベクトルylは、繰り返し回数がl回目のときのn次元ベクトルyを表す。
 ステップS2210では、更新部113は、使用率yiをi番目の要素とするn次元ベクトルylを
Figure JPOXMLDOC01-appb-M000040
により計算する。
 ステップS2220では、更新部113は、各p∈[r]に対して、ステップS2221を繰り返し実行する。ただし、ステップS2240が実行された場合は、Correction処理を終了し、サブルーチンの呼び出し元に戻る。なお、以降のステップS2221では、或るpに関するステップS2221について説明する。
 ステップS2221では、更新部113は、現時点において、pに対応する新たな戦略集合の中でコストが最小になる戦略sl pと、pに対応するアクティブ集合の中でコストが最大になる戦略vl pとを計算する。具体的には、更新部113は、
Figure JPOXMLDOC01-appb-M000041
を計算する。また、このとき、更新部113は、
Figure JPOXMLDOC01-appb-M000042
も計算する。ここで、dl p,FWはxpからコスト最小の戦略sl pに向かう方向を表し、dl p,Aはxpからコスト最大の戦略vl pに向かう方向と逆方向を表す。
 なお、上記のステップS2221において、更新部113は、各戦略に対するコストを計算することで、戦略sl pと戦略vl pとを計算すればよい。これは、pに対応する戦略集合と比較して、pに対応する新たな戦略集合やアクティブ集合の大きさは非常に小さい(せいぜいO(n)程度)ためである。
 ステップS2230では、更新部113は、全てのp∈[r]に対して、戦略vl pのコストと戦略sl pのコストとの差が許容誤差ε以下であるか否かを判定する。すなわち、更新部113は、
Figure JPOXMLDOC01-appb-M000043
を満たすか否かを判定する。なお、上記の内積部分は<∇Φ(yl),vl p>-<∇Φ(yl),sl p>であり、<∇Φ(yl),vl p>は戦略vl pのコスト、<∇Φ(yl),sl p>は戦略sl pのコストをそれぞれ表している。
 全てのp∈[r]に対して戦略vl pのコストと戦略sl pのコストとの差が許容誤差ε以下であると判定された場合(ステップS2230でYES)はステップS2240が実行され、そうでなければ(ステップS2230でNO)ステップS2250~ステップS2280が実行される。
 ステップS2240では、更新部113は、パラメータ
Figure JPOXMLDOC01-appb-M000044
をそれぞれ
Figure JPOXMLDOC01-appb-M000045
として、現時点のパラメータ
Figure JPOXMLDOC01-appb-M000046
を出力(つまり、サブルーチンの呼び出し元に出力)する。そして、更新部113は、Correction処理を終了し、サブルーチンの呼び出し元に戻る。
 ステップS2250では、更新部113は、
Figure JPOXMLDOC01-appb-M000047
によりgl FW及びgl Aを計算する。
 ステップS2260では、更新部113は、gl FW及びgl Aの大小関係に応じて、dl及びγmaxを計算する。すなわち、更新部113は、gl FW≧gl Aである場合は
Figure JPOXMLDOC01-appb-M000048
によりdl及びγmaxを計算し、gl FW<gl Aである場合は
Figure JPOXMLDOC01-appb-M000049
によりdl及びγmaxを計算する。このことは、xlを更新する際に、gl FW≧gl Aであればxpをdl p,FWの方向に進めることとし、そうでなければxpをdl p,Aの方向に進めることを意味する。
 ステップS2270では、更新部113は、
Figure JPOXMLDOC01-appb-M000050
によりγlを計算する。ここで、関数Fは、
Figure JPOXMLDOC01-appb-M000051
である。すなわち、更新部113は、xlからdlの方向に進んだときに関数Fの値が最小になる点γlを求める。これは、例えば、ラインサーチにより求めればよい。
 ステップS2280では、更新部113は、パラメータを更新する。ここで、更新部113は、以下のようにパラメータを更新する。
 まず、xlについては、
Figure JPOXMLDOC01-appb-M000052
と更新する。これは、現時点で関数Fの値が最小となるxをxl+1とすることを意味する。
 また、各戦略を選んだプレイヤーの割合zについては、gl FW≧gl Aである場合は、各p∈[r]に対して、
Figure JPOXMLDOC01-appb-M000053
により更新する。一方で、gl FW<gl Aである場合は、各p∈[r]に対して、
Figure JPOXMLDOC01-appb-M000054
により更新する。なお、各戦略を選んだプレイヤーの割合zの更新方法については、例えば、上記の参考文献4等を参照されたい。
 また、アクティブ集合については、各p∈[r]に対して、
Figure JPOXMLDOC01-appb-M000055
により更新する。つまり、zS p>0となる戦略Sだけを集めて、新たなアクティブ集合としている。
 <パラメータがε-近似Wardrop均衡状態を満たす理由>
 次に、均衡状態計算処理で出力されたパラメータがε-近似Wardrop均衡状態を満たす理由について説明する。コスト関数CSを用いた場合、ε-近似Wardrop均衡状態は、任意のp∈[r]を1つ固定したときに、任意の
Figure JPOXMLDOC01-appb-M000056
に対して、均衡状態計算処理で出力されたパラメータxが
Figure JPOXMLDOC01-appb-M000057
を満たすことと言い換えることができる。
 図2のステップS1430により、
Figure JPOXMLDOC01-appb-M000058
が成り立つことが保証される。一方で、図3のステップS2230により、
Figure JPOXMLDOC01-appb-M000059
が成り立つことが保証される。
 したがって、上記の2つの不等式により、CS(x)-ε≦CS´(x)+εが成り立つ。これにより、均衡状態計算処理で出力されたパラメータはε-近似Wardrop均衡状態を満たす。
 <ハードウェア構成>
 次に、本実施形態に係る均衡状態計算装置10のハードウェア構成について、図4を参照しながら説明する。図4は、本実施形態に係る均衡状態計算装置10のハードウェア構成の一例を示す図である。
 図4に示すように、本実施形態に係る均衡状態計算装置10は一般的なコンピュータ又はコンピュータシステムで実現され、入力装置201と、表示装置202と、外部I/F203と、通信I/F204と、プロセッサ205と、メモリ装置206とを有する。これら各ハードウェアは、それぞれがバス207を介して通信可能に接続されている。
 入力装置201は、例えば、キーボードやマウス、タッチパネル等である。表示装置202は、例えば、ディスプレイ等である。なお、均衡状態計算装置10は、入力装置201及び表示装置202のうちの少なくとも一方を有していなくてもよい。
 外部I/F203は、外部装置とのインタフェースである。外部装置には、例えば、記録媒体203a等がある。均衡状態計算装置10は、外部I/F203を介して、記録媒体203aの読み取りや書き込み等を行うことができる。記録媒体203aには、例えば、均衡状態計算装置10が有する各機能部(入力部101、最適化部102及び出力部103)を実現する1以上のプログラムが格納されていてもよい。
 なお、記録媒体203aには、例えば、CD(Compact Disc)、DVD(Digital Versatile Disk)、SDメモリカード(Secure Digital memory card)、USB(Universal Serial Bus)メモリカード等がある。
 通信I/F204は、均衡状態計算装置10を通信ネットワークに接続するためのインタフェースである。なお、均衡状態計算装置10が有する各機能部を実現する1以上のプログラムは、通信I/F204を介して、所定のサーバ装置等から取得(ダウンロード)されてもよい。
 プロセッサ205は、例えば、CPU(Central Processing Unit)やGPU(Graphics Processing Unit)等の各種演算装置である。均衡状態計算装置10が有する各機能部は、例えば、メモリ装置206等に格納されている1以上のプログラムがプロセッサ205に実行させる処理により実現される。
 メモリ装置206は、例えば、HDD(Hard Disk Drive)やSSD(Solid State Drive)、RAM(Random Access Memory)、ROM(Read Only Memory)、フラッシュメモリ等の各種記憶装置である。均衡状態計算装置10が有する記憶部104は、例えば、メモリ装置206を用いて実現可能である。なお、例えば、記憶部104は、均衡状態計算装置10と通信ネットワークを介して接続される記憶装置等を用いて実現されていてもよい。
 本実施形態に係る均衡状態計算装置10は、図4に示すハードウェア構成を有することにより、上述した均衡状態計算処理を実現することができる。なお、図4に示すハードウェア構成は一例であって、均衡状態計算装置10は、他のハードウェア構成を有していてもよい。例えば、均衡状態計算装置10は、複数のプロセッサ205を有していてもよいし、複数のメモリ装置206を有していてもよい。
 <まとめ>
 以上のように、本実施形態に係る均衡状態計算装置10は、一般の戦略集合からなる混雑ゲームに対しても、近似の精度に保証がある均衡状態(ε-近似Wardrop均衡状態)を実用上高速に求めることができる。これにより、例えば、多点間通信や予算制約付きの2点間通信等の複雑な状況をモデル化した混雑ゲームに対しても、実用上高速に均衡状態を計算することができる。
 したがって、例えば、通信ネットワークや高速ネットワークの設計時に、その設計により各通信路や道路にどの程度の混雑が発生するか、またプレイヤーの実際のコストがどの程度になるかをシミュレーションすることが可能となる。このため、例えば、設計に複数の案が存在する場合等に、そのシミュレーションによる性能を比較することが可能になる。
 なお、本件発明者は、本実施形態に係る均衡状態計算装置10により、戦略集合の中身を全て列挙する場合と比べて均衡状態の計算が1000倍程度高速になるケースや、戦略集合の中身を全て列挙することがメモリや時間の制約上叶わない場合でも均衡状態の計算が数秒で終わるケースがあることを確認している。
 本発明は、具体的に開示された上記の実施形態に限定されるものではなく、請求の範囲の記載から逸脱することなく、種々の変形や変更、既知の技術との組み合わせ等が可能である。
 10    均衡状態計算装置
 101   入力部
 102   最適化部
 103   出力部
 104   記憶部
 111   初期設定部
 112   最短経路計算部
 113   更新部

Claims (8)

  1.  混雑ゲームの均衡状態を計算する均衡状態計算装置であって、
     前記混雑ゲームのプレイヤーが使用する戦略であって、アイテムの組合せで表される戦略の集合をゼロサプレス型二分決定図で表現したグラフ情報を入力する入力手段と、
     前記入力手段により入力されたグラフ情報を用いて、Frank-Wolfeアルゴリズムの変種により前記均衡状態において前記戦略を選んだプレイヤーの割合を含む均衡状態情報を計算する計算手段と、
     ことを特徴とする均衡状態計算装置。
  2.  前記計算手段は、
     前記グラフ情報が表すゼロサプレス型二分決定図の各ノードに対して、前記ノードの0-枝の距離を0、前記ノードの1-枝の距離を、前記ノードに対応するアイテムのコストとして、前記ゼロサプレス型二分決定図の根ノードから終端ノードまでの最短経路を動的計画法により探索することで、前記コストが最小となる戦略を表す第1のコスト最小戦略を計算し、
     前記第1のコスト最小戦略を用いて、前記均衡状態情報を更新する、ことを特徴とする請求項1に記載の均衡状態計算装置。
  3.  前記計算手段は、
     前記第1のコスト最小戦略の計算と前記均衡状態情報の更新とを、前記戦略の集合を使用するプレイヤーの平均コストと前記第1のコスト最小戦略のコストとの差が所定の誤差以下となるまで繰り返し実行する、ことを特徴とする請求項2に記載の均衡状態計算装置。
  4.  前記計算手段は、
     前記第1のコスト最小戦略を用いて、Away-step Frank-Wolfeアルゴリズムに基づくアルゴリズムにより前記均衡状態情報を更新する、ことを特徴とする請求項2又は3に記載の均衡状態計算装置。
  5.  前記計算手段は、
     前記第1のコスト最小戦略と前記プレイヤーが現在選択している戦略の集合を表すアクティブ集合とから作成された新たな戦略集合を用いて、前記新たな戦略集合の中でコストが最小となる戦略を表す第2のコスト最小戦略を計算し、
     前記アクティブ集合の中でコストが最大となる戦略を表すコスト最大戦略を計算し、
     前記コスト最大戦略と前記第2のコスト最小戦略とを用いて前記均衡状態情報を更新する、ことを特徴とする請求項2乃至4の何れか一項に記載の均衡状態計算装置。
  6.  前記計算手段は、
     前記コスト最大戦略と前記第2のコスト最小戦略との差が所定の誤差以下となるまで、前記均衡状態情報の更新を繰り返し実行する、請求項5に記載の均衡状態計算装置。
  7.  混雑ゲームの均衡状態を計算する均衡状態計算装置が、
     前記混雑ゲームのプレイヤーが使用する戦略であって、アイテムの組合せで表される戦略の集合をゼロサプレス型二分決定図で表現したグラフ情報を入力する入力手順と、
     前記入力手順により入力された情報を用いて、Frank-Wolfeアルゴリズムの変種により前記均衡状態において前記戦略を選んだプレイヤーの割合を含む均衡状態情報を計算する計算手順と、
     を実行することを特徴とする均衡状態計算方法。
  8.  コンピュータを、請求項1乃至6の何れか一項に記載の均衡状態計算装置における各手段として機能させるためのプログラム。
PCT/JP2019/050675 2019-12-24 2019-12-24 均衡状態計算装置、均衡状態計算方法及びプログラム Ceased WO2021130867A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
US17/787,859 US20230052060A1 (en) 2019-12-24 2019-12-24 Equilibrium calculation apparatus, equilibrium calculation method and program
JP2021566610A JP7279820B2 (ja) 2019-12-24 2019-12-24 均衡状態計算装置、均衡状態計算方法及びプログラム
PCT/JP2019/050675 WO2021130867A1 (ja) 2019-12-24 2019-12-24 均衡状態計算装置、均衡状態計算方法及びプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2019/050675 WO2021130867A1 (ja) 2019-12-24 2019-12-24 均衡状態計算装置、均衡状態計算方法及びプログラム

Publications (1)

Publication Number Publication Date
WO2021130867A1 true WO2021130867A1 (ja) 2021-07-01

Family

ID=76575811

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2019/050675 Ceased WO2021130867A1 (ja) 2019-12-24 2019-12-24 均衡状態計算装置、均衡状態計算方法及びプログラム

Country Status (3)

Country Link
US (1) US20230052060A1 (ja)
JP (1) JP7279820B2 (ja)
WO (1) WO2021130867A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7222441B1 (ja) 2022-08-05 2023-02-15 富士電機株式会社 分析装置、分析方法及びプログラム

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2013147812A1 (en) * 2012-03-29 2013-10-03 Empire Technology Development, Llc Resource management for data center based gaming
WO2014053149A1 (en) * 2012-10-05 2014-04-10 Andrew Wireless Systems Gmbh Capacity optimization sub-system for distributed antenna system
US11954612B2 (en) * 2017-09-05 2024-04-09 International Business Machines Corporation Cognitive moderator for cognitive instances
US11134026B2 (en) * 2018-06-01 2021-09-28 Huawei Technologies Co., Ltd. Self-configuration of servers and services in a datacenter
CN110472354B (zh) * 2019-08-20 2023-07-18 东南大学 一种新能源汽车渗透的环境影响评估方法
US12387597B2 (en) * 2020-10-13 2025-08-12 The Regents Of The University Of Colorado, A Body Corporate Systems and methods for system optimal traffic routing

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
"Wiley Encyclopedia of Operations Research and Management Science", 13 December 2017, JOHN WILEY & SONS, INC., Hoboken, NJ, USA, ISBN: 978-0-470-40053-1, article COCHRAN JAMES J., COX LOUIS A., KESKINOCAK PINAR, KHAROUFEH JEFFREY P., SMITH J. COLE, CORREA JOSÉ R., STIER-MOSES NICOLÁS E.: "Wardrop Equilibria", XP055893973, DOI: 10.1002/9780470400531.eorms0962 *
ALEX FABRIKANT ; CHRISTOS PAPADIMITRIOU ; KUNAL TALWAR: "The complexity of pure Nash equilibria", THEORY OF COMPUTING, ACM, 2 PENN PLAZA, SUITE 701NEW YORKNY10121-0701USA, 13 June 2004 (2004-06-13) - 16 June 2004 (2004-06-16), 2 Penn Plaza, Suite 701New YorkNY10121-0701USA , pages 604 - 612, XP058394297, ISBN: 978-1-58113-852-8, DOI: 10.1145/1007352.1007445 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7222441B1 (ja) 2022-08-05 2023-02-15 富士電機株式会社 分析装置、分析方法及びプログラム
JP2024022159A (ja) * 2022-08-05 2024-02-16 富士電機株式会社 分析装置、分析方法及びプログラム

Also Published As

Publication number Publication date
JP7279820B2 (ja) 2023-05-23
US20230052060A1 (en) 2023-02-16
JPWO2021130867A1 (ja) 2021-07-01

Similar Documents

Publication Publication Date Title
Durand et al. Complexity and optimality of the best response algorithm in random potential games
Ravichandiran Deep Reinforcement Learning with Python
US10217072B2 (en) Association-based product design
US20240020543A1 (en) Glp-1/gip dual agonists
Griffin et al. In search of lost mixing time: adaptive Markov chain Monte Carlo schemes for Bayesian variable selection with very large p
Amr Hands-On Machine Learning with scikit-learn and Scientific Python Toolkits: A practical guide to implementing supervised and unsupervised machine learning algorithms in Python
US20160132787A1 (en) Distributed, multi-model, self-learning platform for machine learning
CN113646784A (zh) 信息处理装置、信息处理系统、信息处理方法、存储介质及程序
US11763084B2 (en) Automatic formulation of data science problem statements
Analui et al. On distributionally robust multiperiod stochastic optimization
Jain et al. TensorFlow Machine Learning Projects: Build 13 real-world projects with advanced numerical computations using the Python ecosystem
Venturelli et al. A Kriging-assisted multiobjective evolutionary algorithm
Mueller et al. Algorithms for dummies
Hagg et al. Designing air flow with surrogate-assisted phenotypic niching
Zhai et al. Deep q-learning with prioritized sampling
CN120112901A (zh) 优化用于硬件装置的算法
CN117093732B (zh) 多媒体资源的推荐方法及装置
CN120380487A (zh) 特定于上下文的机器学习模型的生成和部署
Pilát et al. Feature extraction for surrogate models in genetic programming
US12327194B2 (en) Active learning using causal network feedback
JP7279820B2 (ja) 均衡状態計算装置、均衡状態計算方法及びプログラム
US20210365822A1 (en) Computer implemented method for generating generalized additive models
Guzdial et al. Classical pcg
Petrowski et al. Evolutionary algorithms
JP2022178099A (ja) 最適化方法、最適化装置、及びプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19958020

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021566610

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19958020

Country of ref document: EP

Kind code of ref document: A1