WO2022123743A1 - パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム - Google Patents

パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム Download PDF

Info

Publication number
WO2022123743A1
WO2022123743A1 PCT/JP2020/046123 JP2020046123W WO2022123743A1 WO 2022123743 A1 WO2022123743 A1 WO 2022123743A1 JP 2020046123 W JP2020046123 W JP 2020046123W WO 2022123743 A1 WO2022123743 A1 WO 2022123743A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
aggregated data
parameters
aggregated
parameter estimation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/046123
Other languages
English (en)
French (fr)
Inventor
佑典 田中
具治 岩田
健 倉島
浩之 戸田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to US18/255,925 priority Critical patent/US20240045921A1/en
Priority to JP2022567985A priority patent/JP7439957B2/ja
Priority to PCT/JP2020/046123 priority patent/WO2022123743A1/ja
Publication of WO2022123743A1 publication Critical patent/WO2022123743A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
    • G06F17/10—Complex mathematical operations
    • G06F17/11—Complex mathematical operations for solving equations, e.g. nonlinear equations, general mathematical optimization problems
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
    • G06F17/10—Complex mathematical operations
    • G06F17/18—Complex mathematical operations for evaluating statistical data, e.g. average values, frequency distributions, probability functions, regression analysis
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00—Machine learning

Definitions

  • the present invention relates to a technique for estimating high resolution data from aggregated data aggregated in coarse particle size.
  • Spatial data refers to data given as a pair of position information (latitude and longitude, etc.) and some value associated with it.
  • Non-Patent Document 3 attempts have been made to utilize data in different domains (city, etc.) (Non-Patent Document 2).
  • the target high-resolution data is predicted by simultaneously modeling a plurality of types of aggregated data based on the multivariate Gaussian process based on the Gaussian process.
  • the technique disclosed in Non-Patent Document 2 even if the domain (city or the like) is different, the data in a plurality of domains is utilized by sharing the parameters of the Gaussian process between the domains.
  • the technique has a problem that the similarity between domains cannot be considered.
  • the present invention has been made in view of the above points, and an object of the present invention is to utilize a wide variety of data in a plurality of domains to realize highly accurate prediction of high-resolution data.
  • the disclosed technique is a parameter estimator that estimates a plurality of parameters used to calculate high resolution data from aggregated data aggregated in coarse particle size.
  • Marginal likelihood assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
  • a parameter estimation unit that estimates multiple parameters that are unknown variables in the model so as to maximize
  • a storage unit for storing the plurality of parameters is provided.
  • the plurality of parameters are provided with a parameter estimation device including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
  • FIG. 1 It is a block diagram of the aggregated data high-resolution apparatus in embodiment of this invention. It is a figure which shows the hardware configuration example of the apparatus. It is a figure for demonstrating the flow of the whole processing. It is a figure which shows the search example to the search part and the output example from an output part in embodiment of this invention.
  • Means 1 Multivariate Gaussian process model for aggregated data that introduces prior distributions for mixing coefficients
  • Means 1 Multivariate Gaussian process model for aggregated data that introduces prior distributions for mixing coefficients
  • Means 1 “spatial scale parameters”, “mixing coefficients”, and “parameters for noise”
  • the value of the aggregated data is modeled by the integrated value of the Gaussian process in the corresponding region. .. Assuming that the aggregated data actually observed was generated from the above model, the unknown variables are estimated so as to maximize the marginal likelihood.
  • Means 2 Efficient parameter estimation by variational Bayesian method
  • means 2 in addition to the multivariate Gaussian process model in means 1, unknown variables are estimated based on the variational Bayesian method.
  • the aggregated data high resolution device 100 incorporating the above-mentioned means is provided.
  • the aggregated data high resolution device 100 can target all the data aggregated in any area (hereinafter referred to as aggregated data for the sake of simplicity), and the type of data used (poor degree, air pollution degree, etc.) It does not depend on the number of dimensions d ⁇ ⁇ 1, ... ⁇ of the data input space (traffic volume, etc.), and can flexibly deal with them and estimate high-resolution data.
  • the aggregated data high resolution device 100 in the present embodiment models the value of the aggregated data by the integrated value of the Gaussian process in the corresponding region based on the multivariate Gaussian process model expressed by the linear mixture of a plurality of latent Gaussian processes. Then, after setting the prior distribution for the mixing coefficient, the unknown variable is estimated based on the variational Bayes method, and the high-resolution data is predicted.
  • data aggregated in a two-dimensional space will be mainly described as aggregated data, but the present invention is applied to data aggregated in a space at an arbitrary number of dimensions. Can be done. For example, when considering a one-dimensional space, it corresponds to time-series data such as sensor data aggregated at arbitrary time intervals.
  • FIG. 1 shows a configuration diagram of the aggregated data high resolution apparatus 100 according to the present embodiment.
  • the aggregated data high resolution apparatus 100 shown in the figure includes an aggregated data storage unit 1, a target division storage unit 2, an operation unit 3, a search unit 4, a high resolution data processing unit 5, a parameter estimation unit 6, and a spatial scale parameter storage unit. 7. It has a mixing coefficient storage unit 8, a parameter storage unit 9 for noise, a prior distribution hyperparameter storage unit 10, a high resolution data calculation unit 11, and an output unit 12. The details of the operation of each part will be described later.
  • the aggregated data high resolution device 100 may be composed of a plurality of devices (computers) or may be composed of one device. Further, the aggregated data high resolution apparatus 100 may be referred to as an aggregated data high resolution system. Further, the aggregated data high resolution device 100 may be called a parameter estimation device. Further, in FIG. 1, a device including a functional unit other than the aggregated data storage unit 1 and the target division storage unit 2 may be referred to as an aggregated data high resolution device 100.
  • a device having a function for estimating parameters that is, a parameter estimation unit 6
  • a parameter estimation device an apparatus having a function of increasing the resolution (that is, the high resolution data calculation unit 11) without including the function of estimating parameters
  • an aggregated data high resolution apparatus an apparatus having a function of increasing the resolution (that is, the high resolution data calculation unit 11) without including the function of estimating parameters.
  • Both the above-mentioned aggregated data high resolution device and parameter estimation device are realized by, for example, causing a computer to execute a program describing the processing contents described in the present embodiment. It is possible.
  • the "computer” may be a physical machine or a virtual machine on the cloud. When using a virtual machine, the “hardware” described here is virtual hardware.
  • the above program can be recorded on a computer-readable recording medium (portable memory, etc.), saved, and distributed. It is also possible to provide the above program through a network such as the Internet or e-mail.
  • FIG. 2 is a diagram showing an example of the hardware configuration of the above computer.
  • the computer of FIG. 2 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, which are connected to each other by a bus B, respectively.
  • the program that realizes the processing on the computer is provided by, for example, a recording medium 1001 such as a CD-ROM or a memory card.
  • a recording medium 1001 such as a CD-ROM or a memory card.
  • the program is installed in the auxiliary storage device 1002 from the recording medium 1001 via the drive device 1000.
  • the program does not necessarily have to be installed from the recording medium 1001, and may be downloaded from another computer via the network.
  • the auxiliary storage device 1002 stores the installed program and also stores necessary files, data, and the like.
  • the memory device 1003 reads and stores the program from the auxiliary storage device 1002 when the program is instructed to start.
  • the CPU 1004 realizes the function related to the device according to the program stored in the memory device 1003.
  • the interface device 1005 is used as an interface for connecting to a network.
  • the display device 1006 displays a GUI (Graphical User Interface) or the like by a program.
  • the input device 1007 is composed of a keyboard, a mouse, buttons, a touch panel, and the like, and is used for inputting various operation instructions.
  • the output device 1008 outputs the calculation result.
  • the parameter estimation unit 6 models the value of the aggregated data by the integral value of the Gaussian process in the corresponding region based on the multivariate Gaussian process model expressed by the linear mixture of a plurality of latent Gaussian processes, and the mixing coefficient is used. After setting the prior distribution, multiple parameters that are unknown variables are estimated based on the Gaussian process.
  • the high-resolution data calculation unit 11 calculates high-resolution data from the aggregated data using the plurality of parameters calculated in S101, and outputs the high-resolution data from the output unit 12.
  • the aggregated data storage unit 1 stores data that can be analyzed by the aggregated data high resolution apparatus 100, reads out the data according to the request from the high resolution data processing unit 5, and converts the corresponding data into the high resolution data processing unit. Send to 5.
  • X v ⁇ R d be the input space in the v-th domain
  • x ⁇ X v be the input variable.
  • P vs. represents the division associated with the corresponding data.
  • the division corresponds to, for example, the division of a city by address or region.
  • N vs. represents the number of regions included in the divided P vs.
  • Area argument n 1, ising , N vs.
  • the nth observation contained in the sth data is expressed as (R vsn , y vsn ) as a set of the region R vsn and the value y vsn ⁇ R.
  • the data stored in the aggregated data storage unit 1 is
  • the aggregated data storage unit 1 stores areas and values for each domain, each type of data, and each area.
  • the aggregated data storage unit 1 can be realized by, for example, a Web server, a database server including a database, or the like.
  • the aggregated data storage unit 1 may be a storage device in one computer.
  • the target division storage unit 2 stores divisions that can be output by the aggregated data high resolution apparatus 100, reads data according to a request from the high resolution data processing unit 5, and inputs the corresponding data to the high resolution data processing unit. Send to 5. It shall represent a division targeting P target .
  • One of the regions included in the P target is referred to as an R target .
  • any target division can be used, and a division based on an address or region, a mesh of an arbitrary size set by the user, or the like can be considered.
  • the target division storage unit 2 is realized by, for example, a Web server, a database server including a database, or the like.
  • the target division storage unit 2 may be a storage device in one computer.
  • the operation unit 3 receives various operations from the user on the data of the aggregated data storage unit 1 and the target division storage unit 2. Various operations are operations such as registering, modifying, and deleting stored information.
  • the input means of the operation unit 3 may be any one such as a keyboard, a mouse, a menu screen, and a touch panel.
  • the operation unit 3 is realized by, for example, a device driver for an input means such as a mouse or control software for a menu screen.
  • the search unit 4 accepts the target data type for high resolution and the target division.
  • the high-resolution data predicted by the aggregated data high-resolution device 100 is output for the data specified by the search unit 4 and the target division.
  • the input means of the search unit 4 may be any one such as a keyboard, a mouse, a menu screen, and a touch panel.
  • the search unit 4 can be realized by a device driver of an input means such as a mouse or control software of a menu screen.
  • the high-resolution data processing unit 5 includes a parameter estimation unit 6, a spatial scale parameter storage unit 7, a mixing coefficient storage unit 8, a parameter storage unit 9 for noise, a prior distribution hyperparameter storage unit 10, and a high resolution. It has a data calculation unit 11.
  • the parameter estimation unit 6 estimates the spatial scale parameters, mixing coefficients, parameters for noise, and hyperparameters of prior distribution using the data stored in the aggregated data storage unit 1 as training data. Then, the predicted high resolution data using these parameters is output.
  • the high-resolution data processing unit 5 after incorporating the above unknown variables, aggregated data in a plurality of domains is modeled based on a multivariate Gaussian process model expressed by a linear mixture of a plurality of latent Gaussian processes, and observation data is observed. Estimates the unknown variable based on the marginal likelihood given that. After that, high-resolution data is calculated based on the predicted distribution of the Gaussian process.
  • ⁇ XV ⁇ l (x, x'): X ⁇ X ⁇ R is the l (L) th correlation function, and any one can be used.
  • the correlation function is the l (L) th correlation function, and any one can be used.
  • ⁇ l is a spatial scale parameter of the l-th correlation function.
  • v 1, ....
  • f vs (x) be the Gaussian process for the sth aggregate data of each domain
  • S variable Gaussian process f v (x) (f v1 (x), ..., f vs. (X)) T as a linear mixture of L independent Gaussian processes
  • g v (x) (g v1 (x), ...., g vL (x)) T , where W v is a mixed matrix of S ⁇ L, and the (s, l) element. W vsl ⁇ R, which is, represents the mixing coefficient. Further, n v (x) represents a noise process for each data, and is a Gaussian process with an average of 0 S variates.
  • the vector 0 is a vector having 0 (zero) in all the elements, and ⁇ v (x, x') is.
  • ⁇ vs (x, x'): X v ⁇ X v ⁇ R is a correlation function of the noise process for the s th data of the v th domain, and any one can be used. here,
  • K v (x, x'): X ⁇ X ⁇ RS ⁇ S represents a correlation matrix.
  • ⁇ (x, x') diag ( ⁇ 1 (x, x'), ...., ⁇ L (x, x')).
  • Prior distribution for mixing coefficient w vsl is w vsl ⁇ p (w sl ) And. Any distribution can be used for p (w sl ), but here the Gaussian distribution is used.
  • -w sl and ⁇ 2 are hyperparameters of prior distribution with respect to the mixing coefficient.
  • the symbol placed at the beginning of the character is described before the character, for example, " -w sl ".
  • y v is a multidimensional Gaussian distribution
  • a vs (x) (a vs1 (x), ...., a vsN vs (x)) T.
  • a vsN vs (x) (a vs1 (x), ...., a vsN vs (x)) T.
  • vsN vs (x) (a vs1 (x), ...., a vsN vs (x)) T.
  • ⁇ 2 vs is a noise dispersion parameter for the sth data.
  • I is an identity matrix
  • O is a matrix in which all elements are 0.
  • C v is a correlation matrix of N v ⁇ N v , and is a correlation matrix.
  • Q (W v ) is called a proposed distribution and is introduced to approximate the posterior distribution of W v .
  • Q ( Wv) can set an arbitrary probability distribution, it is assumed here that each element wvsl of Wv follows an independent Gaussian distribution for the sake of simplicity. Such an approximation is called mean field approximation.
  • the parameter estimation unit 6 performs an operation for estimating various parameters so as to maximize the equation (23), and obtains various parameters. Any continuous optimization technique can be used for maximization.
  • the parameters to be estimated by the parameter estimation unit 6 are summarized below.
  • the spatial scale parameter storage unit 7 stores ⁇ l
  • l 1, ..., L ⁇ obtained by the parameter estimation unit 6.
  • the spatial scale parameter storage unit 6 may be anything as long as this information is stored and can be restored.
  • the spatial scale parameter storage unit 6 may be a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
  • the mixing coefficient storage unit 8 stores ⁇ W v
  • v 1, ..., V ⁇ obtained by the parameter estimation unit 6.
  • the mixing coefficient storage unit 8 may be any as long as this information is stored and can be restored.
  • the mixing coefficient storage unit 8 may be a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
  • the parameter storage unit 9 for the is is ⁇ vs
  • v 1, ..., V ⁇ is stored.
  • the parameter storage unit 9 for noise may be anything as long as this information is stored and can be restored.
  • the parameter storage unit 9 for noise may be a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
  • the hyperparameter storage unit 10 of the prior distribution stores ⁇ -w sl
  • the hyperparameter storage unit 10 of the prior distribution may be anything as long as this information is stored and can be restored.
  • the prior-distributed hyperparameter storage unit 10 may be, for example, a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
  • the high resolution data calculation unit 11 calculates the high resolution data using the parameters stored in each of the above storage units. Hereinafter, the processing contents of the high resolution data calculation unit 11 will be described in detail.
  • the high-resolution data calculation unit 11 calculates high-resolution data using various learned parameters for the aggregated data specified by the search unit 4 and the target division, and passes the high-resolution data to the output unit 12.
  • the calculation method of high resolution data will be described below. Since the derivation of the posterior distribution can be performed in the same way regardless of the domain, the v-th domain will be described below.
  • the post-process of the S-variate Gaussian process f v (x) of the v-th domain is derived.
  • the ex post process f * v (x) is
  • m * v (x): X v ⁇ RS represents an average vector
  • K * v (x, x ′): X v ⁇ X v ⁇ RS ⁇ S represents a correlation matrix
  • the high-resolution data calculation unit 11 calculates the target high-resolution data by integrating the posterior average (30) in each region R target included in the target division P target .
  • the output unit 12 outputs high-resolution data based on the information from the high-resolution data calculation unit 11.
  • the output is a concept including display on a display, printing on a printer, sound output, transmission to an external device, and the like.
  • the output unit 12 may or may not include an output device such as a display or a speaker.
  • the output unit 12 can be realized by the drive software of the output device, the driver software of the output device, the output device, or the like.
  • Figure 4 shows an output example.
  • the target domain argument, the aggregated data argument, and the target division are received from the search unit 3, and the visualization result of the high resolution data is displayed in the output unit 12 accordingly.
  • the shade in the visualization result is determined in proportion to the data value, for example.
  • the shade in the visualization result is determined in proportion to the data value, for example.
  • This specification discloses at least the parameter estimation device, the aggregated data high resolution device, the parameter estimation method, the aggregated data high resolution method, and the program of each of the following items.
  • (Section 1) A parameter estimator that estimates multiple parameters used to calculate high resolution data from aggregated data aggregated to a coarser grain size. Marginal likelihood, assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
  • a parameter estimation unit that estimates multiple parameters that are unknown variables in the model so as to maximize
  • a storage unit for storing the plurality of parameters is provided.
  • the plurality of parameters are parameter estimation devices including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
  • the model is described in the first item, which is a model in which the value of the aggregated data of each region in a plurality of regions obtained by dividing the space of the domain is modeled by the integrated value of the multivariate Gaussian process in the corresponding region.
  • Parameter estimator (Section 3)
  • the parameter estimation unit is the parameter estimation device according to the first or second term, which estimates the plurality of parameters by using the variational Bayes method.
  • (Section 4) Aggregated data height including a high-resolution data calculation unit that calculates high-resolution data from aggregated data using the plurality of parameters estimated by the parameter estimation device according to any one of the items 1 to 3.
  • Resolution device. (Section 5) It is an aggregated data high resolution device that calculates high resolution data from aggregated data aggregated to coarse particle size. Marginal likelihood, assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
  • a parameter estimation unit that estimates multiple parameters that are unknown variables in the model so as to maximize A storage unit that stores the plurality of parameters, and A high-resolution data calculation unit that calculates high-resolution data from aggregated data using the plurality of parameters is provided.
  • the plurality of parameters are aggregate data high resolution devices including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
  • (Section 6) A parameter estimation method performed by a parameter estimator that estimates multiple parameters used to calculate high resolution data from aggregated data aggregated to a coarser grain size.
  • Marginal likelihood assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
  • the step of estimating multiple parameters that are unknown variables in the model so as to maximize A step of storing the plurality of parameters in the storage unit is provided.
  • the plurality of parameters are parameter estimation methods including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
  • (Section 7) It is a method of increasing the resolution of aggregated data executed by the aggregated data high-resolution device that calculates high-resolution data from the aggregated data aggregated in coarse particle size.
  • Marginal likelihood assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
  • the step of estimating multiple parameters that are unknown variables in the model so as to maximize A step of storing the plurality of parameters in the storage unit, and A step of calculating high resolution data from aggregated data using the plurality of parameters is provided.
  • the plurality of parameters are aggregate data high resolution methods including hyperparameters of prior distribution with respect to the mixing coefficient used in the linear mixing.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Theoretical Computer Science (AREA)
  • Computational Mathematics (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Pure & Applied Mathematics (AREA)
  • Software Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • Algebra (AREA)
  • Operations Research (AREA)
  • Computing Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Artificial Intelligence (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Probability & Statistics with Applications (AREA)
  • Complex Calculations (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置であって、実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、前記複数のパラメータを格納する格納部と、を備え、前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含むパラメータ推定装置。

Description

パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム
 本発明は、粗い粒度に集約された集約データから高解像度データを推定する技術に関連するものである。
 近年、政府や企業などが都市環境や事業の改善を目的として、様々な種類の空間データ(貧困度、大気汚染度、犯罪数、人口、交通量など)を収集し、公開している。空間データとは、位置情報(緯度経度など)とそれに紐づく何らかの値とのペアで与えられるデータを指す。
 しかし、このような空間データは、収集コストが高く十分なサンプル数を確保することが難しい。そのため、ある程度粗い粒度の領域(住所や地域など)において集約されて提供されることが多い。このようなデータを以下では「集約データ」と呼ぶ。より効果的な都市環境の改善のためには、できる限り高解像度のデータを入手することが望ましい。例えば、貧困度が高い地域や大気汚染度が高い地域を詳細に絞りこむことで、より適切な介入が可能となる。
 したがって、低解像度の集約データから高解像度なデータを予測するという問題は重要である。従来技術では、ターゲットとする低解像度の集約データとは別に各々の解像度を持つ様々な種類の集約データを用意し、多変量ガウス過程に基づいて同時にモデル化することで低解像度のデータに対して高精度な予測が実現されている(非特許文献3)。更に、異なるドメイン(都市など)におけるデータを活用する試みもなされている(非特許文献2)。
D. Kingma and M. Welling. Auto-encoding variational Bayes. In ICLR, 2014. Y. Tanaka, T. Tanaka, T. Iwata, T. Kurashima, M. Okawa, Y. Akagi, and H. Toda. Spatially aggregated Gaussian processes with multivariate areal outputs. In NeurIPS, pages 3000-3010, 2019. F. Yousefi, M. T. Smith, and M. A. ´Alvarez. Multi-task learning for aggregated data using Gaussian processes. In NeurIPS, pages 15050-15060, 2019. 8
 非特許文献2、3に開示されている従来技術では、ガウス過程を基礎として、複数種類の集約データを多変量ガウス過程に基づいて同時にモデル化することでターゲットとする高解像度データを予測する。このとき、非特許文献2に開示された技術では、ドメイン(都市など)が異なる場合でも、ガウス過程のパラメータをドメイン間で共有することにより、複数ドメインにおけるデータを活用する。しかし、当該技術にはドメイン間の類似度が考慮できないという問題点がある。
 すなわち、従来技術では、異なるドメインにおいて同じ種類のデータの組が存在したとしても、それらのデータ間の依存関係の強さはドメイン毎に独立であることを仮定している。しかし、ドメインによっては、低解像度のデータしか入手できない場合も少なくない。そのような場合に、データ間の依存関係の強さを適切に推定することが難しいという問題が存在した。
 本発明は、上記の点に鑑みてなされたものであり、複数ドメインにおける多種多様なデータを活用して、高解像度データの高精度な予測を実現することを目的とする。
 開示の技術によれば、粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置であって、
 実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
 前記複数のパラメータを格納する格納部と、を備え、
 前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
 パラメータ推定装置が提供される。
 開示の技術によれば、複数ドメインにおける多種多様なデータを活用して、高解像度データの高精度な予測を実現することが可能となる。
本発明の実施の形態における集約データ高解像度化装置の構成図である。 装置のハードウェア構成例を示す図である。 全体の処理の流れを説明するための図である。 本発明の実施の形態における検索部への検索例と,出力部からの出力例を示す図である。
 以下、図面を参照して本発明の実施の形態(本実施の形態)を説明する。以下で説明する実施の形態は一例に過ぎず、本発明が適用される実施の形態は、以下の実施の形態に限られるわけではない。
 (実施の形態の概要)
 本実施の形態では、ドメイン間の類似度が考慮できないという従来技術の問題を解決するために、データ間の依存関係を表すパラメータ(混合係数)に対する事前分布を導入し、それを加味したパラメータ推定手段を導入している。これにより、ドメイン間の類似度を同時に推定しつつ、複数ドメインにおける多種多様なデータを活用して、高解像度データの高精度な予測を実現することが可能となる。より具体的には、下記の手段1と手段2が導入される。
 (1)手段1:混合係数に対する事前分布を導入した、集約データのための多変量ガウス過程モデル
 手段1においては、「空間スケールパラメータ」と、「混合係数」と、「ノイズに対するパラメータ」と、「事前分布のハイパーパラメータ」とを未知変数として、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程に基づいて、集約データの値を該当領域におけるガウス過程の積分値によりモデル化する。実際に観測された集約データが上記のモデルから生成されたものと仮定して、周辺尤度を最大化するように未知変数を推定する。
 手段1に対する効果は下記のとおりである。
 複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程モデルに基づいて、各ドメインにおける複数のデータを同時にモデル化し、更に、混合係数に対する事前分布を導入する。これにより、ドメイン間の類似度を推定しつつ、「空間スケールパラメータ」、及び、「混合係数」を複数ドメインにおけるデータを活用して学習することができる。
 (1)手段2:変分ベイズ法による効率的なパラメータ推定手段
 手段2においては、手段1における多変量ガウス過程モデルに加えて、変分ベイズ法に基づいて、未知変数を推定する。
 手段2に対する効果は下記のとおりである。
 未知変数を推定する際には、周辺尤度を計算する必要があるが、実際には解析的に計算することは難しい。そこで任意の近似的な推定手段を用いることが考えられるが、選択肢の一つとして変分ベイズ法に基づく学習方法を導出する。これにより、より効率的に未知変数を学習することができる。
 本実施の形態では、上述した手段を組み込んだ集約データ高解像度化装置100が提供される。集約データ高解像度化装置100は、任意の領域で集約されたデータ(以下では簡単のため集約データと呼ぶ)全般を対象とすることができ、使用するデータの種類(貧困度、大気汚染度や交通量など)やデータの入力空間の次元数d∈{1,...}に依存せず、それらに対して柔軟に対応し、高解像度データの推定を行うことができる。
 v=1,.....Vをドメインの引数とし、v番目のドメインにおけるデータの種類集合をSvとする。ここで、全ドメインにおけるデータの種類集合の和集合をSとすると、Sv⊂Sである。以下では、簡略化のため全てのドメインにおいてS種類のデータが存在するとし、s=1,....,Sをデータ種類の引数として定式化を行う。ただし、上記の定義のようにドメインによって得られるデータの種類が異なっていても同様にして定式化が可能である。
 本実施の形態における集約データ高解像度化装置100は、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程モデルに基づいて、集約データの値を該当領域におけるガウス過程の積分値によりモデル化し、混合係数に対する事前分布を設定した上で、変分ベイズ法に基づいて未知変数を推定し、高解像度データを予測する。
 以下、集約データ高解像度化装置100の構成と動作について詳細に説明する。なお、以下では、主に集約データとして二次元空間上で集約されたデータを実施例として説明をするが、本発明は、任意の次元数における空間上において集約されたデータに対して適用することができる。例えば、一次元空間を考えた場合、センサデータ等の時系列データが任意の時間間隔毎に集約されたものに対応する。
 (装置構成例)
 図1に、本実施の形態における集約データ高解像度化装置100の構成図を示す。同図に示す集約データ高解像度化装置100は、集約データ格納部1、ターゲット分割格納部2、操作部3、検索部4、高解像度データ処理部5、パラメータ推定部6、空間スケールパラメータ格納部7、混合係数格納部8、ノイズに対するパラメータ格納部9、事前分布のハイパーパラメータ格納部10、高解像度データ算出部11、出力部12を有する。各部の動作詳細については後述する。
 集約データ高解像度化装置100は、複数の装置(コンピュータ)から構成されてもよいし、1つの装置で構成されてもよい。また、集約データ高解像度化装置100を集約データ高解像度化システムと呼んでもよい。また、集約データ高解像度化装置100をパラメータ推定装置と呼んでもよい。また、図1において、集約データ格納部1、ターゲット分割格納部2以外の機能部からなる装置を集約データ高解像度化装置100と呼んでもよい。
 また、高解像度化を行う機能を含まずにパラメータ推定を行う機能(つまり、パラメータ推定部6)を有する装置をパラメータ推定装置と呼んでもよい。また、パラメータ推定を行う機能を含まずに、高解像度化を行う機能(つまり、高解像度データ算出部11)を有する装置を集約データ高解像度化装置と呼んでもよい。
  (ハードウェア構成例)
 上述した集約データ高解像度化装置、パラメータ推定装置(これらを総称して装置と呼ぶ)はいずれも、例えば、コンピュータに、本実施の形態で説明する処理内容を記述したプログラムを実行させることにより実現可能である。なお、この「コンピュータ」は、物理マシンであってもよいし、クラウド上の仮想マシンであってもよい。仮想マシンを使用する場合、ここで説明する「ハードウェア」は仮想的なハードウェアである。
 上記プログラムは、コンピュータが読み取り可能な記録媒体(可搬メモリ等)に記録して、保存したり、配布したりすることが可能である。また、上記プログラムをインターネットや電子メール等、ネットワークを通して提供することも可能である。
 図2は、上記コンピュータのハードウェア構成例を示す図である。図2のコンピュータは、それぞれバスBで相互に接続されているドライブ装置1000、補助記憶装置1002、メモリ装置1003、CPU1004、インタフェース装置1005、表示装置1006、入力装置1007、出力装置1008等を有する。
 当該コンピュータでの処理を実現するプログラムは、例えば、CD-ROM又はメモリカード等の記録媒体1001によって提供される。プログラムを記憶した記録媒体1001がドライブ装置1000にセットされると、プログラムが記録媒体1001からドライブ装置1000を介して補助記憶装置1002にインストールされる。但し、プログラムのインストールは必ずしも記録媒体1001より行う必要はなく、ネットワークを介して他のコンピュータよりダウンロードするようにしてもよい。補助記憶装置1002は、インストールされたプログラムを格納すると共に、必要なファイルやデータ等を格納する。
 メモリ装置1003は、プログラムの起動指示があった場合に、補助記憶装置1002からプログラムを読み出して格納する。CPU1004は、メモリ装置1003に格納されたプログラムに従って、当該装置に係る機能を実現する。インタフェース装置1005は、ネットワークに接続するためのインタフェースとして用いられる。表示装置1006はプログラムによるGUI(Graphical User Interface)等を表示する。入力装置1007はキーボード及びマウス、ボタン、又はタッチパネル等で構成され、様々な操作指示を入力させるために用いられる。出力装置1008は演算結果を出力する。
 (集約データ高解像度化装置100の処理動作)
 まず、図3を参照して、全体の処理の流れを説明する。
 S101において、パラメータ推定部6が、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程モデルに基づいて、集約データの値を該当領域におけるガウス過程の積分値によりモデル化し、混合係数に対する事前分布を設定した上で、変分ベイズ法に基づいて未知変数である複数のパラメータを推定する。
 S102において、高解像度データ算出部11が、S101で算出された複数のパラメータを用いて、集約データから高解像度データを算出し、出力部12から出力する。
 以下、集約データ高解像度化装置100を構成する各部の機能及び処理動作について説明する。
 <集約データ格納部1>
 集約データ格納部1は、集約データ高解像度化装置100によって解析され得るデータを格納しており、高解像度データ処理部5からの要求にしたがって、データを読み出し、該当のデータを高解像度データ処理部5に送信する。
 v番目のドメインにおける入力空間をXv⊂Rdとし、x∈Xvを入力変数とする。例えば、d=2の場合は、Xvはv番目の都市全体に、xは緯度経度等に対応する。v番目のドメインのs番目のデータに対し、Pvsを該当データが紐づく分割を表すものとする。ここで分割とは、例えば住所や地域による都市の分割等に対応する。また、Nvsは分割Pvsに含まれる領域数を表すものとする。領域の引数n=1,.......,Nvsに対して、n番目の領域をRvsn∈Pvsとする。s番目のデータに含まれるn番目の観測を、領域Rvsnと値yvsn∈Rの組として(Rvsn,yvsn)と表す。集約データ格納部1に格納されるデータは、
Figure JPOXMLDOC01-appb-M000001
である。つまり、集約データ格納部1には、ドメイン毎、データの種類毎、領域毎の領域と値が格納されている。集約データ格納部1は、具体的には、例えば、Webサーバや、データベースを具備するデータベースサーバ等によって実現することができる。集約データ格納部1が、1つのコンピュータ内の記憶装置であってもよい。
 <ターゲット分割格納部2>
 ターゲット分割格納部2は、集約データ高解像度化装置100によって出力され得る分割を格納しており、高解像度データ処理部5からの要求にしたがって、データを読み出し、該当のデータを高解像度データ処理部5に送信する。Ptargetをターゲットとする分割を表すものとする。Ptargetに含まれる領域の1つをRtargetと表す。ここで、ターゲット分割は任意のものを使用することが可能であり、住所や地域に基づく分割や使用者が設定した任意のサイズのメッシュなどが考えられる。ターゲット分割格納部2は、例えば、Webサーバや、データベースを具備するデータベースサーバ等により実現される。ターゲット分割格納部2が、1つのコンピュータ内の記憶装置であってもよい。
 <操作部3>
 操作部3は、集約データ格納部1、及び、ターゲット分割格納部2のデータに対するユーザからの各種操作を受け付ける。各種操作とは、格納された情報を登録、修正、削除する操作等である。操作部3の入力手段は、キーボードやマウス、メニュー画面、タッチパネルによるもの等、どのようなものでもよい。操作部3は、例えば、マウス等の入力手段のデバイスドライバや、メニュー画面の制御ソフトウェアで実現される。
 <検索部4>
 検索部4は、高解像度化を行う対象とするデータ種別、及び、ターゲット分割を受け付ける。検索部4で指定されたデータ、及び、ターゲット分割に対して、集約データ高解像度化装置100によって予測された高解像度データを出力する。なお、検索部4の入力手段は、キーボードやマウス、メニュー画面、タッチパネルによるもの等、どのようなものでもよい。検索部4は、マウス等の入力手段のデバイスドライバや、メニュー画面の制御ソフトウェアで実現され得る。
 <高解像度データ処理部5>
 図1に示すとおり、高解像度データ処理部5は、パラメータ推定部6、空間スケールパラメータ格納部7、混合係数格納部8、ノイズに対するパラメータ格納部9、事前分布のハイパーパラメータ格納部10、高解像度データ算出部11を有する。
 高解像度データ処理部5では、パラメータ推定部6により、集約データ格納部1に格納されたデータを学習データとして、空間スケールパラメータ、混合係数、ノイズに対するパラメータ、事前分布のハイパーパラメータを推定する。そして、これらのパラメータを用いて予測された高解像度データを出力する。
 高解像度データ処理部5では、上記の未知変数を組み込んだ上で、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程モデルに基づいて、複数ドメインにおける集約データをモデル化し、観測データが与えられたとしたときの周辺尤度に基づいて未知変数を推定する。その後、ガウス過程の予測分布に基づいて高解像度データを算出する。
 <パラメータ推定部6>
 以下、パラメータ推定部6において処理されるガウス過程モデルとパラメータ推定法について詳細に説明する。まず、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程の定式化を行う。V×L個の独立なガウス過程を
Figure JPOXMLDOC01-appb-M000002
とする。ここで、X=X1∪X2........∪XVとしたとき、γl(x,x′):X×X→Rはl(エル)個目の相関関数であり、任意のものを使用できる。ここでは、相関関数は
Figure JPOXMLDOC01-appb-M000003
を用いるものとする。ここで、βlはl(エル)個目の相関関数の空間スケールパラメータである。v=1,....,Vに対して、fvs(x)を各ドメインのs番目の集約データに対するガウス過程とし、S変量ガウス過程fv(x)=(fv1(x),.....,fvS(x))TをL個の独立なガウス過程の線型混合として
Figure JPOXMLDOC01-appb-M000004
と表す。ここで、gv(x)=(gv1(x),.....,gvL(x))Tを表し、WvはS×Lの混合行列であり、(s,l)要素であるwvsl∈Rは混合係数を表す。また、nv(x)は各データに対するノイズ過程を表し、S変量の平均0のガウス過程であり、
Figure JPOXMLDOC01-appb-M000005
と表される。ここで、ベクトル0は全要素に0(ゼロ)を持つベクトルであり、Λv(x,x′)は
Figure JPOXMLDOC01-appb-M000006
とする。λvs(x,x′):Xv×Xv→Rはv番目のドメインのs番目のデータに対するノイズ過程の相関関数であり、任意のものを使用できる。ここでは、
Figure JPOXMLDOC01-appb-M000007
を用いるものとする。式(4)におけるgv(x)とnv(x)は積分消去を行うことができ、S変量ガウス過程は
Figure JPOXMLDOC01-appb-M000008
と書ける。ここで、Kv(x,x′):X×X→RS×Sは相関行列を表し、
Figure JPOXMLDOC01-appb-M000009
となる。ここで、Γ(x,x′)=diag(γ1(x,x′),.....,γL(x,x′))とした。
 次に、混合係数に対する事前分布の導入を行う。混合係数wvslに対する事前分布を
   wvsl~p(wsl)
とする。p(wsl)は任意の分布を用いることができるが、ここではガウス分布を用いて
Figure JPOXMLDOC01-appb-M000010
とする。ここで、―wsl及びη2は混合係数に対する事前分布のハイパーパラメータである。なお、明細書のテキストの記載の便宜上、文字の頭に置かれる記号は、例えば「―wsl」のように、文字の前に記載している。
 次に、集約データの観測モデルについて説明する。集約データの値を該当領域におけるガウス過程の積分値を平均に持つガウス分布の実現値で表現したものが観測モデルである。v=1,...,Vに対して同様のモデルを用いるため、v番目のドメインに対する定式化について説明する。s番目のガウス過程から生成されるNvs次元の観測ベクトルをyvs=(yvs1,.....,yvsNvs)Tと表し、S個のガウス過程から生成される観測ベクトルをまとめて
Figure JPOXMLDOC01-appb-M000011
と表す。yvは多次元ガウス分布
Figure JPOXMLDOC01-appb-M000012
に従うと仮定する。ここで、Nv=ΣS s=1Nvsとして、A(x):Xv→RNv×Sであり、
Figure JPOXMLDOC01-appb-M000013
とする。ここで、avs(x)=(avs1(x),.....,avsNvs(x))Tとした。なお、明細書での記載の便宜上、avsNvs(x)と記載しているが、avsNvsに関して、これは、「vsNvs」をaに下付きで付けたものを意図している。avsn(x)は任意のものを使用することができ、avsn(x)の設定の仕方によって各領域における集約の仕方を変えることができる。ここでは各領域Rvsnにおいて領域平均をした結果、観測が得られる場合を考える。このときavsn(x)は
Figure JPOXMLDOC01-appb-M000014
と書ける。ここで、1(・)は指示関数であり、Cが真のとき1(C)=1を出力し、そうでないとき1(C)=0を出力する。また、
Figure JPOXMLDOC01-appb-M000015
とし、σ2 vsはs番目のデータに対するノイズ分散パラメータである。ここで、Iは単位行列であり、Oは全要素を0とする行列である。
 次に、各種パラメータを学習する方法について説明する。上記の式(12)のガウス過程モデル(観測モデル)から観測データが生成されたと仮定したとき、対数周辺尤度は
Figure JPOXMLDOC01-appb-M000016
Figure JPOXMLDOC01-appb-M000017
Figure JPOXMLDOC01-appb-M000018
と記述できる。ここで、
Figure JPOXMLDOC01-appb-M000019
と解析的に計算可能である。また、CvはNv×Nvの相関行列であり、
Figure JPOXMLDOC01-appb-M000020
である。式(20)における二重積分の計算について説明する。入力空間Xvの次元数がd=1の場合には、非特許文献3と同様の方法によって解析的に計算が可能である。d>1の場合には、積分を解析的に計算することは困難なので、非特許文献2と同様の方法によって離散近似を行い数値的に計算する。式(18)におけるp(Wv)は
Figure JPOXMLDOC01-appb-M000021
である。式(18)を最大化することで各種パラメータの解を得ることができる。ここで、Wvについての積分は解析的には行えないため、任意の近似方法を用いる。例えば、数値近似、モンテカルロ近似、変分近似等である。近似後の目的関数を最大化するための方法としては、任意の連続最適化手法が使用可能である。例えばBFGS法を用いて解くことができる。
 以下では、式(18)におけるWvについての積分を、変分近似を用いて近似する場合について説明する。本手法は変分ベイズ法と呼ばれる。変分ベイズ法では式(18)の変分下限(Evidence lower bound)を算出し、それを目的関数としてパラメータ推定を行う。イェンセンの不等式を用いて、変分下限は
Figure JPOXMLDOC01-appb-M000022
Figure JPOXMLDOC01-appb-M000023
と記述できる。ここで、Q(Wv)は提案分布と呼ばれ、Wvの事後分布を近似するために導入される。Q(Wv)は任意の確率分布を設定することができるが、ここでは簡単のためWvの各要素wvslが独立なガウス分布に従うものとする。このような近似は平均場近似と呼ばれ、
Figure JPOXMLDOC01-appb-M000024
と表される。ここで、―w′vsl及びη′vsl 2は変分パラメータである。このとき、式(23)の第二項におけるKLダイバージェンスは解析的に計算ができるが、第一項の期待値は解析的に計算が困難である。そこで以下のようにサンプリングを用いた近似を行う。
Figure JPOXMLDOC01-appb-M000025
Figure JPOXMLDOC01-appb-M000026
ここで、^Wv~Q(Wv)である。また、変分パラメータの推定を可能にするために非特許文献1と同様にReparameterization trickを用いて^Wvの各成分を
Figure JPOXMLDOC01-appb-M000027
とする。ここで、ε~N(0,1)とする。
 以上より、パラメータ推定部6は、式(23)を最大化するように各種パラメータを推定する演算を行って、各種パラメータを求める。最大化のためには、任意の連続最適化手法を使用することができる。
 パラメータ推定部6が推定すべきパラメータをまとめると下記のとおりである。
 ・空間スケールパラメータ{βl|l=1,...,L}
 ・混合係数{Wv|v=1,...,V}
 ・ノイズに対するパラメータ{αvs|v=1,...,V;s=1,...,S}と{κvs|v=1,...,V;s=1,...,S}と{Σv|v=1,...,V}
 ・事前分布のハイパーパラメータ{‐wsl|s=1,...,S;l=1,...,L}とη2
 なお、変分パラメータは上記パラメータを推定するための補助変数として機能し、後述する高解像度データ算出部11においては使用されない。推定されたパラメータは、下記の各格納部に格納される。なお、複数のパラメータを格納する格納部が1つだけ備えられることとしてもよい。
 <空間スケールパラメータ格納部7>
 空間スケールパラメータ格納部7は、パラメータ推定部6で求めた{βl|l=1,...,L}を格納する。空間スケールパラメータ格納部6は、この情報が保存され、復元可能なものであればどのようなものでもよい。例えば、空間スケールパラメータ格納部6は、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
 <混合係数格納部8>
 混合係数格納部8は、パラメータ推定部6で求めた{Wv|v=1,...,V}を格納する。混合係数格納部8は、この情報が保存され、復元可能なものであればどのようなものでもよい。例えば、混合係数格納部8は、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
 <ノイズに対するパラメータ格納部9>
 イズに対するパラメータ格納部9は、パラメータ推定部6で求めた{αvs|v=1,...,V;s=1,...,S}、{κvs|v=1,...,V;s=1,...,S}、{Σv|v=1,...,V}を格納する。ノイズに対するパラメータ格納部9は、この情報が保存され、復元可能なものであればどのようなものでもよい。例えば、ノイズに対するパラメータ格納部9は、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
 <事前分布のハイパーパラメータ格納部10>
 事前分布のハイパーパラメータ格納部10は、パラメータ推定部6で求めた{‐wsl|s=1,...,S;l=1,...,L}、η2を格納する。事前分布のハイパーパラメータ格納部10は、この情報が保存され、復元可能なものであればどのようなものでもよい。事前分布のハイパーパラメータ格納部10は、例えば、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
 上記の各格納部に格納されているパラメータを使用して、高解像度データ算出部11が、高解像度データの算出を行う。以下、高解像度データ算出部11の処理内容を詳細に説明する。
 <高解像度データ算出部11>
 高解像度データ算出部11は、検索部4で指定された集約データ、及び、ターゲット分割に対して、学習済みの各種パラメータを用いて高解像度データを算出し、出力部12へ渡す。以下に、高解像度データの算出法について説明する。ドメインによらず事後分布の導出は同様に行うことができるので、以下ではv番目のドメインについて説明をする。まず、v番目のドメインのS変量ガウス過程fv(x)の事後プロセスを導出する。事後プロセスf* v(x)は、
Figure JPOXMLDOC01-appb-M000028
と書ける。ここで、m* v(x):Xv→RSは平均ベクトルを表し、K* v(x,x′):Xv×Xv→RS×Sは相関行列を表す。次に、Hv(x):Xv→RNv×Sを
Figure JPOXMLDOC01-appb-M000029
と置く。Hv(x)を用いて、
Figure JPOXMLDOC01-appb-M000030
Figure JPOXMLDOC01-appb-M000031
と記述できる。高解像度データ算出部11は、事後平均(30)をターゲット分割Ptargetに含まれる各領域Rtargetにおいて積分することによって、目的の高解像度データを算出する。
 <出力部12>
 出力部12は、高解像度データ算出部11からの情報に基づいて、高解像度データを出力する。ここで、出力とは、ディスプレイへの表示、プリンタへの印字、音出力、外部装置への送信等を含む概念である。出力部12は、ディスプレイやスピーカ等の出力デバイスを含むと考えても含まないと考えてもよい。出力部12は、出力デバイスのドライブソフト又は、出力デバイスのドライバソフトと出力デバイス等で実現され得る。
 図4に出力例を示す。検索部3から対象とするドメインの引数、集約データの引数、及び、ターゲット分割を受け取り、それに応じて、出力部12において高解像度データの可視化結果が表示される。
 可視化結果における濃淡は、例えば、データ値に比例して決まる。このような出力例を用いて、例えば、貧困度が高い領域や大気汚染度が高い領域を詳細に絞りこみ、より適切な介入を行うための検討に活用できる。
 (実施の形態の効果)
 以上、説明したとおり、本実施の形態では、データ間の依存関係を表すパラメータ(混合係数)に対する事前分布を導入し、それを加味したパラメータ推定を行うこととしているので、ドメイン間の類似度を同時に推定しつつ、複数ドメインにおける多種多様なデータを活用して、高解像度データの高精度な予測を実現することが可能となる。
 (付記)
 本明細書には、少なくとも下記各項のパラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラムが開示されている。
(第1項)
 粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置であって、
 実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
 前記複数のパラメータを格納する格納部と、を備え、
 前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
 パラメータ推定装置。
(第2項)
 前記モデルは、ドメインの空間を分割することより得られる複数の領域における各領域の集約データの値を、該当領域における多変量ガウス過程の積分値でモデル化したモデルである
 第1項に記載のパラメータ推定装置。
(第3項)
 前記パラメータ推測部は、変分ベイズ法を用いて前記複数のパラメータを推定する
 第1項又は第2項に記載のパラメータ推定装置。
(第4項)
 第1項ないし第3項のうちいずれか1項に記載のパラメータ推定装置により推定された前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部
 を備える集約データ高解像度化装置。
(第5項)
 粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置であって、
 実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
 前記複数のパラメータを格納する格納部と、
 前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部と、を備え、
 前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
 集約データ高解像度化装置。
(第6項)
 粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置が実行するパラメータ推定方法であって、
 実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
 前記複数のパラメータを格納部に格納するステップと、を備え、
 前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
 パラメータ推定方法。
(第7項)
 粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置が実行する集約データ高解像度化方法であって、
 実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
 前記複数のパラメータを格納部に格納するステップと、
 前記複数のパラメータを用いて、集約データから高解像度データを算出するステップと、を備え、
 前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
 集約データ高解像度化方法。
(第8項)
 コンピュータを、第1項ないし第3項のうちいずれか1項に記載のパラメータ推定装置における各部として機能させるためのプログラム、又は、第5項に記載の集約データ高解像度化装置における各部として機能させるためのプログラム。
 以上、本実施の形態について説明したが、本発明はかかる特定の実施形態に限定されるものではなく、特許請求の範囲に記載された本発明の要旨の範囲内において、種々の変形・変更が可能である。
100 集約データ高解像度化装置
1 集約データ格納部
2 ターゲット分割格納部
3 操作部
4 検索部
5 高解像度データ処理部
6 パラメータ推定部
7 空間スケールパラメータ格納部
8 混合係数格納部
9 ノイズに対するパラメータ格納部
10 事前分布のハイパーパラメータ格納部
11 高解像度データ算出部
12 出力部
1000 ドライブ装置
1001 記録媒体
1002 補助記憶装置
1003 メモリ装置
1004 CPU
1005 インタフェース装置
1006 表示装置
1007 入力装置

Claims (8)

  1.  粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置であって、
     実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
     前記複数のパラメータを格納する格納部と、を備え、
     前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
     パラメータ推定装置。
  2.  前記モデルは、ドメインの空間を分割することより得られる複数の領域における各領域の集約データの値を、該当領域における多変量ガウス過程の積分値でモデル化したモデルである
     請求項1に記載のパラメータ推定装置。
  3.  前記パラメータ推測部は、変分ベイズ法を用いて前記複数のパラメータを推定する
     請求項1又は2に記載のパラメータ推定装置。
  4.  請求項1ないし3のうちいずれか1項に記載のパラメータ推定装置により推定された前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部
     を備える集約データ高解像度化装置。
  5.  粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置であって、
     実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
     前記複数のパラメータを格納する格納部と、
     前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部と、を備え、
     前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
     集約データ高解像度化装置。
  6.  粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置が実行するパラメータ推定方法であって、
     実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
     前記複数のパラメータを格納部に格納するステップと、を備え、
     前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
     パラメータ推定方法。
  7.  粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置が実行する集約データ高解像度化方法であって、
     実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
     前記複数のパラメータを格納部に格納するステップと、
     前記複数のパラメータを用いて、集約データから高解像度データを算出するステップと、を備え、
     前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
     集約データ高解像度化方法。
  8.  コンピュータを、請求項1ないし3のうちいずれか1項に記載のパラメータ推定装置における各部として機能させるためのプログラム、又は、請求項5に記載の集約データ高解像度化装置における各部として機能させるためのプログラム。
PCT/JP2020/046123 2020-12-10 2020-12-10 パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム Ceased WO2022123743A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
US18/255,925 US20240045921A1 (en) 2020-12-10 2020-12-10 Parameter estimation apparatus, aggregated data resolution enhancement apparatus, parameter estimation method, aggregated data resolution enhancement method and program
JP2022567985A JP7439957B2 (ja) 2020-12-10 2020-12-10 パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム
PCT/JP2020/046123 WO2022123743A1 (ja) 2020-12-10 2020-12-10 パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2020/046123 WO2022123743A1 (ja) 2020-12-10 2020-12-10 パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム

Publications (1)

Publication Number Publication Date
WO2022123743A1 true WO2022123743A1 (ja) 2022-06-16

Family

ID=81973446

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/046123 Ceased WO2022123743A1 (ja) 2020-12-10 2020-12-10 パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム

Country Status (3)

Country Link
US (1) US20240045921A1 (ja)
JP (1) JP7439957B2 (ja)
WO (1) WO2022123743A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP2530626A1 (en) * 2011-06-01 2012-12-05 BAE Systems Plc. Heterogeneous data fusion using Gaussian processes
JP2017033198A (ja) * 2015-07-30 2017-02-09 日本電信電話株式会社 時空間変数予測装置及びプログラム

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP2530626A1 (en) * 2011-06-01 2012-12-05 BAE Systems Plc. Heterogeneous data fusion using Gaussian processes
JP2017033198A (ja) * 2015-07-30 2017-02-09 日本電信電話株式会社 時空間変数予測装置及びプログラム

Also Published As

Publication number Publication date
US20240045921A1 (en) 2024-02-08
JP7439957B2 (ja) 2024-02-28
JPWO2022123743A1 (ja) 2022-06-16

Similar Documents

Publication Publication Date Title
de Fondeville et al. High-dimensional peaks-over-threshold inference
Zou et al. Empirical Bayes estimates of finite mixture of negative binomial regression models and its application to highway safety
Kazianka et al. Copula-based geostatistical modeling of continuous and discrete data including covariates
Sellers et al. Underdispersion models: Models that are “under the radar”
Middleton et al. Unbiased estimation of the average treatment effect in cluster-randomized experiments
US11593860B2 (en) Method, medium, and system for utilizing item-level importance sampling models for digital content selection policies
Plumlee et al. Building accurate emulators for stochastic simulations via quantile kriging
Schneider et al. Learning stochastic closures using ensemble Kalman inversion
De Jong et al. Multiple imputation of predictor variables using generalized additive models
Wu et al. Uapd: Predicting urban anomalies from spatial-temporal data
Yang et al. A stochastic expectation-maximization algorithm for the analysis of system lifetime data with known signature
Ezzahrioui et al. Asymptotic results of a nonparametric conditional quantile estimator for functional time series
Daziano et al. Computational Bayesian statistics in transportation modeling: from road safety analysis to discrete choice
Lu et al. Estimation of Sobol's sensitivity indices under generalized linear models
Ionides et al. Bagged filters for partially observed interacting systems
Zou et al. Mixture modeling of freeway speed and headway data using multivariate skew-t distributions
Zhou et al. Multiple imputation in two-stage cluster samples using the weighted finite population Bayesian bootstrap
Kuo et al. Estimating the safety impacts in before–after studies using the Naïve Adjustment Method
Hallgren et al. Changepoint detection in non-exchangeable data
Horrigue et al. Non parametric regression quantile estimation for dependent functional data under random censorship: Asymptotic normality
Hiabu et al. Smooth backfitting of proportional hazards with multiplicative components
Beyaztas et al. Spatial function-on-function regression
Kim et al. A neural network-based adaptive cut-off approach to normality testing for dependent data
Browne A comparison of the equivalent weights particle filter and the local ensemble transform Kalman filter in application to the barotropic vorticity equation
JP7439957B2 (ja) パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20965125

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2022567985

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 18255925

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20965125

Country of ref document: EP

Kind code of ref document: A1