WO2022123743A1 - パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム - Google Patents
パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム Download PDFInfo
- Publication number
- WO2022123743A1 WO2022123743A1 PCT/JP2020/046123 JP2020046123W WO2022123743A1 WO 2022123743 A1 WO2022123743 A1 WO 2022123743A1 JP 2020046123 W JP2020046123 W JP 2020046123W WO 2022123743 A1 WO2022123743 A1 WO 2022123743A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- aggregated data
- parameters
- aggregated
- parameter estimation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/11—Complex mathematical operations for solving equations, e.g. nonlinear equations, general mathematical optimization problems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/18—Complex mathematical operations for evaluating statistical data, e.g. average values, frequency distributions, probability functions, regression analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
Definitions
- the present invention relates to a technique for estimating high resolution data from aggregated data aggregated in coarse particle size.
- Spatial data refers to data given as a pair of position information (latitude and longitude, etc.) and some value associated with it.
- Non-Patent Document 3 attempts have been made to utilize data in different domains (city, etc.) (Non-Patent Document 2).
- the target high-resolution data is predicted by simultaneously modeling a plurality of types of aggregated data based on the multivariate Gaussian process based on the Gaussian process.
- the technique disclosed in Non-Patent Document 2 even if the domain (city or the like) is different, the data in a plurality of domains is utilized by sharing the parameters of the Gaussian process between the domains.
- the technique has a problem that the similarity between domains cannot be considered.
- the present invention has been made in view of the above points, and an object of the present invention is to utilize a wide variety of data in a plurality of domains to realize highly accurate prediction of high-resolution data.
- the disclosed technique is a parameter estimator that estimates a plurality of parameters used to calculate high resolution data from aggregated data aggregated in coarse particle size.
- Marginal likelihood assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
- a parameter estimation unit that estimates multiple parameters that are unknown variables in the model so as to maximize
- a storage unit for storing the plurality of parameters is provided.
- the plurality of parameters are provided with a parameter estimation device including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
- FIG. 1 It is a block diagram of the aggregated data high-resolution apparatus in embodiment of this invention. It is a figure which shows the hardware configuration example of the apparatus. It is a figure for demonstrating the flow of the whole processing. It is a figure which shows the search example to the search part and the output example from an output part in embodiment of this invention.
- Means 1 Multivariate Gaussian process model for aggregated data that introduces prior distributions for mixing coefficients
- Means 1 Multivariate Gaussian process model for aggregated data that introduces prior distributions for mixing coefficients
- Means 1 “spatial scale parameters”, “mixing coefficients”, and “parameters for noise”
- the value of the aggregated data is modeled by the integrated value of the Gaussian process in the corresponding region. .. Assuming that the aggregated data actually observed was generated from the above model, the unknown variables are estimated so as to maximize the marginal likelihood.
- Means 2 Efficient parameter estimation by variational Bayesian method
- means 2 in addition to the multivariate Gaussian process model in means 1, unknown variables are estimated based on the variational Bayesian method.
- the aggregated data high resolution device 100 incorporating the above-mentioned means is provided.
- the aggregated data high resolution device 100 can target all the data aggregated in any area (hereinafter referred to as aggregated data for the sake of simplicity), and the type of data used (poor degree, air pollution degree, etc.) It does not depend on the number of dimensions d ⁇ ⁇ 1, ... ⁇ of the data input space (traffic volume, etc.), and can flexibly deal with them and estimate high-resolution data.
- the aggregated data high resolution device 100 in the present embodiment models the value of the aggregated data by the integrated value of the Gaussian process in the corresponding region based on the multivariate Gaussian process model expressed by the linear mixture of a plurality of latent Gaussian processes. Then, after setting the prior distribution for the mixing coefficient, the unknown variable is estimated based on the variational Bayes method, and the high-resolution data is predicted.
- data aggregated in a two-dimensional space will be mainly described as aggregated data, but the present invention is applied to data aggregated in a space at an arbitrary number of dimensions. Can be done. For example, when considering a one-dimensional space, it corresponds to time-series data such as sensor data aggregated at arbitrary time intervals.
- FIG. 1 shows a configuration diagram of the aggregated data high resolution apparatus 100 according to the present embodiment.
- the aggregated data high resolution apparatus 100 shown in the figure includes an aggregated data storage unit 1, a target division storage unit 2, an operation unit 3, a search unit 4, a high resolution data processing unit 5, a parameter estimation unit 6, and a spatial scale parameter storage unit. 7. It has a mixing coefficient storage unit 8, a parameter storage unit 9 for noise, a prior distribution hyperparameter storage unit 10, a high resolution data calculation unit 11, and an output unit 12. The details of the operation of each part will be described later.
- the aggregated data high resolution device 100 may be composed of a plurality of devices (computers) or may be composed of one device. Further, the aggregated data high resolution apparatus 100 may be referred to as an aggregated data high resolution system. Further, the aggregated data high resolution device 100 may be called a parameter estimation device. Further, in FIG. 1, a device including a functional unit other than the aggregated data storage unit 1 and the target division storage unit 2 may be referred to as an aggregated data high resolution device 100.
- a device having a function for estimating parameters that is, a parameter estimation unit 6
- a parameter estimation device an apparatus having a function of increasing the resolution (that is, the high resolution data calculation unit 11) without including the function of estimating parameters
- an aggregated data high resolution apparatus an apparatus having a function of increasing the resolution (that is, the high resolution data calculation unit 11) without including the function of estimating parameters.
- Both the above-mentioned aggregated data high resolution device and parameter estimation device are realized by, for example, causing a computer to execute a program describing the processing contents described in the present embodiment. It is possible.
- the "computer” may be a physical machine or a virtual machine on the cloud. When using a virtual machine, the “hardware” described here is virtual hardware.
- the above program can be recorded on a computer-readable recording medium (portable memory, etc.), saved, and distributed. It is also possible to provide the above program through a network such as the Internet or e-mail.
- FIG. 2 is a diagram showing an example of the hardware configuration of the above computer.
- the computer of FIG. 2 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, which are connected to each other by a bus B, respectively.
- the program that realizes the processing on the computer is provided by, for example, a recording medium 1001 such as a CD-ROM or a memory card.
- a recording medium 1001 such as a CD-ROM or a memory card.
- the program is installed in the auxiliary storage device 1002 from the recording medium 1001 via the drive device 1000.
- the program does not necessarily have to be installed from the recording medium 1001, and may be downloaded from another computer via the network.
- the auxiliary storage device 1002 stores the installed program and also stores necessary files, data, and the like.
- the memory device 1003 reads and stores the program from the auxiliary storage device 1002 when the program is instructed to start.
- the CPU 1004 realizes the function related to the device according to the program stored in the memory device 1003.
- the interface device 1005 is used as an interface for connecting to a network.
- the display device 1006 displays a GUI (Graphical User Interface) or the like by a program.
- the input device 1007 is composed of a keyboard, a mouse, buttons, a touch panel, and the like, and is used for inputting various operation instructions.
- the output device 1008 outputs the calculation result.
- the parameter estimation unit 6 models the value of the aggregated data by the integral value of the Gaussian process in the corresponding region based on the multivariate Gaussian process model expressed by the linear mixture of a plurality of latent Gaussian processes, and the mixing coefficient is used. After setting the prior distribution, multiple parameters that are unknown variables are estimated based on the Gaussian process.
- the high-resolution data calculation unit 11 calculates high-resolution data from the aggregated data using the plurality of parameters calculated in S101, and outputs the high-resolution data from the output unit 12.
- the aggregated data storage unit 1 stores data that can be analyzed by the aggregated data high resolution apparatus 100, reads out the data according to the request from the high resolution data processing unit 5, and converts the corresponding data into the high resolution data processing unit. Send to 5.
- X v ⁇ R d be the input space in the v-th domain
- x ⁇ X v be the input variable.
- P vs. represents the division associated with the corresponding data.
- the division corresponds to, for example, the division of a city by address or region.
- N vs. represents the number of regions included in the divided P vs.
- Area argument n 1, ising , N vs.
- the nth observation contained in the sth data is expressed as (R vsn , y vsn ) as a set of the region R vsn and the value y vsn ⁇ R.
- the data stored in the aggregated data storage unit 1 is
- the aggregated data storage unit 1 stores areas and values for each domain, each type of data, and each area.
- the aggregated data storage unit 1 can be realized by, for example, a Web server, a database server including a database, or the like.
- the aggregated data storage unit 1 may be a storage device in one computer.
- the target division storage unit 2 stores divisions that can be output by the aggregated data high resolution apparatus 100, reads data according to a request from the high resolution data processing unit 5, and inputs the corresponding data to the high resolution data processing unit. Send to 5. It shall represent a division targeting P target .
- One of the regions included in the P target is referred to as an R target .
- any target division can be used, and a division based on an address or region, a mesh of an arbitrary size set by the user, or the like can be considered.
- the target division storage unit 2 is realized by, for example, a Web server, a database server including a database, or the like.
- the target division storage unit 2 may be a storage device in one computer.
- the operation unit 3 receives various operations from the user on the data of the aggregated data storage unit 1 and the target division storage unit 2. Various operations are operations such as registering, modifying, and deleting stored information.
- the input means of the operation unit 3 may be any one such as a keyboard, a mouse, a menu screen, and a touch panel.
- the operation unit 3 is realized by, for example, a device driver for an input means such as a mouse or control software for a menu screen.
- the search unit 4 accepts the target data type for high resolution and the target division.
- the high-resolution data predicted by the aggregated data high-resolution device 100 is output for the data specified by the search unit 4 and the target division.
- the input means of the search unit 4 may be any one such as a keyboard, a mouse, a menu screen, and a touch panel.
- the search unit 4 can be realized by a device driver of an input means such as a mouse or control software of a menu screen.
- the high-resolution data processing unit 5 includes a parameter estimation unit 6, a spatial scale parameter storage unit 7, a mixing coefficient storage unit 8, a parameter storage unit 9 for noise, a prior distribution hyperparameter storage unit 10, and a high resolution. It has a data calculation unit 11.
- the parameter estimation unit 6 estimates the spatial scale parameters, mixing coefficients, parameters for noise, and hyperparameters of prior distribution using the data stored in the aggregated data storage unit 1 as training data. Then, the predicted high resolution data using these parameters is output.
- the high-resolution data processing unit 5 after incorporating the above unknown variables, aggregated data in a plurality of domains is modeled based on a multivariate Gaussian process model expressed by a linear mixture of a plurality of latent Gaussian processes, and observation data is observed. Estimates the unknown variable based on the marginal likelihood given that. After that, high-resolution data is calculated based on the predicted distribution of the Gaussian process.
- ⁇ XV ⁇ l (x, x'): X ⁇ X ⁇ R is the l (L) th correlation function, and any one can be used.
- the correlation function is the l (L) th correlation function, and any one can be used.
- ⁇ l is a spatial scale parameter of the l-th correlation function.
- v 1, ....
- f vs (x) be the Gaussian process for the sth aggregate data of each domain
- S variable Gaussian process f v (x) (f v1 (x), ..., f vs. (X)) T as a linear mixture of L independent Gaussian processes
- g v (x) (g v1 (x), ...., g vL (x)) T , where W v is a mixed matrix of S ⁇ L, and the (s, l) element. W vsl ⁇ R, which is, represents the mixing coefficient. Further, n v (x) represents a noise process for each data, and is a Gaussian process with an average of 0 S variates.
- the vector 0 is a vector having 0 (zero) in all the elements, and ⁇ v (x, x') is.
- ⁇ vs (x, x'): X v ⁇ X v ⁇ R is a correlation function of the noise process for the s th data of the v th domain, and any one can be used. here,
- K v (x, x'): X ⁇ X ⁇ RS ⁇ S represents a correlation matrix.
- ⁇ (x, x') diag ( ⁇ 1 (x, x'), ...., ⁇ L (x, x')).
- Prior distribution for mixing coefficient w vsl is w vsl ⁇ p (w sl ) And. Any distribution can be used for p (w sl ), but here the Gaussian distribution is used.
- -w sl and ⁇ 2 are hyperparameters of prior distribution with respect to the mixing coefficient.
- the symbol placed at the beginning of the character is described before the character, for example, " -w sl ".
- y v is a multidimensional Gaussian distribution
- a vs (x) (a vs1 (x), ...., a vsN vs (x)) T.
- a vsN vs (x) (a vs1 (x), ...., a vsN vs (x)) T.
- vsN vs (x) (a vs1 (x), ...., a vsN vs (x)) T.
- ⁇ 2 vs is a noise dispersion parameter for the sth data.
- I is an identity matrix
- O is a matrix in which all elements are 0.
- C v is a correlation matrix of N v ⁇ N v , and is a correlation matrix.
- Q (W v ) is called a proposed distribution and is introduced to approximate the posterior distribution of W v .
- Q ( Wv) can set an arbitrary probability distribution, it is assumed here that each element wvsl of Wv follows an independent Gaussian distribution for the sake of simplicity. Such an approximation is called mean field approximation.
- the parameter estimation unit 6 performs an operation for estimating various parameters so as to maximize the equation (23), and obtains various parameters. Any continuous optimization technique can be used for maximization.
- the parameters to be estimated by the parameter estimation unit 6 are summarized below.
- the spatial scale parameter storage unit 7 stores ⁇ l
- l 1, ..., L ⁇ obtained by the parameter estimation unit 6.
- the spatial scale parameter storage unit 6 may be anything as long as this information is stored and can be restored.
- the spatial scale parameter storage unit 6 may be a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
- the mixing coefficient storage unit 8 stores ⁇ W v
- v 1, ..., V ⁇ obtained by the parameter estimation unit 6.
- the mixing coefficient storage unit 8 may be any as long as this information is stored and can be restored.
- the mixing coefficient storage unit 8 may be a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
- the parameter storage unit 9 for the is is ⁇ vs
- v 1, ..., V ⁇ is stored.
- the parameter storage unit 9 for noise may be anything as long as this information is stored and can be restored.
- the parameter storage unit 9 for noise may be a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
- the hyperparameter storage unit 10 of the prior distribution stores ⁇ -w sl
- the hyperparameter storage unit 10 of the prior distribution may be anything as long as this information is stored and can be restored.
- the prior-distributed hyperparameter storage unit 10 may be, for example, a specific area of a database or a general-purpose storage device (memory or hard disk device) provided in advance.
- the high resolution data calculation unit 11 calculates the high resolution data using the parameters stored in each of the above storage units. Hereinafter, the processing contents of the high resolution data calculation unit 11 will be described in detail.
- the high-resolution data calculation unit 11 calculates high-resolution data using various learned parameters for the aggregated data specified by the search unit 4 and the target division, and passes the high-resolution data to the output unit 12.
- the calculation method of high resolution data will be described below. Since the derivation of the posterior distribution can be performed in the same way regardless of the domain, the v-th domain will be described below.
- the post-process of the S-variate Gaussian process f v (x) of the v-th domain is derived.
- the ex post process f * v (x) is
- m * v (x): X v ⁇ RS represents an average vector
- K * v (x, x ′): X v ⁇ X v ⁇ RS ⁇ S represents a correlation matrix
- the high-resolution data calculation unit 11 calculates the target high-resolution data by integrating the posterior average (30) in each region R target included in the target division P target .
- the output unit 12 outputs high-resolution data based on the information from the high-resolution data calculation unit 11.
- the output is a concept including display on a display, printing on a printer, sound output, transmission to an external device, and the like.
- the output unit 12 may or may not include an output device such as a display or a speaker.
- the output unit 12 can be realized by the drive software of the output device, the driver software of the output device, the output device, or the like.
- Figure 4 shows an output example.
- the target domain argument, the aggregated data argument, and the target division are received from the search unit 3, and the visualization result of the high resolution data is displayed in the output unit 12 accordingly.
- the shade in the visualization result is determined in proportion to the data value, for example.
- the shade in the visualization result is determined in proportion to the data value, for example.
- This specification discloses at least the parameter estimation device, the aggregated data high resolution device, the parameter estimation method, the aggregated data high resolution method, and the program of each of the following items.
- (Section 1) A parameter estimator that estimates multiple parameters used to calculate high resolution data from aggregated data aggregated to a coarser grain size. Marginal likelihood, assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
- a parameter estimation unit that estimates multiple parameters that are unknown variables in the model so as to maximize
- a storage unit for storing the plurality of parameters is provided.
- the plurality of parameters are parameter estimation devices including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
- the model is described in the first item, which is a model in which the value of the aggregated data of each region in a plurality of regions obtained by dividing the space of the domain is modeled by the integrated value of the multivariate Gaussian process in the corresponding region.
- Parameter estimator (Section 3)
- the parameter estimation unit is the parameter estimation device according to the first or second term, which estimates the plurality of parameters by using the variational Bayes method.
- (Section 4) Aggregated data height including a high-resolution data calculation unit that calculates high-resolution data from aggregated data using the plurality of parameters estimated by the parameter estimation device according to any one of the items 1 to 3.
- Resolution device. (Section 5) It is an aggregated data high resolution device that calculates high resolution data from aggregated data aggregated to coarse particle size. Marginal likelihood, assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
- a parameter estimation unit that estimates multiple parameters that are unknown variables in the model so as to maximize A storage unit that stores the plurality of parameters, and A high-resolution data calculation unit that calculates high-resolution data from aggregated data using the plurality of parameters is provided.
- the plurality of parameters are aggregate data high resolution devices including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
- (Section 6) A parameter estimation method performed by a parameter estimator that estimates multiple parameters used to calculate high resolution data from aggregated data aggregated to a coarser grain size.
- Marginal likelihood assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
- the step of estimating multiple parameters that are unknown variables in the model so as to maximize A step of storing the plurality of parameters in the storage unit is provided.
- the plurality of parameters are parameter estimation methods including hyperparameters of prior distribution to the mixing coefficient used in the linear mixing.
- (Section 7) It is a method of increasing the resolution of aggregated data executed by the aggregated data high-resolution device that calculates high-resolution data from the aggregated data aggregated in coarse particle size.
- Marginal likelihood assuming that the actually observed aggregated data is generated from a model based on a multivariate Gaussian process that is a linear mixture of multiple latent Gaussian processes for multiple types of aggregated data in multiple domains.
- the step of estimating multiple parameters that are unknown variables in the model so as to maximize A step of storing the plurality of parameters in the storage unit, and A step of calculating high resolution data from aggregated data using the plurality of parameters is provided.
- the plurality of parameters are aggregate data high resolution methods including hyperparameters of prior distribution with respect to the mixing coefficient used in the linear mixing.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- Data Mining & Analysis (AREA)
- Theoretical Computer Science (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Software Systems (AREA)
- General Engineering & Computer Science (AREA)
- Databases & Information Systems (AREA)
- Algebra (AREA)
- Operations Research (AREA)
- Computing Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Medical Informatics (AREA)
- Artificial Intelligence (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Probability & Statistics with Applications (AREA)
- Complex Calculations (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
前記複数のパラメータを格納する格納部と、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
パラメータ推定装置が提供される。
本実施の形態では、ドメイン間の類似度が考慮できないという従来技術の問題を解決するために、データ間の依存関係を表すパラメータ(混合係数)に対する事前分布を導入し、それを加味したパラメータ推定手段を導入している。これにより、ドメイン間の類似度を同時に推定しつつ、複数ドメインにおける多種多様なデータを活用して、高解像度データの高精度な予測を実現することが可能となる。より具体的には、下記の手段1と手段2が導入される。
手段1においては、「空間スケールパラメータ」と、「混合係数」と、「ノイズに対するパラメータ」と、「事前分布のハイパーパラメータ」とを未知変数として、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程に基づいて、集約データの値を該当領域におけるガウス過程の積分値によりモデル化する。実際に観測された集約データが上記のモデルから生成されたものと仮定して、周辺尤度を最大化するように未知変数を推定する。
手段2においては、手段1における多変量ガウス過程モデルに加えて、変分ベイズ法に基づいて、未知変数を推定する。
図1に、本実施の形態における集約データ高解像度化装置100の構成図を示す。同図に示す集約データ高解像度化装置100は、集約データ格納部1、ターゲット分割格納部2、操作部3、検索部4、高解像度データ処理部5、パラメータ推定部6、空間スケールパラメータ格納部7、混合係数格納部8、ノイズに対するパラメータ格納部9、事前分布のハイパーパラメータ格納部10、高解像度データ算出部11、出力部12を有する。各部の動作詳細については後述する。
上述した集約データ高解像度化装置、パラメータ推定装置(これらを総称して装置と呼ぶ)はいずれも、例えば、コンピュータに、本実施の形態で説明する処理内容を記述したプログラムを実行させることにより実現可能である。なお、この「コンピュータ」は、物理マシンであってもよいし、クラウド上の仮想マシンであってもよい。仮想マシンを使用する場合、ここで説明する「ハードウェア」は仮想的なハードウェアである。
まず、図3を参照して、全体の処理の流れを説明する。
集約データ格納部1は、集約データ高解像度化装置100によって解析され得るデータを格納しており、高解像度データ処理部5からの要求にしたがって、データを読み出し、該当のデータを高解像度データ処理部5に送信する。
ターゲット分割格納部2は、集約データ高解像度化装置100によって出力され得る分割を格納しており、高解像度データ処理部5からの要求にしたがって、データを読み出し、該当のデータを高解像度データ処理部5に送信する。Ptargetをターゲットとする分割を表すものとする。Ptargetに含まれる領域の1つをRtargetと表す。ここで、ターゲット分割は任意のものを使用することが可能であり、住所や地域に基づく分割や使用者が設定した任意のサイズのメッシュなどが考えられる。ターゲット分割格納部2は、例えば、Webサーバや、データベースを具備するデータベースサーバ等により実現される。ターゲット分割格納部2が、1つのコンピュータ内の記憶装置であってもよい。
操作部3は、集約データ格納部1、及び、ターゲット分割格納部2のデータに対するユーザからの各種操作を受け付ける。各種操作とは、格納された情報を登録、修正、削除する操作等である。操作部3の入力手段は、キーボードやマウス、メニュー画面、タッチパネルによるもの等、どのようなものでもよい。操作部3は、例えば、マウス等の入力手段のデバイスドライバや、メニュー画面の制御ソフトウェアで実現される。
検索部4は、高解像度化を行う対象とするデータ種別、及び、ターゲット分割を受け付ける。検索部4で指定されたデータ、及び、ターゲット分割に対して、集約データ高解像度化装置100によって予測された高解像度データを出力する。なお、検索部4の入力手段は、キーボードやマウス、メニュー画面、タッチパネルによるもの等、どのようなものでもよい。検索部4は、マウス等の入力手段のデバイスドライバや、メニュー画面の制御ソフトウェアで実現され得る。
図1に示すとおり、高解像度データ処理部5は、パラメータ推定部6、空間スケールパラメータ格納部7、混合係数格納部8、ノイズに対するパラメータ格納部9、事前分布のハイパーパラメータ格納部10、高解像度データ算出部11を有する。
以下、パラメータ推定部6において処理されるガウス過程モデルとパラメータ推定法について詳細に説明する。まず、複数の潜在ガウス過程の線形混合で表現された多変量ガウス過程の定式化を行う。V×L個の独立なガウス過程を
wvsl~p(wsl)
とする。p(wsl)は任意の分布を用いることができるが、ここではガウス分布を用いて
・混合係数{Wv|v=1,...,V}
・ノイズに対するパラメータ{αvs|v=1,...,V;s=1,...,S}と{κvs|v=1,...,V;s=1,...,S}と{Σv|v=1,...,V}
・事前分布のハイパーパラメータ{‐wsl|s=1,...,S;l=1,...,L}とη2
なお、変分パラメータは上記パラメータを推定するための補助変数として機能し、後述する高解像度データ算出部11においては使用されない。推定されたパラメータは、下記の各格納部に格納される。なお、複数のパラメータを格納する格納部が1つだけ備えられることとしてもよい。
空間スケールパラメータ格納部7は、パラメータ推定部6で求めた{βl|l=1,...,L}を格納する。空間スケールパラメータ格納部6は、この情報が保存され、復元可能なものであればどのようなものでもよい。例えば、空間スケールパラメータ格納部6は、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
混合係数格納部8は、パラメータ推定部6で求めた{Wv|v=1,...,V}を格納する。混合係数格納部8は、この情報が保存され、復元可能なものであればどのようなものでもよい。例えば、混合係数格納部8は、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
イズに対するパラメータ格納部9は、パラメータ推定部6で求めた{αvs|v=1,...,V;s=1,...,S}、{κvs|v=1,...,V;s=1,...,S}、{Σv|v=1,...,V}を格納する。ノイズに対するパラメータ格納部9は、この情報が保存され、復元可能なものであればどのようなものでもよい。例えば、ノイズに対するパラメータ格納部9は、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
事前分布のハイパーパラメータ格納部10は、パラメータ推定部6で求めた{‐wsl|s=1,...,S;l=1,...,L}、η2を格納する。事前分布のハイパーパラメータ格納部10は、この情報が保存され、復元可能なものであればどのようなものでもよい。事前分布のハイパーパラメータ格納部10は、例えば、データベースや、あらかじめ備えられた汎用的な記憶装置(メモリやハードディスク装置)の特定領域であってもよい。
高解像度データ算出部11は、検索部4で指定された集約データ、及び、ターゲット分割に対して、学習済みの各種パラメータを用いて高解像度データを算出し、出力部12へ渡す。以下に、高解像度データの算出法について説明する。ドメインによらず事後分布の導出は同様に行うことができるので、以下ではv番目のドメインについて説明をする。まず、v番目のドメインのS変量ガウス過程fv(x)の事後プロセスを導出する。事後プロセスf* v(x)は、
出力部12は、高解像度データ算出部11からの情報に基づいて、高解像度データを出力する。ここで、出力とは、ディスプレイへの表示、プリンタへの印字、音出力、外部装置への送信等を含む概念である。出力部12は、ディスプレイやスピーカ等の出力デバイスを含むと考えても含まないと考えてもよい。出力部12は、出力デバイスのドライブソフト又は、出力デバイスのドライバソフトと出力デバイス等で実現され得る。
以上、説明したとおり、本実施の形態では、データ間の依存関係を表すパラメータ(混合係数)に対する事前分布を導入し、それを加味したパラメータ推定を行うこととしているので、ドメイン間の類似度を同時に推定しつつ、複数ドメインにおける多種多様なデータを活用して、高解像度データの高精度な予測を実現することが可能となる。
本明細書には、少なくとも下記各項のパラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラムが開示されている。
(第1項)
粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
前記複数のパラメータを格納する格納部と、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
パラメータ推定装置。
(第2項)
前記モデルは、ドメインの空間を分割することより得られる複数の領域における各領域の集約データの値を、該当領域における多変量ガウス過程の積分値でモデル化したモデルである
第1項に記載のパラメータ推定装置。
(第3項)
前記パラメータ推測部は、変分ベイズ法を用いて前記複数のパラメータを推定する
第1項又は第2項に記載のパラメータ推定装置。
(第4項)
第1項ないし第3項のうちいずれか1項に記載のパラメータ推定装置により推定された前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部
を備える集約データ高解像度化装置。
(第5項)
粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
前記複数のパラメータを格納する格納部と、
前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部と、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
集約データ高解像度化装置。
(第6項)
粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置が実行するパラメータ推定方法であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
前記複数のパラメータを格納部に格納するステップと、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
パラメータ推定方法。
(第7項)
粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置が実行する集約データ高解像度化方法であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
前記複数のパラメータを格納部に格納するステップと、
前記複数のパラメータを用いて、集約データから高解像度データを算出するステップと、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
集約データ高解像度化方法。
(第8項)
コンピュータを、第1項ないし第3項のうちいずれか1項に記載のパラメータ推定装置における各部として機能させるためのプログラム、又は、第5項に記載の集約データ高解像度化装置における各部として機能させるためのプログラム。
1 集約データ格納部
2 ターゲット分割格納部
3 操作部
4 検索部
5 高解像度データ処理部
6 パラメータ推定部
7 空間スケールパラメータ格納部
8 混合係数格納部
9 ノイズに対するパラメータ格納部
10 事前分布のハイパーパラメータ格納部
11 高解像度データ算出部
12 出力部
1000 ドライブ装置
1001 記録媒体
1002 補助記憶装置
1003 メモリ装置
1004 CPU
1005 インタフェース装置
1006 表示装置
1007 入力装置
Claims (8)
- 粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
前記複数のパラメータを格納する格納部と、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
パラメータ推定装置。 - 前記モデルは、ドメインの空間を分割することより得られる複数の領域における各領域の集約データの値を、該当領域における多変量ガウス過程の積分値でモデル化したモデルである
請求項1に記載のパラメータ推定装置。 - 前記パラメータ推測部は、変分ベイズ法を用いて前記複数のパラメータを推定する
請求項1又は2に記載のパラメータ推定装置。 - 請求項1ないし3のうちいずれか1項に記載のパラメータ推定装置により推定された前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部
を備える集約データ高解像度化装置。 - 粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するパラメータ推測部と、
前記複数のパラメータを格納する格納部と、
前記複数のパラメータを用いて、集約データから高解像度データを算出する高解像度データ算出部と、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
集約データ高解像度化装置。 - 粗い粒度に集約された集約データから高解像度データを算出するために使用される複数のパラメータを推定するパラメータ推定装置が実行するパラメータ推定方法であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
前記複数のパラメータを格納部に格納するステップと、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
パラメータ推定方法。 - 粗い粒度に集約された集約データから高解像度データを算出する集約データ高解像度化装置が実行する集約データ高解像度化方法であって、
実際に観測された集約データが、複数ドメインにおける複数の種類の集約データに対する複数の潜在ガウス過程を線形混合で表現した多変量ガウス過程に基づくモデルから生成されたものと仮定して、周辺尤度を最大化するように、前記モデルにおける未知変数である複数のパラメータを推定するステップと、
前記複数のパラメータを格納部に格納するステップと、
前記複数のパラメータを用いて、集約データから高解像度データを算出するステップと、を備え、
前記複数のパラメータは、前記線形混合で用いられる混合係数に対する事前分布のハイパーパラメータを含む
集約データ高解像度化方法。 - コンピュータを、請求項1ないし3のうちいずれか1項に記載のパラメータ推定装置における各部として機能させるためのプログラム、又は、請求項5に記載の集約データ高解像度化装置における各部として機能させるためのプログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/255,925 US20240045921A1 (en) | 2020-12-10 | 2020-12-10 | Parameter estimation apparatus, aggregated data resolution enhancement apparatus, parameter estimation method, aggregated data resolution enhancement method and program |
| JP2022567985A JP7439957B2 (ja) | 2020-12-10 | 2020-12-10 | パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム |
| PCT/JP2020/046123 WO2022123743A1 (ja) | 2020-12-10 | 2020-12-10 | パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/046123 WO2022123743A1 (ja) | 2020-12-10 | 2020-12-10 | パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022123743A1 true WO2022123743A1 (ja) | 2022-06-16 |
Family
ID=81973446
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/046123 Ceased WO2022123743A1 (ja) | 2020-12-10 | 2020-12-10 | パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20240045921A1 (ja) |
| JP (1) | JP7439957B2 (ja) |
| WO (1) | WO2022123743A1 (ja) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP2530626A1 (en) * | 2011-06-01 | 2012-12-05 | BAE Systems Plc. | Heterogeneous data fusion using Gaussian processes |
| JP2017033198A (ja) * | 2015-07-30 | 2017-02-09 | 日本電信電話株式会社 | 時空間変数予測装置及びプログラム |
-
2020
- 2020-12-10 US US18/255,925 patent/US20240045921A1/en active Pending
- 2020-12-10 WO PCT/JP2020/046123 patent/WO2022123743A1/ja not_active Ceased
- 2020-12-10 JP JP2022567985A patent/JP7439957B2/ja active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP2530626A1 (en) * | 2011-06-01 | 2012-12-05 | BAE Systems Plc. | Heterogeneous data fusion using Gaussian processes |
| JP2017033198A (ja) * | 2015-07-30 | 2017-02-09 | 日本電信電話株式会社 | 時空間変数予測装置及びプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| US20240045921A1 (en) | 2024-02-08 |
| JP7439957B2 (ja) | 2024-02-28 |
| JPWO2022123743A1 (ja) | 2022-06-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| de Fondeville et al. | High-dimensional peaks-over-threshold inference | |
| Zou et al. | Empirical Bayes estimates of finite mixture of negative binomial regression models and its application to highway safety | |
| Kazianka et al. | Copula-based geostatistical modeling of continuous and discrete data including covariates | |
| Sellers et al. | Underdispersion models: Models that are “under the radar” | |
| Middleton et al. | Unbiased estimation of the average treatment effect in cluster-randomized experiments | |
| US11593860B2 (en) | Method, medium, and system for utilizing item-level importance sampling models for digital content selection policies | |
| Plumlee et al. | Building accurate emulators for stochastic simulations via quantile kriging | |
| Schneider et al. | Learning stochastic closures using ensemble Kalman inversion | |
| De Jong et al. | Multiple imputation of predictor variables using generalized additive models | |
| Wu et al. | Uapd: Predicting urban anomalies from spatial-temporal data | |
| Yang et al. | A stochastic expectation-maximization algorithm for the analysis of system lifetime data with known signature | |
| Ezzahrioui et al. | Asymptotic results of a nonparametric conditional quantile estimator for functional time series | |
| Daziano et al. | Computational Bayesian statistics in transportation modeling: from road safety analysis to discrete choice | |
| Lu et al. | Estimation of Sobol's sensitivity indices under generalized linear models | |
| Ionides et al. | Bagged filters for partially observed interacting systems | |
| Zou et al. | Mixture modeling of freeway speed and headway data using multivariate skew-t distributions | |
| Zhou et al. | Multiple imputation in two-stage cluster samples using the weighted finite population Bayesian bootstrap | |
| Kuo et al. | Estimating the safety impacts in before–after studies using the Naïve Adjustment Method | |
| Hallgren et al. | Changepoint detection in non-exchangeable data | |
| Horrigue et al. | Non parametric regression quantile estimation for dependent functional data under random censorship: Asymptotic normality | |
| Hiabu et al. | Smooth backfitting of proportional hazards with multiplicative components | |
| Beyaztas et al. | Spatial function-on-function regression | |
| Kim et al. | A neural network-based adaptive cut-off approach to normality testing for dependent data | |
| Browne | A comparison of the equivalent weights particle filter and the local ensemble transform Kalman filter in application to the barotropic vorticity equation | |
| JP7439957B2 (ja) | パラメータ推定装置、集約データ高解像度化装置、パラメータ推定方法、集約データ高解像度化方法、及びプログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20965125 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2022567985 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18255925 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20965125 Country of ref document: EP Kind code of ref document: A1 |






























