WO2022014657A1 - 解析装置、解析方法及びプログラム - Google Patents
解析装置、解析方法及びプログラム Download PDFInfo
- Publication number
- WO2022014657A1 WO2022014657A1 PCT/JP2021/026531 JP2021026531W WO2022014657A1 WO 2022014657 A1 WO2022014657 A1 WO 2022014657A1 JP 2021026531 W JP2021026531 W JP 2021026531W WO 2022014657 A1 WO2022014657 A1 WO 2022014657A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- analysis
- probability
- measure
- measures
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/18—Complex mathematical operations for evaluating statistical data, e.g. average values, frequency distributions, probability functions, regression analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F7/00—Methods or arrangements for processing data by operating upon the order or content of the data handled
- G06F7/38—Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation
- G06F7/48—Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation using non-contact-making devices, e.g. tube, solid state device; using unspecified devices
- G06F7/52—Multiplying; Dividing
- G06F7/523—Multiplying only
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N7/00—Computing arrangements based on specific mathematical models
Definitions
- the present invention relates to an analysis device, an analysis method and a program.
- Kernel mean embedding is known as a framework for dealing with such randomness in data analysis. Randomness is formulated by a probability measure, which is a set function that expresses the likelihood of an event. Kernel average embedding is a method that gives this probability measure the concept of "closeness" such as inner product and norm, and the closeness between probability measures is given by the inner product in a space called RKHS (reproducing kernel Hilbert space). Since many data analysis methods are based on the concept of closeness, they can measure the closeness of data including randomness, estimate the probability measure to generate data with certain randomness, and so on. It is possible to apply the general data analysis of the above to data with randomness.
- RKHM reproducing kernel Hilbert C * -module
- C * -algebra which is a generalization of matrices and linear operators, instead of the inner product, which is usually a complex number. You will be able to perform analysis while doing so. This makes it possible to analyze interacting data with high accuracy and extract interaction information.
- Non-Patent Document 1 and Non-Patent Document 2 are limited to theoretical research, and there is still no framework for using a probability measure having a value for a linear agonist in actual data analysis.
- Recently, research to analyze data appearing from quantum using machine learning methods has also attracted attention, and from such a viewpoint, a probability measure with a value in a linear action element that can handle multiple randomnesses at the same time. It is considered important to use the framework for data analysis.
- One embodiment of the present invention has been made in view of the above points, and an object thereof is to realize data analysis having a plurality of randomnesses.
- the analysis apparatus includes an acquisition unit that acquires a data set of a plurality of data having randomness, and probability measures ⁇ and ⁇ on the data set, and is von Neumann.
- the inner product or norm of ⁇ ( ⁇ ) and ⁇ ( ⁇ ) mapped on the RKHM by the mapping ⁇ that extends the kernel average embedding of the probability measures ⁇ and ⁇ that have values in the ring is the inner product of the probability measures ⁇ and ⁇ .
- it is characterized by having an analysis unit that calculates as a norm.
- the present embodiment describes an analysis device 10 capable of performing data analysis having a plurality of randomnesses.
- analysis of data having a plurality of randomness in particular, visualization of data representing data and a quantum state when a plurality of random data interact with each other, for example. , Abnormality detection, etc. can be performed.
- the analysis device 10 according to the present embodiment in addition to the analysis such as visualization and abnormality detection, for example, the data in which the abnormality is detected based on the analysis result (particularly, the abnormality detection result etc.) can be obtained. You may control the stop of the represented device, device, program, etc.
- the kernel average embedding is extended to give the concept of "closeness" such as inner product and norm to the probability measure having a value in the linear operator.
- the value of the inner product is not a complex numerical value but a linear operator value.
- kernel average embedding using RKHM is used instead of the known kernel average embedding using RKHS.
- positive means that it is a positive-definite value in von Neumann-algebra, and is a generalization of an Hermitian matrix (that is, a Hermitian positive-definite value) in which all eigenvalues are 0 or more.
- This map ⁇ is also called a feature map.
- RKHM a space called RKHM. This space is expressed as M k.
- M k the inner product ⁇ , ⁇ > k of the A value and the magnitude of the A value
- k can be determined.
- An A-value measure on X is a function ⁇ from a subset of X called a measurable set to A, and an infinite number of countable sets E 1 , E 2 , ... ⁇ ⁇ For
- the A-value function f is a sequence of functions called a simple function.
- integral with respect to ⁇ a f is defined by the limit of the integral with respect to ⁇ of s i.
- the single function s some finite number of measurable set E 1, ⁇ , such as the c 1 so that there is no intersection of any two pairs in E n, ⁇ , c n ⁇ A against
- k (x, y) is the complex value positives definite kernel on all components X 2
- Example 2 Measure representing the quantum state In quantum mechanics, let A be the set of all bounded linear operators. Since the quantum state is represented by the linear action element ⁇ and the observation is represented by the A value measure ⁇ , for the linear action elements ⁇ 1 and ⁇ 2 representing the quantum state and the A value measures ⁇ 1 and ⁇ 2 representing the observation. , Observation of each state The closeness of ⁇ 1 ⁇ 1 and ⁇ 2 ⁇ 2 can be expressed by the inner product of ⁇ ( ⁇ 1 ⁇ 1 ) and ⁇ ( ⁇ 2 ⁇ 2).
- the inner product of ⁇ ( ⁇ 1 ) and ⁇ ( ⁇ 2 ) can be calculated by the following equation (2) for the states ⁇ 1 and ⁇ 2 ⁇ C m ⁇ m.
- A C m ⁇ m .
- G be a matrix having ⁇ ( ⁇ i ), ⁇ ( ⁇ j )> k ⁇ A in the (i, j) block for a plurality of A value measures ⁇ 1 , ⁇ , ⁇ n. Since G is an Hermitian positive-definite matrix, there are eigenvalues ⁇ 1 ⁇ ... ⁇ ⁇ mn ⁇ 0 and orthonormal eigenvectors v 1 , ..., v mn corresponding to these eigenvalues, respectively.
- the dimension reduction can be performed so as to keep the information of the covariance between the data.
- the existing method using kernel mean embedding in RKHS for machine learning and statistics is the kernel mean embedding of the probability measure in RKHS in the kernel in RKHM of the covariance measure described in Example 1 above.
- average embedding it can be applied to data with multiple dependent elements. For example, the following examples can be given.
- FIG. 1 is a diagram showing an example of a hardware configuration of the analysis device 10 according to the present embodiment.
- the analysis device 10 is realized by a general computer or a computer system, and as hardware, an input device 11, a display device 12, an external I / F 13, and a communication I / It has an F14, a processor 15, and a memory device 16. Each of these hardware is connected so as to be communicable via the bus 17.
- the input device 11 is, for example, a keyboard, a mouse, a touch panel, or the like.
- the display device 12 is, for example, a display or the like.
- the analysis device 10 does not have to have at least one of the input device 11 and the display device 12.
- the external I / F13 is an interface with an external device.
- the external device includes a recording medium 13a and the like.
- the analysis device 10 can read or write the recording medium 13a via the external I / F 13.
- the recording medium 13a includes, for example, a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), a USB (Universal Serial Bus) memory card, and the like.
- the communication I / F 14 is an interface for connecting the analysis device 10 to the communication network.
- the processor 15 is, for example, various arithmetic units such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).
- the memory device 16 is, for example, various storage devices such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), a RAM (Random Access Memory), a ROM (Read Only Memory), and a flash memory.
- the analysis device 10 can realize the data analysis process described later.
- the hardware configuration shown in FIG. 1 is an example, and the analysis device 10 may have another hardware configuration.
- the analysis device 10 may have a plurality of processors 15 or a plurality of memory devices 16.
- FIG. 2 is a diagram showing an example of the functional configuration of the analysis device 10 according to the present embodiment.
- the analysis device 10 has an acquisition unit 101, an analysis unit 102, and a storage unit 103 as functional units.
- the acquisition unit 101 and the analysis unit 102 are realized, for example, by a process in which one or more programs installed in the analysis device 10 are executed by the processor 15.
- the storage unit 103 can be realized by using, for example, the memory device 16.
- the storage unit 103 may be realized by, for example, a storage device (for example, a database server or the like) connected to the analysis device 10 via a communication network.
- the storage unit 103 stores data to be analyzed (for example, an element of X to be analyzed and its A value measure, and when applied to Example 2 above, a linear operator ⁇ that further represents a quantum state). Will be done.
- the acquisition unit 101 acquires the data to be analyzed from the storage unit 103.
- the analysis unit 102 analyzes the data acquired by the acquisition unit 101 (that is, for example, calculation of the inner product / norm, visualization / abnormality detection using the calculation result, etc.).
- FIG. 3 is a flowchart showing an example of data analysis processing according to the present embodiment.
- the acquisition unit 101 obtains the data to be analyzed (that is, the element of X to be analyzed and its A value measure, and when applied to the above Example 2, the linear operator ⁇ that further represents the quantum state, etc.). Acquired from the storage unit 103 (step S101).
- the analysis unit 102 analyzes the data acquired in the above step S101 (step S102).
- step S102 For data analysis, calculation of inner product / norm described in "2. Application of kernel average embedding using RKHM", visualization / abnormality detection using the calculation result, comparison between data, and data Generation, learning, etc. can be mentioned.
- the above equation (1) is used when the measure represents the covariance between a plurality of data having randomness, and the above equation (1) is used when the measure represents the quantum state. As shown in 2).
- the analysis device 10 is capable of data analysis having a plurality of randomness (particularly, visualization of data when a plurality of random data are interacting with each other and data representing a quantum state, and abnormality detection. Etc.) can be done.
- ⁇ X is a measure whose (i, j) component represents the covariance of X i and X j.
- the A value measure is such that At this time, the inner product of ⁇ ( ⁇ X ) and ⁇ ( ⁇ Y ), the inner product of ⁇ ( ⁇ Y ) and ⁇ ( ⁇ Z ), and the inner product of ⁇ ( ⁇ X ) and ⁇ ( ⁇ Z ) are calculated by the above equation (1). ), And ⁇ X , ⁇ Y , and ⁇ Z were visualized on the first and second spindles by Kernel PCA. The results are shown in FIG. As shown in FIG. 4, the distances between ⁇ Y and ⁇ Z that are related to each other are short, while the distances between ⁇ X and ⁇ Y that are not related and the distance between ⁇ X and ⁇ Z are far. It has become.
- the distance between each data is measured by the analysis device 10 according to the present embodiment (that is, measured by
- the existing distance was measured and compared with the one that was subjected to the two-sample test (conventional method).
- Conventional methods include RKHS described in Reference 1 and Reference 4 “BK Sriperumbudur, K. Fukumizu, A. Gretton, B. Scholkopf, and GRG Lanckriet, On the empirical estimation of integral probability metrics. Electronic Journal. Of Statistics, 6: 1550-1599, 2012. ”, Kantrovich and Dadley were adopted.
- the proposed method and the conventional method were tested 50 times with different data, and the rate at which the two types of samples were determined to follow the same distribution was calculated. The results are shown in Table 1 below.
- ⁇ is defined in the same manner as in Example 2 above. A small amount of noise was added to each of ⁇ 1 and ⁇ 2, and 50 samples were prepared for each.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Algebra (AREA)
- Computing Systems (AREA)
- Bioinformatics & Computational Biology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Evolutionary Biology (AREA)
- Operations Research (AREA)
- Probability & Statistics with Applications (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Complex Calculations (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一実施形態に係る解析装置は、ランダム性を持つ複数のデータのデータ集合を取得する取得部と、前記データ集合上の確率測度μ及びνであって、フォン・ノイマン環に値を持つ確率測度μ及びνを、カーネル平均埋め込みを拡張した写像ΦによってRKHM上にそれぞれ写像したΦ(μ)及びΦ(ν)の内積又はノルムを、前記確率測度μ及びνの内積又はノルムとして計算する解析部と、を有することを特徴とする。
Description
本発明は、解析装置、解析方法及びプログラムに関する。
自然界に現れるデータは基本的にランダム性を含んでおり、ランダム性を考慮したデータ解析技術が従来から研究されている。このようなランダム性をデータ解析において扱う枠組みとして、カーネル平均埋め込み(kernel mean embedding)が知られている。ランダム性は、事象の起こりやすさを表す集合関数である確率測度によって定式化される。カーネル平均埋め込みでは、この確率測度に内積やノルムといった「近さ」の概念を与える手法であり、確率測度同士の近さはRKHS(reproducing kernel Hilbert space)と呼ばれる空間での内積により与えられる。多くのデータ解析手法は近さの概念により成り立っているため、これにより、ランダム性を含むデータ同士の近さを測ったり、或るランダム性のあるデータを生成する確率測度を推定したりする等の一般的なデータ解析を、ランダム性のあるデータに対して適用することが可能となる。
一方で、ランダム性を含まないデータ解析技術で、複数のデータの相互作用を考慮する枠組みとして、RKHM(reproducing kernel Hilbert C*-module)を用いたものが知られている。RKHMはRKHSの拡張であり、通常複素数値である内積の代わりに、行列や線形作用素を一般化したC*-algebraと呼ばれる空間に値を持つ内積を定義することで、相互作用の情報を保存したまま解析を行うことができるようになる。これにより、相互作用のあるデータを精度良く解析したり、相互作用の情報を抽出したりすることが可能となる。
ところで、データの中には、複数のランダムなデータが相互作用し合って生じるものも多い。また、量子計算等の量子を扱う分野においては、量子の状態が各観測の確率という複数の確率により表現される。ランダム性を定式化するためには確率測度が用いられるが、データ解析における既存の枠組みでは、確率測度は複素数値であり、複数のランダム性を同時に扱うことはできない。一方で、量子力学においては、複数の確率により表される量子の状態を定式化するために、Hilbert空間上の線形作用素に値を持つ確率測度が用いられている(例えば、非特許文献1)。また、純粋数学の分野では、これをより一般化したベクトル値測度という概念が理論的に研究されている(例えば、非特許文献2)。
H.E. Brandt, Quantum measurement with a positive operator-valued measure. Acta Phys. Hung. B 20, 95-99, 2004.
C. W. Swartz, Products of vector measures by means of Fubini's theorem, Mathematica Slovaca, 27(4):375-382, 1977.
しかしながら、上記の非特許文献1や非特許文献2は理論的な研究に留まっており、実際のデータ解析において、線形作用素に値を持つ確率測度を利用する枠組みは未だない。最近では量子から現れるデータを、機械学習の手法を用いて解析する研究も注目を集めており、そのような観点からも、複数のランダム性を同時に扱えるような、線形作用素に値を持つ確率測度をデータ解析において利用する枠組みは重要であると考えられる。
本発明の一実施形態は、上記の点に鑑みてなされたもので、複数のランダム性を持つデータ解析を実現することを目的とする。
上記目的を達成するため、一実施形態に係る解析装置は、ランダム性を持つ複数のデータのデータ集合を取得する取得部と、前記データ集合上の確率測度μ及びνであって、フォン・ノイマン環に値を持つ確率測度μ及びνを、カーネル平均埋め込みを拡張した写像ΦによってRKHM上にそれぞれ写像したΦ(μ)及びΦ(ν)の内積又はノルムを、前記確率測度μ及びνの内積又はノルムとして計算する解析部と、を有することを特徴とする。
複数のランダム性を持つデータ解析を実現することができる。
以下、本発明の一実施形態について説明する。本実施形態は、複数のランダム性を持つデータ解析を行うことができる解析装置10について説明する。本実施形態に係る解析装置10を用いることで、複数のランダム性を持つデータの解析、特に、例えば、複数のランダムなデータが相互作用し合っている場合のデータや量子状態を表すデータの可視化、異常検知等を行うことが可能となる。なお、本実施形態に係る解析装置10は、このような可視化や異常検知等の解析に加えて、例えば、その解析結果(特に、異常検知結果等)に基づいて、異常が検知されたデータが表す装置、機器、プログラム等の停止等の制御を行ってもよい。
<理論的構成及びその応用例>
まず、本実施形態の理論的構成及びその応用例について説明する。本実施形態では、カーネル平均埋め込みを拡張し、線形作用素に値を持つ確率測度に内積・ノルム等の「近さ」の概念を与える。ただし、複数のランダム性の情報をできるだけ保った解析を行うために、内積の値は複素数値ではなく、線形作用素値とする。このために、RKHSを用いた既知のカーネル平均埋め込みの代わりに、RKHMを用いたカーネル平均埋め込みとする。
まず、本実施形態の理論的構成及びその応用例について説明する。本実施形態では、カーネル平均埋め込みを拡張し、線形作用素に値を持つ確率測度に内積・ノルム等の「近さ」の概念を与える。ただし、複数のランダム性の情報をできるだけ保った解析を行うために、内積の値は複素数値ではなく、線形作用素値とする。このために、RKHSを用いた既知のカーネル平均埋め込みの代わりに、RKHMを用いたカーネル平均埋め込みとする。
1. RKHMを用いたカーネル平均埋め込み
Xをデータ(ランダム性を持つデータ)の属する空間、Aをフォン・ノイマン環(von Neumann-algebra)とし、A値positive definite kernel k:X×X→Aを考える。ただし、写像k:X×X→AがA値positive definite kernelであるとは、以下の条件1及び条件2を満たすことをいう。なお、フォン・ノイマン間の具体例としては、例えば、線形作用素全体の集合や行列全体の集合等が挙げられる。
Xをデータ(ランダム性を持つデータ)の属する空間、Aをフォン・ノイマン環(von Neumann-algebra)とし、A値positive definite kernel k:X×X→Aを考える。ただし、写像k:X×X→AがA値positive definite kernelであるとは、以下の条件1及び条件2を満たすことをいう。なお、フォン・ノイマン間の具体例としては、例えば、線形作用素全体の集合や行列全体の集合等が挙げられる。
(条件1)任意のx,y∈Xに対して、k(x,y)=k(x,y)* (*は共役を表す)
(条件2)mを任意の自然数として、任意のx0,x1,・・・,xm-1∈Xと任意のc0,c1,・・・,cm-1∈Aに対して、
(条件2)mを任意の自然数として、任意のx0,x1,・・・,xm-1∈Xと任意のc0,c1,・・・,cm-1∈Aに対して、
ここで、positiveとはvon Neumann-algebraで正定値であることを意味し、全ての固有値が0以上であるエルミート行列(つまり、エルミート正定値)等の一般化である。
A値positive definite kernel kが与えられたとき、XからA値関数への写像φを、φ(x)=k(・,x)により定義する。この写像φはfeature mapとも呼ばれる。
自然数mと、x0,x1,・・・,xm-1∈Xと、c0,c1,・・・,cm-1∈Aに対して、
X上のA値測度とは、可測集合と呼ばれるXの部分集合からAへの関数μで、任意の2ペアの交わりがないような可算無限個の可測集合E1,E2,・・・に対して、
A値測度に対して、その測度に関する積分を考えることができる。A値関数fが、単関数と呼ばれる関数の列
このとき、s(x)を左からμで積分した値を
上記の設定の下、有限なA値測度をRKHMの元に移す写像Φを、
例えば、X=Rd、A=Cm×mに対して、k:X×X→Aを、
2. RKHMを用いたカーネル平均埋め込みの応用
2.1 A値測度の間の距離
有限A値測度μ、νのA値距離を以下で定義する。
2.1 A値測度の間の距離
有限A値測度μ、νのA値距離を以下で定義する。
γ(μ,ν)=|Φ(μ)-Φ(ν)|k
このとき、Φが単射であれば、例えば、||γ(μ,ν)||は距離の性質を完全に満たす。つまり、||γ(μ,ν)||=||γ(ν,μ)||、||γ(μ,ν)||=0ならばμ=ν、||γ(μ,ν)||≦||γ(μ,λ)||+||γ(λ,ν)||が任意の有限A値測度μ、ν、λに対して成立する。
このとき、Φが単射であれば、例えば、||γ(μ,ν)||は距離の性質を完全に満たす。つまり、||γ(μ,ν)||=||γ(ν,μ)||、||γ(μ,ν)||=0ならばμ=ν、||γ(μ,ν)||≦||γ(μ,λ)||+||γ(λ,ν)||が任意の有限A値測度μ、ν、λに対して成立する。
以下に有限A値測度の例を2つ挙げる。
例1:ランダム性を持つ複数のデータ間の共分散を表す測度
A=Cm×mとする。Xに値を持つm個の確率変数X1,・・・,XmとY1,・・・,Ymを考える。PをX上の確率測度とし、μXを、(i,j)成分がXiとXjの共分散を表す測度(Xi,Xj)*PになるようなA値測度(又は、その測度を中心化したバージョン
A=Cm×mとする。Xに値を持つm個の確率変数X1,・・・,XmとY1,・・・,Ymを考える。PをX上の確率測度とし、μXを、(i,j)成分がXiとXjの共分散を表す測度(Xi,Xj)*PになるようなA値測度(又は、その測度を中心化したバージョン
実際には、X1,・・・,Xmから得られたデータ{x1,1,x1,2,・・・,x1,N},・・・,{xm,1,xm,2,・・・,xm,N}と、Y1,・・・,Ymから得られたデータ{y1,1,y1,2,・・・,y1,N},・・・,{ym,1,ym,2,・・・,ym,N}とが与えられた際、Φ(μX)とΦ(μY)の内積〈Φ(μX),Φ(μY)〉kの(i,j)成分を以下の式(1)のように近似する。
例2:量子の状態を表す測度
量子力学において、Aを有界線形作用素全体の集合とする。量子の状態は線形作用素ρにより表され、その観測はA値測度μにより表されるため、量子の状態を表す線形作用素ρ1、ρ2、観測を表すA値測度μ1、μ2に対し、各状態の観測μ1ρ1とμ2ρ2の近さはΦ(μ1ρ1)とΦ(μ2ρ2)の内積により表すことができる。
量子力学において、Aを有界線形作用素全体の集合とする。量子の状態は線形作用素ρにより表され、その観測はA値測度μにより表されるため、量子の状態を表す線形作用素ρ1、ρ2、観測を表すA値測度μ1、μ2に対し、各状態の観測μ1ρ1とμ2ρ2の近さはΦ(μ1ρ1)とΦ(μ2ρ2)の内積により表すことができる。
例えば、A=Cm×m、X=Cmとし、i=1,・・・,sに対し、|ψi〉∈Xを正規化されたベクトルとする。これに対して、観測(つまり、X上のA値測度)
A=Cm×mとする。複数のA値測度μ1,・・・,μnに対して、〈Φ(μi),Φ(μj)〉k∈Aを(i,j)ブロックに持つ行列をGとする。Gはエルミート正定値行列になるため、固有値λ1≧・・・≧λmn≧0と、これらの固有値にそれぞれ対応する正規直交な固有ベクトルv1,・・・,vmnとが存在する。第i主軸を
2.3 その他の応用例
機械学習や統計のRKHSにおけるカーネル平均埋め込みを用いる既存の方法は、RKHSにおける確率測度のカーネル平均埋め込みを、上記の例1で記載した共分散を表す測度のRKHMにおけるカーネル平均埋め込みに一般化することで、依存し合う複数の要素を持つデータに対して適用可能となる。例えば、以下のような例が挙げられる。
機械学習や統計のRKHSにおけるカーネル平均埋め込みを用いる既存の方法は、RKHSにおける確率測度のカーネル平均埋め込みを、上記の例1で記載した共分散を表す測度のRKHMにおけるカーネル平均埋め込みに一般化することで、依存し合う複数の要素を持つデータに対して適用可能となる。例えば、以下のような例が挙げられる。
・参考文献1「A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Scholkopf, and A. Smola, A kernel two-sample test, Journal of Machine Learning Research, 13(1):723-773, 2012.」に記載されているtwo-sample testを一般化することで、依存し合う複数の要素を持つデータ同士の比較が可能となる。
・参考文献2「W. Jitkrittum, P. Sangkloy, M. W. Gondal, A. Raj, J. Hays, and B. Scholkopf, Kernel mean matching for content addressability of GANs, In Proceedings of the 36th International Conference on Machine Learning, volume 97, pages 3140-3151, 2019.」に記載されている生成モデルに対するkernel mean matchingを一般化することで、依存し合う複数の要素の共分散の情報を保ったデータを生成できる。
・参考文献3「H. Li, S. J. Pan, S. Wang, and A. C. Kot, Heterogeneous domain adaptation via nonlinear matrix factorization, IEEE Transactions on Neural Networks and Learning Systems, 31:984-996, 2019.」に記載されているMMDを用いたdomain adaptationを一般化することで、ソースドメインとターゲットドメインのデータが依存し合う複数の要素を持つ場合に、その共分散の情報を保って学習を行うことができる。
また、上記の例2で記載した量子の状態を表す測度に対するカーネル平均埋め込みの内積を用いて、量子の状態に対する機械学習や統計の手法を用いた解析が可能となる。
<解析装置10のハードウェア構成>
次に、本実施形態に係る解析装置10のハードウェア構成について、図1を参照しながら説明する。図1は、本実施形態に係る解析装置10のハードウェア構成の一例を示す図である。
次に、本実施形態に係る解析装置10のハードウェア構成について、図1を参照しながら説明する。図1は、本実施形態に係る解析装置10のハードウェア構成の一例を示す図である。
図1に示すように、本実施形態に係る解析装置10は一般的なコンピュータ又はコンピュータシステムで実現され、ハードウェアとして、入力装置11と、表示装置12と、外部I/F13と、通信I/F14と、プロセッサ15と、メモリ装置16とを有する。これらの各ハードウェアは、それぞれがバス17を介して通信可能に接続されている。
入力装置11は、例えば、キーボードやマウス、タッチパネル等である。表示装置12は、例えば、ディスプレイ等である。なお、解析装置10は、入力装置11及び表示装置12のうちの少なくとも一方を有していなくてもよい。
外部I/F13は、外部装置とのインタフェースである。外部装置には、記録媒体13a等がある。解析装置10は、外部I/F13を介して、記録媒体13aの読み取りや書き込み等を行うことができる。なお、記録媒体13aには、例えば、CD(Compact Disc)、DVD(Digital Versatile Disk)、SDメモリカード(Secure Digital memory card)、USB(Universal Serial Bus)メモリカード等がある。
通信I/F14は、解析装置10を通信ネットワークに接続するためのインタフェースである。プロセッサ15は、例えば、CPU(Central Processing Unit)やGPU(Graphics Processing Unit)等の各種演算装置である。メモリ装置16は、例えば、HDD(Hard Disk Drive)やSSD(Solid State Drive)、RAM(Random Access Memory)、ROM(Read Only Memory)、フラッシュメモリ等の各種記憶装置である。
本実施形態に係る解析装置10は、図1に示すハードウェア構成を有することにより、後述するデータ解析処理を実現することができる。なお、図1に示すハードウェア構成は一例であって、解析装置10は、他のハードウェア構成を有していてもよい。例えば、解析装置10は、複数のプロセッサ15を有していてもよいし、複数のメモリ装置16を有していてもよい。
<解析装置10の機能構成>
次に、本実施形態に係る解析装置10の機能構成について、図2を参照しながら説明する。図2は、本実施形態に係る解析装置10の機能構成の一例を示す図である。
次に、本実施形態に係る解析装置10の機能構成について、図2を参照しながら説明する。図2は、本実施形態に係る解析装置10の機能構成の一例を示す図である。
図2に示すように、本実施形態に係る解析装置10は、機能部として、取得部101と、解析部102と、記憶部103とを有する。取得部101及び解析部102は、例えば、解析装置10にインストールされた1以上のプログラムがプロセッサ15に実行させる処理により実現される。また、記憶部103は、例えば、メモリ装置16を用いて実現可能である。なお、記憶部103は、例えば、解析装置10と通信ネットワークを介して接続される記憶装置(例えば、データベースサーバ等)により実現されていてもよい。
記憶部103には、解析対象のデータ(例えば、解析対象となるXの元及びそのA値測度、上記の例2に対して適用する場合は更に量子の状態を表す線形作用素ρ等)が記憶される。
取得部101は、解析対象のデータを記憶部103から取得する。解析部102は、取得部101によって取得されたデータの解析(つまり、例えば、内積・ノルムの計算や、その計算結果を用いた可視化・異常検知等)を行う。
<データ解析処理>
次に、本実施形態に係る解析装置10が実行するデータ解析処理の流れについて、図3を参照しながら説明する。図3は、本実施形態に係るデータ解析処理の一例を示すフローチャートである。
次に、本実施形態に係る解析装置10が実行するデータ解析処理の流れについて、図3を参照しながら説明する。図3は、本実施形態に係るデータ解析処理の一例を示すフローチャートである。
まず、取得部101は、解析対象のデータ(つまり、解析対象となるXの元及びそのA値測度、上記の例2に対して適用する場合は更に量子の状態を表す線形作用素ρ等)を記憶部103から取得する(ステップS101)。
そして、解析部102は、上記のステップS101で取得されたデータの解析を行う(ステップS102)。なお、データの解析としては、上記の「2. RKHMを用いたカーネル平均埋め込みの応用」に記載した内積・ノルムの計算やその計算結果を用いた可視化・異常検知、データ同士の比較、データの生成、学習等が挙げられる。なお、内積の計算方法の具体例は、ランダム性を持つ複数のデータ間の共分散を表す測度である場合は上記の式(1)、量子の状態を表す測度である場合は上記の式(2)に示す通りである。
以上により、本実施形態に係る解析装置10は、複数のランダム性を持つデータ解析(特に、複数のランダムなデータが相互作用し合っている場合のデータや量子状態を表すデータの可視化、異常検知等)を行うことができる。
<実験>
最後に、上記の「2.1 A値測度の間の距離」に記載した例1及び例2に対して、本実施形態に係る解析装置10を適用した場合の実験結果について説明する。
最後に、上記の「2.1 A値測度の間の距離」に記載した例1及び例2に対して、本実施形態に係る解析装置10を適用した場合の実験結果について説明する。
1. ランダム性を持つ複数のデータ間の共分散を表す測度
X=R、Ω=R5とし、Ω上の、Xに値を持つ以下の式(4)~(6)のような確率変数からデータを作成した。
X=R、Ω=R5とし、Ω上の、Xに値を持つ以下の式(4)~(6)のような確率変数からデータを作成した。
(既存手法との比較)
上記の式(4)で定義される[X1,X2,X3]に従う独立なデータと、上記の式(5)で定義される[Y1,Y2,Y3]に従う独立なデータとをそれぞれ用意し、上記の参考文献1に記載されているtwo-sample testを行った。なお、two-sample testは2種類のサンプルが同じ確率分布に従うかどうかを判定するテストである。
上記の式(4)で定義される[X1,X2,X3]に従う独立なデータと、上記の式(5)で定義される[Y1,Y2,Y3]に従う独立なデータとをそれぞれ用意し、上記の参考文献1に記載されているtwo-sample testを行った。なお、two-sample testは2種類のサンプルが同じ確率分布に従うかどうかを判定するテストである。
本実施形態に係る解析装置10により各データ間の距離を測り(つまり、|Φ(μX)-Φ(μY)|kにより測り)、two-sample testを行ったもの(提案手法)と、既存の距離を測り、two-sample testを行ったもの(従来手法)とを比較した。従来手法としては、参考文献1に記載されているRKHSと、参考文献4「B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Scholkopf, and G. R. G. Lanckriet, On the empirical estimation of integral probability metrics. Electronic Journal of Statistics, 6:1550-1599, 2012.」に記載されているKantrovich及びDadleyとを採用した。また、以下のCase1及びCase2のそれぞれの場合で、提案手法及び従来手法の各手法において異なるデータで50回テストを行い、2種類のサンプルが同じ分布に従うと判定された率を計算した。その結果を以下の表1に示す。
・Case1:[X1,X2,X3]に従う独立なデータ10個と[X1,X2,X3]に従う独立なデータ10個
・Case2:[X1,X2,X3]に従う独立なデータ10個と[Y1,Y2,Y3]に従う独立なデータ10個
・Case2:[X1,X2,X3]に従う独立なデータ10個と[Y1,Y2,Y3]に従う独立なデータ10個
2. 量子の状態を表す測度
上記の例2において、m=2、s=4とする。また、
上記の例2において、m=2、s=4とする。また、
このとき、ρ1に関する50個の各サンプルρ1,i(ただし、i=1,・・・,50)について上記の式(3)に示す誤差(再構成誤差)を最小にする第1主軸p1を求め、ρ1に関する50個の各サンプルとρ2に関する50個の各サンプルρj,i(ただし、j=1,2、i=1,・・・,50)それぞれに関してCm×m値の再構成誤差
図5に示すように、ρ1に関するサンプルに比べてρ2に関するサンプルの異常度は高くなっていることから、正常状態であるρ1に対してρ2が乖離している(つまり、異常状態である)ということが精度良く表現できているといえる。
本発明は、具体的に開示された上記の実施形態に限定されるものではなく、特許請求の範囲の記載から逸脱することなく、種々の変形や変更、既知の技術等との組み合わせが可能である。
本願は、日本国に2020年7月16日に出願された基礎出願2020-122352号に基づくものであり、その全内容はここに参照をもって援用される。
10 解析装置
11 入力装置
12 表示装置
13 外部I/F
13a 記録媒体
14 通信I/F
15 プロセッサ
16 メモリ装置
17 バス
101 取得部
102 解析部
103 記憶部
11 入力装置
12 表示装置
13 外部I/F
13a 記録媒体
14 通信I/F
15 プロセッサ
16 メモリ装置
17 バス
101 取得部
102 解析部
103 記憶部
Claims (6)
- ランダム性を持つ複数のデータのデータ集合を取得する取得部と、
前記データ集合上の確率測度μ及びνであって、フォン・ノイマン環に値を持つ確率測度μ及びνを、カーネル平均埋め込みを拡張した写像ΦによってRKHM上にそれぞれ写像したΦ(μ)及びΦ(ν)の内積又はノルムを、前記確率測度μ及びνの内積又はノルムとして計算する解析部と、
を有することを特徴とする解析装置。 - 前記確率測度はランダム性を持つ複数のデータ間の共分散を表す測度を各成分とする行列、前記フォン・ノイマン環はm×mの複素数値行列全体の集合であり、
前記解析部は、
前記データ集合上に値を持つm個の確率変数をそれぞれX1,・・・,Xm及びY1,・・・,Ym、XiとXjの共分散を表す測度を(i,j)成分とする確率測度をμ=μX、YiとYjの共分散を表す測度を(i,j)成分とする確率測度をν=μYとして、前記確率変数X1,・・・,Xmから得られたデータと前記確率変数Y1,・・・,Ymから得られたデータとを用いて、Φ(μX)及びΦ(μY)の内積を、m×mの複素数値行列を値に持つ正定値カーネルにより近似計算する、ことを特徴とする請求項1に記載の解析装置。 - 前記確率測度は量子力学において量子の状態を表す測度、前記フォン・ノイマン環はm×mの複素数値行列全体の集合であり、
前記解析部は、
前記量子の観測を表す前記フォン・ノイマン環上の測度をμ'、前記量子の状態をρ1及びρ2、前記確率測度をμ=ρ1μ'、ν=ρ2μ'として、前記データ集合に含まれるデータを用いて、Φ(ρ1μ')及びΦ(ρ2μ')の内積を、m×mの複素数値行列を値に持つ正定値カーネルにより計算する、ことを特徴とする請求項1に記載の解析装置。 - 前記解析部は、
前記内積又はノルムの計算結果を用いて、前記データ集合の次元削減、前記確率測度の可視化、又は前記確率測度に対する異常検知を行う、ことを特徴とする請求項1又は2に記載の解析装置。 - ランダム性を持つ複数のデータのデータ集合を取得する取得手順と、
前記データ集合上の確率測度μ及びνであって、フォン・ノイマン環に値を持つ確率測度μ及びνを、カーネル平均埋め込みを拡張した写像ΦによってRKHM上にそれぞれ写像したΦ(μ)及びΦ(ν)の内積又はノルムを、前記確率測度μ及びνの内積又はノルムとして計算する解析手順と、
をコンピュータが実行することを特徴とする解析方法。 - コンピュータを、請求項1乃至4の何れか一項に記載の解析装置として機能させるプログラム。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/015,391 US20230237124A1 (en) | 2020-07-16 | 2020-07-14 | Analysis apparatus, analysis method and program |
| JP2022536433A JP7396601B2 (ja) | 2020-07-16 | 2021-07-14 | 解析装置、解析方法及びプログラム |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2020-122352 | 2020-07-16 | ||
| JP2020122352 | 2020-07-16 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022014657A1 true WO2022014657A1 (ja) | 2022-01-20 |
Family
ID=79555649
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2021/026531 Ceased WO2022014657A1 (ja) | 2020-07-16 | 2021-07-14 | 解析装置、解析方法及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20230237124A1 (ja) |
| JP (1) | JP7396601B2 (ja) |
| WO (1) | WO2022014657A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117196414B (zh) * | 2023-11-06 | 2024-04-05 | 南通联润金属制品有限公司 | 一种金属加工质量控制系统 |
-
2020
- 2020-07-14 US US18/015,391 patent/US20230237124A1/en active Pending
-
2021
- 2021-07-14 WO PCT/JP2021/026531 patent/WO2022014657A1/ja not_active Ceased
- 2021-07-14 JP JP2022536433A patent/JP7396601B2/ja active Active
Non-Patent Citations (2)
| Title |
|---|
| SRIPERUMBUDUR B, GRETTON A, FUKUMIZU K, LANCKRIET G, SCHÖLKOPF B: "Injective Hilbert Space Embeddings of Probability Measures", PROCEEDINGS OF THE 21ST ANNUAL CONFERENCE ON LEARNING THEORY, OMNIPRESS, 1 July 2008 (2008-07-01), XP055898622 * |
| YUKA HASHIMOTO; ISAO ISHIKAWA; MASAHIRO IKEDA; FUYUTA KOMURA; TAKESHI KATSURA; YOSHINOBU KAWAHARA: "Analysis via Orthonormal Systems in Reproducing Kernel Hilbert C^*-Modules and Applications", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 2 March 2020 (2020-03-02), 201 Olin Library Cornell University Ithaca, NY 14853 , XP081612255 * |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2022014657A1 (ja) | 2022-01-20 |
| JP7396601B2 (ja) | 2023-12-12 |
| US20230237124A1 (en) | 2023-07-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10726325B2 (en) | Facilitating machine-learning and data analysis by computing user-session representation vectors | |
| Yin et al. | Sequential sufficient dimension reduction for large p, small n problems | |
| Shao et al. | Interactive regression lens for exploring scatter plots | |
| US20170091302A1 (en) | Method and apparatus for representing multidimensional data | |
| US11270221B2 (en) | Unsupervised clustering in quantum feature spaces using quantum similarity matrices | |
| US20220239673A1 (en) | System and method for differentiating between human and non-human access to computing resources | |
| Pathak et al. | An assessment of the missing data imputation techniques for covid-19 data | |
| US20230418700A1 (en) | Real time detection of metric baseline behavior change | |
| Tao et al. | Predicting time series by data-driven spatiotemporal information transformation | |
| Caithness et al. | Can functional characteristics usefully define the cloud computing landscape and is the current reference model correct? | |
| Saleem et al. | Direct feature evaluation in black-box optimization using problem transformations | |
| WO2022014657A1 (ja) | 解析装置、解析方法及びプログラム | |
| Li et al. | Prediction of bidirectional milling forces based on spindle current signals by using deep learning algorithms | |
| Li et al. | A Slicing‐Free Perspective to Sufficient Dimension Reduction: Selective Review and Recent Developments | |
| CN114726581B (zh) | 一种异常检测方法、装置、电子设备及存储介质 | |
| Buzzicotti et al. | Inferring turbulent environments via machine learning | |
| WO2023015142A1 (en) | Principal component analysis | |
| Sakata et al. | Enhancing spectral analysis in nonlinear dynamics with pseudoeigenfunctions from continuous spectra | |
| Senapati et al. | Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing | |
| Liu et al. | Dimension estimation using weighted correlation dimension method | |
| Marques et al. | Gaussian process for radiance functions on the sphere | |
| Yan et al. | Scale-invariant Mexican Hat wavelet descriptor for non-rigid shape similarity measurement | |
| GB2603607A (en) | Interative state detection for molecular dynamics data | |
| Rodriguez Dominguez | Order-Constrained Spectral Causality in Multivariate Time Series | |
| Kittiwachana et al. | Self-organizing map quality control index |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21843295 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2022536433 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21843295 Country of ref document: EP Kind code of ref document: A1 |



























