WO2006093121A1 - 数値計算処理装置 - Google Patents
数値計算処理装置 Download PDFInfo
- Publication number
- WO2006093121A1 WO2006093121A1 PCT/JP2006/303694 JP2006303694W WO2006093121A1 WO 2006093121 A1 WO2006093121 A1 WO 2006093121A1 JP 2006303694 W JP2006303694 W JP 2006303694W WO 2006093121 A1 WO2006093121 A1 WO 2006093121A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- cell
- particle
- calculation
- storage means
- particles
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C10/00—Computational theoretical chemistry, i.e. ICT specially adapted for theoretical aspects of quantum chemistry, molecular mechanics, molecular dynamics or the like
Definitions
- the present invention relates to a numerical calculation processing device.
- the present invention relates to a dedicated calculation processing apparatus that performs numerical operations in calculation technology simulation calculation.
- Patent Document 1 discloses a processing device for calculating an interaction force for a multi-gravity system and a multi-electric force system.
- Patent Document 1 discloses a dedicated computer that calculates gravity and electric force acting between particles and their potentials. This dedicated computer has a feature of accelerating the calculation by pipelined internal calculation processing and continuously giving coordinates, mass, and electric charge (for example, claims of Patent Document 1) (See 1 and 2).
- the force used for the calculation is a force with a short reach (for example, van der Waals force in the case of molecular dynamics simulation)
- a technique for mounting this technique on a dedicated computer as described in Patent Document 1 has already been disclosed in Non-Patent Document 1.
- the space in which the particles exist is divided so that a box called a cell is a unit, and the data about the particles in each cell is stored in a continuous area of the storage device.
- Patent Document 1 Japanese Patent Laid-Open No. 5-89157
- the present invention at least either increases the calculation efficiency when the force reach distance is short, or increases the calculation efficiency even if a process not including the force from a specific particle is included in the calculation.
- the problem is to solve this.
- the present invention is implemented in hardware so that the amount of data to be transferred can be reduced in a computing device that calculates classical mechanical force and potential energy continuously by a pipeline method. That is, in the present invention, the calculation processing device calculates a force or potential energy acting between a plurality of particles arranged in a predetermined space by a cell index method using a plurality of cells obtained by dividing the space.
- a particle storage means for storing coordinate data in space and physical quantity data necessary for calculating force or potential energy in a continuous address storage area for particles in the same cell;
- a plurality of cell index storage means for storing the first address in the particle storage means and the number of particles in the cell or the final address in correspondence with the cell number identifying the cell, for each particle belonging to each cell;
- Coordinate data for each particle included in a central cell including a particle whose force or potential energy is to be calculated and at least one peripheral cell at a predetermined relative position with respect to the central cell.
- physical quantity data are read from the particle storage means, and the force or potential energy acting on the particles in the central cell is calculated and calculated.
- the coordinate data of each peripheral cell is generated by adding the relative coordinate data of the peripheral cells stored at successive addresses of the cell relative position storage means to the coordinate data of the peripheral cell, and from the coordinates of each peripheral cell, the peripheral cell
- the peripheral cell A cell number generating means, and the pipeline calculating means is a head address read from the cell index storing means in correspondence with the cell number from the cell number generating means in the storage area of the particle storing means. Read particle coordinate data and physical quantity data from the storage area addressed by the number of particles or the final address A shall perform the numerical calculations, the computing device is provided.
- the predetermined space in the present invention is a space in which particles to be simulated are arranged, and is a space defined as a three-dimensional region having a predetermined size inside the computer.
- a particle is a target of gravity many-body simulation (such as a star with a mass that causes gravity) or a molecular dynamics simulation, and causes Coulomb or van der Waals forces.
- the atoms and ions having the following charges are generally referred to.
- a plurality of these particles are used in the simulation, and each particle is associated with a coordinate in space and a physical quantity (mass, charge, etc.) according to the purpose of the simulation.
- the cell index method assigns particles existing in a space to multiple cells, and calculates the effect on the force applied to a particle by calculating the effect of the force only on the cells within the reach distance. It is a technique to raise.
- the particle storage means is an arbitrary storage means for storing particle coordinate data and physical quantity data.
- the pipeline calculation means accesses only the minimum frequency necessary for the calculation.
- the nopline calculation means executes the calculation of the classical mechanical force or potential energy between two particles at high speed, and sequentially changes one of the particles to be calculated.
- Line processing is performed to integrate the classical mechanical force and potential energy among many particles.
- This pipeline calculator includes floating point adders, multipliers, and dividers to perform floating point numerical calculations. Further, a plurality of the nopline calculation means are provided, and are configured to increase the throughput of the numerical operation by parallel (parallel) processing.
- the present invention further comprises boundary processing means for determining whether the coordinate data of each neighboring cell is in-space coordinate data or out-of-space coordinate data and outputting a boundary determination signal, It is more preferable that the pipeline calculation means performs different arithmetic processing according to the boundary determination signal.
- the boundary processing means of the present invention is used to appropriately process a point where a peripheral cell may be out of space when the calculation process proceeds by receiving the coordinates of the center cell. Thereby, the boundary processing can be executed by the calculation processing device of the present invention.
- the boundary processing means when determining that the coordinate data of the surrounding cells is outside the space, the size data that gives the size of the space is used.
- the coordinates of surrounding cells translated in space can be calculated by adding or subtracting to the coordinate data.
- each of the nopline calculation means has a plurality of accumulated data storage means, and the calculation result of force or potential is used as the boundary determination signal.
- the accumulated data storage means is, for example, an appropriate register, and separately accumulates whether the force or potential energy value is from a surrounding cell in the space or whether it is from a surrounding cell outside the space. be able to.
- the power and potential energy of the calculation result can be integrated into 8 different integrated data storage means so that the accumulated data can be identified as in-space or out-of-space according to the X, Y, Z coordinate axes. .
- the calculation processing apparatus of the present invention it is possible to perform a process of removing (excluding) predetermined particles from the calculation of force and potential energy.
- the calculation processing device of the present invention excludes the excluded particle address storage means for holding the address in the particle storage means of the particles to be excluded from the accumulation of force or potential energy, and the particles to be excluded. And a plurality of excluded particle list means including an excluded particle flag storage means for holding a flag for designating whether or not to force for each pipeline calculation means.
- storage means power Compare the address of the particle storage means specified by the read head address and the number of particles or the final address with the address stored in the excluded particle address storage means, For particles with the same address, output an integration control signal in combination with the flag stored in the exclusion flag storage. It is preferable that the pipeline calculation means switches between performing and not performing integration according to the integration control signal. As a result, if there is a particle to be excluded, the process of removing the influence of the desired particle force in the calculation of force and potential energy by performing the operation of excluding the accumulated force can be performed while executing the accumulation process. The calculation efficiency of the calculation processing device due to the excluded particles can be prevented from being lowered.
- the calculation processing device of the present invention only transmits the coordinates of the central cell, which is necessary to transmit the entire list of conventional cells from the outside each time the central cell moves as the calculation proceeds.
- the calculation processing efficiency can be improved.
- FIG. 1 is an explanatory diagram showing an overall configuration of a calculation system according to an embodiment of this invention.
- FIG. 2 is a block configuration diagram showing details of a dedicated computer unit in the embodiment of the present invention.
- FIG. 3 is a block configuration diagram showing a detailed configuration of a dedicated processing unit in the embodiment of the present invention.
- FIG. 4 is an explanatory diagram showing an example of a space in which a simulation system is set on a computer and an arrangement of cells in the space in the cell index method used in the embodiment of the present invention.
- FIG. 5 is a block configuration diagram showing a detailed configuration of a boundary condition processing unit in the embodiment of the present invention.
- FIG. 1 is an explanatory diagram conceptually showing the overall configuration of the calculation system 1 of the present embodiment.
- This computer system 1 includes a host computer unit 20 and a dedicated computer unit 10. .
- the host computer unit 20 sends the coordinate data of the particles necessary for calculating the force acting between the particles and the physical quantity data such as charge and mass to the dedicated computer unit 10.
- the dedicated computer unit 10 Based on this data, the dedicated computer unit 10 performs specific calculations such as classical mechanical force calculation between particles and potential energy at high speed, and applies power and potential energy to the host computer unit 20. Returns the calculation result. All other calculations necessary for the simulation are executed by the host computer unit 20.
- the dedicated computer unit 10 functions as an acceleration device for the host computer unit 20 that exclusively executes processing specialized in classical mechanical calculation among the simulation calculations targeted by the calculation system 1.
- the calculation system 1 of the present embodiment greatly reduces the amount of communication between the host computer unit 20 and the dedicated computer unit 10 that has been required in the past, thereby reducing the time required for communication. Improve the calculation efficiency.
- FIG. 2 shows a detailed block configuration when the dedicated computer unit 10 is implemented as a computer.
- the host computer unit 20 includes a central processing unit (host CPU) 220, and storage means (not shown).
- Various programs and data for operating the dedicated computer unit 10 to execute the target simulation are stored in the storage means.
- the dedicated computer unit 10 is generally composed of dedicated computer boards 10 to 10. Exclusive
- the computer boards 10 to 10 have high-speed setting data and data with the host computer unit 20.
- An interface LSI 320 for exchanging calculation result data is provided.
- a dedicated processing unit 100 indicated as MDG3-LSI is configured to be able to communicate with the interface LSI 320.
- the dedicated processing unit 100 is an integrated circuit in which the present invention is implemented.
- the dedicated processing unit 100 is mounted with all the elements necessary for the processing of the present invention, such as a particle memory (particle storage unit) 150 and a pipeline unit (pipeline calculation unit) 160 described later.
- the dedicated computer unit 10 is composed of one dedicated computer board 10 and its
- the dedicated computer board 10 has 12 dedicated processing unit 100 LSIs. Dedicated for each
- the processing unit 100 is connected in a ring shape and connected to the first and last power S interface LSIs 320.
- the host computer section 20 has an interface
- the host computer unit 280 is communicably connected to the host CPU 220 through an appropriate bus, whereby the host computer unit 20 is connected to the interface LSI 320 of the dedicated computer unit 10 by the high-speed interface line 300.
- the interface LSI 320 delivers the data to the dedicated processing unit 100, and the dedicated processing unit 100 The data is stored in an internal memory or register.
- the interface LSI 320 reads the calculation result data from each dedicated processing unit 100 and transmits it to the host computer unit 20.
- FIG. 3 is a block configuration diagram showing a detailed configuration of the dedicated processing unit 100.
- the dedicated processing unit 100 used in the present embodiment uses a main memory (particle memory, particle storage means) 150 that stores information on each particle.
- This particle memory 150 has a total of 160 bits of 120 bits for each particle (40 bits for each of X, y, and z) and 40 bits of data necessary for calculation of force or potential such as charge 'particle species' mass. Stored as one word.
- the particle memory 150 has a capacity of 32,768 words, and the number of particles that can be handled simultaneously is 32,768. To address the particles, a 15-bit wide address can be used.
- FIG. 4 is an explanatory diagram showing an example of a space in which a simulation system is set on a computer and an arrangement of cells in the space when the cell index method is used.
- a space 70 having a predetermined size is set on a computer, and particles to be simulated are placed in this space as a simulation system.
- each position in the space 70 can be designated by the coordinate value of (X, ⁇ , Z), and typically has an extension to the boundary of the boundary surfaces B to B. Where B and B are X
- 3 4 is the boundary determined by the minimum and maximum values of the Y coordinate
- B and B are the boundaries determined by the minimum and maximum values of the Z coordinate.
- a periodic boundary condition For example, a periodic boundary condition
- space 70 internal force S is further divided into cells and the calculation proceeds.
- Each of these cells It is a three-dimensional space area divided into an appropriate size suitable for calculations involving short forces. From the cell (center cell c) to be focused on and the surrounding cells set around it
- the cell Since the cell has an area within the reach of the force from the apex of the center cell, the force that normally extends to the cell adjacent to the nearest neighbor cell (second nearest neighbor cell), and further subdivision can be performed. it can.
- the cell index method a large number of particles (not shown) scattered in the space 70 are set, the cells in the space 70 are sequentially set as the central cell, and the range that can be reached by the force with a short reach is included.
- a peripheral cell is set, and force and potential energy are calculated in the local space including the central cell and the peripheral cell.
- the calculation under the setting of a certain center cell is completed for one step, the same calculation is performed by moving the center cell and surrounding cells.
- the next calculation is executed in the same way for one step, and a multi-step calculation is executed.
- the center cell can be specified by coordinates (XO, YO, ZO). These coordinates are expressed as integers each 12 bits wide.
- a list of relative positions ( ⁇ , ⁇ , ⁇ ) of the peripheral cells to be calculated with respect to the central cell is recorded in the relative cell position memory 104.
- This memory has a capacity of 1024 words, with each word being 12 bits wide in the dedicated processing unit 100 of the present embodiment. In each word, the assignment of ⁇ , ⁇ , and ⁇ can be changed within a 12-bit width, usually depending on the purpose of the force calculation that assigns 4 bits to ⁇ , ⁇ , and ⁇ .
- the coordinates of the center cell are the forces input from the host computer unit 20 to the dedicated processing unit 100 each time the center cell is changed.
- the relative position ( ⁇ , ⁇ , ⁇ ) is written from the host computer unit 20 at an appropriate time, such as at the start of calculation, and is retained even if the coordinates (XO, YO, Z0) of the center cell are changed.
- the dedicated processing unit 100 includes a calculation cell counter 102.
- the calculation cell counter 102 generates an address of the relative cell position memory 104 to be used for calculation. For example, calculation When the total number of peripheral cells to be used for the calculation is 6 as in the example shown in FIG. 4, the calculation cell counter 102 generates 6 addresses, which are stored in the relative cell position memory 104 according to the addresses. Six relative positions are selected from the list of relative positions ( ⁇ , ⁇ , ⁇ ⁇ ).
- the calculation cell counter 102 has a width of 10 bits, and an initial value can be set. Also, there is a separate end value register (not shown). When the end value is reached, the calculation ends.
- the dedicated processing unit 100 is provided with a cell number generation unit 130.
- the cell number generation unit 130 includes a first part 130A and a second part 130B.
- the first part 130A of the cell number generation unit 130 is provided with adders 108a to 108c.
- the relative position ( ⁇ , ⁇ , ⁇ ⁇ ) of the cell output from the relative cell position memory 104 includes the coordinate (XO, Y0, Z0) position of the central cell transmitted from the host computer unit 20. Addition, the absolute position (X, ⁇ , Z) of each cell is calculated. The calculation result is 13 bits with a sign.
- FIG. 5 shows a detailed configuration of the boundary condition processing unit 110.
- the X data input to the boundary condition processing unit 110 is compared by the comparators 112a and 112b against the value of Xmin that gives the boundary surface B of space 70 and the value of Xmax that gives the boundary surface B.
- the logical sum of the comparison result bits of the comparators 112a and 112b is output as Xover to each of the pipeline sections 160 described later. Note that the output of this Xover is the boundary determination signal of the present invention.
- the logical sum is logically ANDed with an input Cyclic that controls whether to set the periodic boundary mode, and is output to the cell translation unit 152 as a translation flag.
- the boundary condition processing unit 110 adds Xsize to X when X is less than Xmin, and subtracts Xsize from X when X is greater than or equal to Xmax.
- Xsize represents the size of the space 70 in the X direction, and is calculated by Xmax-Xmin.
- the adder / subtractor 114 is configured to receive the outputs from the comparators 112a and 112b and change the operation. The output of the adder / subtractor 114 is input to the multiplexer 116 together with the input X.
- the multiplexer 116 is controlled by the parallel movement flag. When the parallel movement flag is 1, the multiplexer 116 selects the output of the adder / subtractor 114.
- the multiplexer 116 selects the input X, and the parallel movement processing is performed. Is output as the absolute position XT where is performed.
- the output of comparator 112b In combination with the movement flag, it becomes an output indicating the direction of translation, and is output to the cell translation unit 152.
- the flag (Ignore, not shown) that indicates that the case that exceeds the boundary is set to 1, X is less than Xmin or X is greater than or equal to Xmax. Remove the cell from the calculation. The above processing is also performed for ⁇ and ⁇ .
- the input Cyclic and Ignore flags can be set independently for X, ⁇ , and ⁇ ⁇ so that the dedicated processing unit 100 can be operated flexibly according to the calculation purpose.
- Xmin, Xmax, and Xsize are all 12-bit values
- Cyclic and Ignore are 1-bit inputs.
- Xmin, Xmax, Xsize, Cyclic, and Ignore are the values that are set as the initial settings of the calculations, so there is no need to change them by moving the center cell.
- a pipeline unit 160 is provided.
- the pipeline section 160 is provided with pipelines PL to PL.
- Each pipeline PL-PL has
- Eight registers F (0) to F (7) are provided so that the force can be divided and accumulated. Therefore, whether the force is divided and integrated is also switched by a flag (not shown) that can be provided in the pipelines PL to PL for each XYZ component.
- each output Xover, Yover, Zover (indicating a value of 1 or
- the cell translation unit 152 translates the particle coordinates (X, y, z) in accordance with the application of the periodic boundary condition. Refer to the translation flag for each x, y, z component and the flag for the direction flag to translate (6 bits for 2 bits and 3 directions) output from the boundary condition processing unit 110, and the direction flag to translate Depending on the contents of, the physical size of the basic cell is added to or subtracted from x, y, z respectively.
- Xd and Yd are numbers corresponding to the maximum number of cells in the X and Y directions, respectively, and are given by separate registers.
- the width is 12 bits.
- This cell number CID is input to the cell index memory 142.
- the cell index memory 142 has a 15-bit start address indicating from which address in the particle memory 150 the particles in the cell are stored at the address corresponding to the cell number CID, and the number of particles in the cell (15 bits). ) Is stored. Since the cell number CID is 12 bits, the cell index memory 142 is composed of 4096 words and a 30-bit width memory.
- the start address is set as the initial value of the particle memory address counter 144, and reading of the particle memory 150 is started.
- the number of particles in the cell is set as the initial value of the particle number counter 146.
- the width of particle memory address counter 144 and particle number counter 146 is 15 bits. This process is continued until the particle number counter 146 becomes zero.
- a signal for incrementing the calculation cell counter 102 is output in order to shift to the processing of the next cell. In the actual processing, the processing of the next cell is advanced as much as possible while reading the particle memory 150 for a certain cell. Also, when the number of particles in the cell to be calculated is very small, the reading of the particle memory 150 ends before the next cell is ready. In this case, the calculation is suspended and the next particle memory 150 is read out.
- the exclusive processing unit 100 of the present embodiment includes an excluded particle processing unit 170.
- the excluded particle processing unit 170 has 16 excluded particle lists EPL to EPL.
- Each of the excluded particle lists EPL to EPL is the particle memory of the particles to be excluded 150
- the dedicated processing unit 100 of this embodiment has 20 pipelines PL to PL,
- each pipeline is logically identified as two pipelines by time division operation. Therefore, the exclusion particle flag register 174 is composed of a 40-bit width register. Excretion The particle removal address register 172 is a 15-bit register. Each of the excluded particle lists EPL to EPL is a particle memory output from the particle memory address counter 144 1
- the output of the EPL output is further input to the logical sum circuit 178 to calculate the logical sum for each bit.
- the output from the OR circuit 178 becomes an integration control signal of the present invention.
- the corresponding pipeline operates so that the calculated force is not accumulated in the registers F (0) to F (8) when the flag power of each bit is satisfied, and when the flag of each bit is 0
- the calculated force is accumulated in the registers F (0) to F (8).
Landscapes
- Engineering & Computer Science (AREA)
- Computing Systems (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
セルインデックス法の計算を効率よく行う専用計算機を実現する。 座標データと物理量データとを格納する粒子記憶手段150と、その先頭アドレスと粒子数または最終アドレスをセル毎に格納するセルインデックス記憶手段142と、各粒子に対して座標データおよび物理量データを粒子記憶手段から読み込んで、中心セルの粒子の力またはポテンシャルエネルギーを並行して計算する複数のパイプライン計算手段160と、周辺セルの相対座標データを格納するセル相対位置記憶手段104と、周辺セルのセル番号を生成するセル番号生成手段130A、Bとを備える計算処理装置。パイプライン計算手段は、セル番号生成手段からのセル番号に対応させてセルインデックス記憶手段から読み出される先頭アドレスと粒子数または最終アドレスとによって、粒子記憶手段の記憶領域からデータを読み込んで数値計算を実行する。
Description
明 細 書
数値計算処理装置
技術分野
[0001] 本発明は数値計算処理装置に関する。特に、本発明は、計算技術シミュレーション 計算において数値演算を行なう専用計算処理装置に関する。
背景技術
[0002] 従来より、コンピュータの高速ィ匕が要望されている。特に計算負荷の高い科学技術 シミュレーションなどの分野にぉ 、て、数値計算を高速に処理する演算処理装置が 求められている。
[0003] このような科学技術シミュレーションのうち、分子動力学シミュレーション、重力多体 問題シミュレーションなどの古典粒子シミュレーションは、物質の性質や宇宙の成り立 ちを研究する上で重要な計算手法である。これらのシミュレーションにおいては、粒 子間に働く古典的な力の計算が計算時間の大部分を占める。このため、この力の計 算を高速に実行するための専用計算機が開発されている。
[0004] 特許文献 1には、重力多体系および電気力多体系用相互作用力計算用処理装置 が開示されている。特許文献 1は、粒子間に働く重力および電気力およびそれらのポ テンシャルを計算する専用計算機を開示している。この専用計算機は、内部の計算 処理をパイプラインィ匕してこれに連続して座標 ·質量 ·電荷を与えることにより計算を 加速するという特徴を有している(例えば、特許文献 1の請求項 1および 2を参照)。
[0005] 上記従来の専用計算機においては、計算処理の過程において、多数存在する粒 子それぞれの座標 ·質量 ·電荷を連続して与えて計算が行われて 、る。この従来の 専用計算機では、分子動力学シミュレーションを例とすれば、互いに化学結合によつ て結合して 、る粒子の間も、互いに化学結合によっては結合して 、な 、粒子の間も 同様に力を計算する。化学結合によって結合している粒子間における力を除いて力 の和を正しく計算する場合のように、連続する粒子の座標 ·質量 ·電荷から本来含め るべきではない力を取り除いて計算したい場合には、上記従来の専用計算機におい ては、計算途中で記憶装置上の座標 '質量'電荷を書き換えるか、計算終了後別の
手法でこれを差し引くかのいずれかを行う必要がある。これらの処理は、いずれも計 算時間を要し、従来技術における速度低下の大きな要因となっている。
[0006] また、計算に用いる力が到達距離の短い力(例えば、分子動力学シミュレーション の場合であれば、ファンデルワールス力)である場合には、近くの粒子からの影響の みを計算する「セルインデックス法」と呼ばれる手法によって高 、計算効率が実現さ れることが知られている。そして、この手法を上記特許文献 1に記載されているような 専用計算機上に実装する手法についても既に非特許文献 1に公開されている。しか し、セルインデックス法の従来の実装方法では、粒子の存在する空間をセルと呼ばれ る箱を単位となるように分割し、各セル内の粒子に関するデータを記憶装置の連続し た領域に記憶させる。これを呼び出すためには、各セル内の粒子をおさめた先頭ァ ドレスとセル内の粒子数との組について、全ての計算対象セル分のリストを作成し、こ のリストに基づいて粒子を読み出す必要がある。この実装手法では、計算の対象とな るセル(中心セル)を変更するたびにこのリストを演算装置へと転送する必要があり、 計算処理の効率を低下させて!/、る。
特許文献 1:特開平 5— 89157号公報
特干文献 1 : T. Fukusnige, M. Taiji, J. Makino, T. Ebisuzaki, D. bugimoto, Astrop hysical J. 468 (1996) 51.
発明の開示
発明が解決しょうとする課題
[0007] 本発明は、力の到達距離が短い場合の計算効率を高めること、または、特定の粒 子からの力を計算に含めない処理を行っても計算効率を高めることの少なくともいず れかを解決することを課題とする。
課題を解決するための手段
[0008] 本発明は、古典力学的な力やポテンシャルエネルギーの計算をパイプライン方式 により連続して計算する計算装置において、転送するデータ量を小さくできるように ハードウ アに実装する。すなわち、本発明においては、所定の空間内に配置される 複数の粒子間に作用する力またはポテンシャルエネルギーを、該空間を分割した複 数のセルを用いるセルインデックス法によって計算する計算処理装置であって、粒子
のそれぞれにつ 、て、空間内での座標データと力またはポテンシャルエネルギーの 算出に必要な物理量データとを、同一のセル内にある粒子について連続するァドレ スの記憶領域に格納する粒子記憶手段と、各セルに属する粒子にっ 、ての該粒子 記憶手段における先頭アドレスと該セル内の粒子数または最終アドレスとを、セルを 特定するセル番号に対応させて格納するセルインデックス記憶手段と、複数のセル のうち、力またはポテンシャルエネルギーを算出すべき粒子を含む中心セルと該中 心セルに対して所定の相対位置にある少なくとも一つの周辺セルとに含まれる各粒 子に対して、座標データおよび物理量データを粒子記憶手段から読み込んで、該中 心セルの粒子に作用する力またはポテンシャルエネルギーを数値計算して積算する 少なくとも 1つのパイプライン計算手段と、中心セルに対する周辺セルの相対座標デ ータを連続するアドレスの記憶領域に格納するセル相対位置記憶手段と、中心セル の座標データを受け付けて、中心セルの該座標データに、セル相対位置記憶手段 の連続するアドレスに格納されている周辺セルの相対座標データを加算して各周辺 セルの座標を生成し、各周辺セルの該座標から、該周辺セルのセル番号を生成する セル番号生成手段とを備え、パイプライン計算手段は、粒子記憶手段の記憶領域の うち、セル番号生成手段からのセル番号に対応させてセルインデックス記憶手段から 読み出される先頭アドレスと粒子数または最終アドレスとによってアドレス指定される 記憶領域から、粒子の座標データと物理量データとを読み込んで数値計算を実行す るものである、計算処理装置が提供される。
[0009] 本発明における所定の空間は、シミュレーションされる粒子が配置される空間であり 、コンピュータの内部において所定の大きさの立体領域として定義される空間である
[0010] 粒子とは、重力多体シミュレーションの対象となる質点(重力の原因となる質量を有 する星などの天体)や、分子動力学シミュレーションの対象となり、クーロン力やファン デルワールス力の原因となる電荷を有する原子やイオンなどを一般にいう。この粒子 は、シミュレーションにおいては複数用いられて、個々の粒子には、空間における座 標と、そのシミュレーション目的に応じた物理量 (質量、電荷など)が対応付けされて いる。
[0011] セルインデックス法は、空間に存在する粒子を、複数のセルに帰属させ、ある粒子 に与える力を、その到達距離内にあるセルのみ力 の影響を計算することにより、計 算効率を高める手法である。
[0012] 粒子記憶手段とは、粒子の座標データと物理量データとを格納する任意の記憶手 段である。本発明の計算処理装置においては、パイプライン計算手段によって、計算 に必要な最低限の頻度だけアクセスされる。
[0013] ノ ィプライン計算手段は、シミュレーションにおける計算のうち、 2粒子間の古典力 学的な力またはポテンシャルエネルギーの計算を高速に実行し、計算対象となる粒 子の一方を順次変更しながらパイプライン処理を行って、多数の粒子間での古典力 学的な力やポテンシャルエネルギーを積算する処理を行う。このパイプライン計算手 段は、浮動小数点数値計算を行うために、浮動小数点による加算器、乗算器、除算 器を備えている。また、このノ ィプライン計算手段は複数備えられていて、並行 (並列 )処理によって数値演算のスループットを高めるように構成されて 、る。
[0014] 本発明においては、各周辺セルの座標データが空間内の座標データであるか空 間外の座標データであるかの判定を行って境界判定信号を出力する境界処理手段 をさらに備え、パイプライン計算手段は、該境界判定信号に応じて異なる演算処理を 行うこととするとさらに好適である。本発明の境界処理手段は、中心セルの座標を受 け取って計算処理が進む際に、周辺セルが空間外になる可能性がある点を適切に 処理するために用いられる。これにより、本発明の計算処理装置によって境界処理を 実行することができる。
[0015] 境界処理手段を用いる場合の好適な一例としては、境界処理手段は、周辺セルの 座標データを空間外のものであると判定した場合には、空間の大きさを与える大きさ データを座標データに加算または減算して、空間内に平行移動された周辺セルの座 標を算出するものとすることができる。これにより、シミュレーションで多用される周期 境界条件の処理も本発明の計算処理装置によって実行することができる。
[0016] また、境界処理手段を用いる場合の好適な一例としては、ノ ィプライン計算手段そ れぞれは、複数の積算データ記憶手段を有し、力またはポテンシャルの計算結果を 、境界判定信号に応じて異なる積算データ記憶手段に積算させるものとすることがで
きる。積算データ記憶手段は、例えば適当なレジスタであり、力やポテンシャルエネ ルギ一の値が空間内の周辺セルからのものである力、空間外の周辺セルのものであ るかを別々に積算することができる。例えば、積算されたデータが X, Y, Zの座標軸 別に空間内か空間外かが識別できるように、 8種類の別々の積算データ記憶手段に 計算結果の力やポテンシャルエネルギーを積算することができる。
[0017] 本発明の計算処理装置では、さらに、所定の粒子について、力やポテンシャルエネ ルギ一の計算から取り除く(排除する)処理を行うことができる。これを行うために、本 発明の計算処理装置は、力またはポテンシャルエネルギーの積算から排除すべき粒 子の粒子記憶手段におけるアドレスを保持する排除粒子アドレス記憶手段と、該排 除すべき粒子を排除すべき力否かをパイプライン計算手段毎に指定するフラグを保 持する排除粒子フラグ記憶手段とを含む排除粒子リスト手段をさらに複数含み、該排 除粒子リスト手段は、セル番号生成手段力ゝらのセル番号に対応してセルインデックス 記憶手段力 読み出される先頭アドレスと粒子数または最終アドレスとによって指定 される粒子記憶手段のアドレスを、排除粒子アドレス記憶手段に格納されるアドレスと 比較して、両アドレスが一致する粒子について、排除フラグ記憶手段に格納されるフ ラグと組み合わせて積算制御信号を出力するものであり、パイプライン計算手段は、 該積算制御信号に応じて積算を実行するか実行しないかを切り替えるものとされると 好適である。これにより、排除される粒子がある場合には積算力 除外する動作を行 うことにより、力やポテンシャルエネルギーの計算において所望の粒子力もの影響を 取り除く処理が、積算処理を実行しながら行うことができ、排除粒子による計算処理 装置の計算効率の低下が防止できる。
発明の効果
[0018] 本発明の計算処理装置では、計算の進行に伴って中心セルが移動するたびに従 来セルのリスト全てを外部から送信する必要があったものを、中心セルの座標を送信 するのみで計算できるようになり、計算処理効率が向上する。
また、本発明の計算処理装置で排除粒子リストを含む場合には、排除したい粒子 のアドレスと、パイプラインごとのフラグを送るだけで特定の粒子力ものカゃポテンシ ャルを計算力 取り除くことができるようになり、計算処理効率が向上する。
図面の簡単な説明
[0019] [図 1]本発明の実施の形態の計算システムの全体構成を示す説明図である。
[図 2]本発明の実施の形態における専用計算機部の詳細を示すブロック構成図であ る。
[図 3]本発明の実施の形態における専用処理部の詳細な構成を示すブロック構成図 である。
[図 4]本発明の実施の形態において用いられるセルインデックス法において、計算機 上にシミュレーション系が設定される空間と、その空間内におけるセルの配置の例を 示す説明図である。
[図 5]本発明の実施の形態における境界条件処理部の詳細な構成を示すブロック構 成図である。
符号の説明
[0020] 1 計算システム
10 専用計算機部
10 -10 専用計算機ボード
1 N
70 空間
100専用処理部
102計算セルカウンタ
104相対セル位置メモリ
110境界条件処理部
112a, 112b 比較器
114加減算器
116マルチプレクサ
130A 130B セノレ番号生成咅
132a 乗算器
134加算器
142セノレインデックスメモリ
144粒子メモリ番地カウンタ
146粒子数カウンタ
150粒子メモリ
152セル平行移動ユニット
160パイプライン計算手段
170排除粒子処理部
172排除粒子アドレスレジスタ
174排除粒子フラグレジスタ
176a 一致検出器
176b 論理積回路
178論理和回路
20 ホスト計算機部
220ホスト CPU
280インターフェース咅
300高速インターフェース回線
320インターフェース: LSI
B -B 境界面
1 6
EPL -EPL 排除粒子リスト
1 16
PL -PL パイプライン
1 20
C 中心セル
0
C - C 最隣接セル
1 6
CIDセル番号
発明を実施するための最良の形態
[0021] 以下、図面を参照して、本発明の専用計算機を含む計算システムの実施の形態に ついて説明し、さらに、そのシステムにおいて用いられている専用計算機部 10を構 成する集積回路の構成について説明する。
[0022] 本実施の形態は、分子動力学シミュレーションを行うための計算システムの実施の 形態である。図 1は、本実施の形態の計算システム 1の全体の構成を概念的に示す 説明図である。この計算システム 1は、ホスト計算機部 20と専用計算機部 10からなる
。ホスト計算機部 20は粒子の間に働く力の計算に必要な粒子の座標データや、電荷 、質量などの物理量データを専用計算機部 10に送る。専用計算機部 10はこのデー タに基づ 、て、例えば粒子間の古典力学的な力の計算やポテンシャルエネルギー の計算のなどの特定の計算を高速に行ない、ホスト計算機部 20に力、ポテンシャル エネルギーなどの計算結果を返す。シミュレーションに必要なその他の部分の計算 は全てホスト計算機部 20で実行する。つまり、専用計算機部 10は、計算システム 1が 目的としているシミュレーション計算のうち、古典力学的な計算に特化した処理を専ら 実行するような、ホスト計算機部 20の加速装置として機能する。本実施の形態の計 算システム 1は、従来必要であったホスト計算機部 20と専用計算機部 10の間の通信 量を大幅に削減して通信に要する時間を削減することによって、計算システム 1全体 の計算効率を向上させる。
[0023] この専用計算機部 10を計算機として実装した場合の詳細なブロック構成を図 2に 示す。図 2に示したように、ホスト計算機部 20には、中央演算処理装置 (ホスト CPU) 220が備えられており、図示しない記憶手段が備えられている。専用計算機部 10を 動作させて目的のシミュレーションを実行するための各種プログラムやデータは、そ の記憶手段に格納されている。
[0024] 専用計算機部 10は、一般には専用計算機ボード 10〜10によって構成される。専
1 N
用計算機ボード 10〜10には、ホスト計算機部 20との間で高速に設定データや計
1 N
算結果データをやり取りするためのインターフェース LSI320が備えられている。そし て、 MDG3-LSIと記されている専用処理部 100がこのインターフェース LSI320と通 信可能に構成されている。この専用処理部 100は、本発明を実装した集積回路であ る。
[0025] 専用処理部 100には、後述する粒子メモリ(粒子記憶手段) 150、パイプライン部( ノ ィプライン計算手段) 160等、本発明の処理に必要な要素全てが実装されている。 実装例では、専用計算機部 10は、 1つの専用計算機ボード 10により構成され、その
1
専用計算機ボード 10には 12個の専用処理部 100の LSIが搭載されている。各専用
1
処理部 100は、図 2に示したように、リング状に連結されて、最初のものと最後のもの 力 Sインターフェース LSI320に接続されている。ホスト計算機部 20には、インターフエ
ース部 280が適当なバスを通じてホスト CPU220と通信可能に接続されており、これ により、ホスト計算機部 20は、専用計算機部 10のインターフェース LSI320と高速ィ ンターフェース回線 300によって接続されて 、る。
[0026] 本実施形態の計算システム 1にお 、ては、ホスト計算機部 20が計算のためのデー タを送ると、インターフェース LSI320は専用処理部 100に向けてデータを受け渡し、 専用処理部 100は、内部のメモリやレジスタにそのデータを格納する。専用処理部 1 00が計算を終えると、インターフェース LSI320は計算結果データを各専用処理部 1 00から読み出してホスト計算機部 20に送信する。
[0027] 図 3は、専用処理部 100の詳細な構成を示すブロック構成図である。本実施の形態 において用いる専用処理部 100では、各粒子の情報を格納するメインメモリ(粒子メ モリ、粒子記憶手段) 150を用いる。この粒子メモリ 150は、各粒子について、座標を 120ビット (X, y, z各 40ビット)、電荷'粒子種'質量などの力またはポテンシャルの計 算に必要なデータ 40ビットの合計 160ビットを 1ワードとして格納している。そして、粒 子メモリ 150は、 32, 768ワードの容量を有しており、同時に扱える粒子数が 32、 76 8個である。粒子をアドレス指定するためには、 15ビット幅のアドレスを用いることがで きる。
[0028] 図 4を用いて分子動力学シミュレーションにおけるセルインデックス法について説明 する。図 4は、セルインデックス法を用いる場合において、計算機上にシミュレーショ ン系が設定される空間と、その空間内におけるセルの配置の例を示す説明図である 。分子動力学シミュレーションは、所定の大きさの空間 70を計算機上に設定し、この 空間にシミュレーションの対象の粒子を配置させて、シミュレーション系として行う。こ の空間 70は、その中の各位置が (X, Υ, Z)の座標値により指定可能であり、典型的 には、境界面 B〜Bの境界までの広がりを有している。ここで、 B、 Bはそれぞれ、 X
1 6 1 2
座標の最小値と最大値によって定まる境界であり、同様に、 B、 B
3 4は Y座標の最小値 と最大値によって、 B、 Bは Z座標の最小値と最大値によって定まる境界である。シミ
5 6
ユレーシヨンにおいては、典型的には、各境界面 B〜Bにおいて適当な境界条件(
1 6
たとえば、周期境界条件)が設定される。セルインデックス法においては、空間 70内 力 Sさらにセルに区分されて計算が進められる。このセルのそれぞれは、到達距離の
短い力を対象とする計算に適するような適当な大きさに区切られた立体空間領域で あり、計算の際に注目するセル(中心セル c )と、その周りに設定される周辺セルから
0
なる。図 4では、図示の都合で立方体形状の中心セル Cの各面に隣接する最隣接
0
セル c〜cを周辺セルとして明示している力 計算に考慮する周辺セルの範囲は、
1 6
中心セルの頂点から力の到達距離内に領域を持つセルであるので、通常最隣接セ ルに隣接するセル (第 2最隣接セル)まで広がっている力、また、さらに細かい分割を 行うこともできる。セルインデックス法の計算においては、空間 70中に散らばる粒子( 図示しない)を多数設定し、その空間 70のセルを順次中心セルとして定め、さらに、 到達距離の短い力が及びうる範囲を含むように周辺セルを設定して、その中心セル と周辺セルとを含めた局所空間にお 、て力やポテンシャルエネルギーの計算を実行 する。ある中心セルの設定のもとでの計算が 1ステップ分終了すると、次に、中心セル および周辺セルを移動させて同様の計算を行う。空間 70内の全セルを中心セルとし て計算し終わると、次の計算を 1ステップ分同様に実行し、多数ステップの計算を実 行してゆく。
[0029] 再び、図 3に戻って、本実施の形態の計算システムに用いる専用処理部 100の構 成を説明する。中心セルの指定は、座標 (XO, YO, ZO)によって行うことができる。こ の座標は、それぞれ各 12ビット幅の整数で表現する。中心セルに対し、計算する対 象となる周辺セルの相対位置(ΔΧ, ΔΥ, ΔΖ)のリストは相対セル位置メモリ 104に 記録されている。このメモリは、本実施の形態の専用処理部 100では、各ワードが 12 ビット幅であり、 1024ワードの容量とされている。各ワードにおいては、通常は、 ΔΧ , ΔΥ, ΔΖに 4ビットずつ割り当てられる力 計算の目的に応じて、 12ビットの幅内で ΔΧ, ΔΥ, ΔΖの割り当てを変更することができる。中心セルの座標(XO, YO, Z0) は、中心セルが変更されるたびにホスト計算機部 20から専用処理部 100に入力され る力 相対セル位置メモリ 104における周辺セルの相対位置(ΔΧ, ΔΥ, ΔΖ)のリス トは、計算の開始時等の適当な時点でホスト計算機部 20から書き込まれており、中 心セルの座標(XO, YO, Z0)が変更されても保持されている。
[0030] 専用処理部 100には計算セルカウンタ 102が備えられている。計算セルカウンタ 10 2は、計算に用いるべき相対セル位置メモリ 104のアドレスを生成する。例えば、計算
に用いるべき周辺セルの数が図 4に示した例のように合計 6である場合には、計算セ ルカウンタ 102は 6つのアドレスを生成し、そのアドレスによって、相対セル位置メモリ 104に格納されている相対位置(ΔΧ, ΔΥ, Δ Ζ)のリストから、 6つの相対位置が選 択される。計算セルカウンタ 102の幅は 10ビットであり、初期値は設定可能である。ま た終了値のレジスタ(図示しない)が別にあり、終了値に到達すると計算を終了する。 専用処理部 100には、セル番号生成部 130が備えられている。このセル番号生成 部 130は、第 1部分 130Aと第 2部分 130Bにより構成されている。セル番号生成部 1 30の第 1部分 130Aには加算器 108a〜108cが備えられている。これにより、相対セ ル位置メモリ 104から出力されたセルの相対位置(ΔΧ, ΔΥ, Δ Ζ)には、ホスト計算 機部 20から送信される中心セルの座標 (XO, Y0, Z0)位置が加算されて、各セルの 絶対位置 (X, Υ, Z)が計算される。計算結果は符号つきで 13ビットとなる。
[0031] 次に、各セルの絶対位置 (X, Υ, Z)は境界条件処理部 110に入力されて、境界条 件の処理が行われる。図 5に、境界条件処理部 110の詳細な構成を示す。境界条件 処理部 110に入力された Xのデータは、比較器 112aおよび 112bによって、空間 70 の境界面 Bを与える Xminの値および境界面 Bを与える Xmaxの値に対して比較さ
1 1
れる。各比較器 112aおよび 112bの比較結果ビットの論理和の出力は、 Xoverとして 後述するパイプライン部 160のそれぞれに出力される。なお、この Xoverの出力は、 本発明の境界判定信号となる。また、この論理和の出力は、周期境界モードに設定 するかどうかを制御する入力 Cyclicと論理積がとられて、平行移動フラグとしてセル 平行移動ユニット 152に出力される。
[0032] 境界条件処理部 110では、 Xが Xmin未満の場合には Xに Xsizeを加算し、 Xが Xm ax以上の場合には Xから Xsizeを減算する。なお、 Xsizeは、空間 70の X方向のサイ ズを現し、 Xmax— Xminにより算出される。このために、加減算器 114は、比較器 11 2aおよび 112bからの出力を受け取って動作を変更するように構成される。加減算器 114の出力は、入力された Xとともにマルチプレクサ 116に入力される。マルチプレク サ 116は、平行移動フラグによって制御されて、平行移動フラグが 1のときには加減 算器 114の出力を選択し、平行移動フラグが 0のときには入力された Xを選択し、平 行移動処理が行われた絶対位置 XTとして出力する。比較器 112bの出力は、平行移
動フラグと組み合わせて、平行移動方向を示す出力となるため、セル平行移動ュ- ット 152に出力される。なお、境界からはみ出た場合を処理することを示すフラグ (Ign ore,図示しない)が 1にされている際には、 Xが Xmin未満となったり、 Xが Xmax以 上となったりした場合にそのセルを計算から除外する。以上の処理は、 Υ、 Ζについて も同様に行われる。ここで、計算目的に応じて柔軟に専用処理部 100を動作させるこ とができるように、 X, Υ, Ζに対して独立に入力 Cyclicや Ignoreフラグが設定できる。 Xmin, Xmax, Xsizeは全て 12ビットの値であり、 Cyclic, Ignoreは 1ビットの入力で ある。 Xmin, Xmax, Xsize, Cyclic, Ignoreは、計算の初期設定として設定される 値であるために、中心セルの移動によって変更する必要がない。
[0033] 再度図 3に戻って専用処理部 100の処理を説明する。本実施の形態に用いる専用 処理部 100においては、パイプライン部 160が備えられている。このパイプライン部 1 60には、パイプライン PL〜PL が備えられている。各パイプライン PL〜PL には、
1 20 1 20 セルの絶対位置 (X, Υ, Z)から、そのセルが境界面 B〜Bをまたいだかどうかの判
1 6
定を行い、力を分割して積算することができるように、 8個のレジスタ F (0)〜F (7) (図 示しない)が備えられている。したがって、力を分割して積算するかどうかも、 XYZの 成分ごとにパイプライン PL〜PL に設けられ得るフラグ(図示しない)によって切り替
1 20
えることができる。具体的には、これらのフラグがすべて 1とされている場合には、 X, Υ, Z方向で境界をまたいだかどうかを示す各出力 Xover, Yover, Zover (またいだ ときに値 1、またがないときに値 0となる本発明の境界判定信号)を用いて、力の番号 Nfdiv=Xover+ 2 *Yover + 4 * Zoverを算出して、このセルからの寄与を、力を 積算するレジスタ F (Nfdiv)に加える。
[0034] セル平行移動ユニット 152は、粒子の座標 (X, y, z)を周期境界条件の適用に応じ て平行移動させる。境界条件処理部 110より出力された、 x、 y、 zの成分ごとの平行 移動フラグおよび平行移動する方向フラグ(2ビット 3方向で計 6ビット)のフラグを参 照し、平行移動する方向フラグの内容に応じて、基本セルの物理的大きさを x、 y、 z それぞれに加算または減算する。
[0035] 専用処理部 100には、セルの絶対位置 (X, Υ, Z)から各セルに固有のセル番号( CID: CelllD)を生成するセル番号生成部 130の第 2部分 130Bが備えられて 、る。
セル番号生成部 130の第 2部分 130Bには、乗算器 132aおよび 132bと加算器 134 が備えられ、これらによって、 CID=X+Y*Xd + Z *Xd *Ydという演算を行って、 CIDが算出される。ここで、 Xd, Ydはそれぞれ X, Y方向の最大セル数に対応する 数であり、別々のレジスタで与える。幅は 12ビットである。
[0036] このセル番号 CIDは、セルインデックスメモリ 142に入力される。セルインデックスメ モリ 142は、セル番号 CIDに対応するアドレスに、セル内の粒子が粒子メモリ 150内 のどのアドレスから格納されているかを示す 15ビットの開始アドレスと、セル内の粒子 数(15ビット)とを記憶しておく。セル番号 CIDは 12ビットであるので、セルインデック スメモリ 142は 4096ワード、 30ビット幅のメモリにより構成される。
[0037] 開始アドレスを粒子メモリ番地カウンタ 144の初期値とし、粒子メモリ 150の読み出 しを始める。セル内の粒子数を粒子数カウンタ 146の初期値とする。サイクルごとに 粒子メモリ番地カウンタ 144を 1増やし、粒子数カウンタ 146を 1減らす。粒子メモリ番 地カウンタ 144、粒子数カウンタ 146の幅は 15ビットである。この処理を粒子数カウン タ 146が 0になるまで続ける。粒子数カウンタ 146が 0になったら、次のセルの処理に 移行するために、計算セルカウンタ 102をインクリメントするための信号が出力される 。なお、実際の処理では、あるセルについての粒子メモリ 150の読み出し中に、次の セルの処理をできるだけ先に進めておくように構成されている。また、計算するセル 内の粒子数が非常に少ないときには、次のセルの準備ができる前に粒子メモリ 150 の読み出しが終わってしまう。この場合には、計算を一時停止して次の粒子メモリ 15 0の読み出しが始まるのを待つ。
[0038] 本実施の形態の専用処理部 100には、排除粒子処理部 170が備えられている。そ の排除粒子処理部 170には、 16個の排除粒子リスト EPL〜EPL が備えられている
1 16
。排除粒子リスト EPL〜EPL のそれぞれは、排除されるべき粒子の粒子メモリ 150
1 16
におけるアドレスを格納する排除粒子アドレスレジスタ 172と、各パイプラインに対し て粒子を排除するかどうかのフラグを格納する排除粒子フラグレジスタ 174を有して いる。本実施の形態の専用処理部 100には 20本のノ ィプライン PL〜PL があり、
1 20 各パイプラインは時分割動作させて論理的に 2本のパイプラインとして識別される。し たがって、排除粒子フラグレジスタ 174は 40ビット幅のレジスタにより構成される。排
除粒子アドレスレジスタ 172は 15ビット幅のレジスタにより構成される。排除粒子リスト EPL〜EPL のそれぞれは、粒子メモリ番地カウンタ 144から出力される粒子メモリ 1
1 16
50のアドレス値と、排除粒子アドレスレジスタ 172の出力とが等しい場合には、排除 粒子フラグレジスタ 174からの出力を出力し、そうでなければ 0を出力するように、一 致検出器 176aと論理積回路 176bが備えられている。全ての排除粒子リスト EPL〜
1
EPL 力もの出力は、さらに論理和回路 178に入力されてビットごとに論理和が計算
16
されて、各パイプラインに出力される。この論理和回路 178からの出力は、本発明の 積算制御信号となる。対応するパイプラインは、各ビットのフラグ力 である場合には 、計算した力をレジスタ F (0)〜F (8)に積算しないように動作し、各ビットのフラグが 0 である場合には、計算した力をレジスタ F (0)〜F (8)に積算するように動作する。 以上、本発明の実施の形態につき述べたが、これらの実施の形態および実施例は 本発明の思想を具体ィ匕する一例に過ぎない。すなわち、本発明は既述の実施の形 態に限定されるものではなぐ本発明の技術的思想に基づいて各種の変形、変更お よび組み合わせが可能である。
Claims
[1] 所定の空間内に配置される複数の粒子間に作用する力またはポテンシャルェネル ギーを、該空間を分割した複数のセルを用いるセルインデックス法によって計算する 計算処理装置であって、
前記粒子のそれぞれにつ 、て、前記空間内での座標データと前記力またはポテン シャルエネルギーの算出に必要な物理量データとを、同一のセル内にある粒子につ いて連続するアドレスの記憶領域に格納する粒子記憶手段と、
各セルに属する粒子についての該粒子記憶手段における先頭アドレスと該セル内 の粒子数または最終アドレスとを、セルを特定するセル番号に対応させて格納するセ ルインデックス記憶手段と、
前記複数のセルのうち、力またはポテンシャルエネルギーを算出すべき粒子を含む 中心セルと該中心セルに対して所定の相対位置にある少なくとも一つの周辺セルと に含まれる各粒子に対して、座標データおよび物理量データを前記粒子記憶手段 力 読み込んで、該中心セルの粒子に作用する力またはポテンシャルエネルギーを 数値計算して積算する少なくとも 1つのパイプライン計算手段と、
中心セルに対する前記周辺セルの相対座標データを連続するアドレスの記憶領域 に格納するセル相対位置記憶手段と、
中心セルの座標データを受け付けて、中心セルの該座標データに、前記セル相対 位置記憶手段の連続するアドレスに格納されている周辺セルの相対座標データをカロ 算して各周辺セルの座標を生成し、各周辺セルの該座標から、該周辺セルの前記セ ル番号を生成するセル番号生成手段と
を備え、
前記パイプライン計算手段は、前記粒子記憶手段の記憶領域のうち、前記セル番 号生成手段からのセル番号に対応させて前記セルインデックス記憶手段力 読み出 される先頭アドレスと粒子数または最終アドレスとによってアドレス指定される記憶領 域から、粒子の座標データと物理量データとを読み込んで前記数値計算を実行する ものである、計算処理装置。
[2] 各周辺セルの座標データが前記空間内の座標データであるか前記空間外の座標
データであるかの判定を行って境界判定信号を出力する境界処理手段をさらに備え 前記パイプライン計算手段は、該境界判定信号に応じて異なる演算処理を行う、請 求項 1に記載の計算処理装置。
[3] 前記境界処理手段は、前記周辺セルの座標データを前記空間外のものであると判 定した場合には、前記空間の大きさを与える大きさデータを前記座標データに加算 または減算して、前記空間内に平行移動された前記周辺セルの座標を算出するもの である、請求項 2に記載の計算処理装置。
[4] 前記パイプライン計算手段それぞれは、複数の積算データ記憶手段を有し、前記 力またはポテンシャルの計算結果を、前記境界判定信号に応じて異なる積算データ 記憶手段に積算させるものである、請求項 2に記載の計算処理装置。
[5] 前記力またはポテンシャルエネルギーの積算から排除すべき粒子の前記粒子記憶 手段におけるアドレスを保持する排除粒子アドレス記憶手段と、該排除すべき粒子を 排除すべき力否かをパイプライン計算手段毎に指定するフラグを保持する排除粒子 フラグ記憶手段とを含む排除粒子リスト手段をさらに複数含み、
該排除粒子リスト手段は、前記セル番号生成手段からのセル番号に対応して前記 セルインデックス記憶手段力 読み出される先頭アドレスと粒子数または最終アドレ スとによって指定される前記粒子記憶手段のアドレスを、前記排除粒子アドレス記憶 手段に格納されるアドレスと比較して、両アドレスが一致する粒子について、前記排 除フラグ記憶手段に格納されるフラグと組み合わせて積算制御信号を出力するもの であり、
前記パイプライン計算手段は、該積算制御信号に応じて積算を実行するか実行し ないかを切り替える、請求項 1に記載の計算処理装置。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2005-053628 | 2005-02-28 | ||
| JP2005053628A JP4740610B2 (ja) | 2005-02-28 | 2005-02-28 | 数値計算処理装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2006093121A1 true WO2006093121A1 (ja) | 2006-09-08 |
Family
ID=36941150
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2006/303694 Ceased WO2006093121A1 (ja) | 2005-02-28 | 2006-02-28 | 数値計算処理装置 |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP4740610B2 (ja) |
| WO (1) | WO2006093121A1 (ja) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104636632A (zh) * | 2015-03-10 | 2015-05-20 | 中国人民解放军国防科学技术大学 | 高精度相位小存储量查表计算方法 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8089485B2 (en) * | 2007-07-17 | 2012-01-03 | Prometech Software, Inc. | Method for constructing data structure used for proximate particle search, program for the same, and storage medium for storing program |
| JP7133843B2 (ja) * | 2018-10-19 | 2022-09-09 | 国立研究開発法人理化学研究所 | 処理装置、処理システム、処理方法、プログラム、及び、記録媒体 |
| GB201904340D0 (en) * | 2019-03-28 | 2019-05-15 | Univ Dublin | Configurational energy calculation and crystal structure prediction |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH09251449A (ja) * | 1996-03-15 | 1997-09-22 | Fuji Xerox Co Ltd | 多体問題用計算装置 |
| JP2000003352A (ja) * | 1998-06-15 | 2000-01-07 | Hitachi Ltd | 分子動力学法計算装置 |
| JP2001236342A (ja) * | 2000-02-23 | 2001-08-31 | Gazo Giken:Kk | 多体問題解析装置に用いるアドレス装置 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0477853A (ja) * | 1990-07-13 | 1992-03-11 | Fujitsu Ltd | 粒子運動シミュレーション方式 |
| JPH05173787A (ja) * | 1991-12-25 | 1993-07-13 | Hitachi Ltd | 超並列計算機及びそれを用いた計算手法 |
| JPH05324581A (ja) * | 1992-05-22 | 1993-12-07 | Hitachi Ltd | 並列多体問題シミュレータ及び当該シミュレータにおける負荷分散方法 |
| JP3528990B2 (ja) * | 1995-04-14 | 2004-05-24 | 富士ゼロックス株式会社 | 多体問題用計算装置 |
| JPH09160898A (ja) * | 1995-12-05 | 1997-06-20 | Hitachi Ltd | 粒子座標及び時系列粒子座標圧縮方法 |
-
2005
- 2005-02-28 JP JP2005053628A patent/JP4740610B2/ja not_active Expired - Fee Related
-
2006
- 2006-02-28 WO PCT/JP2006/303694 patent/WO2006093121A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH09251449A (ja) * | 1996-03-15 | 1997-09-22 | Fuji Xerox Co Ltd | 多体問題用計算装置 |
| JP2000003352A (ja) * | 1998-06-15 | 2000-01-07 | Hitachi Ltd | 分子動力学法計算装置 |
| JP2001236342A (ja) * | 2000-02-23 | 2001-08-31 | Gazo Giken:Kk | 多体問題解析装置に用いるアドレス装置 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104636632A (zh) * | 2015-03-10 | 2015-05-20 | 中国人民解放军国防科学技术大学 | 高精度相位小存储量查表计算方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP4740610B2 (ja) | 2011-08-03 |
| JP2006236256A (ja) | 2006-09-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7469407B2 (ja) | ニューラルネットワーク計算ユニットにおける入力データのスパース性の活用 | |
| US11907726B2 (en) | Systems and methods for virtually partitioning a machine perception and dense algorithm integrated circuit | |
| US20210357732A1 (en) | Neural network accelerator hardware-specific division of inference into groups of layers | |
| CN102999313B (zh) | 一种基于蒙哥马利模乘的数据处理方法 | |
| CN114242179B (zh) | 精炼合金元素收得率预测方法、系统、电子设备及介质 | |
| CN110990063B (zh) | 一种用于基因相似性分析的加速装置、方法和计算机设备 | |
| CN109863476A (zh) | 动态变量精度计算 | |
| CN101604261B (zh) | 超级计算机的任务调度方法 | |
| CN102064783B (zh) | 一种概率假设密度粒子滤波器的设计方法及滤波器 | |
| Ray et al. | Hardcaml msm: A high-performance split cpu-fpga multi-scalar multiplication engine | |
| CN108733347B (zh) | 一种数据处理方法及装置 | |
| JP4477959B2 (ja) | ブロードキャスト型並列処理のための演算処理装置 | |
| CN112711440A (zh) | 用于转换数据类型的转换器、芯片、电子设备及其方法 | |
| JP4740610B2 (ja) | 数値計算処理装置 | |
| JP2008217061A (ja) | Simd型マイクロプロセッサ | |
| Sheng et al. | VMAgent: Scheduling simulator for reinforcement learning | |
| CN120723486A (zh) | 一种规约结果的获取方法及计算设备 | |
| CN113031913B (zh) | 乘法器、数据处理方法、装置及芯片 | |
| CN113031915B (zh) | 乘法器、数据处理方法、装置及芯片 | |
| CN106407137A (zh) | 基于邻域模型的协同过滤推荐算法的硬件加速器和方法 | |
| CN109582279B (zh) | 数据运算装置及相关产品 | |
| CN109754107A (zh) | 预测方法和预测装置 | |
| JP3778489B2 (ja) | プロセッサ、演算装置及び演算方法 | |
| RU2858252C1 (ru) | Устройство для вычисления дистанции между элементами беспроводного кластерного устройства по координатам | |
| Ma et al. | Hardware-accelerated parallel genetic algorithm for fitness functions with variable execution times |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 06714831 Country of ref document: EP Kind code of ref document: A1 |