WO2004099990A1 - 計算機システム及び同システムに適用される故障計算機代替制御方法 - Google Patents
計算機システム及び同システムに適用される故障計算機代替制御方法 Download PDFInfo
- Publication number
- WO2004099990A1 WO2004099990A1 PCT/JP2004/006500 JP2004006500W WO2004099990A1 WO 2004099990 A1 WO2004099990 A1 WO 2004099990A1 JP 2004006500 W JP2004006500 W JP 2004006500W WO 2004099990 A1 WO2004099990 A1 WO 2004099990A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- computer
- computers
- failed
- unit
- boot
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/16—Error detection or correction of the data by redundancy in hardware
- G06F11/20—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements
- G06F11/202—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements where processing functionality is redundant
- G06F11/2038—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements where processing functionality is redundant with a single idle spare processing component
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/16—Error detection or correction of the data by redundancy in hardware
- G06F11/20—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements
- G06F11/202—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements where processing functionality is redundant
- G06F11/2023—Failover techniques
- G06F11/2028—Failover techniques eliminating a faulty processor or activating a spare
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/14—Error detection or correction of the data by redundancy in operations
- G06F11/1402—Saving, restoring, recovering or retrying
- G06F11/1415—Saving, restoring, recovering or retrying at system level
- G06F11/1417—Boot up procedures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/14—Error detection or correction of the data by redundancy in operations
- G06F11/1402—Saving, restoring, recovering or retrying
- G06F11/1415—Saving, restoring, recovering or retrying at system level
- G06F11/142—Reconfiguring to eliminate the error
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/16—Error detection or correction of the data by redundancy in hardware
- G06F11/20—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements
- G06F11/202—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements where processing functionality is redundant
- G06F11/2046—Error detection or correction of the data by redundancy in hardware using active fault-masking, e.g. by switching out faulty elements or by switching in spare elements where processing functionality is redundant where the redundant components share persistent storage
Definitions
- the present invention relates to a computer system including a plurality of computers including a preliminary computer called a provisioning node.
- the present invention provides a computer system suitable for enabling a boot computer to be executed on a spare computer when a running computer fails, and a fault computer applied to the system. It relates to an alternative control method.
- a high-density computer system including tens to hundreds of computer nodes in one housing has been marketed.
- Such a computer system often includes a spare computer called a provisioning node.
- a spare computer is not normally used, but is used as a substitute computer (alternate node) when a computer normally used fails. To do so, an operator operation is required to set the boot image that was running on the failed computer as the boot image for the standby computer. By starting the spare computer after this setting, the spare computer can be used as a substitute for the failed computer.
- a spare computer provisioning node
- the spare computer can be used in place of the failed computer.
- conventional techniques require the intervention of an operator to make a spare computer usable as a substitute for a failed computer.
- the present invention has been made in consideration of the above circumstances, and has as its object to operate a computer included in a computer system in the event of a failure. —To make a spare computer available to replace a failed computer without the intervention of an evening.
- a computer system including a plurality of computers including a standby computer.
- the computer system includes a plurality of storage devices that individually store boot images for booting the plurality of computers, respectively, and a first storage device that stores statuses of the plurality of computers.
- a second unit for storing information indicating a correspondence relationship between a storage unit and each of the plurality of storage devices and a computer booted by a boot image stored in the storage device;
- a storage unit; a faulty computer search unit configured to search for a faulty computer in the computer system; and a faulty computer when the faulty computer search unit finds a faulty computer.
- a spare computer for use as an alternative to the plurality of computers according to the status of the plurality of computers stored in the first storage unit.
- a storage device in which a spare computer search unit and a boot image for booting the selected spare computer when the failed computer search unit finds a failed computer are stored by the spare computer search unit.
- a boot computer image selection unit configured to select the boot computer image according to the information stored in the second storage unit, and the boot computer image selected by the standby computer search unit.
- a boot unit configured to boot using the boot image stored in the storage device selected by the selection unit.
- FIG. 1 is a block diagram showing a configuration of a computer system according to the first embodiment of the present invention.
- Figure 3 is a state transition diagram showing the state transition of the computer.
- FIG. 4 is a diagram showing an example of the data structure of the database DBB in FIG.
- Fig. 5 shows an example of the data structure of the data base HD in Fig. 1.
- FIG. 6 is a flowchart showing a processing procedure of a failure computer search process F # mainly in the first embodiment.
- FIG. 7 is a diagram showing a state in which the failure computer search process F # is failed over from the computer C 1 to the computer C 2 in the first embodiment.
- FIG. 8 is a diagram showing the contents of the database DBD # after execution of step S5 in FIG.
- FIG. 9 is a diagram showing the contents of the database CDDBi after execution of step S5 in FIG.
- FIG. 10 is a diagram illustrating a state in which a computer image that was being executed by the failed computer C 1 can now be executed by the computer C 5.
- FIG. 11 is a block diagram showing the configuration of a computer system according to the second embodiment of the present invention.
- Fig. 12 shows the failure computer search program of the second embodiment. 2004/006500
- FIG. 13 is a block diagram showing a configuration of a computer system according to the third embodiment of the present invention.
- FIG. 14 is a flowchart showing a processing procedure of a failure computer search process FP 1 according to the third embodiment.
- FIG. 15 is a block diagram showing a configuration of a computer system according to the fourth embodiment of the present invention.
- FIG. 16 is a flowchart showing a processing procedure of a failure computer search process FP1 mainly in the fourth embodiment.
- FIG. 1 is a block diagram showing a configuration of a computer system according to the first embodiment of the present invention.
- the computer system in Fig. 1 is composed of five computers C1 to C5.
- the computers C 1 to C 5 are interconnected by a network N.
- Computers C1 to C5 are connected to the storage area network SAN.
- the storage device S S is also connected to the storage area network S AN.
- the storage device S S is provided with disks (disk drives) D 1 to D 4. That is, the computers C1 to C5 and the disks D1 to D4 in the storage device SS are connected by the storage area network SAN.
- Hosts "host-1”, “host-2”, “host-3” and “host-4" are assigned to disks D1, D2, D3 and D4, respectively. Pre-stored image for computer Have been.
- the “host-i” boot image is Includes operating system OS i and application programs that operate under the operating system OS i. That is, the disk Di is used as the "host-i” boot disk.
- FIG. 1 shows that the computers C 1 to C 4 of the computers C 1 to C 5 are started and operated by using the boot images stored in the disks D 1 to D 4, respectively.
- the state is shown.
- FIG. 1 also shows a state in which the operating systems ⁇ S 1 to OS 4 are operating on the operating computers C 1 to C 4, respectively.
- These running computers C1 to C4 recognize their own host names as "host-1" to "host-4" recorded in the boot images of disks D1 to D4, respectively.
- the remaining computer C5 among the computers C1 to C5 is arranged as a spare computer called a provisioning node in the state of FIG. That is, the boot image including the operating system is not loaded on the computer C5, and the computer C5 is not booted.
- the number of spare computers does not need to be one, and a configuration in which a plurality of spare computers are arranged may be used.
- the storage device SS has a database DBDB.
- the data base DBDB uses the disks D1 to D4 and the disks D1 to D4 in the storage device SS as boot disks.
- management table 7 holds a management table (management information) for managing the correspondence with the computer (boot computer).
- the preliminary computer mining unit PPi has a database CDDBi.
- the database CDDB i holds a management table (management information) for managing the relationship between the status of each computer C 1 to C 5 and the disks D 1 to D 4 in the computer system of FIG. I do.
- the boot image stored on the disk D i in the storage device SS is harmed by “host-i” as the host name.
- the corresponding relationship between the boot image and the computer so that it is executed on the computer
- Boot unit B Bi is the host name
- the computer to which "host-i" is assigned is booted by the boot image set by the boot image setting unit BSi.
- the faulty computer search process FP runs on one of the running computers C1 to C4, for example, the computer C1.
- Failure computer search process FP uses database HDB It functions as a fault computer search unit for searching for fault computers.
- the database HDB holds a management table (management information) for managing whether the computer to which "host-i" is assigned is running (live).
- the preliminary computer search unit P Pi, the boot image setting unit B Si, and the boot unit B Bi, which operate on the computer C i, are, for example, program units.
- each program unit is included in the fault computer alternative control program.
- the units PP i, BS i and B BI are realized by the computer C i (not shown C PU) reading and executing the fault computer alternative control program.
- the fault computer search process FP running on computer C 1 is also a program unit included in the fault computer substitution control program.
- the search process F P is also realized by the computer C 1 reading and executing the faulty computer alternative control program. Therefore, the search process FP may operate on a computer other than the computer C1.
- the operating computers C1 to C4 also operate the class control units CC1 to CC4.
- the cluster control units CC1 to CC4 detect a failed computer (failed computer) by communicating with each other via the network N.
- the cluster control units CC1 to CC4 mutually and periodically transmit a signal called a hard beat signal.
- the class control units CC1 to CC4 allow the heartbeat signal to be transmitted for a predetermined period of time (timeout time). Based on the interruption, the failure occurrence is determined for the computer on which the corresponding cluster control unit exists.
- the cluster control units CC1 to CC4 constitute one virtual class control system CC.
- the cluster control system CC performs a control for transferring the service executed on the computer in which the failure was detected (that is, the failure computer) to another computer.
- a failure computer search process FP is defined in advance as one of the services controlled by the class control system CC.
- the search process FP is designed so that, when the computer on which the process FP is operating (computer C1 in FIG. 1) is stopped due to a failure or the like, the search process FP is started by another computer. Controlled by control system CC.
- the class controller C C i is realized by the computer C i reading and executing a cluster control software program (cluster software).
- cluster software cluster software
- the class software and the fault computer alternative control program are independent programs.
- code information corresponding to a fault computer alternative control program into class software in advance.
- each record of the database CDDB i includes items for setting information of a computer, the status of the computer, and each disk.
- the computer identifier used to identify the computer is used as the computer information.
- Status is the status of the corresponding computer And indicates “Run” (S), “Provisioning” (P), “Down” (D) or “Reserve” (R). If the status is "Active” (S), the disk item in the corresponding record contains the boot image used to start the computer indicated by the computer identifier in the record.
- a disk identifier is set to identify the stored disk (storage device).
- the database C DDB i in FIG. 2 shows the state of the computer system in FIG. First, the computers C 1 to C 4 are "operating"
- Figure 3 is a state transition diagram showing the state transition of the computer.
- a state (1) that is not physically incorporated into the system is shown.
- the computer (preliminary computer) in the “provisioning” (P) state passes through the “reserved” (R) state and becomes the “operating” (S) state. It can be switched to a state.
- Fig. 4 shows an example of the data structure of the database DBDB in Fig. 1.
- each record in the database DBDB includes information (for example, a disk identifier) of a disk (boot disk) storing a boot image, and a record of the disk.
- Fig. 5 shows an example of the data structure of the database HDB in Fig. 1.
- each record of the database HDB includes items for setting information indicating a host (for example, a host name) and information of a counter (counter value).
- a host for example, a host name
- a counter counter value
- FIG. 1 is a flowchart mainly showing the processing procedure of the fault computer search process FP.
- an example is given of an operation in which a failure computer is detected, and a backup computer is started using the boot image applied to the failure computer.
- the computers C 1 to C 4 are operating according to the boot images of “host-1” to “host-4”, respectively, that is, “operating” (S ).
- Boot images of "hos1" to "host-4" are stored on disks D1 to D4 in the storage device SS, respectively.
- the computer C5 is arranged in the computer system as a standby computer and is in the "provisioning" (P) state.
- the contents of the databases CDDB1 to CDDB4 of the preliminary computer search units PP1 to PP4 in the operating computers C1 to C4 are all as shown in FIG. .
- the contents of the database DBDB in the storage device SS are shown in Fig. 4. It's getting up.
- the unit CC i transmits and receives a heartbeat signal to and from each other.
- the computer Cj determined to have a fault by the class control system CC is assigned to the computer Cl on which the faulty computer search process FP is operating, that is, the host name "host-1". It is assumed that it is the assigned computer C 1.
- the search process FP is defined as a service that should be taken over (failover) to another computer when the computer on which the process FP is running is determined to be a faulty computer. I have.
- the computers C2 to C4 are positioned as standby computers.
- the search process FP is executed by the computer C 1. 2 to C4.
- the search process FP is carried over from the fault occurrence computer C1 to the computer C2 as shown in FIG. Thereafter, the search process FP runs on the computer C2.
- the failure computer search process FP executes the following procedure to detect such a state and switch from the computer C j to the spare computer.
- the search process FP refers to the database HDB to search for a host whose count has not changed during the above-mentioned fixed time (step S2). If there are no hosts whose counters have not changed, that is, if none of the counters corresponding to each host has changed, the search process FP determines that no host has failed. to decide. In this case, the mining process FP returns to step S1 and sleeps. Thereafter, the search process FP repeats the processing of steps S1 and S2 unless there is no host whose power has not changed.
- the search process FP determines the cause of the interruption of the heartbeat signal from the cluster control unit of the host. This is because the host has failed and the host has been rebooted. It is one of the things to be done.
- the search process FP sleeps for a certain period of time necessary for the reboot to determine this factor (step S3). After that, the search process FP determines whether the count of the host previously determined to be unchanged has not changed by referring to the de-night base HDB again. (Step S4). If the above counter has not changed yet, the search process FP determines that the corresponding host is not alive and has therefore failed.
- the failure of the computer C 1 to which the host name “host-1” is assigned is determined.
- the search process FP is controlled by the class control system C C and is operating on the computer C 2 instead of the computer C 1 (see FIG. 7).
- the standby computer PP2 is used to search for a standby computer (provisioning node).
- the search for the preliminary computer by the preliminary computer search unit PP2 is performed as follows.
- the preliminary computer search unit PP2 refers to the database CDDB2. Then, the search unit P P2 will be able to see the status “Provisioning”
- the computer identifier of the computer in the state of (P) is C5.
- the computer C5 is detected (selected) as a standby computer.
- the search unit PP 2 detects (selects) the computer C 5 as a spare computer, it operates the database CDDB 2 as follows. That is, when the computer C 1 is a failure computer and the computer C 5 is a standby computer, the search unit PP 2 sets the status of the failure computer C 1 and the standby computer C 5 to “down” (D) and “ To "Reserve” (R).
- the database operation contents of this search unit PP 2 are reflected in the database base CDDB 3 and CDDB 4 of the search units PP 3 and PP 4 of the other computers C 3 and C 4 in operation. .
- the contents of the database CDDB 3 and CDDB 4 will be changed to match the contents of the database CDDB 2.
- FIG. 9 shows the contents of the databases CDDB 2 to CDDB 4 (CDDB i) after this change.
- the failure computer search process FP performs a control for causing the spare computer C5 to use the boot image used by the failed computer C1. Do. In this case, the search process FP executes the boot image used by the computer C1 on the spare computer C5, so that the database DBDB in the storage device SS is set to the boot image setting unit. G is operated by BS2 (step S5).
- the boot image setting unit BS2 is a record in which the failure computer C1 is set as the boot computer among the records in the database DBDB, that is, the information (calculation) of the computer C1. Select the record that contains the (machine identifier). This selected record As is apparent from FIG. 4, the information of the disk D1 (disk identifier) is set in the code in pairs with the information of the computer C1 (computer identifier). Therefore, the selected record indicates that the boot image used to boot the failure computer C1 is stored in the disk D1.
- the boot image setting unit BS 2 operates the selected record as follows in order to execute the boot image stored in the disk D 1 on the preliminary computer C 5.
- the update operation of the record by the setting unit BS2, that is, the update operation of the database DBDB, is performed by replacing the boot image used for booting the failure computer C1 with the backup computer C5. This is equivalent to selecting the boot image as the boot image for booting. That is, the setting unit BS2 functions as a boot image selection unit.
- the boot image selection operation (operation of the database DBDB) by the setting unit BS2 allows the boot image (disk D1) of “host-1” running on the failure computer C1 to be executed. Is indirectly set in the spare computer C5.
- Figure 8 shows the contents of the database DBB at this time.
- the failure computer search process FP executes the boot image by executing the step S5 using the boot image setting unit BS2.
- the preliminary computer C5 is booted with the selected boot image (step S6).
- the spare computer C5 has an interface circuit (not shown).
- This interface circuit operates as a control circuit for connecting to the network N.
- This interface circuit is on standby to receive a special packet transmitted from the network N to itself. Therefore, the standby current is always supplied to the interface circuit.
- the interface circuit receives a special packet addressed to itself via the network N, the interface circuit has a function to start (boot) a computer having the interface circuit (in this case, the preliminary computer 5). Have.
- the standby computer C5 having such an interface circuit is set in a state in which it can always be started (standby).
- the boot nit BB 2 sends a special packet to the backup computer C 5 via the network N in order to activate the backup computer C 5.
- the interface circuit of the preparatory computer C5 causes the preparatory computer C5 to start startup processing (processing such as boot boot loader) based on the reception of the special packet.
- the preliminary computer C5 that has started the boot process refers to the database DBDB to search for a disk on which a booty image for booting itself is recorded.
- the standby computer C5 refers to the database DBDB to search for a record in which the identifier of the standby computer C5 is set.
- the preliminary computer C 5 According to the identifier of the disk D1 recorded in the record in which the identifier of the found preliminary computer C5 is set, the "host-11" block stored in the disk D1 in the storage device SS is recorded. — Boot with toy image.
- the technique of transmitting a special bucket to a specific computer via a network and activating the specific computer as described above is widely and generally known under the name of Wakeon LAN (trademark).
- the boot image of host-1 "that was being executed by the failed computer C1 can be executed by the computer C5.
- the computer C 5 is started as “host-1.”
- the database CDDB 2 is operated, and the status of the computer C 5 changes from “reserved” (R) to “active” (S).
- the operating system ⁇ S 1 the standby computer search unit PP 1, and the boot image setting unit that were operating on the computer C 1 until the computer C 1 failed.
- the BS1, the boot unit BB1 and the class control unit CC1 operate on the computer C5 as shown in FIG.
- the process FP is taken over by another computer (here, the computer C2). For this reason, the failure of the computer C 1 can be reliably determined by the search process FP.
- the standby computer search unit, boot image setting unit, and boot unit (here, standby computer) that operate on the succeeding computer Using the computer search unit PP 2, the boot image setting unit BS 2 and the boot unit BB 2), search for the spare computer and the boot image used by the fault computer C 1, respectively. It is possible to realize the setting for making the standby computer usable and to automate the boot of the standby computer.
- a counter that is incremented each time a heartbeat signal is received within the timeout period is used to detect a host failure.
- the time when the heartbeat signal is still not received even after the timeout time elapses that is, by monitoring the elapsed time from the timeout time (first timeout time)
- the failure of the corresponding host may be determined.
- FIG. 11 is a block diagram showing a configuration of a computer system according to the second embodiment of the present invention.
- components equivalent to those of the computer system of FIG. 1 are denoted by the same reference numerals.
- the configuration of the computer system in Fig. 11 will be described focusing on the differences from the computer in Fig. 1.
- a remote distribution server RDS is connected to the network N.
- the remote distribution server RDS is connected to the network N.
- disk drive R1, R2, R3 and R4 are provided.
- the disks R 1, R 2, R 3 and R 4 contain
- the computers C 1, C 2, C 3, C 4 and C 5 have disks (local disk drives) D 1, D 2, D 3, D 4 and D 5, respectively.
- Disks Dl, D2, D3, and D4 contain disks Rl, R2, R of remote distribution server RDS.
- FIG. 11 is a flowchart mainly showing the processing procedure of the fault computer search process FP.
- an operation of detecting a failure computer, setting the boot image applied to the failure computer as a standby computer, and starting the standby computer will be described as an example.
- the computers C 1 to C 4 are connected to the host computers “host-l” to “host-4” copied to the disks D 1 to D 4, respectively.
- Operating state that is, the state of “operating” (S).
- Computer C5 is located in the computer system as a standby computer, and is called "Provisioning".
- the search process FP is a process corresponding to steps S I, S 2, S 3 and S 4 in FIG.
- Steps Sl, S12, S13, and S14 determine the failure of the computer C1.
- the search process FP is taken over by any of the standby computers C2 to C4, for example, the computer C2, in the following stage.
- This stage means that the heartbeat signal from the cluster control unit CC1 operating on the computer C1 is interrupted beyond the time-out time, and as a result, the computer C1 is controlled by the cluster control system CC. This is the stage at which the occurrence of a failure has been detected.
- the spare computer search unit PP2 is used to search for a spare computer.
- the search process FP causes the backup computer C5 to use the boot image of "host-1" used by the failed computer C1. Control for the operation.
- the search process FP executes the boot image by the boot image setting unit BS2 to execute the boot image. Copy it to the local disk D5 of C5 (step S15).
- the boot image of "host-1" is stored on disk R1 in the remote distribution server RDS.
- boot image setup unit BS2 selects disk R1 from remote distribution server RDS, and boot image setup unit BS2 is stored on disk R1.
- the boot image of "host-1” is copied to the local disk D5 of the preliminary computer C5 (step S15).
- the boot image of “host-1” executed on the failure computer C 1 is directly set to the spare computer C 5 detected (selected) by the failure computer search process FP. .
- This is different from the first embodiment in which the boot image of "host-1" executed on the failure computer C1 is indirectly set on the spare computer C5.
- the search process FP causes the boot unit BB2 to boot the computer C5 according to the boot image of "host-1" copied to the disk D5 of the preliminary computer C5.
- Step S16 The boot operation of the computer C5 by the boot unit BB2 is performed in the same manner as the boot operation in the first embodiment.
- the computer C5 is started as "host-1".
- the operating system OS 1, the standby computer search unit PP 1, the boot image setting unit BS 1, the unit BB 1, and the cluster control unit CC 1 are configured to operate on the computer C 5. Become.
- FIG. 13 is a block diagram showing a configuration of a computer system according to the third embodiment of the present invention.
- components equivalent to those of the computer system of FIG. 1 are denoted by the same reference numerals.
- computers C1, C2, C3, and C4 are stored on disks D1, D2, D3, and D4 in the storage device SS, respectively.
- the feature of the fault computer search process FP1, FP2, FP3 and FP4 is that the host names are "host-1", "host-2”,
- the search processes FP 1, FP 2, FP 3 and FP 4 correspond to the search process FP, Is different from the search process FP, which operates only on one of the computers.
- Another feature of the search processes FP 1, FP 2, FP 3 and FP 4 is that they can recognize the computer on which they should operate.
- the search processes FP1, FP2, FP3, and FP4 determine the host name (own host name) assigned to the computer on which they are running by using the storage device.
- the search processes FP1, FP2, FP3, and FP4 search for faulty computers by using the host name recognition function. Therefore, unlike the search process FP, the search processes FP1, FP2, FP3, and FP4 do not require the database HDB to search for a faulty computer.
- failure computer search processes FP1, FP2, FP3, and FP4 are defined in advance as services controlled by the cluster control system CC.
- the search processes FP1, FP2, FP3, and FP4 are executed by the computers on which the processes FP1, FP2, FP3, and FP4 are operating (the computers C1, C2, When C3 and C4) are stopped due to a failure or the like, they are controlled by the class control system CC so that they are started by other computers.
- FIG. 14 is a flowchart mainly showing the processing procedure of the fault computer search process FPl (FPi).
- FPi fault computer search process
- TM fault computer search process
- the fault computer search process FP 1 determines whether or not it is running on the computer which should operate itself, that is, the computer C 1 whose host name is “host-1” (Ste S21). Normally, the search process F P1 is running on the computer C 1 as shown in FIG. In this case, the search process FP1 sleeps until the next time it is started (step S28). Now, it is assumed that a failure has occurred in the computer C 1 on which the search process F P 1 is running. Then, the search process FP1 is transferred (failed over) from the faulty computer C1 to another computer in the computer system under the control of the class controller CC. That is, the location where the search process FP1 is activated is changed from the faulty computer C1 to another computer.
- the search process FP1 operates on a computer different from the computer (C1) having the host name "host-1" on which the search process FP1 should operate.
- the search process FP1 has been started on the computer C2 whose host name is "host-2".
- the search process FP 1 When the search process FP 1 is started on the computer C 2, in step S 21 above, the search process FP 1 is not a computer C 1 whose own host name is “host-1” and is not a computer (host). It is determined that it is running on the computer C2) whose name is "hosts.” Then, the search process FP1 determines that the host name on which it was originally running is It recognizes that a failure has occurred in the computer C1 on "host-1". In this case, the search process FP1 searches for a spare computer by using, for example, the spare computer search unit PP2 on the computer C2 on which it is currently operating (steps S22, S23).
- the search for the preliminary computer by the search unit PP 2 is performed by referring to the database CDDB 2 and determining that the computer is in the “provisioning” (P) state. This is achieved by acquiring the computer identifier of If there is no computer in the “provisioning” (P) state, the faulty computer search process FP1 sleeps for a certain period of time (step S24), and then enters the standby computer search unit. A backup computer is searched again using PP2 (steps S22 and S23).
- the computer C5 is detected as a standby computer.
- the failure computer search process FP1 executes processing (steps S25 and S26) corresponding to steps S5 and S6 in FIG. 6 when the standby computer C5 is detected (selected).
- the search process FP 1 controls the boot computer of “host-1” used by the failed computer C 1 to be used by the standby computer C 5.
- the search process F P 1
- Step S25 By operating this database DBDB, the first embodiment Similarly to the state, the boot image used to boot the failure computer C1 is selected as the boot image for booting the spare computer C5.
- the search process FP1 uses the boot unit BB2 on the computer C2 on which it is currently operating to switch the standby computer C5 to the above-mentioned “host-1” booth.
- the boot image of "host-1" that was being executed by the failed computer C1 can be executed by the computer C5. That is, the computer C5 was started as "host-1", and was operating on the computer C1 until the computer C1 failed.
- the operating system OS1 and the standby computer search unit PP1 Then, the boot image setting unit BS1, the boot unit BB1, and the class control unit CC1 operate on the computer C5.
- step S28 Sleep until activated on step 28 (step S28).
- the spare computer search unit PP 1 (PP i), the boot image setting unit BS 1 (BS i), and the boot unit BB 1 (BB i) are connected to the fault computer search process FP 1 (FP i).
- the preliminary computer search unit PPI (PPi), the boot image setting unit BSI (BSi) and the boot unit BB1 (BBi) are attached to the fault computer search process FP1 (FPi). It is also possible to adopt a configuration that includes (includes). In this configuration, when the fault computer search process FP1 is transferred from the computer C1 to the computer C2 under the control of the class control system CC, the standby computer search unit PP1 and the boot image setting are performed.
- the unit BS1 and the boot unit BB1 are also moved to the computer C2.
- the fault computer search process FP1 uses the spare computer search unit PP1, the boot image setting unit BS1 and the boot unit BB1 to perform the processing shown in the flowchart of FIG. Run.
- the cluster control system CC uses a time-out time longer than the time-out time used for detecting a computer failure with respect to the failure computer search process FP 1 (FP i). Movement control from a faulty computer to another computer It is good to do aileo.
- FIG. 15 is a block diagram showing a configuration of a computer system according to the fourth embodiment of the present invention.
- components equivalent to those of the computer system of FIG. 11 or FIG. 13 are denoted by the same reference numerals.
- the difference between the computer system of FIG. 15 and the computer system of FIG. 13 applied in the third embodiment is that the computer system of FIG. 11 applied in the second embodiment differs from the computer system of FIG. This differs from the computer system of FIG. 1 applied in the embodiment of FIG.
- the processing procedure of the faulty computer search process FP1 (FPi) in the computer system of Fig. 15 is shown in the flowchart of Fig. 16. As is clear from FIG.
- the failure computer search process FPI corresponds to steps S21 to S28 of the flow chart of FIG. 14 applied in the third embodiment.
- Step S31 to S38 step S37 is executed by the cluster control system CC similarly to step S27 in FIG.
- the flow chart of FIG. 16 differs from the flow chart of FIG. 14 in the process (step S35) corresponding to step S25 in FIG. That is, the process for causing the spare computer C5 to use the boot image used by the failure computer C1 is different.
- step S35 the boot image used by the failure computer C1, that is, the boot image stored in the disk R1 in the remote distribution server RDS, is immediately copied to the local disk of the standby computer C5. Processing to copy to D5 is performed. This process is shown in Figure 12 This is the same as step S15.
- the failure detection of the computer is used as a trigger to execute the boot image executed by the computer.
- the technique of starting up the standby computer in the printer is applied. Applying this technology, for example, a configuration in which a standby computer is started with the boot image executed by the computer, as a trigger when the load on the computer becomes larger than the reference value, for example, is triggered. Is also possible. In this case, it is not always necessary to take over the fault computer search process. It is also possible to adopt a configuration in which the arrival of a predetermined time triggers the same computer to be booted with different boot images depending on the time zone.
- the same computer is started in the first boot image including the first operating system during the day, for example, and is started in the second boot image including the second operating system during the night.
- This configuration is particularly suitable for the computer system shown in Fig. 1 or Fig. 13 which does not require copying of the booty image.
- the present invention is not limited to the above-described embodiments as they are, and can be embodied by modifying constituent elements without departing from the scope of the invention in the implementation stage.
- various inventions can be formed by appropriately combining a plurality of constituent elements disclosed in the above embodiments. For example, some components may be deleted from all the components shown in the embodiment. Further, constituent elements of different embodiments may be appropriately combined.
- a boot image is set in the spare computer so that the spare computer can be used as a substitute for the failed computer without the intervention of an operator. Can be activated.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Quality & Reliability (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Hardware Redundancy (AREA)
- Stored Programmes (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US10/556,051 US7478230B2 (en) | 2003-05-09 | 2004-05-07 | Computer system and failed computer replacing method to the same system |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2003132117A JP3737810B2 (ja) | 2003-05-09 | 2003-05-09 | 計算機システム及び故障計算機代替制御プログラム |
| JP2003-132117 | 2003-05-09 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2004099990A1 true WO2004099990A1 (ja) | 2004-11-18 |
Family
ID=33432153
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2004/006500 Ceased WO2004099990A1 (ja) | 2003-05-09 | 2004-05-07 | 計算機システム及び同システムに適用される故障計算機代替制御方法 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US7478230B2 (ja) |
| JP (1) | JP3737810B2 (ja) |
| CN (1) | CN100382041C (ja) |
| WO (1) | WO2004099990A1 (ja) |
Families Citing this family (18)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2006043309A1 (ja) * | 2004-10-18 | 2006-04-27 | Fujitsu Limited | 運用管理プログラム、運用管理方法および運用管理装置 |
| DE602004027424D1 (de) * | 2004-10-18 | 2010-07-08 | Fujitsu Ltd | Operationsverwaltungsprogramm, operationsverwaltun |
| EP1814027A4 (en) * | 2004-10-18 | 2009-04-29 | Fujitsu Ltd | PROGRAM, METHOD AND INSTALLATION FOR OPERATIONAL MANAGEMENT |
| JP4462024B2 (ja) * | 2004-12-09 | 2010-05-12 | 株式会社日立製作所 | ディスク引き継ぎによるフェイルオーバ方法 |
| JP2008533573A (ja) * | 2005-03-10 | 2008-08-21 | テレコム・イタリア・エッセ・ピー・アー | 障害回復アーキテクチャー |
| JP4710518B2 (ja) | 2005-09-28 | 2011-06-29 | 株式会社日立製作所 | 計算機システムとそのブート制御方法 |
| JP4544146B2 (ja) * | 2005-11-29 | 2010-09-15 | 株式会社日立製作所 | 障害回復方法 |
| US8209417B2 (en) * | 2007-03-08 | 2012-06-26 | Oracle International Corporation | Dynamic resource profiles for clusterware-managed resources |
| JP2008269352A (ja) * | 2007-04-20 | 2008-11-06 | Toshiba Corp | アドレス変換装置及びプロセッサシステム |
| JP2010170351A (ja) * | 2009-01-23 | 2010-08-05 | Hitachi Ltd | 計算機システムのブート制御方法 |
| US8161142B2 (en) | 2009-10-26 | 2012-04-17 | International Business Machines Corporation | Addressing node failure during a hyperswap operation |
| JP5150696B2 (ja) * | 2010-09-28 | 2013-02-20 | 株式会社バッファロー | 記憶処理装置及びフェイルオーバ制御方法 |
| JP2012175574A (ja) * | 2011-02-23 | 2012-09-10 | Toshiba Corp | 送信装置及び送信システム |
| JP5484434B2 (ja) * | 2011-12-19 | 2014-05-07 | 株式会社日立製作所 | ネットワークブート計算機システム、管理計算機、及び計算機システムの制御方法 |
| JP5307223B2 (ja) * | 2011-12-22 | 2013-10-02 | テレコム・イタリア・エッセ・ピー・アー | 障害回復アーキテクチャ |
| CN103973470A (zh) * | 2013-01-31 | 2014-08-06 | 国际商业机器公司 | 用于无共享集群的集群管理方法和设备 |
| JP6123375B2 (ja) * | 2013-03-14 | 2017-05-10 | 日本電気株式会社 | 監視制御装置及び方法、組み込み制御装置、並びにコンピュータ・プログラム |
| US9880859B2 (en) * | 2014-03-26 | 2018-01-30 | Intel Corporation | Boot image discovery and delivery |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH06231101A (ja) * | 1993-01-29 | 1994-08-19 | Natl Aerospace Lab | 受信タイムアウト検出機構 |
| JPH10105423A (ja) * | 1996-09-27 | 1998-04-24 | Nec Corp | ネットワークサーバの障害監視方式 |
| JP2001256071A (ja) * | 2000-03-13 | 2001-09-21 | Fuji Electric Co Ltd | 冗長化システム |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5796934A (en) * | 1996-05-31 | 1998-08-18 | Oracle Corporation | Fault tolerant client server system |
| US5852724A (en) * | 1996-06-18 | 1998-12-22 | Veritas Software Corp. | System and method for "N" primary servers to fail over to "1" secondary server |
| FR2752631B1 (fr) * | 1996-08-22 | 1999-01-22 | Schneider Automation | Procede de chargement d'un systeme d'exploitation |
| US5903717A (en) * | 1997-04-02 | 1999-05-11 | General Dynamics Information Systems, Inc. | Fault tolerant computer system |
| US6363497B1 (en) * | 1997-05-13 | 2002-03-26 | Micron Technology, Inc. | System for clustering software applications |
| US5996086A (en) * | 1997-10-14 | 1999-11-30 | Lsi Logic Corporation | Context-based failover architecture for redundant servers |
| EP1035465A3 (en) * | 1999-03-05 | 2006-10-04 | Hitachi Global Storage Technologies Japan, Ltd. | Disk storage apparatus and computer system using the same |
| US6609213B1 (en) * | 2000-08-10 | 2003-08-19 | Dell Products, L.P. | Cluster-based system and method of recovery from server failures |
| JP2002222160A (ja) * | 2001-01-26 | 2002-08-09 | Fujitsu Ltd | 中継装置 |
| CN1319237C (zh) * | 2001-02-24 | 2007-05-30 | 国际商业机器公司 | 超级计算机中通过动态重新划分的容错 |
| GB0112781D0 (en) * | 2001-05-25 | 2001-07-18 | Global Continuity Plc | Method for rapid recovery from a network file server failure |
| US6874103B2 (en) * | 2001-11-13 | 2005-03-29 | Hewlett-Packard Development Company, L.P. | Adapter-based recovery server option |
-
2003
- 2003-05-09 JP JP2003132117A patent/JP3737810B2/ja not_active Expired - Lifetime
-
2004
- 2004-05-07 CN CNB2004800159436A patent/CN100382041C/zh not_active Expired - Lifetime
- 2004-05-07 US US10/556,051 patent/US7478230B2/en not_active Expired - Lifetime
- 2004-05-07 WO PCT/JP2004/006500 patent/WO2004099990A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH06231101A (ja) * | 1993-01-29 | 1994-08-19 | Natl Aerospace Lab | 受信タイムアウト検出機構 |
| JPH10105423A (ja) * | 1996-09-27 | 1998-04-24 | Nec Corp | ネットワークサーバの障害監視方式 |
| JP2001256071A (ja) * | 2000-03-13 | 2001-09-21 | Fuji Electric Co Ltd | 冗長化システム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2004334698A (ja) | 2004-11-25 |
| CN1802636A (zh) | 2006-07-12 |
| US20070067613A1 (en) | 2007-03-22 |
| JP3737810B2 (ja) | 2006-01-25 |
| CN100382041C (zh) | 2008-04-16 |
| US7478230B2 (en) | 2009-01-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP4462024B2 (ja) | ディスク引き継ぎによるフェイルオーバ方法 | |
| US10853056B2 (en) | System and method for supporting patching in a multitenant application server environment | |
| US7444502B2 (en) | Method for changing booting configuration and computer system capable of booting OS | |
| US7953831B2 (en) | Method for setting up failure recovery environment | |
| US7287186B2 (en) | Shared nothing virtual cluster | |
| US8010827B2 (en) | Method and computer system for failover | |
| US6134673A (en) | Method for clustering software applications | |
| JP3737810B2 (ja) | 計算機システム及び故障計算機代替制御プログラム | |
| EP1397744B1 (en) | Recovery computer for a plurality of networked computers | |
| US20010056554A1 (en) | System for clustering software applications | |
| JP4572250B2 (ja) | 計算機切り替え方法、計算機切り替えプログラム及び計算機システム | |
| JP2008097276A (ja) | 障害回復方法、計算機システム及び管理サーバ | |
| JP2007293422A (ja) | ネットワークブート計算機システムの高信頼化方法 | |
| JP5316616B2 (ja) | 業務引き継ぎ方法、計算機システム、及び管理サーバ | |
| JP2003099146A (ja) | 計算機システムの起動制御方式 | |
| JP5285045B2 (ja) | 仮想環境における故障復旧方法及びサーバ及びプログラム | |
| US7437445B1 (en) | System and methods for host naming in a managed information environment | |
| US7657734B2 (en) | Methods and apparatus for automatically multi-booting a computer system | |
| JP5131336B2 (ja) | ブート構成変更方法 | |
| CN111427721B (zh) | 异常恢复方法及装置 | |
| JP5484434B2 (ja) | ネットワークブート計算機システム、管理計算機、及び計算機システムの制御方法 | |
| JP5267544B2 (ja) | ディスク引き継ぎによるフェイルオーバ方法 | |
| JP2011086316A (ja) | 引継方法、計算機システム及び管理サーバ | |
| JP4877368B2 (ja) | ディスク引き継ぎによるフェイルオーバ方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| AK | Designated states |
Kind code of ref document: A1 Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW |
|
| AL | Designated countries for regional patents |
Kind code of ref document: A1 Designated state(s): BW GH GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| WWE | Wipo information: entry into national phase |
Ref document number: 20048159436 Country of ref document: CN |
|
| 122 | Ep: pct application non-entry in european phase | ||
| WWE | Wipo information: entry into national phase |
Ref document number: 2007067613 Country of ref document: US Ref document number: 10556051 Country of ref document: US |
|
| WWP | Wipo information: published in national office |
Ref document number: 10556051 Country of ref document: US |