WO2005017736A1 - ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム - Google Patents

ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム Download PDF

Info

Publication number
WO2005017736A1
WO2005017736A1 PCT/JP2004/011780 JP2004011780W WO2005017736A1 WO 2005017736 A1 WO2005017736 A1 WO 2005017736A1 JP 2004011780 W JP2004011780 W JP 2004011780W WO 2005017736 A1 WO2005017736 A1 WO 2005017736A1
Authority
WO
WIPO (PCT)
Prior art keywords
exceeds
period
threshold
response time
average response
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2004/011780
Other languages
English (en)
French (fr)
Inventor
Tadaomi Kato
Juichi Sakai
Naoki Hirabayashi
Takaaki Yamato
Tomonari Horikoshi
Yutaka Hiyoshi
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hiyoshi Keiko
Fujitsu Ltd
Original Assignee
Hiyoshi Keiko
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hiyoshi Keiko, Fujitsu Ltd filed Critical Hiyoshi Keiko
Priority to JP2005513194A priority Critical patent/JPWO2005017736A1/ja
Publication of WO2005017736A1 publication Critical patent/WO2005017736A1/ja
Priority to US11/321,578 priority patent/US20060106926A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/30Monitoring
    • G06F11/34Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment
    • G06F11/3409Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment for performance assessment
    • G06F11/3419Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment for performance assessment by assessing time
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/30Monitoring
    • G06F11/34Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment
    • G06F11/3452Performance evaluation by statistical analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F2201/00Indexing scheme relating to error detection, to error correction, and to monitoring
    • G06F2201/81Threshold

Definitions

  • the present invention relates to a system including a disk array device and a server that inputs and outputs data to and from the disk array device.
  • a business system is a system in which a server that provides services to client terminals via a network and a disk array device that stores various data used by application programs running on the server are connected. Is used everywhere.
  • performance information various information related to the performance of the system (performance information) is monitored so that the time required for processing the application is equal to or greater than a certain standard, and points (bottlenecks) that may cause the processing of the application to be delayed are identified.
  • performance information information related to the performance of the system
  • points points (bottlenecks) that may cause the processing of the application to be delayed are identified.
  • a process for detecting whether or not a bottleneck has occurred is performed. If a bottleneck is detected, a process for identifying the bottleneck and eliminating the bottleneck is performed for the bottleneck.
  • the bottleneck of the disk array device includes resources such as a CPU and a physical disk in the disk array device.
  • resources such as a CPU and a physical disk in the disk array device.
  • detection and identification of bottlenecks in a disk array device are performed as one unit, and the resource usage rate calculated by dividing the accumulated value of the time that resources have been used for a predetermined time by the predetermined time is used. If the resource usage rate exceeds the threshold, it was identified that the resource was a bottleneck.
  • FIG. 1 is a diagram for explaining a disk usage rate and the occurrence of a bottleneck accompanying the processing of an application.
  • the vertical axis represents elapsed time 11, and the horizontal axis represents input / output such as writing and reading issued by the server during application processing.
  • FIG. 1A shows a case where ten requests arrive in a concentrated manner at a certain time
  • FIG. 1B shows a case where ten requests arrive relatively evenly.
  • FIG. 1A shows an example in which a bottleneck occurs as a result of intensively arriving 10 requests exceeding the processing capacity of the disk array device in a short time. Before 10 requests have been processed, 10 requests arrive one after another, so it takes more time to process 1 ⁇ requests that arrive later. In Figure 1B, request 1 is being processed smoothly, and no bottleneck has been seen.
  • Disk usage which is the ratio of the average response time obtained by dividing the cumulative value of the response time by the number of 10 requests arriving at a predetermined time, and the total time the disk has been used in the predetermined time
  • Figure 1A shows that the average response time is 35 ms (ms) and the disk usage is 53%
  • Figure 1B is that the average response time is 14 ms and the disk usage is 67% become.
  • Patent Document 1 As a related art related to the reason, there is a disk array device that resolves 10 conflicts (Patent Document 1) and the like.
  • Patent Document 1 JP 2000-215007 A
  • an object of the present invention is to provide a system and a program capable of appropriately detecting the occurrence of a bottleneck.
  • the object is to provide a server that provides a service to a client terminal via a network, a disk array device connected to the server and the network and storing data used by the server, A monitoring terminal connected to the disk array device to detect a bottleneck of the disk array device, wherein the disk array device or the server is issued from the server to the disk array device.
  • the performance information including the number of 10 requests, the time required to process each 10 requests, and the resource usage rate of each resource included in the disk array device is calculated and periodically notified to the monitoring terminal.
  • the monitoring terminal may calculate the processing time included in the performance information that is periodically notified.
  • the above object is provided in claim 1, wherein the monitoring terminal sets a time when the average response time exceeds the first threshold for a period continuously exceeding the first predetermined period.
  • the monitoring terminal is configured to determine that a result of accumulating a period in which the average response time exceeds the first threshold for a third predetermined period is the first period. This is achieved by providing the system according to claim 3, wherein a time point exceeding a predetermined period is set as a reference point.
  • the above-mentioned object is the system according to claim 4, wherein in claim 3, the monitoring terminal obtains the accumulation result every third predetermined period. Is achieved by providing
  • the above object is provided in claim 3, wherein the monitoring terminal obtains the accumulation result at intervals shorter than the third predetermined period. Is achieved by providing a system. [0017] In addition, the above object is provided in claim 3, wherein the monitoring terminal is configured to determine that the average response time falls below a third threshold lower than the first threshold within the third predetermined period. In this case, this is achieved by providing the system according to claim 6, wherein the accumulated period is reset to zero once.
  • the above object is achieved in the claim 1 in which the monitoring terminal is in a fourth predetermined period before the reference point and further in a period in which the average response time exceeds a fourth threshold value. If the ratio of the period in which the resource usage rate exceeds the second threshold set for each resource exceeds the predetermined ratio, the resource is identified as a bottleneck. This is achieved by providing a system according to claim 7.
  • the above object is to provide a system having a server for providing a service to a client terminal via a network, and a disk array device connected to the server and the network and storing data used by the server.
  • a program executed by a terminal connected to the disk array device via the network, wherein the program is periodically notified to the terminal by the server or the disk array device.
  • the object is to provide a server that provides a service to a client terminal via a network, a disk array device connected to the server and the network and storing data used by the server, A monitoring terminal that is connected to the disk array device through a disk array device and detects a bottleneck of the disk array device, wherein the disk array device or the server is issued from the server to the disk array device.
  • the number of 10 requests to be made and the cost of processing each 1 ⁇ request Performance information including time and resource usage for each resource included in the disk array device is calculated and periodically notified to the monitoring terminal, and the monitoring terminal is included in the periodically notified performance information.
  • the reference point is a time at which the period during which the average response time exceeds the first threshold continuously exceeds the second predetermined period. Further, the reference point may be a time at which the sum of the periods in which the average response time exceeds the first threshold for the third predetermined period exceeds the second predetermined period.
  • the reference point is a waveform that can be obtained by plotting the average response time against time by placing the time on the horizontal axis and the average response time on the vertical axis during the period when the average response time continuously exceeds the first threshold. And the time at which the area of the part surrounded by the horizontal line indicating the first threshold value and the average response time exceeds a predetermined area.
  • the reference point is defined as a waveform obtained by plotting the average response time with respect to time, and the sum of the area of the portion surrounded by the horizontal line indicating the first threshold value accumulated for the third predetermined period. It may be a time that exceeds the area of.
  • the above object is to provide a system having a server for providing a service to a client terminal via a network, and a disk array device connected to the server and the network and storing data used by the server.
  • a program executed by a terminal connected to the disk array device via the network, wherein the program is periodically notified to the terminal by the server or the disk array device.
  • the average response time obtained by dividing the processing time included in the received performance information by the 1 ⁇ number of requests is the first
  • the reference point time is determined based on the period exceeding the threshold, and the ratio of the period in which the resource usage rate exceeds the second threshold set for each resource to the first predetermined period before the reference point Is a certain percentage This is achieved by providing a program characterized by specifying the resource as a bottleneck when the number exceeds the threshold.
  • the bottleneck can be specified based on two criteria. Neck detection can be performed appropriately.
  • FIG. 1 is a diagram for explaining the disk usage rate and the occurrence of a bottleneck accompanying the processing of an application.
  • FIG. 2 is a diagram showing a configuration example of an entire system according to an embodiment of the present invention.
  • FIG. 3 is a diagram showing a configuration example of a server.
  • FIG. 4 is a diagram showing a configuration example of a disk array device.
  • FIG. 5 is a flowchart illustrating a bottleneck detection method according to the embodiment of the present invention.
  • FIG. 6 is a diagram illustrating reference point conditions (1).
  • FIG. 7 is a diagram illustrating reference point conditions (No. 2).
  • FIG. 8 is a modified example of the method of calculating the accumulation period.
  • FIG. 9 is a diagram illustrating an example of an interval at which an accumulation period is calculated.
  • FIG. 10 is a diagram for explaining conditions (part 1) for specifying a bottleneck.
  • FIG. 11 is a diagram for explaining conditions (part 2) for specifying a bottleneck.
  • FIG. 12 is a view for explaining reference point conditions (No. 3).
  • FIG. 13 is a view for explaining reference point conditions (No. 4).
  • the resource utilization rate is monitored as A reference point for detecting a bottleneck is determined based on a condition set for a response time that does not detect a bottleneck based on the usage rate. It refers to the history of the performance information before the reference point, and identifies the bottleneck based on the specific conditions set for the resource usage.
  • FIG. 2 is a diagram showing a configuration example of a general system in the embodiment of the present invention.
  • the server 22 provides a service to the client terminal 24 via the network 21.
  • Various services such as a web server, a mail server, and a database server are provided depending on the application running on the server 22.
  • the monitoring terminal 25 is a terminal for monitoring the operation state of the server 22 ⁇ disk array device 23.
  • a SAN (Storage Channel) having a configuration including an FC (Fiber Channel) switch, etc.
  • the disk array device 23 connected to the server 22 via the (Area Network) 26 stores various data used for the above applications.
  • the server 22 accesses the data stored in the disk array device 23 and responds to the client terminal 24 with a processing result based on the application.
  • FIG. 3 is a diagram showing a configuration example of the server 22.
  • the basic configuration is the same for the client terminal 24 and the monitoring terminal 25.
  • the server 22 includes a network interface 36 (network IF) for processing communication via a network, and an input / output IF 38 for processing data exchange with a disk array device 23 connected to the server 22 and peripheral devices such as an FC switch.
  • OS internal storage 37 to be installed, memory 35 to store OS and applications read out for execution and data necessary for processing, and server 22 And a CPU 34 for controlling each of the devices according to a program stored in the memory.
  • Each device in the server 22 is connected by an internal bus 39.
  • FIG. 4 is a diagram showing a configuration example of the disk array device 23.
  • the disk array device 23 includes a network IF 43 for processing communication via a network, an input / output IF 45 for processing data exchange with the server 22 connected to the disk array device 23, and peripheral devices 40 such as FC switches, and a data interface.
  • a disk group 46 including a plurality of disks 47 for storing data, a memory 42 for storing firmware which is a program for controlling the disk array device 23, and for storing data necessary for processing, Control the device according to the firmware And a CPU 41 for controlling.
  • Each device in the disk array device 23 is connected by the internal bus 44.
  • a bottleneck detection method according to the embodiment of the present invention will be described.
  • a reference point for detecting a bottleneck is determined based on a condition set for a response time. Then, by referring to the history of the performance information before the reference point, the bottleneck is identified based on the specific conditions set for the resource utilization.
  • FIG. 5 is a flowchart illustrating a bottleneck detection method according to the embodiment of the present invention.
  • the bottleneck detection method of the present invention is implemented by executing a program stored in the memory 36 of the monitoring terminal 25.
  • the monitoring terminal shown in FIG. 2 is used to detect a bottleneck of a disk array device will be described with reference to the configuration examples of the respective devices shown in FIGS.
  • a condition (reference point condition) regarding a response time when a reference point for detecting a bottleneck is set is set in the monitoring terminal 25 of FIG. 2 (S1).
  • the reference point condition for example, a period in which the average response time continuously exceeds the predetermined threshold reaches a predetermined period, or a cumulative period in which the average response time exceeds the first threshold within the first predetermined period. May reach a second predetermined period.
  • the reference point conditions will be described later with reference to FIGS.
  • These conditions are stored in advance in storage means such as the memory 35 and the built-in disk 37 included in the monitoring terminal 25.
  • storage means such as the memory 35 and the built-in disk 37 included in the monitoring terminal 25.
  • a number specifying a reference point condition is associated with each of a plurality of conditions, and the number is stored in a variable corresponding to the reference point condition. Then, the condition can be determined by reading the number stored in the variable corresponding to the reference point condition. If there is only one condition, the condition is automatically used.
  • a condition (specifying condition) for specifying a bottleneck is set in the monitoring terminal 25 for each resource included in the disk array device 23 (S2).
  • the specific condition is, for example, a ratio of a period in which the usage rate of a certain resource exceeds a predetermined threshold set for the resource in a predetermined period. It can be set to exceed a predetermined value.
  • these conditions may be stored as variables in storage means such as the memory 35 or the built-in disk 37 included in the monitoring terminal 25, and the specific condition may be determined by reading out the variables. .
  • the specific conditions will be described later with reference to FIGS.
  • performance information on the disk array device 23 is acquired by the monitoring terminal 25 (S3).
  • the CPU 41 periodically executes the firmware to acquire performance information including at least 1 request number, 1 response time, and resource usage rate of the resources included in the disk array device 23, It can be stored in storage means such as the memory 42.
  • the performance information stored in the server 22 and disk array device 23 can be periodically updated via the network. It can be acquired by the monitoring terminal 25 and stored in a storage means such as the built-in disk 37 included in the monitoring terminal 25. Thus, in step S3, the performance information on the disk array device 23 can be acquired by the monitoring terminal 25.
  • Agent Network Management Protocol
  • SNMP manager an SNMP manager
  • the monitoring terminal 25 determines whether a bottleneck is detected based on the acquired performance information, and determines a reference point when bottleneck detection is performed (S4).
  • the bottleneck detection is determined by determining the force whose response time included in the performance information acquired in step S3 satisfies the reference point condition set in step S1. Specific examples of this determination will be described later with reference to FIGS.
  • step S4 If the reference point condition is not satisfied in step S4, the bottleneck detection process is not performed, so the process proceeds to step S8, waits for a certain period of time, acquires performance information again (S3), and detects the bottleneck. The process of determining whether to perform (S4) is repeated. If the reference point condition is satisfied in step S4, the time that satisfies the condition is determined as the reference point, and the monitoring terminal 25 determines whether the resource is a bottleneck for each resource based on the performance information acquired in step S3. (S5). In step S5, it is sufficient to determine the power of the resource use rate of each resource included in the acquired performance information that satisfies the specific condition set in step S2. This size Specific examples are described later in FIGS. 10 and 11.
  • step S5 the resource is specified as a bottleneck by the monitoring terminal 25 (S6).
  • the processing after the resource that is the bottleneck is identified varies.
  • the system administrator can be notified by e-mail, a display device (not shown) connected to the monitoring terminal 25 can indicate that the resource is a bottleneck, and can perform automatic processing. You can also. More specifically, the automatic processing is, for example, disconnecting the CPU or the disk from the system configuration, stopping the disk, or increasing the cooling fan speed of the CPU.
  • step S5 determines whether or not the determination in step S5 has been completed for all resources included in the disk array device 23 (S7). If there is a resource that has not been determined yet (No in step S7), the process returns to step S5 and continues. If the determination in step S5 has been completed for all resources (Yes in step S7), the process proceeds to step S8, and after a certain period of time, performance information is acquired again (S3), and it is determined whether a bottleneck is detected. A determination is made (S4).
  • the monitoring terminal 25 can periodically acquire performance information and detect a bottleneck.
  • the response time used to determine the ability to detect a bottleneck is the response time, which increases in time with the occurrence of the bottleneck. It is possible to detect bottlenecks more appropriately than in the examples.
  • the resource usage rate is used as a condition to identify a bottleneck, and by using response time as a condition for performing bottleneck detection (reference point condition), a single piece of performance information (resource usage) can be obtained. It is possible to identify the bottleneck more appropriately than in the conventional example using only the rate.
  • the monitoring terminal 25 executes the bottleneck detection process.
  • any terminal that is connected to the disk array device 23 via the network 21 is described. Can also be executed. Therefore, the method can be executed by the server 22, in which case the method of the present invention can be applied without introducing new hardware.
  • a reference point condition a period in which the average response time continuously exceeds a certain threshold can be set to reach a predetermined period.
  • FIG. 6 is a diagram illustrating reference point conditions (No. 1). Based on the graph of FIG. 6 showing an example of the average response time that changes with the period, a case where the bottleneck detection process is executed by applying the conditions will be described.
  • the first continuous average response time exceeds 30 ms in section 61.
  • the total period (cumulative period) of section 61 is less than the predetermined period of 600 seconds. Therefore, in section 61, bottleneck detection was not performed.
  • the state in which the average response time exceeds the threshold for 600 seconds or more continues, so the time 63 where the accumulation period exceeds 600 seconds is determined as the reference point, and Neck detection is performed.
  • FIG. 7 is a diagram illustrating a reference point condition (No. 2). Based on the graph of FIG. 7 showing an example of the average response time that changes with the period, a case in which the bottleneck detection process is executed by applying the conditions will be described.
  • 3600 seconds are used as the first predetermined period
  • 600 seconds are used as the second predetermined period
  • 30 ms is used as the threshold value.
  • the total period during which the average response time exceeds 30 ms is less than the second predetermined period of 600 seconds. Therefore, at block 71, The detection of the torneck is not performed. In the next 3600 seconds (block 72), bottleneck detection is performed when the cumulative period exceeds 600 seconds.
  • FIG. 8 shows a modification of the method of calculating the cumulative period in FIG. In Fig. 7, the period in which the average response time exceeds the threshold is simply added.In Fig. 8, a second threshold lower than the first threshold is prepared, and when the average response time is lower than the second threshold, The cumulative period is calculated so that the cumulative period up to that point is zero.
  • FIG. 8 is a graph showing an example of an average response time that changes with a period in a certain block divided into 3600 seconds. Adopt 5ms as the second threshold. Other conditions are the same as in Fig. 7. Now, 400 seconds are accumulated in the section 81 where the average response time exceeds the first threshold (30 ms). However, when the average response time falls below the second threshold, the previous cumulative period is reset to zero. Thereafter, again, the interval 82 in which the average response time exceeds the first threshold value continues for 200 seconds, but does not reach the second predetermined period because the accumulated value S has been reset (the accumulated period has not been reset. If not, this point is determined as the reference point and bottleneck detection is performed).
  • FIG. 9 is a diagram illustrating an example of an interval at which the accumulation period is calculated. In other words, it is a diagram for explaining a modification of how to take the first predetermined period in FIG. In FIG. Forces with blocks that are separated every 3600 seconds, assuming that one predetermined period (3600 seconds) does not overlap each other. .
  • FIG. 9A illustrates the same method as in FIG. 3600 sec block 91 Force S Position so that they do not overlap each other.
  • the 3600 second block 91 is slightly shifted.
  • the amount of displacement may be uniform or non-uniform.
  • step S2 the specific conditions set in step S2 will be described using examples and examples.
  • the ratio (impact) of the total time during which the resource utilization exceeds the first threshold within the predetermined period to the predetermined time is calculated, and the ratio is equal to or more than the predetermined value. And can be set.
  • the predetermined period As an example of the predetermined period, it is simply set to a time range from the reference point to a period before the predetermined period. Based on the graph of FIG. 10 showing an example of the average response time that changes with the period, a case where the bottleneck detection process is specified by applying the conditions will be described.
  • 3600 seconds is adopted as the predetermined period.
  • a CPU usage threshold of 80% and a disk usage threshold of 60% are adopted.
  • 80% is adopted as the predetermined value for the degree of influence.
  • the CPU bottlenecks if the total period during which the CPU usage exceeds 80% is 80% or more of the entire range of the effect.
  • a disk is identified as a bottleneck if the total period during which the disk usage rate exceeds 60% is 80% or more of the entire range in which the impact is monitored.
  • the section 102 where the CPU usage rate exceeds 80% from the reference point to 3600 seconds before is 20% of the range 101 in which the degree of impact is viewed, and the disk usage rate is 60%. It can be seen that 95% of the section 103, which exceeds the threshold, occupies 95% of the range 101 where the degree of impact is viewed. Therefore, a disk exceeding a predetermined value (80%) set for the degree of influence is identified as a bottleneck.
  • the average The response time is set to a time range exceeding the second threshold value. Based on the graph of FIG. 11 showing an example of the average response time that changes with the period, a case where a bottleneck is specified by applying the conditions will be described.
  • FIG. 11 30 ms is adopted as the second threshold. Otherwise, the procedure is the same as in Fig. 10.
  • the time range in which the average response time exceeds the second threshold (30 ms) up to 3600 seconds before the reference point is further extracted as a range in which the degree of influence is viewed. Then, two sections 111 and 1 12 correspond.
  • the resource identified as the bottleneck has a state where the response time is high at the reference point and the resource utilization rate is high before the reference point. Resources.
  • bottleneck detection is performed based on the response time, and by using a resource usage rate different from the response time as the specific condition, the bottleneck can be identified based on two criteria. Neck detection can be performed appropriately.
  • connection method between the disk array device 23 and the server 22 is not limited to a method via a SAN, and the present invention can be applied to a direct connection using a SCSKSmall Computer System Interface (CCS) cable or the like.
  • CCS SCSKSmall Computer System Interface
  • the performance information stored in the disk array device 23 is used to detect a bottleneck in the disk array device 23.
  • the server 22 is also provided in the OS.
  • the CPU 34 periodically executes commands and the like to obtain at least 1 ⁇ ⁇ ⁇ ⁇ ⁇ number of requests, 10 response times, and performance information including the resource usage rate of the resources included in the disk array device 23, and stores the performance information in the internal disk 37. Etc. can be stored in the storage means. Therefore, it is possible to use the performance information stored in the server.
  • the bottleneck detection method of the present invention can be implemented as a program executed in the monitoring terminal 25 or the server 22.
  • the reference point condition which is a condition for starting detection of a bottleneck
  • the period during which the average response time continuously exceeds the predetermined threshold reaches the predetermined period, or the average response time within the first predetermined period reaches the first period.
  • An example is given in which the cumulative period of the period exceeding the threshold reaches the second predetermined period.
  • the bottleneck Detection starts.
  • FIG. 12 is a diagram illustrating reference point conditions (No. 3). Based on the graph of Fig. 12 showing an example of the average response time that changes with the period, the case where the bottleneck detection process is executed when the area of the portion where the average response time continuously exceeds a certain threshold reaches a predetermined area. explain.
  • the area surrounded by the average response time and the horizontal line indicating the threshold of 30 ms is the average response time when the average response time can be represented by a function (including the case where the average response time is approximated by an approximation model). It can be calculated as the integrated value from the beginning to the end of the period where the time exceeds 30 ms. Further, as shown in FIG. 12, the area may be obtained by approximation using a rectangle for each minute section.
  • the section 121 has the first continuous average response time exceeding 30ms.
  • the area calculated from the force section 121 is less than the predetermined area S. Therefore, in section 121, bottleneck detection is not performed.
  • the area calculated continuously from the section 122 where the average response time exceeds 30 ms exceeds the predetermined area. Therefore, the last time of the period in which the average response time exceeds 30 ms is determined as the reference point, and the bottleneck is detected.
  • the reference point may be selected at any time during a period in which the average response time exceeds 30 ms.
  • FIG. 13 is a diagram illustrating reference point conditions (No. 4). Based on the graph of Fig. 13 showing an example of the average response time that changes with the period, the case where the bottleneck detection is executed when the area of the portion where the average response time exceeds the threshold within the predetermined period reaches the predetermined area. explain.
  • 3600 seconds is used as the predetermined period, and 30 ms is used as the threshold.
  • the area surrounded by the average response time in the period in which the average response time exceeds 30 ms and the horizontal line indicating the threshold of 30 ms in the period in which the average response time exceeds 30 ms in 3600 seconds is the predetermined area. If so, the processing from step S5 in FIG. 5 is started.
  • the first block 131 divided into 3600 seconds in FIG. 13 there are two periods in which the average response time exceeds 30 ms, and the portion surrounded by the average response time and the horizontal line indicating the threshold of 30 ms is shown in FIG.
  • the areas are Sl 1 and S12, respectively. And the sum (S11 + S12) does not exceed a predetermined area. Therefore, in the block 131, the detection of the bottleneck is not performed.
  • the total area (S21 + S22) calculated from the period in which the average response time exceeds 30 ms is equal to or larger than the predetermined area. Therefore, the last time of the period when the average response time exceeds 30 ms is determined as the reference point, and the bottleneck is detected.
  • the reference point may be selected at any time during the period when the average response time exceeds 30 ms.
  • a second threshold (5 ms) lower than the first threshold (for example, 30 ms) is prepared as a method of calculating the cumulative area in FIG. If the value is smaller than the threshold value, the cumulative area may be calculated so that the cumulative area up to that point is zero. Further, as an interval for calculating the accumulated area, as shown in FIG. 9B, a block of a predetermined length (for example, 360 seconds) can be slightly shifted to take a predetermined period.
  • the bottleneck detection method provides, for example, a server that provides services to client terminals via a network, and a disk array device that stores various data used by application programs running on the server. It can be applied to a system etc. to which is connected.

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Computer Hardware Design (AREA)
  • Quality & Reliability (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Debugging And Monitoring (AREA)

Abstract

 資源使用率だけを基にボトルネックを検出・特定する従来の方法では、本来解消すべきボトルネックを見逃し、未発生のボトルネックに対してボトルネック解消処理を行う場合があるという課題を有していた。そこで、クライアント端末にサービスを提供するサーバと、サーバが使用するデータが格納されるディスクアレイ装置と、ディスクアレイ装置のボトルネックを検出する監視端末とがネットワークを介して接続されるシステムを提供する。ディスクアレイ装置あるいはサーバは、サーバが発行するIO要求の数と各IO要求を処理するのに要した時間とディスクアレイ装置に含まれる資源毎の資源使用率を含むパフォーマンス情報を算出する。監視端末は、パフォーマンス情報に含まれる処理時間をIO要求数で割った平均応答時間に基づき基準点を定める。そして、基準点以前の所定期間における資源使用率に基づき、資源をボトルネックと特定することを特徴とする。

Description

明 細 書
ディスクアレイ装置におけるボトルネックを検出'
ラム
技術分野
[0001] 本発明は、ディスクアレイ装置とそのディスクアレイ装置に対しデータの入出力を行 うサーバを含むシステムに関する。
背景技術
[0002] 現在業務システムとして、ネットワークを介してクライアント端末にサービスを提供す るサーバと、そのサーバにて稼動するアプリケーションプログラムが使用する各種デ ータを格納するディスクアレイ装置とが接続されたシステムが随所で使用されている。 このようなシステムでは、アプリケーションの処理に伴う時間が増大するとクライアント 端末に提供するサービスを低下させてしまう。そこで、アプリケーションの処理に伴う 時間が一定の基準以上となるよう、システムの'性能に関する様々な情報 (パフォーマ ンス情報)を監視し、アプリケーションの処理を遅らせる原因になり得る箇所 (ボトルネ ック)が発生していないか検出する処理が実行され、ボトルネックが検出された場合、 ボトルネックを特定し、そのボトノレネックに対してボトノレネックを解消する処理が行わ れている。
[0003] ディスクアレイ装置に関するボトノレネックとしては、デイスアレイ装置内の CPU、物理 ディスク等の資源がある。従来は、ディスクアレイ装置におけるボトルネックの検出 '特 定が一体として実行され、所定時間に資源が使用された時間の累積値を、その所定 時間で割ることにより算出される資源使用率を利用し、資源使用率が閾値を超える場 合、その資源がボトルネックであると特定してレ、た。
[0004] し力しながら、資源使用率の上昇とボトルネックの発生は必ずしも対応しない場合 がある。一例として、資源としてディスクが選択された場合を説明する。
[0005] 図 1は、アプリケーションの処理に伴うディスク使用率とボトルネックの発生を説明す るための図である。縦軸が経過時間 11を表し、横軸がアプリケーションの処理に伴つ てサーバにより発行される書き込み、読み込み等の入出力 (10)要求を処理するのに 要する時間 12 (応答時間)を表す。図 1Aは、 10要求がある時間に集中して到着する 場合であり、図 1Bは、 10要求が比較的均等に到着する場合である。
[0006] 図 1Aでは、ディスクアレイ装置における処理能力以上の 10要求が短時間に集中し て到着した結果、ボトルネックが発生する例である。 10要求の処理が済まないうちに、 次々と 10要求が到着するため、後から到着した 1〇要求ほど処理に時間を要している 。図 1Bでは、 1〇要求が順調に処理されており、ボトルネックの発生は見られない。
[0007] 応答時間の累積値を所定時間に到着した 10要求数で割った平均応答時間と、そ の所定時間に占めるディスクが使用された時間を合計した累積時間の割合であるデ イスク使用率をそれぞれ算出してみると、図 1Aでは、平均応答時間が 35ミリ秒 (ms) 、ディスク使用率が 53%であるのに対し、図 1Bでは、平均応答時間 14ms、ディスク 使用率が 67%になる。
[0008] ところ力 従来の資源使用率を監視してボトルネックを検出する方法では、ディスク 使用率の閾値を 60%とした場合、ディスクがボトノレネックとして検出されるのは、図 1 Bの場合である。しかし、実際は図 1Bの場合ボトルネック解消処理を行う必要はなぐ ボトルネック解消処理が必要なのは図 1 Aの場合である。資源としてディスク以外の CPUや他の資源を監視する場合にも資源使用率と応答時間に関して図 1と同じこと 力 s言える。
[0009] 因みに関連する従来技術としては、 10競合を解消するディスクアレイ装置(特許文 献 1)等がある。
特許文献 1 :特開 2000 - 215007号公報
発明の開示
発明が解決しょうとする課題
[0010] このように、資源使用率だけを基にボトルネックを検出'特定する従来の方法では、 本来解消すべきボトルネックを見逃し、未発生のボトルネックに対してボトルネック解 消処理を行う場合があるとレ、う課題を有してレ、た。
課題を解決するための手段
[0011] そこで本発明の目的は、ボトルネックの発生を適切に検出することが可能なシステ ムおよびプログラムを提供することにある。 [0012] 上記目的は、ネットワークを介してクライアント端末にサービスを提供するサーバと、 前記サーバおよび前記ネットワークに接続され、前記サーバが使用するデータが格 納されるディスクアレイ装置と、前記ネットワークを介して前記ディスクアレイ装置に接 続され、前記ディスクアレイ装置のボトルネックを検出する監視端末を有するシステム であって、前記ディスクアレイ装置あるいは前記サーバは、前記サーバから前記ディ スクアレイ装置に対して発行される 10要求の数と各 10要求を処理するのに要した時 間と該ディスクアレイ装置に含まれる資源毎の資源使用率を含むパフォーマンス情 報を算出して前記監視端末に定期的に通知し、前記監視端末は、前記定期的に通 知されるパフォーマンス情報に含まれる前記処理時間を前記 1〇要求数で割った平均 応答時間が第一の閾値を超える期間が、第一の所定期間を超える時刻を基準点とし 、前記基準点以前の第二の所定期間に占める、前記資源使用率が前記資源毎に設 定された第二の閾値を超える期間の割合が、所定の割合を超える場合に、該資源を ボトルネックと特定することを特徴とする請求の範囲第 1項に記載のシステムを提供 することにより達成される。
[0013] また上記目的は、請求の範囲第 1項において、前記監視端末は、前記平均応答時 間が前記第一の閾値を越える期間が、連続して前記第一の所定期間を超える時刻 を基準点とすることを特徴とする請求の範囲第 2項に記載のシステムを提供すること により達成される。
[0014] また上記目的は、請求の範囲第 1項において、前記監視端末は、前記平均応答時 間が前記第一の閾値を超える期間を第三の所定期間累積した結果が、前記第一の 所定期間を超える時刻を基準点とすることを特徴とする請求の範囲第 3項に記載の システムを提供することにより達成される。
[0015] また上記目的は、請求の範囲第 3項において、前記監視端末は、前記第三の所定 期間毎に前記累積結果を求めることを特徴とする請求の範囲第 4項に記載のシステ ムを提供することにより達成される。
[0016] また上記目的は、請求の範囲第 3項において、前記監視端末は、前記第三の所定 期間より短い間隔で前記累積結果を求めることを特徴とする請求の範囲第 5項に記 載のシステムを提供することにより達成される。 [0017] また上記目的は、請求の範囲第 3項において、前記監視端末は、前記第三の所定 期間内に前記平均応答時間が、前記第一の閾値より低い第三の閾値を下回った場 合、累積された期間を一旦ゼロにリセットすることを特徴とする請求の範囲第 6項に記 載のシステムを提供することにより達成される。
[0018] また上記目的は、請求の範囲第 1項において、前記監視端末は、前記基準点以前 であって、更に前記平均応答時間が第四の閾値を超えた期間である第四の所定期 間に占める、前記資源使用率が前記資源毎に設定された前記第二の閾値を超える 期間の割合が、前記所定の割合を超える場合に、該資源をボトルネックと特定するこ とを特徴とする請求の範囲第 7項に記載のシステムを提供することにより達成される。
[0019] また上記目的は、ネットワークを介してクライアント端末にサービスを提供するサー バと、前記サーバおよび前記ネットワークに接続され、前記サーバが使用するデータ が格納されるディスクアレイ装置とを有するシステムに含まれ、該ネットワークを介して 前記ディスクアレイ装置に接続された端末にて実行されるプログラムであって、前記 端末に、前記サーバあるいは前記ディスクアレイ装置により定期的に通知される、前 記ディスクアレイ装置に対して前記サーバから発行される 10要求の数と各 10要求の 処理に要した時間と該ディスクアレイ装置に含まれる資源毎の資源使用率を含むパ フォーマンス情報を受信させ、前記受信したパフォーマンス情報に含まれる前記処 理時間を前記 10要求数で割った平均応答時間が第一の閾値を超える期間が、第一 の所定期間を超える時刻を基準点とし、前記基準点以前の第二の所定期間に占め る、前記資源使用率が前記資源毎に設定された第二の閾値を超える期間の割合が 、所定の割合を超える場合に、該資源をボトルネックと特定させることを特徴とする請 求の範囲第 8項に記載のプログラムを提供することにより達成される。
[0020] また上記目的は、ネットワークを介してクライアント端末にサービスを提供するサー バと、前記サーバおよび前記ネットワークに接続され、前記サーバが使用するデータ が格納されるディスクアレイ装置と、前記ネットワークを介して前記ディスクアレイ装置 に接続され、前記ディスクアレイ装置のボトルネックを検出する監視端末を有するシス テムであって、前記ディスクアレイ装置あるいは前記サーバは、前記サーバから前記 ディスクアレイ装置に対して発行される 10要求の数と各 1〇要求を処理するのに要した 時間と該ディスクアレイ装置に含まれる資源毎の資源使用率を含むパフォーマンス 情報を算出して前記監視端末に定期的に通知し、前記監視端末は、前記定期的に 通知されるパフォーマンス情報に含まれる前記処理時間を前記 10要求数で割った平 均応答時間が第一の閾値を超える期間に基づき基準点となる時間を決定し、前記基 準点以前の第一の所定期間に占める、前記資源使用率が前記資源毎に設定された 第二の閾値を超える期間の割合が、所定の割合を超える場合に、該資源をボトルネ ックと特定することを特徴とするシステムを提供することにより達成される。
[0021] より好ましい実施例によれば、基準点は、平均応答時間が第一の閾値を超える期 間が、連続して第二の所定期間を超える時刻である。また、基準点は、平均応答時 間が第一の閾値を超える期間を第三の所定期間累積した合計が第二の所定期間を 超える時刻でもよレ、。更に、基準点は、平均応答時間が連続して第一の閾値を超え る期間において、時間を横軸に、平均応答時間を縦軸に配置し、時間に対する平均 応答時間をプロットしてできる波形と、平均応答時間が第一の閾値を示す横線とで囲 まれる部分の面積が、所定の面積を超える時刻とすることもできる。また、基準点は、 時間に対する平均応答時間をプロットしてできる波形と、平均応答時間が第一の閾 値を示す横線とで囲まれる部分の面積を第三の所定期間累積した合計が、所定の 面積を超える時刻であってもよい。
[0022] また上記目的は、ネットワークを介してクライアント端末にサービスを提供するサー バと、前記サーバおよび前記ネットワークに接続され、前記サーバが使用するデータ が格納されるディスクアレイ装置とを有するシステムに含まれ、該ネットワークを介して 前記ディスクアレイ装置に接続された端末にて実行されるプログラムであって、前記 端末に、前記サーバあるいは前記ディスクアレイ装置により定期的に通知される、前 記ディスクアレイ装置に対して前記サーバから発行される 10要求の数と各 1〇要求の 処理に要した時間と該ディスクアレイ装置に含まれる資源毎の資源使用率を含むパ フォーマンス情報を受信させ、前記受信したパフォーマンス情報に含まれる前記処 理時間を前記 1〇要求数で割った平均応答時間が第一の閾値を超える期間に基づき 基準点となる時間を決定させ、前記基準点以前の第一の所定期間に占める、前記資 源使用率が前記資源毎に設定された第二の閾値を超える期間の割合が、所定の割 合を超える場合に、該資源をボトルネックと特定させることを特徴とするプログラムを 提供することにより達成される。
発明の効果
[0023] 応答時間を基にボトルネックの検出を実施し、特定条件として応答時間とは異なる 資源使用率を用いることで、 2つの基準によってボトルネックの特定を行うことができ、 従来よりもボトルネックの検出を適切に行うことが可能である。
図面の簡単な説明
[0024] [図 1]アプリケーションの処理に伴うディスク使用率とボトルネックの発生を説明するた めの図である。
[図 2]本発明の実施形態におけるシステム全体の構成例を示す図である。
[図 3]サーバの構成例を示す図である。
[図 4]ディスクアレイ装置の構成例を示す図である。
[図 5]本発明の実施形態におけるボトルネック検出方法を説明するフローチャートで める。
[図 6]基準点条件 (その 1)を説明する図である。
[図 7]基準点条件 (その 2)を説明する図である。
[図 8]累積期間の算出法の変形例である。
[図 9]累積期間が算出される間隔の例を説明する図である。
[図 10]ボトルネックを特定する条件(その 1)を説明するための図である。
[図 11]ボトルネックを特定する条件(その 2)を説明するための図である。
[図 12]基準点条件 (その 3)を説明する図である。
[図 13]基準点条件 (その 4)を説明する図である。
発明を実施するための最良の形態
[0025] 以下、本発明の実施の形態について図面に従って説明する。し力 ながら、本発 明の技術的範囲はかかる実施の形態に限定されるものではない。
[0026] 図 1に示されるように、ボトルネックが発生すると、 1〇要求の処理に要する応答時間 が増大する。従って、ボトルネックの発生を検出するには応答時間を監視するのがよ レ、。そこで本発明の実施形態においては、従来のように資源使用率を監視し、資源 使用率によりボトルネックを検出するのではなぐ応答時間に対して設定された条件 に基づき、ボトルネックを検出する基準点を決定する。そして、基準点以前のパフォ 一マンス情報の履歴を参照し、資源使用率に対して設定された特定条件に基づき、 ボトルネックを特定するものである。
[0027] 図 2は、本発明の実施形態における一般的なシステムの構成例を示す図である。サ ーバ 22は、ネットワーク 21を介してクライアント端末 24に対しサービスを提供する。サ ーバ 22上で稼動するアプリケーションに応じて、ウェブサーバ、メールサーバ、デー タベースサーバ等さまざまなサービスが提供される。監視端末 25は、サーバ 22ゃデ イスクアレイ装置 23の動作状態を監視するための端末である。
[0028] FC(Fiber Channel)スィッチ等を含む構成の SAN(Storage
Area Network)26を介してサーバ 22に接続されたディスクアレイ装置 23には、上記 のアプリケーションに使用されるさまざまなデータが格納される。クライアント端末から の要求に応じてサーバ 22は、ディスクアレイ装置 23に格納されたデータにアクセスし 、アプリケーションに基づく処理結果をクライアント端末 24に応答する。
[0029] 図 3は、サーバ 22の構成例を示す図である。基本的な構成は、クライアント端末 24 、監視端末 25でも同様である。サーバ 22は、ネットワークを介した通信を処理するネ ットワークインタフェース 36 (ネットワーク IF)と、サーバ 22に接続するディスクアレイ装 置 23、 FCスィッチ等の周辺機器とのデータ交換を処理する入出力 IF38と、 OSゃァ プリケーシヨン力 Sインストールされる内蔵ディスク 37と、実行のために読み出された OS やアプリケーションが格納され、また処理に必要なデータが格納されるメモリ 35と、サ ーバ 22内の各装置をメモリに格納されたプログラムに従って制御する CPU34とを有 する。サーバ 22内の各装置は内部バス 39により接続される。
[0030] 図 4は、ディスクアレイ装置 23の構成例を示す図である。ディスクアレイ装置 23は、 ネットワークを介した通信を処理するネットワーク IF43と、ディスクアレイ装置 23に接 続するサーバ 22、 FCスィッチ等の周辺機器 40とのデータ交換を処理する入出力 IF 45と、データを格納するディスク 47を複数含むディスク群 46と、ディスクアレイ装置 2 3を制御するプログラムであるファームウェアが格納され、また処理に必要なデータが 格納されるメモリ 42と、ディスクアレイ装置 23内の各装置をファームウェアに従って制 御する CPU41とを有する。ディスクアレイ装置 23内の各装置は内部バス 44により接 糸冗 れる。
[0031] 続いて本発明の実施形態におけるボトルネック検出方法を説明する。本発明の実 施形態においては、応答時間に対して設定された条件に基づき、ボトルネックを検出 する基準点を決定する。そして、基準点以前のパフォーマンス情報の履歴を参照し、 資源使用率に対して設定された特定条件に基づき、ボトルネックを特定するものであ る。
[0032] 図 5は、本発明の実施形態におけるボトルネック検出方法を説明するフローチヤ一 トである。例えば、監視端末 25のメモリ 36に格納されたプログラムを実行することによ り、本発明のボトルネック検出方法が実施される。ここでは、図 2の監視端末を用いて ディスクアレイ装置のボトルネックを検出する様子を、図 3、図 4に示される各装置の 構成例を参照して説明する。
[0033] まず、ボトルネックを検出する基準点を設定する際の応答時間に関する条件 (基準 点条件)を図 2の監視端末 25に設定する(Sl)。本実施形態においては、応答時間 が基準点条件を満たすことにより、ボトルネックの検出が実行され、基準点以前のパ フォーマンス情報の履歴を参照し、ボトルネックが特定される。基準点条件としては、 例えば、平均応答時間が連続して所定の閾値を超える期間が所定期間に達すること や、第一の所定期間内に平均応答時間が第一の閾値を超える期間の累積期間が第 二の所定期間に達すること等と設定することができる。なお基準点条件については、 図 6から図 9にて後述する。
[0034] これらの条件は、監視端末 25に含まれるメモリ 35や内蔵ディスク 37等の記憶手段 に予め格納される。例えば、複数の条件にそれぞれ、基準点条件を特定する数字を 対応させ、基準点条件に対応する変数にその数字を格納する。すると、基準点条件 に対応する変数に格納された数字を読み出すことにより、条件を決定することができ る。条件力 S1つのみであれば、自動的にその条件が使用される。
[0035] 次に、ボトルネックを特定する条件(特定条件)をディスクアレイ装置 23に含まれる 資源毎に監視端末 25に設定する(S2)。特定条件としては、例えば、所定期間に占 める、ある資源の使用率がその資源に設定された所定の閾値を超える期間の割合が 所定値を越えること等と設定することができる。基準点条件同様これらの条件は、監 視端末 25に含まれるメモリ 35や内蔵ディスク 37等の記憶手段に変数として格納され 、その変数を読み出すことにより特定条件が決定されるよう構成してもよい。なお特定 条件については、図 9、図 10にて後述する。
[0036] 次に、監視端末 25にてディスクアレイ装置 23に関するパフォーマンス情報を取得 する(S3)。ディスクアレイ装置 23においては、定期的にファームウェアを CPU41が 実行することにより、少なくとも 1〇要求数、 1〇応答時間、ディスクアレイ装置 23に含ま れる資源の資源使用率を含むパフォーマンス情報を取得し、メモリ 42等の記憶手段 に蓄積することができる。
[0037] また、サーバ 22やディスクアレイ装置 23に SNMP(Simple
Network Management Protocol)エージェント機能を持つプログラムを組み込み、監視 端末 25に SNMPマネージャ機能を持つプログラムを組み込むことで、ネットワークを介 して、サーバ 22やディスクアレイ装置 23に蓄積されたパフォーマンス情報を定期的 に監視端末 25にて取得し、監視端末 25に含まれる内蔵ディスク 37等の記憶手段に 格納すること力できる。こうして、ステップ S3において、監視端末 25にてディスクァレ ィ装置 23に関するパフォーマンス情報を取得することができる。
[0038] そして、監視端末 25にて、取得したパフォーマンス情報を基にボトルネックを検出 するか判定し、ボトルネックの検出を実行する場合は基準点を決定する(S4)。ステツ プ S4のボトルネック検出判定は、ステップ S3で取得したパフォーマンス情報に含ま れる応答時間がステップ S1で設定された基準点条件を満たす力を判定すればょレ、 。この判定の具体例については図 6から図 9に後述する。
[0039] ステップ S4で基準点条件を満たさない場合、ボトルネック検出処理は行われないの で、ステップ S8に進み、一定時間待機した後、再びパフォーマンス情報を取得し(S3 )、ボトルネックを検出するかを判定する(S4)処理を繰り返す。ステップ S4で基準点 条件を満たす場合、条件を満たす時刻を基準点と決定し、監視端末 25にて、ステツ プ S3で取得したパフォーマンス情報を基に資源毎にその資源がボトルネックかを判 定する(S5)。ステップ S5では、取得したパフォーマンス情報に含まれる資源毎の資 源使用率がステップ S2で設定された特定条件を満たす力、を判定すればよい。この判 定の具体例については図 10および図 1 1に後述する。
[0040] ステップ S5で条件を満たす場合、監視端末 25にてその資源をボトルネックと特定 する(S6)。ボトルネックである資源が特定された後の処理はさまざまである。例えば、 メールでシステム管理者に通知することもできるし、監視端末 25に接続された図示し ないディスプレイ装置にその資源がボトルネックであることを表示することもできるし、 自動的な処理をさせることもできる。 自動的な処理をより具体的に述べると、例えば、 CPUやディスクをシステム構成から切り離したり、ディスクを停止させたり、 CPUの冷去口 ファン速度を上昇させたりすることである。
[0041] ステップ S5で条件を満たさない場合、監視端末にてディスクアレイ装置 23に含まれ るすべての資源についてステップ S5の判定が完了したかを判定する(S7)。未だ、判 定の行われていない資源がある場合 (ステップ S7で Noの場合)、ステップ S5に戻り 処理が続行する。すべての資源についてステップ S5の判定が完了すれば(ステップ S7で Yesの場合)、ステップ S8に進み、一定時間経過した後、再びパフォーマンス 情報を取得し (S 3)、ボトルネックを検出するかを判定する(S4)。
[0042] 以上のボトルネック検出処理により、監視端末 25にて、定期的にパフォーマンス情 報を取得し、ボトルネックの検出を行うことができる。ボトルネックを検出する力を判定 するのに使用されるのは、ボトルネックの発生に連動して時間が増大する応答時間で あり、ボトルネックの発生とは必ずしも連動しない資源使用率を利用する従来例よりも ボトルネックの検出を適切に行うことが可能となる。またボトルネックを特定する条件と して使用されるのは資源使用率であり、ボトルネック検出を実施する条件 (基準点条 件)として応答時間を用いることにより、単一のパフォーマンス情報(資源使用率)の みを用いる従来例よりも、ボトルネックの特定をより適切に行うことが可能となる。
[0043] なお、本発明の実施形態においては、監視端末 25にて、ボトルネック検出処理を 実行する様子を説明したが、ネットワーク 21を介してディスクアレイ装置 23に接続さ れていればどの端末においても実行することが可能である。従ってサーバ 22にて実 行することもでき、その場合新たなハードウェアを導入することなく本発明の方法を適 用すること力 Sできる。
[0044] 続いて、ステップ S 1で設定される基準点条件について、レ、くつかの例を用いて説 明する。まず、基準点条件として、平均応答時間が連続してある閾値を超える期間が 所定期間に達することと設定することができる。
[0045] 図 6は、基準点条件 (その 1)を説明する図である。期間と共に変化する平均応答時 間の一例を示す図 6のグラフを基に、その条件を適用してボトルネック検出処理が実 行される場合を説明する。
[0046] 図 6では、閾値として 30ms、所定期間として 600秒を採用する。つまり、平均応答 時間が 30msを超える期間が 600秒連続した場合、図 5のステップ S5以降の処理が 開始される。
[0047] 図 6で最初に連続して平均応答時間が 30msを超えるのは、区間 61である。しかし 区間 61の期間合計 (累積期間)は、所定期間の 600秒に満たない。そこで、区間 61 では、ボトルネックの検出は実施されなレ、。次に連続して平均応答時間が 30msを超 える区間 62では、 600秒以上平均応答時間が閾値を超える状態が連続するため、 累積期間が 600秒を超える時刻 63が基準点と決定され、ボトルネックの検出が実行 される。
[0048] 連続して平均応答時間が閾値を超えた期間の合計が所定期間に達するのは、平 均応答時間の高い状態が持続していることを意味し、ボトルネックが発生している可 能性が高い。従って、基準点条件をこのように設定することでボトルネックをより適切 に検出することができる。
[0049] 基準点条件の別の条件として、第一の所定期間内に平均応答時間がある閾値を 超える期間の合計 (累積期間)が、第二の所定期間に達することと設定することがで きる。図 7は、基準点条件 (その 2)を説明する図である。期間と共に変化する平均応 答時間の一例を示す図 7のグラフを基に、その条件を適用してボトルネック検出処理 が実行される場合を説明する。
[0050] 図 7では、第一の所定期間として 3600秒、第二の所定期間として、 600秒、閾値と して 30msを採用する。つまり、 3600秒の内、平均応答時間が 30msを超える期間の 合計が 600秒に達した場合、図 5のステップ S5以降の処理が開始される。
[0051] 図 7で 3600秒に区切られた最初のブロック 71では、平均応答時間が 30msを超え る期間の合計は、第二の所定期間の 600秒に満たない。そこで、ブロック 71では、ボ トルネックの検出は実行されなレ、。次の 3600秒(ブロック 72)では、累積期間が 600 秒を超える時、ボトルネックの検出が実行される。
[0052] ある期間内に平均応答時間が閾値を超えた期間の合計が(第二の)所定期間に達 するのは、平均応答時間の高い状態が持続していることを意味し、ボトルネックの発 生の可能性が高い。従って、基準点条件をこのように設定することでボトルネックをよ り検出しやすくすることができる。更に、図 7の設定にすると、連続して平均応答時間 が閾値を超える区間が短レ、ため、図 6の設定ではボトルネックの検出が行われなレヽ 場合でも、ボトルネックの検出が実行されることがあり、よりボトルネックの検出精度を 上げ'ること力 Sできる。
[0053] 図 8は、図 7における累積期間の算出法の変形例である。図 7においては、単純に 平均応答時間が閾値を超える期間を加算するが、図 8では、第一の閾値より低い第 二の閾値を用意し、平均応答時間が第二の閾値を下回る場合、それまでの累積期 間をゼロにするようにして累積期間を算出するものである。
[0054] 図 8は、 3600秒に区切られたあるブロックにおける、期間と共に変化する平均応答 時間の一例を示すグラフである。第二の閾値として 5msを採用する。他の条件は図 7 と同様とする。今、平均応答時間が第一の閾値(30ms)を越える区間 81で 400秒が 累積される。しかし、その後平均応答時間が第二の閾値を下回るとき、それまでの累 積期間がゼロにリセットされる。その後再び、平均応答時間が第一の閾値を超える区 間 82が 200秒連続するが、累積値力 Sリセットされているため、第二の所定期間には 達しない(ちなみに累積期間がリセットされていなければこの時点が基準点と決定さ れ、ボトルネックの検出が実施される)。
[0055] 図 8において平均応答時間が第二の閾値を下回る場合、平均応答時間が変動して レ、ることを意味する。ディスクアレイ装置 23においてボトルネックが発生する場合であ れば、平均応答時間が高い状態が維持されるため、平均応答時間に変動が生じて いる場合、ディスクアレイ装置 23以外でボトルネックが発生している可能性を意味し、 図 8の累積期間算出法にはこれを除外する効果がある。
[0056] 図 9は、累積期間が算出される間隔の例を説明する図である。言い換えると、図 7に おける第一の所定期間の取り方の変形例を説明する図である。図 7においては、第 一の所定期間(3600秒)を互いに重ならない範囲として、 3600秒ごとに区切ったブ ロックが現れた力 図 9では、 3600秒のブロックを少しずつずらして第一の所定時間 を取るものである。
[0057] 図 9Aは、図 7と同じ方法を図に表したものである。 3600秒のブロック 91力 S互レヽに 重ならないように位置する。図 9Bは、 3600秒のブロック 91が少しずつずれて位置す る。ずれの量は、均一でも不均一でも構わない。図 9Bのようにブロックを取ることで、 ボトルネックの検出処理が行われる回数を増やすことができ、よりボトルネックの検出 精度を上げることができる。
[0058] 次に、ステップ S2で設定される特定条件について、レ、くつか例を用いて説明する。
ボトルネックを特定する条件としては、所定期間内に資源使用率が第一の閾値を越 える期間の合計時間が、その所定時間に占める割合 (影響度)を算出し、その割合が 所定値以上であることと設定することができる。
[0059] まず、所定期間の一例としては、単純に基準点から所定期間前までの時間範囲と することである。期間と共に変化する平均応答時間の一例を示す図 10のグラフに基 づき、その条件を適用してボトルネック検出処理が特定される場合を説明する。
[0060] 図 10では、所定期間として 3600秒を採用する。資源毎に設定される資源使用率 の閾値としては、 CPU使用率の閾値として 80%、ディスク使用率の閾値として 60%を 採用する。そして、影響度に対する所定値として 80%を採用する。つまり、基準点か ら 3600秒前までの期間(影響度を見る範囲)において、 CPU使用率が 80%を超えた 期間の合計が影響度を見る範囲全体の 80%以上であれば CPUがボトルネックと特 定され、同様にディスク使用率が 60%を越えた期間の合計が影響度を見る範囲全 体の 80%以上であればディスクがボトルネックと特定される。
[0061] 図 10では、基準点から 3600秒前までにおいて、 CPU使用率が 80%を超えた区間 102が、影響度を見る範囲 101に占める割合が 20%であり、ディスク使用率が 60% を超えた区間 103が、影響度を見る範囲 101に占める割合が 95%であることがわか る。従って、影響度に対して設定された所定値(80%)を超えるディスクがボトルネッ クであると特定される。
[0062] 所定期間の別の一例としては、基準点から所定期間前までの履歴において、平均 応答時間が第二の閾値を超える時間範囲とすることである。期間と共に変化する平 均応答時間の一例を示す図 11のグラフに基づき、その条件を適用してボトルネック が特定される場合を説明する。
[0063] 図 11では、第二の閾値として 30msを採用する。それ以外は図 10の場合と同様と する。図 11では、基準点から 3600秒前までにおいて、更に、平均応答時間が第二 の閾値(30ms)を超える時間範囲を影響度を見る範囲として抜き出す。すると 2つの 区間 111、 1 12が該当する。
[0064] そして、影響度を見る範囲(区間 111、 112)にて、 CPU使用率が 80%を超えた区 間 113が、影響度を見る範囲(区間 111、 112)に占める割合が 20%であり、ディスク 使用率が 60%を超えた時間(区間 114、 115)の合計が、影響度を見る範囲(区間 1 11、 112)に占める割合が 85%であることがわかる。従って、影響度に対して設定さ れた所定値(80%)を超えるディスクがボトルネックであると特定される。
[0065] 以上、本発明の実施形態をまとめると、ボトルネックと特定される資源は、基準点で 応答時間が高い状態が継続しており、基準点以前に資源使用率も高い状態であつ た資源である。こうして、応答時間を基にボトルネックの検出を実施し、特定条件とし て応答時間とは異なる資源使用率を用いることで、 2つの基準によってボトルネックの 特定を行うことができ、従来よりもボトルネックの検出を適切に行うことが可能である。
[0066] なお、上記図 6から図 11にて使用される数値は一例に過ぎず、実施の形態に合わ せて自由に設定することが可能である。また、ディスクアレイ装置 23とサーバ 22間の 接続法は SANを介す方法に限定されず、 SCSKSmall Computer System Interface)ケ 一ブル等を用いたダイレクト接続でも本発明の適用が可能である。
[0067] また、本発明の実施形態においては、ディスクアレイ装置 23におけるボトルネックを 検出するために、ディスクアレイ装置 23に蓄積されるパフォーマンス情報を用いたが 、サーバ 22でも、 OSに備えられたコマンド等を定期的に CPU34が実行することにより 、少なくとも 1〇要求数、 10応答時間、ディスクアレイ装置 23に含まれる資源の資源使 用率を含むパフォーマンス情報を取得し、パフォーマンス情報を内蔵ディスク 37等の 記憶手段に蓄積することができる。従って、サーバに蓄積されるパフォーマンス情報 を利用することも可能である。 [0068] 更に、本発明のボトルネック検出方法は、監視端末 25、あるいはサーバ 22にて実 行されるプログラムとして実施することも可能である。
[0069] ここで更に、ボトルネックの検出を開始するための条件である、基準点条件の変形 例について説明する。図 6から図 9に説明した基準点条件においては、平均応答時 間が連続して所定の閾値を超える期間が所定期間に達することや、第一の所定期間 内に平均応答時間が第一の閾値を超える期間の累積期間が第二の所定期間に達 することを一例として挙げた。ここでは、平均応答時間が閾値を超える部分の面積が 所定面積に達する場合や、所定期間内に平均応答時間が閾値を超える部分の面積 (累積面積)が所定面積に達する場合に、ボトルネックの検出が開始される。
[0070] 図 12は、基準点条件 (その 3)を説明する図である。期間と共に変化する平均応答 時間の一例を示す図 12のグラフを基に、平均応答時間が連続してある閾値を超える 部分の面積が所定面積に達すると、ボトルネック検出処理が実行される場合を説明 する。
[0071] 図 12では、閾値として 30msを採用する。つまり、平均応答時間が 30msを超える 期間の平均応答時間と、閾値である 30msを示す横線とで囲まれる部分の面積が所 定面積に達する場合、図 5のステップ S 5以降の処理が開始される。
[0072] 平均応答時間と、閾値である 30msを示す横線とで囲まれる部分の面積は、平均応 答時間を関数により表せる場合 (近似モデルにより近似される場合も含む)には、平 均応答時間が 30msを超える期間の最初から最後までの積分値として求めることがで きる。また、図 12に示されるように、微小区間毎の長方形による近似により面積を求 めても良い。
[0073] 図 12で最初に連続して平均応答時間が 30msを超えるのは、区間 121である。し 力 区間 121から算出される面積は、所定面積 Sに満たない。そこで、区間 121では 、ボトルネックの検出は実施されない。
[0074] 次に連続して平均応答時間が 30msを超える区間 122から算出される面積は、所 定面積を超える。従って、平均応答時間が 30msを超える期間の最後の時刻が基準 点と決定され、ボトルネックの検出が実行される。なお、基準点は平均応答時間が 30 msを超える期間のどの時刻が選択されてもよい。 [0075] 平均応答時間が所定の閾値を超える期間は短いが、その応答遅延の程度が大き い場合には、ボトルネックが発生している可能性が高レ、。この面積方式を使用すると 、平均応答時間が所定の閾値を超える期間が短いため、図 6から図 9に示す方式で はボトルネックの検出が行われない場合にも、ボトルネックの検出を開始することがで きる。つまり、短い時間帯で応答時間が極端に遅い場合であってもボトノレネックの検 出を開始することができ、基準点条件をこのように設定することでボトルネックをより適 切に検出することができる。
[0076] 図 13は、基準点条件 (その 4)を説明する図である。期間と共に変化する平均応答 時間の一例を示す図 13のグラフを基に、所定期間内に平均応答時間が閾値を超え る部分の面積が所定面積に達するとボトルネックの検出が実行される場合を説明す る。
[0077] 図 13では、所定期間として 3600秒、閾値として 30msを採用する。つまり、 3600秒 の内、平均応答時間が 30msを超える期間における、平均応答時間が 30msを超え る期間の平均応答時間と、閾値である 30msを示す横線とで囲まれる部分の面積が 所定面積に達する場合、図 5のステップ S 5以降の処理が開始される。
[0078] 図 13で 3600秒に区切られた最初のブロック 131では、平均応答時間が 30msを超 える期間が 2箇所あり、平均応答時間と、閾値である 30msを示す横線とで囲まれる 部分の面積は、それぞれ Sl l、 S12であるとする。そして、その合計(S11 + S12)は 所定面積を超えない。そこで、ブロック 131では、ボトノレネックの検出は実行されない
[0079] 次の 3600秒(ブロック 132)では、平均応答時間が 30msを超える期間から算出さ れる面積の合計(S21 + S22)が所定面積以上となる。従って、平均応答時間が 30 msを超える期間の最後の時刻が基準点と決定され、ボトルネックの検出が実行され る。なお、基準点は平均応答時間が 30msを超える期間のどの時刻が選択されてもよ レ、。
[0080] ある期間内に平均応答時間が閾値を超えた期間から算出される面積の合計が所 定面積に達するのは、短い時間帯で応答時間が極端に遅い場合が発生している可 能性を示唆し、ボトルネックが発生している可能性が高い。従って、基準点条件をこ のように設定することでボトルネックをより検出しやすくすることができる。更に、図 13 の設定にすると、連続して平均応答時間が閾値を超える区間が短いため、図 12の設 定ではボトルネックの検出が行われない場合でも、ボトルネックの検出が実行されるこ とがあり、よりボトルネックの検出精度を上げることができる。
[0081] 図 6から図 9に示した基準点条件では、閾値 (例えば 30ms)を大きく超える現象に 対する配慮を行っていない。つまり、所定の閾値を超える期間は短いが、その応答遅 延の程度が大きい場合には、ボトルネックが発生している可能性が高いものの、それ を適切に検出できない事態も起こりうる。一方、図 12、図 13に示される基準点条件に よれば、短い時間帯で応答時間が極端に遅い場合であってもボトルネックの検出を 開始することができ、より適切にボトルネックを検出することができるようになる。
[0082] また、図 13における累積面積の算出法として、図 8に示されるように、第一の閾値( 例えば 30ms)より低い第二の閾値(5ms)を用意し、平均応答時間が第二の閾値を 下回る場合、それまでの累積面積をゼロにするようにして累積面積を算出してもよい 。また、累積面積を算出する間隔として、図 9Bに示されるように、所定長(例えば 360 0秒)のブロックを少しずつずらして所定期間を取ることもできる。
[0083] 図 12、図 13に示すような、面積に基づくボトルネック検出の開始法を採用しても、 その後の処理は図 5に示される場合と変わらずに行うことができる。つまり、ボトルネッ クの判断は、図 10、図 11に示されるように行って良レ、。また、図 12、図 13に示される 変形例であっても、図 1一図 11に示される実施形態同様の効果を得ることができる。 産業上の利用可能性
[0084] 本発明のボトルネック検出方法は、例えば、ネットワークを介してクライアント端末にサ 一ビスを提供するサーバと、そのサーバにて稼動するアプリケーションプログラムが使 用する各種データを格納するディスクアレイ装置とが接続されたシステム等に適用が 可能である。
[0085] 本発明の保護範囲は、上記の実施の形態に限定されず、特許請求の範囲に記載 された発明とその均等物に及ぶものである。

Claims

請求の範囲
[1] ネットワークを介してクライアント端末にサービスを提供するサーバと、前記サーバ および前記ネットワークに接続され、前記サーバが使用するデータが格納されるディ スクアレイ装置と、前記ネットワークを介して前記ディスクアレイ装置に接続され、前記 ディスクアレイ装置のボトルネックを検出する監視端末を有するシステムであって、 前記ディスクアレイ装置あるいは前記サーバは、前記サーバから前記ディスクアレイ 装置に対して発行される 10要求の数と各 10要求を処理するのに要した時間と該ディ スクアレイ装置に含まれる資源毎の資源使用率を含むパフォーマンス情報を算出し て前記監視端末に定期的に通知し、
前記監視端末は、前記定期的に通知されるパフォーマンス情報に含まれる前記処 理時間を前記 10要求数で割った平均応答時間が第一の閾値を超える期間が、第一 の所定期間を超える時刻を基準点とし、前記基準点以前の第二の所定期間に占め る、前記資源使用率が前記資源毎に設定された第二の閾値を超える期間の割合が 、所定の割合を超える場合に、該資源をボトルネックと特定することを特徴とするシス テム。
[2] 請求項 1において、
前記監視端末は、前記平均応答時間が前記第一の閾値を越える期間が、連続し て前記第一の所定期間を超える時刻を基準点とすることを特徴とするシステム。
[3] 請求項 1において、
前記監視端末は、前記平均応答時間が前記第一の閾値を超える期間を第三の所 定期間累積した結果が、前記第一の所定期間を超える時刻を基準点とすることを特 徴とするシステム。
[4] 請求項 3において、
前記監視端末は、前記第三の所定期間毎に前記累積結果を求めることを特徴とす るシステム。
[5] 請求項 3において、
前記監視端末は、前記第三の所定期間より短い間隔で前記累積結果を求めること を特徴
[6] 請求項 3において、
前記監視端末は、前記第三の所定期間内に前記平均応答時間が、前記第一の閾 値より低い第三の閾値を下回った場合、累積された期間を一旦ゼロにリセットすること を特徴とするシステム。
[7] 請求項 1において、
前記監視端末は、前記基準点以前であって、更に前記平均応答時間が第四の閾 値を超えた期間である第四の所定期間に占める、前記資源使用率が前記資源毎に 設定された前記第二の閾値を超える期間の割合が、前記所定の割合を超える場合 に、該資源をボトルネックと特定することを特徴とするシステム。
[8] ネットワークを介してクライアント端末にサービスを提供するサーバと、前記サーバ および前記ネットワークに接続され、前記サーバが使用するデータが格納されるディ スクアレイ装置とを有するシステムに含まれ、該ネットワークを介して前記ディスクァレ ィ装置に接続された端末にて実行されるプログラムであって、
HU記端末に、
前記サーバあるいは前記ディスクアレイ装置により定期的に通知される、前記ディス クアレイ装置に対して前記サーバ力 発行される 10要求の数と各 10要求の処理に要 した時間と該ディスクアレイ装置に含まれる資源毎の資源使用率を含むパフオーマン ス情報を受信させ、
前記受信したパフォーマンス情報に含まれる前記処理時間を前記 10要求数で割つ た平均応答時間が第一の閾値を超える期間が、第一の所定期間を超える時刻を基 準点とし、前記基準点以前の第二の所定期間に占める、前記資源使用率が前記資 源毎に設定された第二の閾値を超える期間の割合が、所定の割合を超える場合に、 該資源をボトルネックと特定させることを特徴とするプログラム。
[9] 請求項 8において、
前記基準点は、前記平均応答時間が前記第一の閾値を越える期間が、連続して 前記第一の所定期間を超える時刻であることを特徴とするプログラム。
[10] 請求項 8において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間を第三の所定 期間累積した結果が、前記第一の所定期間を超える時刻であることを特徴とするプロ グラム。
[11] 請求項 10において、
前記第三の所定期間毎に前記累積結果が求められることを特徴とするプログラム。
[12] 請求項 10において、
前記第三の所定期間より短い間隔で前記累積結果を求めることを特徴とするプログ ラム。
[13] 請求項 10において、
前記第三の所定期間内に前記平均応答時間が、前記第一の閾値より低い第三の 閾値を下回った場合、累積された期間が一旦ゼロにリセットされることを特徴とするプ ログラム。
[14] 請求項 8において、
前記基準点以前の第二の所定期間に占める、前記資源使用率が前記資源毎に設 定された第二の閾値を超える期間の割合が、所定の割合を超える場合の代わりに、 前記基準点以前であって、更に前記平均応答時間が第四の閾値を超えた期間であ る第四の所定期間に占める、前記資源使用率が前記資源毎に設定された前記第二 の閾値を超える期間の割合が、前記所定の割合を超える場合に、該資源をボトルネ ックと特定させることを特徴とするプログラム。
[15] ネットワークを介してクライアント端末にサービスを提供するサーバと、前記サーバ および前記ネットワークに接続され、前記サーバが使用するデータが格納されるディ スクアレイ装置と、前記ネットワークを介して前記ディスクアレイ装置に接続され、前記 ディスクアレイ装置のボトルネックを検出する監視端末を有するシステムであって、 前記ディスクアレイ装置あるいは前記サーバは、前記サーバから前記ディスクアレイ 装置に対して発行される 1〇要求の数と各 10要求を処理するのに要した時間と該ディ スクアレイ装置に含まれる資源毎の資源使用率を含むパフォーマンス情報を算出し て前記監視端末に定期的に通知し、
前記監視端末は、前記定期的に通知されるパフォーマンス情報に含まれる前記処 理時間を前記 1〇要求数で割った平均応答時間が第一の閾値を超える期間に基づき 基準点となる時間を決定し、前記基準点以前の第一の所定期間に占める、前記資源 使用率が前記資源毎に設定された第二の閾値を超える期間の割合が、所定の割合 を超える場合に、該資源をボトルネックと特定することを特徴とするシステム。
[16] 請求項 15において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間が、連続して 第二の所定期間を超える時刻であることを特徴とするシステム。
[17] 請求項 15において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間を第三の所定 期間累積した合計が第二の所定期間を超える時刻であることを特徴とするシステム。
[18] 請求項 15において、
前記基準点は、前記平均応答時間が連続して前記第一の閾値を超える期間にお いて、時間を横軸に、前記平均応答時間を縦軸に配置し、前記時間に対する前記 平均応答時間をプロットしてできる波形と、前記平均応答時間が前記第一の閾値を 示す横線とで囲まれる部分の面積が、所定の面積を超える時刻であることを特徴とす るシステム。
[19] 請求項 15において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間において、時 間を横軸に、前記平均応答時間を縦軸に配置し、前記時間に対する前記平均応答 時間をプロットしてできる波形と、前記平均応答時間が前記第一の閾値を示す横線と で囲まれる部分の面積を第三の所定期間累積した合計が、所定の面積を超える時 刻であることを特徴とするシステム。
[20] 請求項 17又は 19において、
前記第三の所定期間毎に前記累積合計が求められることを特徴とするシステム。
[21] 請求項 17又は 19において、
前記第三の所定期間より短い間隔で前記累積合計が求められることを特徴とする システム。
[22] 請求項 17又は 19において、
前記監視端末は、前記第三の所定期間内に前記平均応答時間が、前記第一の閾 値より低い第三の閾値を下回った場合、前記累積合計が一旦ゼロにリセットされるこ とを特徴とするシステム。
[23] 請求項 15において、
前記監視端末は、前記基準点以前であって、更に前記平均応答時間が第四の閾 値を超えた期間である第四の所定期間に占める、前記資源使用率が前記資源毎に 設定された前記第二の閾値を超える期間の割合が、前記所定の割合を超える場合 に、該資源をボトルネックと特定することを特徴とするシステム。
[24] ネットワークを介してクライアント端末にサービスを提供するサーバと、前記サーバ および前記ネットワークに接続され、前記サーバが使用するデータが格納されるディ スクアレイ装置とを有するシステムに含まれ、該ネットワークを介して前記ディスクァレ ィ装置に接続された端末にて実行されるプログラムであって、
HU記端末に、
前記サーバあるいは前記ディスクアレイ装置により定期的に通知される、前記ディス クアレイ装置に対して前記サーバ力 発行される 10要求の数と各 10要求の処理に要 した時間と該ディスクアレイ装置に含まれる資源毎の資源使用率を含むパフオーマン ス情報を受信させ、
前記受信したパフォーマンス情報に含まれる前記処理時間を前記 10要求数で割つ た平均応答時間が第一の閾値を超える期間に基づき基準点となる時間を決定させ、 前記基準点以前の第一の所定期間に占める、前記資源使用率が前記資源毎に設 定された第二の閾値を超える期間の割合が、所定の割合を超える場合に、該資源を ボトルネックと特定させることを特徴とするプログラム。
[25] 請求項 24において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間が、連続して 第二の所定期間を超える時刻であることを特徴とするプログラム。
[26] 請求項 24において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間を第三の所定 期間累積した合計が第二の所定期間を超える時刻であることを特徴とするプログラム
[27] 請求項 24において、
前記基準点は、前記平均応答時間が連続して前記第一の閾値を超える期間にお いて、時間を横軸に、前記平均応答時間を縦軸に配置し、前記時間に対する前記 平均応答時間をプロットしてできる波形と、前記平均応答時間が前記第一の閾値を 示す横線とで囲まれる部分の面積が、所定の面積を超える時刻であることを特徴とす るプログラム。
[28] 請求項 24において、
前記基準点は、前記平均応答時間が前記第一の閾値を超える期間において、時 間を横軸に、前記平均応答時間を縦軸に配置し、前記時間に対する前記平均応答 時間をプロットしてできる波形と、前記平均応答時間が前記第一の閾値を示す横線と で囲まれる部分の面積を第三の所定期間累積した合計が、所定の面積を超える時 刻であることを特徴とするプログラム。
[29] 請求項 26又は 28において、
前記第三の所定期間毎に前記累積合計が求められることを特徴とするプログラム。
[30] 請求項 26又は 28において、
前記第三の所定期間より短い間隔で前記累積結果が求められることを特徴とする
[31] 請求項 26又は 28において、
前記第三の所定期間内に前記平均応答時間が、前記第一の閾値より低い第三の 閾値を下回った場合、前記累積合計が一旦ゼロにリセットされることを特徴とするプロ グラム。
[32] 請求項 24において、
前記基準点以前であって、更に前記平均応答時間が第四の閾値を超えた期間で ある第四の所定期間に占める、前記資源使用率が前記資源毎に設定された前記第 二の閾値を超える期間の割合が、前記所定の割合を超える場合に、該資源をボトノレ ネックと特定させることを特徴とするプログラム。
PCT/JP2004/011780 2003-08-19 2004-08-17 ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム Ceased WO2005017736A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2005513194A JPWO2005017736A1 (ja) 2003-08-19 2004-08-17 ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム
US11/321,578 US20060106926A1 (en) 2003-08-19 2005-12-29 System and program for detecting disk array device bottlenecks

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JPPCT/JP03/10425 2003-08-19
PCT/JP2003/010425 WO2005017735A1 (ja) 2003-08-19 2003-08-19 ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US11/321,578 Continuation US20060106926A1 (en) 2003-08-19 2005-12-29 System and program for detecting disk array device bottlenecks

Publications (1)

Publication Number Publication Date
WO2005017736A1 true WO2005017736A1 (ja) 2005-02-24

Family

ID=34179399

Family Applications (2)

Application Number Title Priority Date Filing Date
PCT/JP2003/010425 Ceased WO2005017735A1 (ja) 2003-08-19 2003-08-19 ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム
PCT/JP2004/011780 Ceased WO2005017736A1 (ja) 2003-08-19 2004-08-17 ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム

Family Applications Before (1)

Application Number Title Priority Date Filing Date
PCT/JP2003/010425 Ceased WO2005017735A1 (ja) 2003-08-19 2003-08-19 ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム

Country Status (3)

Country Link
US (1) US20060106926A1 (ja)
JP (1) JPWO2005017736A1 (ja)
WO (2) WO2005017735A1 (ja)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008108084A1 (ja) * 2007-03-02 2008-09-12 Panasonic Corporation 再生装置、システムlsi、初期化方法
JP2009187324A (ja) * 2008-02-06 2009-08-20 Nec Corp ファイル保管装置、ファイル保管方法およびプログラム
CN106354590A (zh) * 2015-07-17 2017-01-25 中兴通讯股份有限公司 磁盘检测方法和装置

Families Citing this family (21)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8645922B2 (en) * 2008-11-25 2014-02-04 Sap Ag System and method of implementing a concurrency profiler
JP2012068880A (ja) * 2010-09-22 2012-04-05 Fujitsu Ltd 管理プログラム、管理装置および管理方法
US9251032B2 (en) * 2011-11-03 2016-02-02 Fujitsu Limited Method, computer program, and information processing apparatus for analyzing performance of computer system
CN103379041B (zh) 2012-04-28 2018-04-20 国际商业机器公司 一种系统检测方法和装置以及流量控制方法和设备
US8954546B2 (en) 2013-01-25 2015-02-10 Concurix Corporation Tracing with a workload distributor
US20130283281A1 (en) 2013-02-12 2013-10-24 Concurix Corporation Deploying Trace Objectives using Cost Analyses
US8997063B2 (en) 2013-02-12 2015-03-31 Concurix Corporation Periodicity optimization in an automated tracing system
US8924941B2 (en) 2013-02-12 2014-12-30 Concurix Corporation Optimization analysis using similar frequencies
US20130219372A1 (en) 2013-03-15 2013-08-22 Concurix Corporation Runtime Settings Derived from Relationships Identified in Tracer Data
US9575874B2 (en) 2013-04-20 2017-02-21 Microsoft Technology Licensing, Llc Error list and bug report analysis for configuring an application tracer
US9495199B2 (en) * 2013-08-26 2016-11-15 International Business Machines Corporation Management of bottlenecks in database systems
US9292415B2 (en) 2013-09-04 2016-03-22 Microsoft Technology Licensing, Llc Module specific tracing in a shared module environment
CN103500143B (zh) * 2013-09-27 2016-08-10 华为技术有限公司 硬盘参数调整方法及装置
WO2015071778A1 (en) 2013-11-13 2015-05-21 Concurix Corporation Application execution path tracing with configurable origin definition
US9471375B2 (en) 2013-12-19 2016-10-18 International Business Machines Corporation Resource bottleneck identification for multi-stage workflows processing
CN103810062B (zh) * 2014-03-05 2015-12-30 华为技术有限公司 慢盘检测方法和装置
US20160080229A1 (en) * 2014-03-11 2016-03-17 Hitachi, Ltd. Application performance monitoring method and device
CN106407051B (zh) * 2015-07-31 2019-01-11 华为技术有限公司 一种检测慢盘的方法及装置
CN106407052B (zh) * 2015-07-31 2019-09-13 华为技术有限公司 一种检测磁盘的方法及装置
CN107832202A (zh) * 2017-11-06 2018-03-23 郑州云海信息技术有限公司 一种检测硬盘的方法、装置及计算机可读存储介质
CN114116659B (zh) * 2020-08-26 2024-11-22 白腊梅 一种数据库服务端运行状态的检测方法及检测装置

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5851362A (ja) * 1981-09-24 1983-03-26 Hitachi Ltd 計算機システムの性能予測方式
JP2002082926A (ja) * 2000-09-06 2002-03-22 Nippon Telegr & Teleph Corp <Ntt> 分散アプリケーション試験・運用管理システム
JP2003177963A (ja) * 2001-12-12 2003-06-27 Hitachi Ltd ストレージ装置

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6859783B2 (en) * 1995-12-29 2005-02-22 Worldcom, Inc. Integrated interface for web based customer care and trouble management
US6314465B1 (en) * 1999-03-11 2001-11-06 Lucent Technologies Inc. Method and apparatus for load sharing on a wide area network
US7441045B2 (en) * 1999-12-13 2008-10-21 F5 Networks, Inc. Method and system for balancing load distribution on a wide area network
AU7001701A (en) * 2000-06-21 2002-01-02 Concord Communications Inc Liveexception system
US20010054097A1 (en) * 2000-12-21 2001-12-20 Steven Chafe Monitoring and reporting of communications line traffic information
US6961794B2 (en) * 2001-09-21 2005-11-01 International Business Machines Corporation System and method for analyzing and optimizing computer system performance utilizing observed time performance measures
US20030135609A1 (en) * 2002-01-16 2003-07-17 Sun Microsystems, Inc. Method, system, and program for determining a modification of a system resource configuration

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5851362A (ja) * 1981-09-24 1983-03-26 Hitachi Ltd 計算機システムの性能予測方式
JP2002082926A (ja) * 2000-09-06 2002-03-22 Nippon Telegr & Teleph Corp <Ntt> 分散アプリケーション試験・運用管理システム
JP2003177963A (ja) * 2001-12-12 2003-06-27 Hitachi Ltd ストレージ装置

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008108084A1 (ja) * 2007-03-02 2008-09-12 Panasonic Corporation 再生装置、システムlsi、初期化方法
JPWO2008108084A1 (ja) * 2007-03-02 2010-06-10 パナソニック株式会社 再生装置、システムlsi、初期化方法
CN101589369B (zh) * 2007-03-02 2013-01-23 松下电器产业株式会社 再生装置、系统lsi、初始化方法
US8522339B2 (en) 2007-03-02 2013-08-27 Panasonic Corporation Reproducing apparatus, system LSI, and initialization method
KR101430279B1 (ko) 2007-03-02 2014-08-14 파나소닉 주식회사 재생장치, 시스템 lsi, 초기화방법
JP2009187324A (ja) * 2008-02-06 2009-08-20 Nec Corp ファイル保管装置、ファイル保管方法およびプログラム
CN106354590A (zh) * 2015-07-17 2017-01-25 中兴通讯股份有限公司 磁盘检测方法和装置

Also Published As

Publication number Publication date
US20060106926A1 (en) 2006-05-18
WO2005017735A1 (ja) 2005-02-24
JPWO2005017736A1 (ja) 2007-11-01

Similar Documents

Publication Publication Date Title
WO2005017736A1 (ja) ディスクアレイ装置におけるボトルネックを検出するシステムおよびプログラム
US7653725B2 (en) Management system selectively monitoring and storing additional performance data only when detecting addition or removal of resources
US8645185B2 (en) Load balanced profiling
US20170155560A1 (en) Management systems for managing resources of servers and management methods thereof
CN102356388B (zh) web前端节流
US8977908B2 (en) Method and apparatus for detecting a suspect memory leak
US9027025B2 (en) Real-time database exception monitoring tool using instance eviction data
US10305974B2 (en) Ranking system
CN107659431A (zh) 接口处理方法、装置、存储介质和处理器
CN104778111A (zh) 一种进行报警的方法和装置
CN106685752B (zh) 一种信息处理方法及终端
KR102456150B1 (ko) 실제 환경에서 대용량 시스템에 대한 포괄적 성능평가를 수행하는 방법 및 이를 지원하는 장치
CN112052088A (zh) 自适应的进程cpu资源限制方法、装置、终端及存储介质
CN110674149B (zh) 业务数据处理方法、装置、计算机设备和存储介质
JP4387970B2 (ja) データ入出力プログラム,装置,および方法
CN106441349B (zh) 基于计步器消息的伪造消息判定方法及装置
WO2012087104A1 (en) Intelligent load handling in cloud infrastructure using trend analysis
CN105357026B (zh) 一种资源信息收集方法和计算节点
CN112579396A (zh) 软件系统动态限流方法、装置及设备
CN115150460B (zh) 一种节点安全注册方法、装置、设备及可读存储介质
CN106951318A (zh) 一种电子设备后台进程的管理方法及电子设备
CN108804152B (zh) 配置参数的调节方法及装置
US8976803B2 (en) Monitoring resource congestion in a network processor
CN110727518B (zh) 一种数据处理方法及相关设备
CN113791950A (zh) 一种服务程序的信息处理方法、装置、服务器及存储介质

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A1

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): BW GH GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 2005513194

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 11321578

Country of ref document: US

WWP Wipo information: published in national office

Ref document number: 11321578

Country of ref document: US

122 Ep: pct application non-entry in european phase