WO2025035892A1 - 故障定位方法、设备、装置、存储介质及电子设备 - Google Patents
故障定位方法、设备、装置、存储介质及电子设备 Download PDFInfo
- Publication number
- WO2025035892A1 WO2025035892A1 PCT/CN2024/095818 CN2024095818W WO2025035892A1 WO 2025035892 A1 WO2025035892 A1 WO 2025035892A1 CN 2024095818 W CN2024095818 W CN 2024095818W WO 2025035892 A1 WO2025035892 A1 WO 2025035892A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- preset
- link
- target
- port
- hard disk
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/0706—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation the processing taking place on a specific hardware platform or in a specific software environment
- G06F11/0727—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation the processing taking place on a specific hardware platform or in a specific software environment in a storage system, e.g. in a DASD or network based storage system
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/079—Root cause analysis, i.e. error or fault diagnosis
Definitions
- the present application relates to the field of computers, and in particular, to a fault location method, device, apparatus, non-volatile readable storage medium, and electronic device.
- Hardware link failure is a common hardware failure, which can be manifested as a failure in one or several sections of a very long hardware link.
- the existing technology cannot accurately identify the fault location or range on a very long link, so intelligently repairing the entire link will increase the maintenance scope and reduce the maintenance efficiency.
- the embodiments of the present application provide a fault location method, device, apparatus, non-volatile readable storage medium and electronic device to at least solve the technical problem that the prior art cannot accurately locate the fault position in a link.
- a fault location method comprising: obtaining the login status of at least one target port in a target link with a fault, wherein the target link is used to connect to the target port of at least one hard disk in a preset link order; querying an impact factor corresponding to each login status in a preset set, wherein the preset set is used to record a plurality of login statuses and an impact factor corresponding to each login status; accumulating the impact factor corresponding to each target port one by one in the preset link order, and comparing each accumulated result with a preset location threshold; when the accumulated result is greater than the preset location threshold, determining that the target port corresponding to the last accumulated impact factor is a faulty port.
- the method before obtaining the login status of at least one target port in the target link with a fault, also includes: obtaining the login status of at least one preset port in a preset link, wherein the preset link is used to connect to a preset port of at least one hard disk; querying an impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and an impact factor corresponding to each login status; accumulating the impact factors of all preset ports in the preset link to determine an evaluation result of the preset link; when the evaluation result is greater than a preset fault threshold, determining that the preset link is a target link with a fault.
- the method before obtaining the login status of at least one preset port in the preset link, the method also includes: sending broadcast information to the hard disk set to be detected, wherein the hard disk set to be detected includes at least one hard disk, and each hard disk includes two preset ports; receiving location information returned by the hard disk set to be detected, wherein the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis; determining the preset link based on the preset chassis information and the preset controller information, wherein the preset link is used to cascade controllers with the same identifier in multiple chassis.
- obtaining the login status of at least one preset port in a preset link includes: detecting frame loss information of each preset port, wherein the frame loss information includes at least the number of frame losses; determining a preset frame loss interval corresponding to the number of frame losses in a preset interval set as a target frame loss interval, wherein the preset interval set includes multiple preset frame loss intervals, and each preset frame loss interval is pre-set with a corresponding preset frame loss status; determining a target frame loss status corresponding to the target frame loss area as the login status.
- the method after accumulating the influence factors of all preset ports in the preset link and determining the evaluation result of the preset link, the method also includes: judging whether the evaluation result is greater than a preset redundancy threshold; when the evaluation result is greater than the preset redundancy threshold, determining that the preset link is an abnormal link; determining the total link of the abnormal link, wherein the total link includes two preset links; adjusting the redundant mark of the total link, wherein the redundant mark includes: a first redundant mark indicating that an abnormal link exists in the total link, and a second redundant mark indicating that no abnormal link exists in the total link.
- the method further includes: obtaining target controller information of the faulty port, wherein the target controller information is a unique identifier of the controller connected to the faulty port; obtaining target chassis information of the hard disk where the faulty port is located, wherein the target chassis information is a unique identifier of the chassis of the hard disk where the faulty port is located; and generating alarm information based on the target controller information and the target chassis information.
- the method further includes: The invention includes: determining the upstream port of the faulty port, wherein the upstream port is used to transmit data for the faulty port; obtaining the upstream controller information of the upstream port, wherein the upstream controller information is a unique identifier of the controller connected to the upstream port; obtaining the upstream chassis information of the hard disk where the upstream port is located, wherein the upstream chassis information is a unique identifier of the chassis of the hard disk where the upstream port is located; and generating alarm information according to the target controller information, the target chassis information, the upstream controller information and the upstream chassis information.
- the method further includes: displaying the alarm information through a graphical user interface.
- the method further includes:
- the hard disk where the target port is located is isolated.
- isolating the hard disk where the target port is located includes: determining the hard disk where the target port is located as the target hard disk; setting the target hard disk to a closed state, wherein when the target hard disk is in the closed state, all ports of the target hard disk are disconnected.
- the method further includes: obtaining a redundant mark of the total link where the target link is located, wherein the redundant mark includes: a first redundant mark indicating that there is an abnormal link in the total link, and a second redundant mark indicating that there is no abnormal link in the total link; when the second redundant mark exists in the total link, performing fault isolation on the target controller that controls the faulty port.
- a fault location device includes: a hard disk, a serial bus and a chassis manager; the serial bus is configured to connect to the target port of at least one hard disk in the target link according to a preset link order, wherein the target link is a link with a fault; the chassis manager is configured to obtain the login status of the target port, query the impact factor corresponding to each login status in a preset set, accumulate the impact factor corresponding to each target port one by one according to the preset link order, and compare each accumulated result with a preset positioning threshold; when the accumulated result is greater than the preset positioning threshold, determine that the target port corresponding to the last accumulated impact factor is a faulty port, wherein the preset set is used to record multiple login states and the impact factor corresponding to each login state.
- the device further includes: a controller configured to connect the serial bus to a preset port of the hard disk.
- the device also includes: a bus control chip, configured to publish broadcast events to at least one hard disk and a chassis manager; the chassis manager is also configured to initiate a discovery mechanism to obtain the location information of each hard disk after receiving the broadcast, wherein the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis.
- the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis.
- the chassis manager is further configured to, after initiating the discovery mechanism, subscribe to the frame loss information of each preset port, wherein the frame loss information includes at least the number of frame losses; determine a preset frame loss interval corresponding to the number of frame losses in a preset interval set as a target frame loss interval, wherein the preset interval set includes multiple preset frame loss intervals, each preset frame loss interval is pre-set with a corresponding preset frame loss state; and determine the target frame loss state corresponding to the target frame loss area as a login state.
- the preset port of the hard disk includes: a first preset port and a second preset port;
- the serial bus includes: a first serial bus and a second serial bus; wherein the first serial bus is configured to connect to the first preset port; and the second serial bus is configured to connect to the second preset port.
- the chassis manager is further configured to obtain the login status of at least one preset port in a preset link, wherein the preset link is a link of the first serial bus or a link of the second serial bus; query the impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; accumulate the impact factors of all preset ports in the preset link to determine an evaluation result of the preset link; when the evaluation result is greater than a preset redundancy threshold, determine that the preset link is an abnormal link; adjust the redundant mark of the total link where the abnormal link is located, wherein the total link includes the first serial bus and the second serial bus, and the redundant mark includes: a first redundant mark indicating that an abnormal link exists in the total link, and a second redundant mark indicating that no abnormal link exists in the total link.
- the device also includes: a timer, configured to start when the login status indicates that there is an abnormality in the target port, and set to determine the target time according to a preset time interval; a chassis manager, further configured to obtain the login status of the target port at the target time; and when the login status of the target port at the target time still indicates that there is an abnormality in the target port, isolating the hard disk where the target port is located.
- a timer configured to start when the login status indicates that there is an abnormality in the target port, and set to determine the target time according to a preset time interval
- a chassis manager further configured to obtain the login status of the target port at the target time; and when the login status of the target port at the target time still indicates that there is an abnormality in the target port, isolating the hard disk where the target port is located.
- the chassis manager is further configured to, after determining the faulty port in the target link, obtain a redundant mark of the overall link in which the target link is located, wherein the redundant mark includes: a first redundant mark indicating that there is an abnormal link in the overall link, and a second redundant mark indicating that there is no abnormal link in the overall link; when the second redundant mark exists in the overall link, the target controller that controls the faulty port is fault isolated.
- a fault location device including: an acquisition module, configured to acquire the login status of at least one target port in a target link with a fault, wherein the target link is configured to connect the target port of at least one hard disk according to a preset link sequence; a query module, configured to query the impact factor corresponding to each login status in a preset set, wherein the preset set is configured to record multiple login statuses and the impact factor corresponding to each login status; a processing module, configured to accumulate the impact factor corresponding to each target port one by one according to the preset link sequence, and compare each accumulation result with a preset location threshold; a determination module, configured to determine that the target port corresponding to the last accumulated impact factor is a faulty port when the accumulation result is greater than the preset location threshold.
- a non-volatile readable storage medium is further provided.
- the non-volatile readable storage medium is configured to store a program, wherein when the program is running, the device where the non-volatile readable storage medium is located is controlled to execute the above-mentioned fault location method.
- an electronic device including: a memory and a processor, wherein the processor is configured to run a program stored in the processor, wherein the above-mentioned fault location method is executed when the program is run.
- the login status of at least one target port in a target link having a fault is obtained, wherein the target link is used to The target port of at least one hard disk is connected in a preset link sequence; the impact factor corresponding to each login state is queried in a preset set, wherein the preset set is used to record multiple login states and the impact factor corresponding to each login state; the impact factor corresponding to each target port is accumulated one by one in the preset link sequence, and each accumulated result is compared with a preset positioning threshold; when the accumulated result is greater than the preset positioning threshold, the target port corresponding to the last accumulated impact factor is determined to be a faulty port, and the faulty port where the fault occurs can be determined in the target link where the fault exists, thereby achieving the technical effect of accurately locating the fault position of the faulty target link, thereby solving the technical problem that the prior art cannot accurately locate the fault position in the link.
- FIG1 is a flow chart of a fault location method according to an embodiment of the present application.
- FIG. 2 is a schematic diagram of a storage hardware link status monitoring method architecture based on hard disk command frame loss according to an embodiment of the present application
- FIG3 is a schematic diagram of a storage hardware link according to an embodiment of the present application.
- FIG4 is a schematic diagram of a fault location device according to an embodiment of the present application.
- FIG5 is a schematic diagram of a fault location device according to an embodiment of the present application.
- FIG6 is a structural block diagram of a computer terminal according to an embodiment of the present application.
- SAS Serial Attached SCSI is a computer hub technology, its main function is to transmit data of peripheral parts, such as hard disk, CD-ROM and other devices designed for the interface.
- SAS Expander is an expander that complies with the SAS protocol and can be used for chassis management.
- SCSISmall Computer System Interface SCSI protocol is mainly used to transmit commands, status and block data between the host and the storage device.
- SCSI protocol is the most important backbone.
- RAID generally refers to disk arrays. Disk arrays (Redundant Arrays of Independent Disks, RAID) means "an array composed of several independent disks with redundancy capabilities.”
- an embodiment of a fault location method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
- FIG. 1 is a flow chart of a fault location method according to an embodiment of the present application. As shown in FIG. 1 , the method includes the following steps:
- Step S102 obtaining a login status of at least one target port in a target link having a fault, wherein the target link is used to connect to a target port of at least one hard disk according to a preset link sequence;
- Step S104 querying the impact factor corresponding to each login state in a preset set, wherein the preset set is used to record multiple login states and the impact factor corresponding to each login state;
- Step S106 accumulating the impact factor corresponding to each target port one by one according to the preset link sequence, and comparing each accumulation result with the preset positioning threshold;
- Step S108 When the accumulated result is greater than the preset positioning threshold, the target port corresponding to the last accumulated impact factor is determined to be a faulty port.
- the login status of at least one target port in a target link with a fault is obtained, wherein the target link is used to connect to the target port of at least one hard disk in a preset link sequence; the impact factor corresponding to each login status is queried in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; the impact factor corresponding to each target port is accumulated one by one in the preset link sequence, and each accumulated result is compared with a preset positioning threshold; when the accumulated result is greater than the preset positioning threshold, the target port corresponding to the last accumulated impact factor is determined to be the faulty port, and the faulty port where the fault occurs can be determined in the target link with the fault, thereby achieving the technical effect of accurately locating the fault position of the faulty target link, thereby solving the problem that the prior art cannot The technical problem of accurately locating the fault location in the link.
- the above fault location method may be used in a storage system having multiple hard disks, in which the multiple hard disks may store data in a disk array manner.
- each hard disk may include two preset ports, which are respectively recorded as a first preset port and a second preset port.
- there may be two preset links namely a first preset link and a second preset link, wherein the first preset link is used to connect a first preset port of each hard disk, and the second preset link is used to establish a second preset port of each hard disk.
- the first preset link may be used to transmit uplink data
- the second preset link is used to transmit downlink data
- the target link is a preset link with a fault
- the target port is a port connected to the preset link with a fault
- the login status may indicate the frame loss situation of the corresponding target port or the preset port.
- the login status includes: GOOD, indicating no frame loss exception; IN DOUBT, indicating that there is at least one frame loss, but it has not reached the level of DEGRADED; DEGRADED, indicating that there is a certain number of frame losses, which is greater than IN DOUBT but has not reached the level of EXCLUSION; indicating that there is a large amount of frame loss in the EXCLUDSION disk.
- the impact factor is used to represent the frame loss degree of the target port or the preset port in a numerical manner.
- the preset positioning threshold is used to locate the position of the fault in the target link.
- the impact factor corresponding to each target port in the target link is accumulated one by one, and each accumulation result is compared with the preset positioning threshold. If the accumulated impact factor corresponding to a certain target port is greater than the preset positioning threshold, it means that there is a fault at the target port accumulated later, and the faulty port is located, which is convenient for subsequent maintenance or isolation of the faulty port.
- the location information of each hard disk and preset port can be recorded in advance, and the location information includes at least: preset chassis information and preset controller information.
- the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located
- the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis. After determining the faulty port, the preset controller information of the faulty port can be queried to locate the faulty port.
- the preset port of each hard disk is connected to the preset link through the controller. Since the controller and the preset port correspond one to one, the faulty port can be located by determining the controller of the faulty port.
- the method before obtaining the login status of at least one target port in a target link with a fault, the method also includes: obtaining the login status of at least one preset port in a preset link, wherein the preset link is used to connect to a preset port of at least one hard disk; querying an impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and an impact factor corresponding to each login status; accumulating the impact factors of all preset ports in the preset link to determine an evaluation result of the preset link; and when the evaluation result is greater than a preset fault threshold, determining that the preset link is a target link with a fault.
- the impact factor corresponding to each preset port in each preset link can be accumulated to obtain the evaluation result of the preset link, and then the evaluation result of the preset link is compared with the preset fault threshold.
- the evaluation result of the preset link is greater than the preset fault threshold, it indicates that the preset link has a fault, and the preset link is determined to be the target link, thereby achieving the preliminary positioning of the target link with the fault.
- the method before obtaining the login status of at least one preset port in a preset link, the method also includes: sending broadcast information to a set of hard disks to be detected, wherein the set of hard disks to be detected includes at least one hard disk, and each hard disk includes two preset ports; receiving location information returned by the set of hard disks to be detected, wherein the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate a unique identifier of a chassis where each hard disk is located, and the preset controller information is used to indicate a unique identifier of a controller in each chassis that is connected to a preset port of the hard disk; determining a preset link based on the preset chassis information and the preset controller information, wherein the preset link is used to cascade controllers with the same identifier in multiple chassis.
- the location information of each hard disk and the preset port of each hard disk can be obtained through broadcast information, and after the faulty port is determined, the location of the faulty port and the hard disk on which it is located can be determined from the location information, thereby facilitating the isolation or repair of the faulty port.
- obtaining the login status of at least one preset port in a preset link includes: detecting frame loss information of each preset port, wherein the frame loss information includes at least the number of frame losses; determining a preset frame loss interval corresponding to the number of frame losses in a preset interval set as a target frame loss interval, wherein the preset interval set includes multiple preset frame loss intervals, and each preset frame loss interval is pre-set with a corresponding preset frame loss status; determining a target frame loss status corresponding to the target frame loss area as a login status.
- the login status represents the frame loss situation of the corresponding target port or preset port. Therefore, the login status can be determined according to the number of frame losses, by pre-dividing multiple preset frame loss intervals and setting a corresponding preset frame loss status for each preset frame loss interval, and then determining the preset frame loss interval corresponding to the number of frame losses of the preset port or the target port as the target frame loss interval, and using the preset frame loss status of the target frame loss interval as the login status of the preset port or the target port, thereby realizing the determination of the login status.
- the method after obtaining the login status of at least one preset port in the preset link, the method also includes: when the login status indicates that there is an abnormality in the preset port, obtaining the login status of the preset port at a target time according to a preset time interval; when the login status of the preset port at the target time still indicates that there is an abnormality in the preset port, isolating the hard disk where the preset port is located.
- the hard disk occasionally has a small amount of frame loss, it can be tolerated by the storage system; however, if the hard disk frequently experiences frame loss, the hard disk needs to be processed. Therefore, when it is detected that the login status of the preset port is abnormal, the login status of the preset port can be obtained again at the target time according to the preset time interval. If the login status of the preset port at the target time still indicates that the preset port is abnormal, it means that the preset port frequently experiences frame loss, and the hard disk where the preset port is located needs to be separated from the storage system to isolate the faulty hard disk.
- isolating the hard disk where the preset port is located includes: determining the hard disk where the preset port is located as the target hard disk; setting the target hard disk to a closed state, wherein when the target hard disk is in the closed state, all ports of the target hard disk are disconnected.
- multiple hard disks can establish a disk array in a RAID manner, and then when a hard disk in the disk array fails, the failed hard disk can be set to a closed state Fail. If a hard disk in the closed state Fail is detected in the RAID array, the hard disk is exited, thereby achieving the technical effect of isolating the failed hard disk.
- the method after accumulating the influence factors of all preset ports in the preset link and determining the evaluation result of the preset link, the method also includes: judging whether the evaluation result is greater than a preset redundancy threshold; when the evaluation result is greater than the preset redundancy threshold, determining that the preset link is an abnormal link; determining the total link of the abnormal link, wherein the total link includes two preset links; adjusting the redundant mark of the total link, wherein the redundant mark includes: a first redundant mark indicating that the total link has an abnormal link, and a second redundant mark indicating that the total link does not have an abnormal link.
- each general link in the storage system can be set with two preset links, and the two preset links are redundant backups of each other. If one of the preset links is abnormal, then the other preset link can be used to continue data transmission. Therefore, it is necessary to understand the redundancy of the total link in a timely manner, and then the evaluation results of the corresponding impact factors of all preset ports in each link in the total link can be compared with the preset redundancy threshold.
- the evaluation result is greater than the preset redundancy threshold, it means that the preset link is abnormal, and it is necessary to call another redundant preset link in the total link to meet the data transmission demand, and add a first redundant mark to the total link; and then if both of the two preset links in the total link are abnormal, it means that the total link can no longer meet the data transmission demand, so it is necessary to shut down and alarm, waiting for maintenance.
- the method after determining that the last accumulated target port is a faulty port, the method also includes: obtaining target controller information of the faulty port, wherein the target controller information is a unique identifier of the controller connected to the faulty port; obtaining target chassis information of the hard disk where the faulty port is located, wherein the target chassis information is a unique identifier of the chassis of the hard disk where the faulty port is located; and generating alarm information based on the target controller information and the target chassis information.
- the method also includes: determining the upstream port of the faulty port, wherein the upstream port is used to transmit data for the faulty port; obtaining the upstream controller information of the upstream port, wherein the upstream controller information is a unique identifier of the upstream port connected to the controller; obtaining the upstream chassis information of the hard disk where the upstream port is located, wherein the upstream chassis information is a unique identifier of the chassis of the hard disk where the upstream port is located; and generating alarm information based on the target controller information, the target chassis information, the upstream controller information and the upstream chassis information.
- the alarm information includes the enclosure_index of the chassis where the controller is located (that is, the target chassis information), the canister_index of the controller (that is, the target controller information), the enclosure_index of the upstream chassis connected to the faulty controller (that is, the upstream chassis information) and the canister_index of the upstream controller (that is, the upstream controller information).
- a line is limited by two points to determine the fault location, thereby achieving accurate positioning of the fault location.
- the method further includes: displaying the alarm information through a graphical user interface.
- the method after obtaining the login status of at least one target port in the faulty target link, the method also includes: when the login status indicates that there is an abnormality in the target port, obtaining the login status of the target port at a target time at a preset time interval; when the login status of the target port at the target time still indicates that there is an abnormality in the target port, isolating the hard disk where the target port is located.
- the hard disk occasionally has a small amount of frame loss, it can be tolerated by the storage system; however, if the hard disk frequently experiences frame loss, the hard disk needs to be processed. Therefore, when it is detected that the login status of the target port is abnormal, the login status of the target port can be obtained again at the target time according to the preset time interval. If the login status of the target port at the target time still indicates that the target port is abnormal, it means that the target port frequently experiences frame loss, and the hard disk where the target port is located needs to be separated from the storage system to isolate the faulty hard disk.
- isolating the hard disk where the target port is located includes: determining that the hard disk where the target port is located is the target hard disk; setting the target hard disk to a closed state, wherein when the target hard disk is in the closed state, all ports of the target hard disk are disconnected.
- multiple hard disks can establish a disk array in a RAID manner, and then when a hard disk in the disk array fails, the failed hard disk can be set to a closed state Fail. If a hard disk in the closed state Fail is detected in the RAID array, the hard disk is exited, thereby achieving the technical effect of isolating the failed hard disk.
- the method further includes: obtaining a redundant mark of the total link where the target link is located, wherein the redundant mark includes: a first redundant mark indicating that there is an abnormal link in the total link, and a second redundant mark indicating that there is no abnormal link in the total link; when the second redundant mark exists in the total link, fault isolation is performed on the target controller that controls the faulty port.
- each main link in the storage system can be set with two preset links, and the two preset links are mutually redundant backups. If an abnormality occurs in one of the preset links, another preset link can be used to continue data transmission. Then, in the process of isolating the fault of the faulty port, it is possible to first determine whether the total link has the redundancy condition that meets the fault isolation based on the redundancy mark of the total link, and then if there is a second redundant mark in the total link indicating that there is no abnormal link in the total link, the fault isolation of the faulty port can be achieved.
- the present application also provides an optional embodiment, which provides a storage hardware link status monitoring method based on hard disk command frame loss. Based on the analysis of the number of hard disks with frame loss and the severity of single-block frame loss, the location of the faulty link can be identified, and corresponding fault isolation measures can be taken according to different link failures, effectively ensuring the stability of the system. When a failure occurs, the cause of the failure can be accurately identified and corresponding processing can be performed, and the fault location or range can be accurately identified on a very long link.
- FIG. 2 is a schematic diagram of the architecture of a storage hardware link status monitoring method based on hard disk command frame loss according to an embodiment of the present application. As shown in Figure 2, it includes: a chassis management layer (EN), a protocol layer, a driver layer, and a physical link layer.
- EN chassis management layer
- protocol layer protocol layer
- driver layer driver layer
- physical link layer physical link layer
- the physical link layer is at the lowest layer and is a physical channel for command or data transmission between the storage system and the hard disk.
- the protocol layer and the driver layer are data channels for communication between the storage system and the hard disk.
- the chassis management layer is located at the top layer and its role is to determine the fault location according to the frame loss degree (ie, login status) of each hard disk returned by the protocol layer, and trigger the fault isolation action.
- the frame loss degree ie, login status
- each hard disk has two preset ports, and the frame loss degree (ie, login status) of each preset port is divided into four levels:
- the disk has frame loss: a certain amount of frame loss, greater than IN DOUBT but less than EXCLUSION, affects EN by a factor of 4.
- EXCLUDSION The disk has a large number of frame drops, which affects EN by a factor of 10.
- the chassis management layer records two attributes (i.e., location information) for each hard disk, the chassis where the hard disk is located and the controller where the hard disk is located.
- location information i.e., location information
- the chassis management layer performs corresponding fault isolation operations for the detected fault location and reports the alarm in a timely manner.
- the storage device can promptly detect whether there is a link failure in the frame, and then perform corresponding troubleshooting on the abnormal link, thereby ensuring the security and reliability of the storage system and guaranteeing product quality and production services.
- the implementation process of the storage hardware link status monitoring method based on hard disk command frame loss includes the following steps:
- Step S1 The storage system starts and runs normally.
- Step S2 SAS Expander chip is a programmable chip that can publish broadcast events.
- the chassis management layer EN initiates Discovery (also known as the discovery mechanism).
- EN can obtain the chassis information (enclosure_index) where the hard disk is located and the controller information (canister_index) corresponding to each port preset port of each hard disk, that is, obtain the location information.
- Step S3 EN subscribes to the frame loss information of each port preset by each hard disk in the protocol layer, and checks the frame loss information after each Discover is completed.
- Step S4 the protocol layer returns the command execution status through the SCSI protocol rules and the driver, and sets the frame loss status of each port preset port of each disk.
- the frame loss status is recorded as the login status of the disk.
- the login status (also known as the login status) includes: GOOD: no frame loss exception; IN DOUBT: frame loss exists: at least one frame loss exists, but it has not reached the level of DEGRADED; DEGRADED: frame loss exists: a certain number of frame losses, which is greater than IN DOUBT but has not reached the level of EXCLUSION; and EXCLUDSION: a large number of frame losses exist.
- Step S6 determining a preset link Strand and a total link Chain, wherein each total link Chain includes two preset links Strand.
- FIG3 is a schematic diagram of a storage hardware link according to an embodiment of the present application.
- enclosure0 is the main cabinet
- the hardware link from the same SAS port of the main cabinet to the bottom of the cascade is recorded as a preset link Strand, and each disk corresponds to a preset port port.
- Two preset link Strands symmetrical between the upper and lower controls are recorded as a total link Chain.
- Step S7 when EN finds that multiple disks on a strand have abnormal logins, it needs to accumulate the impact factors corresponding to each disk login. This requires setting three thresholds, namely, preset fault threshold T1, preset location threshold T2, and preset redundancy threshold T3.
- the preset fault threshold T1 is the cumulative value of the impact factors of the login status (ie, login status) of all corresponding hard disks on the entire preset link Strand and the color number node port, which must be at least greater than 30 to meet the simultaneous EXCLUDSION of at least three disks.
- the preset positioning threshold T2 is the cumulative value of the influencing factor of the login status (i.e., login status) of the preset port port of the hard disk corresponding to a controller on the current preset link Strand.
- T2 is smaller than T1 and can be calculated as 6 points to find the location of the first fault point on a single preset link Strand.
- the preset redundancy threshold T3 is the total link Chain threshold, and T3 is smaller than T1, because the total link Chain represents a pair of preset links.
- Strand when the login status (i.e. login status) of the two preset links Strand is abnormal, it means that there is no redundant link, and the storage system is extremely unsafe at this time. Therefore, T3 is lower than T1, and the total link Chain abnormal alarm is reported first.
- the cumulative value of the impact factor of the disk login status (i.e. login status) of each preset link Strand on the two preset links Strand on a total link Chain is greater than T3, the variable (i.e. redundant mark) bad_chain+1 (the initial value of bad_chain is 0).
- Step S8 After the statistics of the accumulated values of the impact factors of login (ie, login status) on each preset link Strand are completed, link fault abnormal point judgment begins.
- Step S9 when the login status (i.e. login status) of the disk on the preset link Strand is greater than T1, it means that multiple disks have login (i.e. login status) abnormalities.
- the probability of multiple disks (at least 3 disks) losing frames at the same time is very small, so it can be considered that there is a problem with a certain section on the preset link Strand.
- Step S10 traverse the login status (i.e. login status) of the disks corresponding to all controllers on the preset link Strand found to be faulty in step S9 and compare them with T2. Find the first controller whose cumulative factor of the corresponding disk is greater than T2, and this controller is the controller that first fails (corresponding to the faulty port).
- login status i.e. login status
- Step S11 After the first faulty controller is found, the controller needs to be isolated for fault.
- step S7 if the value of bad_chain in step S7 is less than 1, the physical link of the controller's uplink port is found. At this time, the controller and the downstream controller on the same preset link Strand as the controller will no longer have the hard disk frame loss problem, thereby achieving the purpose of fault isolation. Then, an alarm of frame loss of the controller is reported.
- the alarm information must include the enclosure_index of the chassis where the controller is located, the canister_index of the controller, the enclosure_index of the upstream chassis connected to the faulty controller, and the canister_index of the upstream controller. A line is limited by two points in the alarm information to determine the fault location.
- step S7 if the value of bad_chain in step S7 is greater than 1, it means that the storage link has no redundancy at this time, and only the above alarm needs to be reported.
- Step S12 when EN finds that only one disk on a preset link Strand has a login exception (a single disk does not distinguish login status), it considers that the disk with login exception is a problem of itself.
- the storage system allows occasional disk login exceptions, but cannot tolerate disk login exceptions all the time. Therefore, when a single disk login exception is found for the first time, EN will perform anti-shake processing, start a 1h timer, and record the disk with login exception in the context of the EN module.
- the latest login status i.e., login status
- the latest login status i.e., login status
- Step S13 through the above steps, the storage system can automatically identify whether it is a point failure (disk) or a line failure (disk, controller) when a link failure occurs, and perform fault isolation on the corresponding hardware to prevent more software anomalies from occurring and ensure storage stability.
- a point failure disk
- a line failure disk, controller
- Step S14 All the above-mentioned alarms are displayed to the user through a GUI interface.
- Figure 4 is a schematic diagram of a fault location device according to an embodiment of the present application.
- the device may include: a hard disk 42, a serial bus 44 and a chassis manager 46; the serial bus is used to connect the target port of at least one hard disk in the target link according to a preset link order, wherein the target link is a link with a fault; the chassis manager is used to obtain the login status of the target port, query the impact factor corresponding to each login status in the preset set, accumulate the impact factor corresponding to each target port one by one according to the preset link order, and compare each accumulated result with the preset location threshold; when the accumulated result is greater than the preset location threshold, determine that the target port corresponding to the last accumulated impact factor is a faulty port, wherein the preset set is used to record multiple login states and the impact factor corresponding to each login state.
- the login status of at least one target port in a target link with a fault is obtained, wherein the target link is used to connect to the target port of at least one hard disk in a preset link order; the impact factor corresponding to each login status is queried in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; the impact factor corresponding to each target port is accumulated one by one in the preset link order, and each accumulated result is compared with a preset positioning threshold; when the accumulated result is greater than the preset positioning threshold, the target port corresponding to the last accumulated impact factor is determined to be the faulty port, and the faulty port where the fault occurs can be determined in the target link with the fault, thereby achieving the technical effect of accurately locating the fault position of the faulty target link, thereby solving the technical problem that the prior art cannot accurately locate the fault position in the link.
- the device further includes: a controller, configured to connect the serial bus to a preset port of the hard disk.
- the device also includes: a bus control chip, used to publish broadcast events to at least one hard disk and a chassis manager; the chassis manager is also used to initiate a discovery mechanism to obtain the location information of each hard disk after receiving the broadcast, wherein the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis.
- the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis.
- the chassis manager is also used to subscribe to the frame loss information of each preset port after initiating the discovery mechanism, wherein the frame loss information includes at least the number of frame losses; determine the preset frame loss interval corresponding to the number of frame losses in the preset interval set as the target frame loss interval, wherein the preset interval set includes multiple preset frame loss intervals, and each preset frame loss interval is pre-set with a corresponding preset frame loss state; determine the target frame loss state corresponding to the target frame loss area as the login state.
- the preset port of the hard disk includes: a first preset port and a second preset port;
- the serial bus includes: a first serial bus and a second serial bus; wherein the first serial bus is used to connect to the first preset port; and the second serial bus is used to connect to the second preset port.
- the chassis manager is also used to obtain the login status of at least one preset port in a preset link, wherein the preset link is a link of a first serial bus or a link of a second serial bus; query the impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login states and the impact factor corresponding to each login state; accumulate the impact factors of all preset ports in the preset link to determine an evaluation result of the preset link; when the evaluation result is greater than a preset redundancy threshold, determine that the preset link is an abnormal link; adjust the redundant mark of the total link where the abnormal link is located, wherein the total link includes the first serial bus and the second serial bus, and the redundant mark includes: a first redundant mark indicating that an abnormal link exists in the total link, and a second redundant mark indicating that no abnormal link exists in the total link.
- the device also includes: a timer, which is started when the login status indicates that there is an abnormality in the target port, and is set to determine the target time according to a preset time interval; a chassis manager, which is also used to obtain the login status of the target port at the target time; and when the login status of the target port at the target time still indicates that there is an abnormality in the target port, the hard disk where the target port is located is isolated.
- the chassis manager is further used to obtain a redundant mark of the overall link where the target link is located after determining the faulty port in the target link, wherein the redundant mark includes: a first redundant mark indicating that there is an abnormal link in the overall link, and a second redundant mark indicating that there is no abnormal link in the overall link; when the second redundant mark exists in the overall link, the target controller that controls the faulty port is fault isolated.
- a fault location device embodiment is also provided. It should be noted that the fault location device can be used to execute the fault location method in the embodiment of the present application, and the fault location method in the embodiment of the present application can be executed in the fault location device.
- Figure 5 is a schematic diagram of a fault location device according to an embodiment of the present application.
- the device may include: an acquisition module 52, used to obtain the login status of at least one target port in a target link with a fault, wherein the target link is used to connect to the target port of at least one hard disk in a preset link order; a query module 54, used to query the impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; a processing module 56, used to accumulate the impact factor corresponding to each target port one by one in a preset link order, and compare each accumulation result with a preset location threshold; a determination module 58, used to determine that the target port corresponding to the last accumulated impact factor is a faulty port when the accumulation result is greater than the preset location threshold.
- an acquisition module 52 used to obtain the login status of at least one target port in a target link with a fault, wherein the target link is used to
- the acquisition module 52 in this embodiment can be used to execute step S102 in the embodiment of the present application
- the query module 54 in this embodiment can be used to execute step S104 in the embodiment of the present application
- the processing module 56 in this embodiment can be used to execute step S106 in the embodiment of the present application
- the determination module 58 in this embodiment can be used to execute step S108 in the embodiment of the present application.
- the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments.
- the login status of at least one target port in a target link with a fault is obtained, wherein the target link is used to connect to the target port of at least one hard disk in a preset link order; the impact factor corresponding to each login status is queried in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; the impact factor corresponding to each target port is accumulated one by one in the preset link order, and each accumulated result is compared with a preset positioning threshold; when the accumulated result is greater than the preset positioning threshold, the target port corresponding to the last accumulated impact factor is determined to be the faulty port, and the faulty port where the fault occurs can be determined in the target link with the fault, thereby achieving the technical effect of accurately locating the fault position of the faulty target link, thereby solving the technical problem that the prior art cannot accurately locate the fault position in the link.
- the embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group.
- the computer terminal may also be replaced by a terminal device such as a mobile terminal.
- the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
- the above-mentioned computer terminal can execute the program code of the following steps in the fault location method: obtaining the login status of at least one target port in the target link with the fault, wherein the target link is used to connect to the target port of at least one hard disk according to a preset link order; querying the impact factor corresponding to each login status in the preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; accumulating the impact factor corresponding to each target port one by one according to the preset link order, and comparing each accumulation result with the preset location threshold; when the accumulation result is greater than the preset location threshold, determining that the target port corresponding to the last accumulated impact factor is the faulty port.
- FIG. 6 is a structural block diagram of a computer terminal according to an embodiment of the present application.
- the computer terminal 60 may include: one or more (only one is shown in the figure) processors 62 and a memory 64 .
- the memory can be used to store software programs and modules, such as program instructions/modules corresponding to the fault location method and device in the embodiment of the present application.
- the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, the above-mentioned fault location method is realized.
- the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
- the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal 60 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
- the processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain the login status of at least one target port in the target link with a fault, wherein the target link is used to connect to the target port of at least one hard disk according to a preset link sequence; query the impact factor corresponding to each login status in the preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; accumulate the impact factor corresponding to each target port one by one according to the preset link sequence, and compare each accumulation result with the preset positioning threshold; when the accumulation result is greater than the preset positioning threshold, determine the target port corresponding to the last accumulated impact factor The port is a faulty port.
- the processor may also execute program code for the following steps: before obtaining the login status of at least one target port in a target link with a fault, obtaining the login status of at least one preset port in a preset link, wherein the preset link is used to connect a preset port of at least one hard disk; querying an impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and an impact factor corresponding to each login status; accumulating the impact factors of all preset ports in the preset link to determine an evaluation result of the preset link; and when the evaluation result is greater than a preset fault threshold, determining that the preset link is a target link with a fault.
- the processor may also execute the program code of the following steps: before obtaining the login status of at least one preset port in a preset link, sending broadcast information to the hard disk set to be detected, wherein the hard disk set to be detected includes at least one hard disk, and each hard disk includes two preset ports; receiving location information returned by the hard disk set to be detected, wherein the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis; determining the preset link according to the preset chassis information and the preset controller information, wherein the preset link is used to cascade controllers with the same identifier in multiple chassis.
- the processor may also execute program code of the following steps: detecting frame loss information of each preset port, wherein the frame loss information includes at least the number of frame losses; determining a preset frame loss interval corresponding to the number of frame losses in a preset interval set as a target frame loss interval, wherein the preset interval set includes multiple preset frame loss intervals, and each preset frame loss interval is pre-set with a corresponding preset frame loss state; determining a target frame loss state corresponding to the target frame loss area as a login state.
- the processor may also execute the program code of the following steps: after accumulating the influence factors of all preset ports in the preset link and determining the evaluation result of the preset link, judging whether the evaluation result is greater than a preset redundancy threshold; when the evaluation result is greater than the preset redundancy threshold, determining that the preset link is an abnormal link; determining the total link of the abnormal link, wherein the total link includes two preset links; adjusting the redundant mark of the total link, wherein the redundant mark includes: a first redundant mark indicating that an abnormal link exists in the total link, and a second redundant mark indicating that no abnormal link exists in the total link.
- the processor may also execute the program code of the following steps: after determining that the last accumulated target port is a faulty port, obtaining target controller information of the faulty port, wherein the target controller information is a unique identifier of the controller connected to the faulty port; obtaining target chassis information of the hard disk where the faulty port is located, wherein the target chassis information is a unique identifier of the chassis of the hard disk where the faulty port is located; and generating alarm information based on the target controller information and the target chassis information.
- the processor may also execute the program code of the following steps: after obtaining the target controller information of the faulty port and the target chassis information of the hard disk where the faulty port is located, determine the upstream port of the faulty port, wherein the upstream port is used to transmit data for the faulty port; obtain the upstream controller information of the upstream port, wherein the upstream controller information is a unique identifier of the upstream port connected to the controller; obtain the upstream chassis information of the hard disk where the upstream port is located, wherein the upstream chassis information is a unique identifier of the chassis of the hard disk where the upstream port is located; and generate alarm information according to the target controller information, the target chassis information, the upstream controller information and the upstream chassis information.
- the processor may also execute program code of the following steps: after generating the alarm information, display the alarm information through a graphical user interface.
- the processor may also execute program code for the following steps: after obtaining the login status of at least one target port in a faulty target link, if the login status indicates that the target port is abnormal, obtaining the login status of the target port at a target time at a preset time interval; if the login status of the target port at the target time still indicates that the target port is abnormal, isolating the hard disk where the target port is located.
- the processor may also execute program code of the following steps: determining that the hard disk where the target port is located is the target hard disk; setting the target hard disk to a closed state, wherein when the target hard disk is in the closed state, all ports of the target hard disk are disconnected.
- the processor may also execute the program code of the following steps: after determining that the last accumulated target port is a faulty port, obtaining a redundant mark of the total link where the target link is located, wherein the redundant mark includes: a first redundant mark indicating that there is an abnormal link in the total link, and a second redundant mark indicating that there is no abnormal link in the total link; when the second redundant mark exists in the total link, performing fault isolation on the target controller that controls the faulty port.
- a fault location scheme By obtaining the login status of at least one target port in a target link with a fault, wherein the target link is used to connect to the target port of at least one hard disk according to a preset link sequence; querying the impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; accumulating the impact factor corresponding to each target port one by one according to the preset link sequence, and comparing each accumulation result with the preset location threshold; when the accumulation result is greater than the preset location threshold, determining that the target port corresponding to the last accumulated impact factor is the fault port, the faulty port that causes the fault can be determined in the target link with the fault, achieving the technical effect of accurately locating the fault position of the faulty target link, thereby solving the technical problem that the prior art cannot accurately locate the fault position in the link.
- the structure shown in FIG. 6 is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices.
- FIG. 6 does not limit the structure of the above-mentioned electronic devices.
- the computer terminal 60 may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG. 6, or have a configuration different from that shown in FIG. 6.
- the non-volatile readable storage medium can include Including: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
- the embodiment of the present application further provides a non-volatile readable storage medium.
- the non-volatile readable storage medium can be used to store the program code executed by the fault location method provided in the above embodiment.
- the non-volatile readable storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: obtaining the login status of at least one target port in a faulty target link, wherein the target link is used to connect to the target port of at least one hard disk in a preset link order; querying an impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and the impact factor corresponding to each login status; accumulating the impact factor corresponding to each target port one by one in the preset link order, and comparing each accumulated result with a preset positioning threshold; when the accumulated result is greater than the preset positioning threshold, determining that the target port corresponding to the last accumulated impact factor is a faulty port.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: before obtaining the login status of at least one target port in a target link with a fault, obtaining the login status of at least one preset port in a preset link, wherein the preset link is used to connect a preset port of at least one hard disk; querying an impact factor corresponding to each login status in a preset set, wherein the preset set is used to record multiple login statuses and the impact factors corresponding to each login status; accumulating the impact factors of all preset ports in the preset link to determine an evaluation result of the preset link; when the evaluation result is greater than a preset fault threshold, determining that the preset link is a target link with a fault.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: before obtaining the login status of at least one preset port in a preset link, sending broadcast information to the hard disk set to be detected, wherein the hard disk set to be detected includes at least one hard disk, and each hard disk includes two preset ports; receiving location information returned by the hard disk set to be detected, wherein the location information includes at least: preset chassis information and preset controller information, the preset chassis information is used to indicate the unique identifier of the chassis where each hard disk is located, and the preset controller information is used to indicate the unique identifier of the controller connected to the preset port of the hard disk in each chassis; determining the preset link based on the preset chassis information and the preset controller information, wherein the preset link is used to cascade controllers with the same identifier in multiple chassis.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: detecting frame loss information of each preset port, wherein the frame loss information includes at least the number of frame losses; determining a preset frame loss interval corresponding to the number of frame losses in a preset interval set as a target frame loss interval, wherein the preset interval set includes multiple preset frame loss intervals, and each preset frame loss interval is pre-set with a corresponding preset frame loss state; determining a target frame loss state corresponding to the target frame loss area as a login state.
- the non-volatile readable storage medium is configured to store program codes for executing the following steps: after accumulating the influence factors of all preset ports in the preset link and determining the evaluation result of the preset link, judging whether the evaluation result is greater than a preset redundancy threshold; when the evaluation result is greater than the preset redundancy threshold, determining that the preset link is an abnormal link; determining the total link of the abnormal link, wherein the total link includes two preset links; adjusting the redundant mark of the total link, wherein the redundant mark includes: a first redundant mark indicating that an abnormal link exists in the total link, and a second redundant mark indicating that no abnormal link exists in the total link.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: after determining that the last accumulated target port is a faulty port, obtaining target controller information of the faulty port, wherein the target controller information is a unique identifier of the controller connected to the faulty port; obtaining target chassis information of the hard disk where the faulty port is located, wherein the target chassis information is a unique identifier of the chassis of the hard disk where the faulty port is located; and generating alarm information based on the target controller information and the target chassis information.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: after obtaining the target controller information of the faulty port and obtaining the target chassis information of the hard disk where the faulty port is located, determining the upstream port of the faulty port, wherein the upstream port is used to transmit data for the faulty port; obtaining the upstream controller information of the upstream port, wherein the upstream controller information is a unique identifier of the upstream port connection controller; obtaining the upstream chassis information of the hard disk where the upstream port is located, wherein the upstream chassis information is a unique identifier of the chassis of the hard disk where the upstream port is located; and generating alarm information based on the target controller information, the target chassis information, the upstream controller information and the upstream chassis information.
- the non-volatile readable storage medium is configured to store program codes for executing the following steps: after generating the alarm information, the method further includes: displaying the alarm information through a graphical user interface.
- the non-volatile readable storage medium is configured to store program code for executing the following steps: after obtaining the login status of at least one target port in a faulty target link, when the login status indicates that there is an abnormality in the target port, obtaining the login status of the target port at a target time at a preset time interval; when the login status of the target port at the target time still indicates that there is an abnormality in the target port, isolating the hard disk where the target port is located.
- the non-volatile readable storage medium is configured to store program codes for executing the following steps: determining that the hard disk where the target port is located is the target hard disk; setting the target hard disk to a closed state, wherein when the target hard disk is in the closed state, all ports of the target hard disk are disconnected.
- the non-volatile readable storage medium is configured to store a program code for executing the following steps: after determining that the last accumulated target port is a faulty port, obtaining a redundant tag of the total link where the target link is located, wherein the redundant tag includes: a first redundant tag indicating that an abnormal link exists in the total link, and a second redundant tag indicating that no abnormal link exists in the total link; when the second redundant tag exists in the total link, In the case of a redundant flag, fault isolation is performed on the target controller that controls the failed port.
- the disclosed technical content can be implemented in other ways.
- the device embodiments described above are only schematic.
- the division of units can be a logical function division. There may be other division methods in actual implementation.
- multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
- Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
- the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
- each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
- the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile readable storage medium.
- the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product.
- the computer software product is stored in a non-volatile readable storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.
- the aforementioned non-volatile readable storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Quality & Reliability (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Debugging And Monitoring (AREA)
Abstract
一种故障定位方法、设备、装置、非易失性可读存储介质电子设备,方法包括:获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路被配置为按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,得到每次累加后的累加结果,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,解决了现有技术无法在链路中进行故障位置的准确定位的技术问题。
Description
相关申请的交叉引用
本申请要求于2023年08月15日提交中国专利局,申请号为202311026712.6,申请名称为“故障定位方法、设备、装置、存储介质及电子设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及计算机领域,具体而言,涉及一种故障定位方法、设备、装置、非易失性可读存储介质及电子设备。
当今是一个数据爆炸的时代,每天都会有海量的数据产生,专业的存储设备为存储这些庞大的数据提可能性,在这些海量的数据中有些数据是非常重要的如金融数据,实验数据等等,这就要求存储设备不仅仅要有功能性还需要有可靠性。
但是,存储设备在长时间的持续运行过程中,硬件难免会出现损耗,硬件的损耗就会引起硬件的异常,间接的带来软件上的异常。硬件链路故障是一种常见的硬件故障,可以表现为在一条很长的硬件的链路上的某一段或某几段发生故障。而现有技术无法在一段很长的链路上准确识别出故障位置或范围,因此智能对该链路进行全部整修,会增加检修范围,降低检修效率。
针对上述现有技术无法在链路中进行故障位置的准确定位的问题,目前尚未提出有效的解决方案。
发明内容
本申请实施例提供了一种故障定位方法、设备、装置、非易失性可读存储介质及电子设备,以至少解决现有技术无法在链路中进行故障位置的准确定位的技术问题。
根据本申请实施例的一个方面,提供了一种故障定位方法,包括:获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口。
可选地,在获取存在故障的目标链路中至少一个目标端口的登录状态之前,方法还包括:获取预设链路中至少一个预设端口的登录状态,其中,预设链路用于连接至少一个硬盘的预设端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果;在评价结果大于预设故障阈值的情况下,确定预设链路为存在故障的目标链路。
可选地,在获取预设链路中至少一个预设端口的登录状态之前,方法还包括:向待检测硬盘集合发送广播信息,其中,待检测硬盘集合包括至少一个硬盘,每个硬盘包括两个预设端口;接收待检测硬盘集合返回的位置信息,其中,位置信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识;根据预设机箱信息和预设控制器信息,确定预设链路,其中,预设链路用于级联多个机箱中具有同一标识的控制器。
可选地,获取预设链路中至少一个预设端口的登录状态包括:检测每个预设端口的丢帧信息,其中,丢帧信息至少包括丢帧数量;在预设区间集合中确定丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,预设区间集合包括多个预设丢帧区间,每个预设丢帧区间预先设有对应的预设丢帧状态;确定目标丢帧区域对应的目标丢帧状态为登录状态。
可选地,在累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果之后,方法还包括:判断评价结果是否大于预设冗余阈值;在评价结果大于预设冗余阈值的情况下,确定预设链路为异常链路;确定异常链路的总链路,其中,总链路包括两条预设链路;调整总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记。
可选地,在确定最后累加的目标端口为故障端口之后,方法还包括:获取故障端口的目标控制器信息,其中,目标控制器信息为故障端口连接控制器的唯一标识;获取故障端口所在硬盘的目标机箱信息,其中,目标机箱信息为故障端口所在硬盘的机箱的唯一标识;根据目标控制器信息和目标机箱信息,生成告警信息。
可选地,在获取故障端口的目标控制器信息和获取故障端口所在硬盘的目标机箱信息之后,方法还包
括:确定故障端口的上游端口,其中,上游端口用于为故障端口传输数据;获取上游端口的上游控制器信息,其中,上游控制器信息为上游端口连接控制器的唯一标识;获取上游端口所在硬盘的上游机箱信息,其中,上游机箱信息为上游端口所在硬盘的机箱的唯一标识;根据目标控制器信息、目标机箱信息、上游控制器信息和上游机箱信息,生成告警信息。
可选地,在生成告警信息之后,方法还包括:将告警信息通过图形用户界面进行展示。
可选地,在获取存在故障的目标链路中至少一个目标端口的登录状态之后,方法还包括:
在登录状态表示目标端口存在异常的情况下,按照预设时间间隔获取目标端口在目标时刻的登录状态;
在目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,对目标端口所在硬盘进行隔离。
可选地,对目标端口所在硬盘进行隔离包括:确定目标端口所在硬盘为目标硬盘;将目标硬盘的设置为关闭状态,其中,在目标硬盘处于关闭状态的情况下,断开目标硬盘的全部端口。
可选地,在确定最后累加的目标端口为故障端口之后,方法还包括:获取目标链路所在总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记;在总链路存在第二冗余标记的情况下,对控制故障端口的目标控制器进行故障隔离。
根据本申请实施例的另一方面,还提供了一种故障定位设备,其特征在于,包括:硬盘、串行总线和机箱管理器;串行总线,被配置为按照预设链路顺序连接目标链路中至少一个硬盘的目标端口,其中,目标链路为存在故障的链路;机箱管理器,被配置为获取目标端口的登录状态,在预设集合中查询每个登录状态对应的影响因子,按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子。
可选地,设备还包括:控制器,被配置为将串行总线与硬盘的预设端口连接。
可选地,设备还包括:总线控制芯片,被配置为向至少一个硬盘和机箱管理器发布广播事件;机箱管理器,还被配置为在收到广播后,发起发现机制获取每个硬盘的位置信息,其中,位置信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识。
可选地,机箱管理器,还被配置为在发起发现机制后,订阅每个预设端口的丢帧信息,其中,丢帧信息至少包括丢帧数量;在预设区间集合中确定丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,预设区间集合包括多个预设丢帧区间,每个预设丢帧区间预先设有对应的预设丢帧状态;确定目标丢帧区域对应的目标丢帧状态为登录状态。
可选地,硬盘的预设端口包括:第一预设端口和第二预设端口;串行总线包括:第一串行总线和第二串行总线;其中,第一串行总线被配置为连接第一预设端口;第二串行总线被配置为连接第二预设端口。
可选地,机箱管理器,还被配置为获取预设链路中至少一个预设端口的登录状态,其在,预设链路为第一串行总线的链路或第二串行总线的链路;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果;在评价结果大于预设冗余阈值的情况下,确定预设链路为异常链路;调整异常链路所在总链路的冗余标记,其中,总链路包括第一串行总线和第二串行总线,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记。
可选地,设备还包括:定时器,被配置为在登录状态表示目标端口存在异常的情况下启动,设定按照预设时间间隔确定目标时刻;机箱管理器,还被配置为获取目标端口在目标时刻的登录状态;并在目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,对目标端口所在硬盘进行隔离。
可选地,机箱管理器,还被配置为在确定目标链路中的故障端口之后,获取目标链路所在总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记;在总链路存在第二冗余标记的情况下,对控制故障端口的目标控制器进行故障隔离。
根据本申请实施例的另一方面,还提供了一种故障定位装置,包括:获取模块,被配置为获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路被配置为按照预设链路顺序连接至少一个硬盘的目标端口;查询模块,被配置为在预设集合中查询每个登录状态对应的影响因子,其中,预设集合被配置为记录多种登录状态,和每种登录状态对应的影响因子;处理模块,被配置为按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;确定模块,被配置为在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口。
根据本申请实施例的另一方面,还提供了一种非易失性可读存储介质,非易失性可读存储介质被配置为存储程序,其中,在程序运行时控制非易失性可读存储介质所在设备执行上述故障定位方法。
根据本申请实施例的另一方面,还提供了一种电子设备,包括:存储器和处理器,处理器被配置为运行存储在处理器中的程序,其中,程序运行时执行上述故障定位方法。
在本申请实施例中,获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于
按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,可以在存在故障的目标链路中确定产生故障的故障端口,实现了对故障的目标链路进行故障位置的准确定位的技术效果,进而解决了现有技术无法在链路中进行故障位置的准确定位技术问题。
此处所说明的附图用来提供对本申请的进一步理解,构成本申请的一部分,本申请的示意性实施例及其说明用于解释本申请,并不构成对本申请的不当限定。在附图中:
图1是根据本申请实施例的一种故障定位方法的流程图;
图2是根据本申请实施例的一种基于硬盘命令丢帧的存储硬件链路状态监控方法架构的示意图;
图3是根据本申请实施例的一种存储硬件链路的示意图;
图4是根据本申请实施例的一种故障定位设备的示意图;
图5是根据本申请实施例的一种故障定位装置的示意图;
图6是根据本申请实施例的一种计算机终端的结构框图。
为了使本技术领域的人员更好地理解本申请方案,下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分的实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都应当属于本申请保护的范围。
需要说明的是,本申请的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
为了便于描述,以下对本申请实施例涉及的部分名词或术语进行说明:
SAS Serial Attached SCSI是一种电脑集线的技术,其功能主要是做周边零件的数据传输,如:硬盘、CD-ROM等设备而设计的接口。
SAS Expander遵循SAS协议的扩展器,可用于机箱管理。
SESSCSI Enclosure Services T10技术委员会制定的用于机箱管理的标准。
EN:全称为Enclosure Mangement机箱管理。
SCSISmall Computer System Interface SCSI协议主要是在主机和存储设备之间传送命令、状态和块数据。在各类存储技术中,SCSI协议可谓是最重要的脊梁
RAID一般指磁盘阵列。磁盘阵列(Redundant Arrays of Independent Disks,RAID),有"数块独立磁盘构成具有冗余能力的阵列”之意。
根据本申请实施例,提供了一种故障定位方法实施例,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图1是根据本申请实施例的一种故障定位方法的流程图,如图1所示,该方法包括如下步骤:
步骤S102,获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;
步骤S104,在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;
步骤S106,按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;
步骤S108,在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口。
在本申请实施例中,获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,可以在存在故障的目标链路中确定产生故障的故障端口,实现了对故障的目标链路进行故障位置的准确定位的技术效果,进而解决了现有技术无法
在链路中进行故障位置的准确定位技术问题。
可选地,上述故障定位方法可以用在具有多个硬盘的存储系统中,该存储系统中多个硬盘可以采用磁盘阵列的方式存储数据。
可选地,每个硬盘可以包括两个预设端口,分别记为第一预设端口和第二预设端口。
可选地,预设链路可以为两条,分别为第一预设链路和第二预设链路,其中,第一预设链路用于连接每个硬盘的第一预设端口,第二预设链路用于建立每个硬盘的第二预设端口。
可选地,第一预设链路可以用于传输上行数据,第二预设链路用于传输下行数据。
可选地,目标链路为存在故障的预设链路,目标端口为存在故障的预设链路中连接的端口。
在上述步骤S102中,登录状态可以表示对应目标端口或预设端口的丢帧情况。
可选地,登录状态包括:GOOD,表示无丢帧异常;IN DOUBT,表示存在至少存在一次丢帧,但是还未到达DEGRADED的程度;DEGRADED,表示存在一定数量的丢帧,大于IN DOUBT但是未达到EXCLUSION的程度;表示EXCLUDSION盘存在大量的丢帧。
在上述步骤S104中,影响因子用于通过数值的方式表示目标端口或预设端口的丢帧程度。
可选地,不同登录状态存在不同的影响因子,例如,GOOD=0,IN DOUBT=1,DEGRADED=4,EXCLUDSION=10。
在上述步骤S106中,预设定位阈值用于对目标链路中存在故障的位置进行定位,在存在故障的目标链路中,逐个累加该目标链路中每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比对,若累加了某一目标端口对应的影响因子后大于预设定位阈值,则说明之后累加的目标端口处存在故障,实现了对故障端口的定位,便于后续对故障端口进行检修或隔离。
可选地,可以预先记录每个硬盘和预设端口的位置信息,该位置信息信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识,在确定故障端口后可以查询该故障端口的预设控制器信息,实现对故障端口的定位。
需要说明的是,每个硬盘的预设端口通过控制器与预设链路连接,由于控制器和预设端口一一对应,因此,可以通过确定故障端口的控制器,实现对故障端口的定位。
作为一种可选的实施例,在获取存在故障的目标链路中至少一个目标端口的登录状态之前,方法还包括:获取预设链路中至少一个预设端口的登录状态,其中,预设链路用于连接至少一个硬盘的预设端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果;在评价结果大于预设故障阈值的情况下,确定预设链路为存在故障的目标链路。
本申请上述实施例,在从故障的目标链路中进行故障定位之前,还需要逐个确定每条预设链路是否属于存在故障的目标链路,进而在确定目标链路的过程中,可以累加每条预设链路中各预设端口对应的影响因子,得到该条预设链路的评价结果,然后将该条预设链路的评价结果与预设故障阈值进行比较,在该条预设链路的评价结果大于预设故障阈值的情况下,说明该条预设链路存在故障,则确定该预设链路为目标链路,实现了对存在故障的目标链路的初步定位。
作为一种可选的实施例,在获取预设链路中至少一个预设端口的登录状态之前,方法还包括:向待检测硬盘集合发送广播信息,其中,待检测硬盘集合包括至少一个硬盘,每个硬盘包括两个预设端口;接收待检测硬盘集合返回的位置信息,其中,位置信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识;根据预设机箱信息和预设控制器信息,确定预设链路,其中,预设链路用于级联多个机箱中具有同一标识的控制器。
本申请上述实施例,通过广播信息可以获取每个硬盘和每个硬盘的预设端口的位置信息,进而在确定故障端口后,可以从该位置信息中确定故障端口所在位置和所在硬盘,进而便于对该故障端口进行隔离或检修。
作为一种可选的实施例,获取预设链路中至少一个预设端口的登录状态包括:检测每个预设端口的丢帧信息,其中,丢帧信息至少包括丢帧数量;在预设区间集合中确定丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,预设区间集合包括多个预设丢帧区间,每个预设丢帧区间预先设有对应的预设丢帧状态;确定目标丢帧区域对应的目标丢帧状态为登录状态。
本申请上述实施例,登录状态表示对应目标端口或预设端口的丢帧情况,因此,登录状态可以根据丢帧数量确定,通过预先划分多个预设丢帧区间,并为每个预设丢帧区间设置相应的预设丢帧状态,进而确定与预设端口或目标端口的丢帧数量对应的预设丢帧区间为目标丢帧区间,并将该目标丢帧区间的预设丢帧状态作为该预设端口或目标端口的登录状态,实现了对登录状态的确定。
作为一种可选的实施例,在获取预设链路中至少一个预设端口的登录状态之后,方法还包括:在登录状态表示预设端口存在异常的情况下,按照预设时间间隔获取预设端口在目标时刻的登录状态;在预设端口在目标时刻的登录状态仍然表示预设端口存在异常的情况下,对预设端口所在硬盘进行隔离。
本申请上述实施例,若硬盘偶尔出现了少量的丢帧可以被存储系统容忍,但是若硬盘频繁出现丢帧就需要对该硬盘进行处理,所以检测到预设端口的登录状态出现异常的情况下,可以按照预设时间间隔,在目标时刻再次获取预设端口的登录状态,若预设端口在目标时刻的登录状态仍然表示预设端口存在异常的情况下,说明该预设端口出现频繁丢帧的情况,则需要将该预设端口所在硬盘从存储系统中分离出来,实现对故障硬盘的隔离。
作为一种可选的实施例,对预设端口所在硬盘进行隔离包括:确定预设端口所在硬盘为目标硬盘;将目标硬盘的设置为关闭状态,其中,在目标硬盘处于关闭状态的情况下,断开目标硬盘的全部端口。
本申请上述实施例,多个硬盘可以采用RAID方式建立磁盘阵列,进而在该磁盘阵列中的某一硬盘出现故障的情况下,可以将出现故障的硬盘设置为关闭状态Fail,则RAID阵列中检测到处于关闭状态Fail的硬盘,则让该硬盘退出,从而实现了对出现故障的硬盘进行隔离的技术效果。
作为一种可选的实施例,在累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果之后,方法还包括:判断评价结果是否大于预设冗余阈值;在评价结果大于预设冗余阈值的情况下,确定预设链路为异常链路;确定异常链路的总链路,其中,总链路包括两条预设链路;调整总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记。
本申请上述实施例,存储系统中每个总链路可以设置两条预设链路,且两条预设链路互为冗余备份,若其中一条预设链路出现异常,那么还可以使用另一条预设链路继续进行数据传输,因此,需要适时了解总链路的冗余情况,进而可以将总链路中每条链路中全部预设端口对应影响因子的评价结果与预设冗余阈值进行比较,若评价结果大于预设冗余阈值,则说明该预设链路存在异常,需要调用总链路中冗余的另一条预设链路满足数据传输需求,并为总链路添加第一冗余标记;进而若总链路中的两条预设链路都出现了异常,则说明总链路已经无法满足数据传输需求,故需要停机报警,等待检修。
作为一种可选的实施例,在确定最后累加的目标端口为故障端口之后,方法还包括:获取故障端口的目标控制器信息,其中,目标控制器信息为故障端口连接控制器的唯一标识;获取故障端口所在硬盘的目标机箱信息,其中,目标机箱信息为故障端口所在硬盘的机箱的唯一标识;根据目标控制器信息和目标机箱信息,生成告警信息。
作为一种可选的实施例,在获取故障端口的目标控制器信息和获取故障端口所在硬盘的目标机箱信息之后,方法还包括:确定故障端口的上游端口,其中,上游端口用于为故障端口传输数据;获取上游端口的上游控制器信息,其中,上游控制器信息为上游端口连接控制器的唯一标识;获取上游端口所在硬盘的上游机箱信息,其中,上游机箱信息为上游端口所在硬盘的机箱的唯一标识;根据目标控制器信息、目标机箱信息、上游控制器信息和上游机箱信息,生成告警信息。
本申请上述实施例,告警信息中包含本控制器所在机箱的enclosure_index(也即目标机箱信息)、控制器的canister_index(也即目标控制器信息),与该故障控制器相连的上游机箱的enclosure_index(也即上游机箱信息)和上游控制器的canister_index(也即上游控制器信息),告警信息中通过两个点限制了一条线,确定了故障位置,实现了对故障位置的准确定位。
作为一种可选的实施例,在生成告警信息之后,方法还包括:将告警信息通过图形用户界面进行展示。
作为一种可选的实施例,在获取存在故障的目标链路中至少一个目标端口的登录状态之后,方法还包括:在登录状态表示目标端口存在异常的情况下,按照预设时间间隔获取目标端口在目标时刻的登录状态;在目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,对目标端口所在硬盘进行隔离。
本申请上述实施例,若硬盘偶尔出现了少量的丢帧可以被存储系统容忍,但是若硬盘频繁出现丢帧就需要对该硬盘进行处理,所以检测到目标端口的登录状态出现异常的情况下,可以按照预设时间间隔,在目标时刻再次获取目标端口的登录状态,若目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,说明该目标端口出现频繁丢帧的情况,则需要将该目标端口所在硬盘从存储系统中分离出来,实现对故障硬盘的隔离。
作为一种可选的实施例,对目标端口所在硬盘进行隔离包括:确定目标端口所在硬盘为目标硬盘;将目标硬盘的设置为关闭状态,其中,在目标硬盘处于关闭状态的情况下,断开目标硬盘的全部端口。
本申请上述实施例,多个硬盘可以采用RAID方式建立磁盘阵列,进而在该磁盘阵列中的某一硬盘出现故障的情况下,可以将出现故障的硬盘设置为关闭状态Fail,则RAID阵列中检测到处于关闭状态Fail的硬盘,则让该硬盘退出,从而实现了对出现故障的硬盘进行隔离的技术效果。
作为一种可选的实施例,在确定最后累加的目标端口为故障端口之后,方法还包括:获取目标链路所在总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记;在总链路存在第二冗余标记的情况下,对控制故障端口的目标控制器进行故障隔离。
本申请上述实施例,存储系统中每个总链路可以设置两条预设链路,且两条预设链路互为冗余备份,
若其中一条预设链路出现异常,那么还可以使用另一条预设链路继续进行数据传输,进而在对故障端口进行故障隔离的过程中,可以先根据总链路的冗余标记,确定该总链路是否具有满足故障隔离的冗余条件,然后在总链路存在指示总链路不存在异常链路的第二冗余标记,可以实现对故障端口的故障隔离。
本申请还提供了一种可选实施例,该可选实施例提供了一种基于硬盘命令丢帧的存储硬件链路状态监控方法。可以基于分析发生丢帧的硬盘数量和单块丢帧的严重程度,识别出故障链路的位置,根据不同的链路故障采取对应的故障隔离措施,有效的保障了系统的稳定性,当故障发生时可以准确的识别出故障原因并进行相应的处理,能够在一段很长的链路上准确的识别出故障位置或范围。
图2是根据本申请实施例的一种基于硬盘命令丢帧的存储硬件链路状态监控方法架构的示意图,如图2所示,包括:机箱管理层(EN),协议层,驱动层,物理链路层。
可选地,物理链路层在最低层,为存储系统和硬盘之间命令或数据传输的物理通道。
可选地,协议层和驱动层是存储系统和硬盘之间通信的数据通道。
可选地,机箱管理层位于最上层作用是根据协议层返回的每块硬盘的丢帧程度(也即登录状态)判断故障位置,触发故障隔离动作。
可选地,每块硬盘有两个预设端口port,每个预设端口port的丢帧程度(也即登录状态)分为四个等级:
GOOD:无丢帧异常,对EN的影响因子0。
IN DOUBT:存在丢帧:至少存在一次丢帧,但是还未到达DEGRADED的程度,对EN的影响因子1。
DEGRADED:盘存在丢帧:一定数量的丢帧,大于IN DOUBT但是未达到EXCLUSION的程度,对EN的影响因子4。
EXCLUDSION:盘存在大量的丢帧,对EN的影响因子10。
可选地,机箱管理层中记录着每块硬盘都有两条属性(也即位置信息),硬盘所在机箱和硬盘所在控制器。当单块盘出现丢帧时,认为是盘的单体故障,但是需要做防抖处理,避免误操作;当多快盘丢帧时,因为多块盘同时丢帧的概率很小,因此故障可能不在盘上,这时需要设置多个阈值,将多块盘的影响因子进行叠加,判断是否达到阈值和这些盘所在的机箱和控制器,进一步判断出故障发生在硬件链路的哪个位置,进而机箱管理层针对检测出的故障位置,进行对应的故障隔离操作,并及时上报告警。
本申请上述实施例,存储设备能够及时的发现机框是否存在链路故障的情况,进而对异常链路进行相应的排障处理,保证了存储系统的安全性和可靠性,使产品质量和生产业务得到了保障。
作为一种可选的实施例,基于硬盘命令丢帧的存储硬件链路状态监控方法的实施过程包括步骤如下:
步骤S1,存储系统正常启动运行。
步骤S2,SAS Expander芯片是一种可编程芯片,可以发布广播事件,机箱管理层(EN)收到广播后,发起Discovery(也即发现机制),Discovery完成后EN可以拿到硬盘所在机箱信息(enclosure_index)和每块硬盘的每个port预设端口对应的控制器信息(canister_index),也即获取位置信息。
步骤S3,EN通过订阅协议层里每块硬盘的每个port预设端口的丢帧信息,每次Discover完成后检查丢帧信息。
步骤S4,协议层通过SCSI协议规则和驱动返回命令执行状态,对每个盘的每个port预设端口的丢帧状态进行设置,丢帧状态记作盘的login状态(也即登录状态)。
可选地,login状态(也即登录状态)包括:GOOD:无丢帧异常;IN DOUBT:存在丢帧:至少存在一次丢帧,但是还未到达DEGRADED的程度;DEGRADED:盘存在丢帧:一定数量的丢帧,大于IN DOUBT但是未达到EXCLUSION的程度;和EXCLUDSION:盘存在大量的丢帧。
步骤S5,协议层对盘描述的不同login状态(也即登录状态)对EN有不同的影响因子,GOOD=0,IN DOUBT=1,DEGRADED=4,EXCLUDSION=10。
步骤S6,确定预设链路Strand和总链路Chain,其中,每条总链路Chain包括两条预设链路Strand。
图3是根据本申请实施例的一种存储硬件链路的示意图,如图3所示,enclosure0是主柜,enclosure1-3挂的是硬盘框(也即机箱)对应enclosure_index=1-3,每个机箱分上下两部分对应canister_index=0和1,其中,硬盘分布在enclosure1-3中,每个机箱直线通过SAS线缆级联起来,从主柜同一各SASport出来的,一直到级联最底部的这条硬件链路记为预设链路Strand,盘的每个与预设端口port对应一条预设链路Strand。上下控对称的两条预设链路Strand记为一条总链路Chain。
步骤S7,当EN发现一条strand上有多块盘login异常,此时需要根据每块盘login对应的影响因子进行累加。这是需要设置三个阈值,预设故障阈值T1,预设定位阈值T2和预设冗余阈值T3。
可选地,预设故障阈值T1为整个预设链路Strand上所有与之对应硬盘的与色号节点port的login状态(即登录状态)的影响因子的累加值,至少要大于30,满足至少三块盘同时EXCLUDSION。
可选地,预设定位阈值T2为当前预设链路Strand上某个控制器所对应硬盘的预设端口port的login状态(即登录状态)的影响因子的累加值T2要小于T1,可以按6分来算,用于寻找单条预设链路Strand上首个故障点的位置。
可选地,预设冗余阈值T3为总链路Chain阈值,T3要小于T1,因为总链路Chain表示一对预设链路
Strand,当两条预设链路Strand上都有login状态(即登录状态)异常时,说明没有冗余链路,此时存储系统极其不安全。因此T3比T1低,优先上报总链路Chain异常告警,当一个总链路Chain上的两条预设链路Strand每出现一条预设链路Strand的盘login状态(即登录状态)的影响因子累加值大于T3时,将变量(也即冗余标记)bad_chain+1(bad_chain初始值为0)。
步骤S8,每条预设链路Strand上的login(即登录状态)的影响因子累加值统计完成之后,开始进行链路故障异常点判断。
步骤S9,当预设链路Strand上盘的login状态(即登录状态)大于T1时代表有多块盘出现了login(即登录状态)异常,多块盘(至少3块)同时出现丢帧的概率是很小的,因此可以认为是预设链路Strand上的某一段出了问题。
步骤S10,再遍历步骤S9中发现故障的预设链路Strand上所有控制器所对应盘的login状态(即登录状态)的影响因子累加值与T2比较。找到首个对应的盘的累加因子大于T2的控制器,这个控制器就是首先发生故障的控制器(对应故障端口)。
步骤S11,找到首个故障控制器后需要对控制器进行故障隔离。
可选地,故障隔离时,如果步骤S7中bad_chain的值小于1,则找到该控制器的上行口物理链路,此时该控制器和与该控制器在同一条预设链路Strand上的下游控制器将不会再发生硬盘丢帧问题,达到故障隔离的目的,然后上报该控制器出现丢帧的告警,告警信息中要包含本控制器所在机箱的enclosure_index、控制器的canister_index,与该故障控制器相连的上游机箱的enclosure_index和上游控制器的canister_index,告警信息中通过两个点限制了一条线,确定了故障位置。
可选地,如果步骤S7中bad_chain的值大于1,表示此时存储链路无冗余,此时仅需要上报上述告警。
步骤S12,当EN发现在一条预设链路Strand上只有一块盘出现login异常(单盘不区分login状态),则认为login异常盘是自身的问题,存储系统是允许盘偶尔login异常,但不能容忍盘一直login异常,因此,在首次发现单盘login异常时,EN要进行防抖处理,启动1h定时器,同时记录此login异常的盘到EN模块的上下文中。
可选地,在定时器的时间到后(也即达到目标时刻),获取盘最新的login状态(即登录状态),如果此时有盘login异常且与1h前login异常的盘是同一块盘,则对这块盘上报告警,RAID模块检测到这块盘有告警后可以将这块盘置为Fail状态,让这块盘退出RAID,EN检测到盘Fail则需要将这块盘的物理预设端口port关闭,达到故障隔离的作用。用户根据告警信息找到对应的故障盘,进行换盘操作。
步骤S13,通过上述步骤存储系统能够在链路故障发生时,自动识别是点故障(盘),还是线故障(盘,控制器),并进行对应硬件上的故障隔离,防止引发更多的软件异常,保障存储的稳定性。
步骤S14,上述提到的所有告警通过GUI界面对用户进行展示。
图4是根据本申请实施例的一种故障定位设备的示意图,如图4所示,该设备可以包括:硬盘42、串行总线44和机箱管理器46;串行总线,用于按照预设链路顺序连接目标链路中至少一个硬盘的目标端口,其中,目标链路为存在故障的链路;机箱管理器,用于获取目标端口的登录状态,在预设集合中查询每个登录状态对应的影响因子,按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子。
在本申请实施例中,获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,可以在存在故障的目标链路中确定产生故障的故障端口,实现了对故障的目标链路进行故障位置的准确定位的技术效果,进而解决了现有技术无法在链路中进行故障位置的准确定位技术问题。
作为一种可选的实施例,设备还包括:控制器,用于将串行总线与硬盘的预设端口连接。
作为一种可选的实施例,设备还包括:总线控制芯片,用于向至少一个硬盘和机箱管理器发布广播事件;机箱管理器,还用于在收到广播后,发起发现机制获取每个硬盘的位置信息,其中,位置信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识。
作为一种可选的实施例,机箱管理器,还用于在发起发现机制后,订阅每个预设端口的丢帧信息,其中,丢帧信息至少包括丢帧数量;在预设区间集合中确定丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,预设区间集合包括多个预设丢帧区间,每个预设丢帧区间预先设有对应的预设丢帧状态;确定目标丢帧区域对应的目标丢帧状态为登录状态。
作为一种可选的实施例,硬盘的预设端口包括:第一预设端口和第二预设端口;串行总线包括:第一串行总线和第二串行总线;其中,第一串行总线用于连接第一预设端口;第二串行总线用于连接第二预设端口。
作为一种可选的实施例,机箱管理器,还用于获取预设链路中至少一个预设端口的登录状态,其在,预设链路为第一串行总线的链路或第二串行总线的链路;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果;在评价结果大于预设冗余阈值的情况下,确定预设链路为异常链路;调整异常链路所在总链路的冗余标记,其中,总链路包括第一串行总线和第二串行总线,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记。
作为一种可选的实施例,设备还包括:定时器,用于在登录状态表示目标端口存在异常的情况下启动,设定按照预设时间间隔确定目标时刻;机箱管理器,还用于获取目标端口在目标时刻的登录状态;并在目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,对目标端口所在硬盘进行隔离。
作为一种可选的实施例,机箱管理器,还用于在确定目标链路中的故障端口之后,获取目标链路所在总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记;在总链路存在第二冗余标记的情况下,对控制故障端口的目标控制器进行故障隔离。
根据本申请实施例,还提供了一种故障定位装置实施例,需要说明的是,该故障定位装置可以用于执行本申请实施例中的故障定位方法,本申请实施例中的故障定位方法可以在该故障定位装置中执行。
图5是根据本申请实施例的一种故障定位装置的示意图,如图5所示,该装置可以包括:获取模块52,用于获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;查询模块54,用于在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;处理模块56,用于按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;确定模块58,用于在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口。
需要说明的是,该实施例中的获取模块52可以用于执行本申请实施例中的步骤S102,该实施例中的查询模块54可以用于执行本申请实施例中的步骤S104,该实施例中的处理模块56可以用于执行本申请实施例中的步骤S106,该实施例中的确定模块58可以用于执行本申请实施例中的步骤S108。上述模块与对应的步骤所实现的示例和应用场景相同,但不限于上述实施例所公开的内容。
在本申请实施例中,获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,可以在存在故障的目标链路中确定产生故障的故障端口,实现了对故障的目标链路进行故障位置的准确定位的技术效果,进而解决了现有技术无法在链路中进行故障位置的准确定位技术问题。
本申请的实施例可以提供一种计算机终端,该计算机终端可以是计算机终端群中的任意一个计算机终端设备。可选地,在本实施例中,上述计算机终端也可以替换为移动终端等终端设备。
可选地,在本实施例中,上述计算机终端可以位于计算机网络的多个网络设备中的至少一个网络设备。
在本实施例中,上述计算机终端可以执行故障定位方法中以下步骤的程序代码:获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口。
可选地,图6是根据本申请实施例的一种计算机终端的结构框图。如图6所示,该计算机终端60可以包括:一个或多个(图中仅示出一个)处理器62、和存储器64。
其中,存储器可用于存储软件程序以及模块,如本申请实施例中的故障定位方法和装置对应的程序指令/模块,处理器通过运行存储在存储器内的软件程序以及模块,从而执行各种功能应用以及数据处理,即实现上述的故障定位方法。存储器可包括高速随机存储器,还可以包括非易失性存储器,如一个或者多个磁性存储装置、闪存、或者其他非易失性固态存储器。在一些实例中,存储器可进一步包括相对于处理器远程设置的存储器,这些远程存储器可以通过网络连接至终端60。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
处理器可以通过传输装置调用存储器存储的信息及应用程序,以执行下述步骤:获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端
口为故障端口。
可选的,上述处理器还可以执行如下步骤的程序代码:在获取存在故障的目标链路中至少一个目标端口的登录状态之前,获取预设链路中至少一个预设端口的登录状态,其中,预设链路用于连接至少一个硬盘的预设端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果;在评价结果大于预设故障阈值的情况下,确定预设链路为存在故障的目标链路。
可选的,上述处理器还可以执行如下步骤的程序代码:在获取预设链路中至少一个预设端口的登录状态之前,向待检测硬盘集合发送广播信息,其中,待检测硬盘集合包括至少一个硬盘,每个硬盘包括两个预设端口;接收待检测硬盘集合返回的位置信息,其中,位置信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识;根据预设机箱信息和预设控制器信息,确定预设链路,其中,预设链路用于级联多个机箱中具有同一标识的控制器。
可选的,上述处理器还可以执行如下步骤的程序代码:检测每个预设端口的丢帧信息,其中,丢帧信息至少包括丢帧数量;在预设区间集合中确定丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,预设区间集合包括多个预设丢帧区间,每个预设丢帧区间预先设有对应的预设丢帧状态;确定目标丢帧区域对应的目标丢帧状态为登录状态。
可选的,上述处理器还可以执行如下步骤的程序代码:在累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果之后,判断评价结果是否大于预设冗余阈值;在评价结果大于预设冗余阈值的情况下,确定预设链路为异常链路;确定异常链路的总链路,其中,总链路包括两条预设链路;调整总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记。
可选的,上述处理器还可以执行如下步骤的程序代码:在确定最后累加的目标端口为故障端口之后,获取故障端口的目标控制器信息,其中,目标控制器信息为故障端口连接控制器的唯一标识;获取故障端口所在硬盘的目标机箱信息,其中,目标机箱信息为故障端口所在硬盘的机箱的唯一标识;根据目标控制器信息和目标机箱信息,生成告警信息。
可选的,上述处理器还可以执行如下步骤的程序代码:在获取故障端口的目标控制器信息和获取故障端口所在硬盘的目标机箱信息之后,确定故障端口的上游端口,其中,上游端口用于为故障端口传输数据;获取上游端口的上游控制器信息,其中,上游控制器信息为上游端口连接控制器的唯一标识;获取上游端口所在硬盘的上游机箱信息,其中,上游机箱信息为上游端口所在硬盘的机箱的唯一标识;根据目标控制器信息、目标机箱信息、上游控制器信息和上游机箱信息,生成告警信息。
可选的,上述处理器还可以执行如下步骤的程序代码:在生成告警信息之后,将告警信息通过图形用户界面进行展示。
可选的,上述处理器还可以执行如下步骤的程序代码:在获取存在故障的目标链路中至少一个目标端口的登录状态之后,在登录状态表示目标端口存在异常的情况下,按照预设时间间隔获取目标端口在目标时刻的登录状态;在目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,对目标端口所在硬盘进行隔离。
可选的,上述处理器还可以执行如下步骤的程序代码:确定目标端口所在硬盘为目标硬盘;将目标硬盘的设置为关闭状态,其中,在目标硬盘处于关闭状态的情况下,断开目标硬盘的全部端口。
可选的,上述处理器还可以执行如下步骤的程序代码:确定最后累加的目标端口为故障端口之后,获取目标链路所在总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记;在总链路存在第二冗余标记的情况下,对控制故障端口的目标控制器进行故障隔离。
采用本申请实施例,提供了一种故障定位方案。通过获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口,可以在存在故障的目标链路中确定产生故障的故障端口,实现了对故障的目标链路进行故障位置的准确定位的技术效果,进而解决了现有技术无法在链路中进行故障位置的准确定位技术问题。
本领域普通技术人员可以理解,图6所示的结构仅为示意,计算机终端也可以是智能手机(如Android手机、iOS手机等)、平板电脑、掌声电脑以及移动互联网设备(Mobile Internet Devices,MID)、PAD等终端设备。图6其并不对上述电子装置的结构造成限定。例如,计算机终端60还可包括比图6中所示更多或者更少的组件(如网络接口、显示装置等),或者具有与图6所示不同的配置。
本领域普通技术人员可以理解上述实施例的各种方法中的全部或部分步骤是可以通过程序来指令终端设备相关的硬件来完成,该程序可以存储于一非易失性可读存储介质中,非易失性可读存储介质可以包
括:闪存盘、只读存储器(Read-Only Memory,ROM)、随机存取器(Random Access Memory,RAM)、磁盘或光盘等。
本申请的实施例还提供了一种非易失性可读存储介质。可选地,在本实施例中,上述非易失性可读存储介质可以用于保存上述实施例所提供的故障定位方法所执行的程序代码。
可选地,在本实施例中,上述非易失性可读存储介质可以位于计算机网络中计算机终端群中的任意一个计算机终端中,或者位于移动终端群中的任意一个移动终端中。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:获取存在故障的目标链路中至少一个目标端口的登录状态,其中,目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;按照预设链路顺序逐个累加每个目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在累加结果大于预设定位阈值的情况下,确定最后累加的影响因子对应的目标端口为故障端口。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在获取存在故障的目标链路中至少一个目标端口的登录状态之前,获取预设链路中至少一个预设端口的登录状态,其中,预设链路用于连接至少一个硬盘的预设端口;在预设集合中查询每个登录状态对应的影响因子,其中,预设集合用于记录多种登录状态,和每种登录状态对应的影响因子;累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果;在评价结果大于预设故障阈值的情况下,确定预设链路为存在故障的目标链路。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在获取预设链路中至少一个预设端口的登录状态之前,向待检测硬盘集合发送广播信息,其中,待检测硬盘集合包括至少一个硬盘,每个硬盘包括两个预设端口;接收待检测硬盘集合返回的位置信息,其中,位置信息至少包括:预设机箱信息和预设控制器信息,预设机箱信息用于表示每个硬盘所在机箱的唯一标识,预设控制器信息用于表示每个机箱中与硬盘的预设端口连接控制器的唯一标识;根据预设机箱信息和预设控制器信息,确定预设链路,其中,预设链路用于级联多个机箱中具有同一标识的控制器。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:检测每个预设端口的丢帧信息,其中,丢帧信息至少包括丢帧数量;在预设区间集合中确定丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,预设区间集合包括多个预设丢帧区间,每个预设丢帧区间预先设有对应的预设丢帧状态;确定目标丢帧区域对应的目标丢帧状态为登录状态。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在累加预设链路中全部预设端口的影响因子,确定预设链路的评价结果之后,判断评价结果是否大于预设冗余阈值;在评价结果大于预设冗余阈值的情况下,确定预设链路为异常链路;确定异常链路的总链路,其中,总链路包括两条预设链路;调整总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在确定最后累加的目标端口为故障端口之后,获取故障端口的目标控制器信息,其中,目标控制器信息为故障端口连接控制器的唯一标识;获取故障端口所在硬盘的目标机箱信息,其中,目标机箱信息为故障端口所在硬盘的机箱的唯一标识;根据目标控制器信息和目标机箱信息,生成告警信息。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在获取故障端口的目标控制器信息和获取故障端口所在硬盘的目标机箱信息之后,确定故障端口的上游端口,其中,上游端口用于为故障端口传输数据;获取上游端口的上游控制器信息,其中,上游控制器信息为上游端口连接控制器的唯一标识;获取上游端口所在硬盘的上游机箱信息,其中,上游机箱信息为上游端口所在硬盘的机箱的唯一标识;根据目标控制器信息、目标机箱信息、上游控制器信息和上游机箱信息,生成告警信息。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在生成告警信息之后,方法还包括:将告警信息通过图形用户界面进行展示。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在获取存在故障的目标链路中至少一个目标端口的登录状态之后,在登录状态表示目标端口存在异常的情况下,按照预设时间间隔获取目标端口在目标时刻的登录状态;在目标端口在目标时刻的登录状态仍然表示目标端口存在异常的情况下,对目标端口所在硬盘进行隔离。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:确定目标端口所在硬盘为目标硬盘;将目标硬盘的设置为关闭状态,其中,在目标硬盘处于关闭状态的情况下,断开目标硬盘的全部端口。
可选地,在本实施例中,非易失性可读存储介质被设置为存储用于执行以下步骤的程序代码:在确定最后累加的目标端口为故障端口之后,获取目标链路所在总链路的冗余标记,其中,冗余标记包括:指示总链路存在异常链路的第一冗余标记,和指示总链路不存在异常链路的第二冗余标记;在总链路存在第二
冗余标记的情况下,对控制故障端口的目标控制器进行故障隔离。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
在本申请的上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其他实施例的相关描述。
在本申请所提供的几个实施例中,应该理解到,所揭露的技术内容,可通过其它的方式实现。其中,以上所描述的装置实施例仅仅是示意性的,例如单元的划分,可以为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,单元或模块的间接耦合或通信连接,可以是电性或其它的形式。
作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个非易失性可读存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个非易失性可读存储介质中,包括若干指令用以使得一台计算机设备(可为个人计算机、服务器或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的非易失性可读存储介质包括:U盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、移动硬盘、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述仅是本申请的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本申请原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本申请的保护范围。
Claims (22)
- 一种故障定位方法,其特征在于,包括:获取存在故障的目标链路中至少一个目标端口的登录状态,其中,所述目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;在预设集合中查询每个所述登录状态对应的影响因子,其中,所述预设集合用于记录多种登录状态,和每种所述登录状态对应的影响因子;按照所述预设链路顺序逐个累加每个所述目标端口对应的影响因子,得到每次累加后的累加结果,并将每次累加结果与预设定位阈值进行比较;在所述累加结果大于所述预设定位阈值的情况下,确定最后累加的影响因子对应的所述目标端口为故障端口。
- 根据权利要求1所述的方法,其特征在于,在获取存在故障的目标链路中至少一个目标端口的登录状态之前,所述方法还包括:获取预设链路中至少一个预设端口的登录状态,其中,所述预设链路用于连接至少一个硬盘的预设端口;在预设集合中查询每个所述登录状态对应的影响因子,其中,所述预设集合用于记录多种登录状态,和每种所述登录状态对应的影响因子;累加所述预设链路中全部预设端口的影响因子,确定所述预设链路的评价结果;在所述评价结果大于预设故障阈值的情况下,确定所述预设链路为存在故障的所述目标链路。
- 根据权利要求2所述的方法,其特征在于,在获取预设链路中至少一个预设端口的登录状态之前,所述方法还包括:向待检测硬盘集合发送广播信息,其中,所述待检测硬盘集合包括至少一个硬盘,每个硬盘包括两个预设端口;接收所述待检测硬盘集合返回的位置信息,其中,所述位置信息至少包括:预设机箱信息和预设控制器信息,所述预设机箱信息用于表示每个硬盘所在机箱的唯一标识,所述预设控制器信息用于表示每个机箱中控制器的唯一标识,所述控制器与所述硬盘的预设端口连接;根据所述预设机箱信息和所述预设控制器信息,确定所述预设链路,其中,所述预设链路用于级联多个机箱中具有同一标识的控制器。
- 根据权利要求2所述的方法,其特征在于,获取预设链路中至少一个预设端口的登录状态包括:检测每个所述预设端口的丢帧信息,其中,所述丢帧信息至少包括丢帧数量;在预设区间集合中确定所述丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,所述预设区间集合包括多个所述预设丢帧区间,每个所述预设丢帧区间预先设有对应的预设丢帧状态;确定所述目标丢帧区间对应的目标丢帧状态为所述登录状态。
- 根据权利要求2所述的方法,其特征在于,在累加所述预设链路中全部预设端口的影响因子,确定所述预设链路的评价结果之后,所述方法还包括:判断所述评价结果是否大于预设冗余阈值;在所述评价结果大于所述预设冗余阈值的情况下,确定所述预设链路为异常链路;确定所述异常链路的总链路,其中,所述总链路包括两条所述预设链路;调整所述总链路的冗余标记,其中,所述冗余标记包括:指示所述总链路存在异常链路的第一冗余标记,和指示所述总链路不存在所述异常链路的第二冗余标记。
- 根据权利要求1所述的方法,其特征在于,在确定本次最后累加的影响因子对应的所述目标端口为故障端口之后,所述方法还包括:获取所述故障端口的目标控制器信息,其中,所述目标控制器信息为所述故障端口连接控制器的唯一标识;获取所述故障端口所在硬盘的目标机箱信息,其中,所述目标机箱信息为所述故障端口所在硬盘的机箱的唯一标识;根据所述目标控制器信息和所述目标机箱信息,生成告警信息。
- 根据权利要求6所述的方法,其特征在于,在获取所述故障端口的目标控制器信息和获取所述故障端口所在硬盘的目标机箱信息之后,所述方法还包括:确定所述故障端口的上游端口,其中,所述上游端口用于为所述故障端口传输数据;获取所述上游端口的上游控制器信息,其中,所述上游控制器信息为所述上游端口连接控制器的唯一标识;获取所述上游端口所在硬盘的上游机箱信息,其中,所述上游机箱信息为所述上游端口所在硬盘的机箱的唯一标识;根据所述目标控制器信息、所述目标机箱信息、所述上游控制器信息和所述上游机箱信息,生成告警信息。
- 根据权利要求6或7所述的方法,其特征在于,在生成告警信息之后,所述方法还包括:将所述告警信息通过图形用户界面进行展示。
- 根据权利要求1所述的方法,其特征在于,在获取存在故障的目标链路中至少一个目标端口的登录状态之后,所述方法还包括:在所述登录状态表示所述目标端口存在异常的情况下,按照预设时间间隔获取所述目标端口在目标时刻的登录状态;在所述目标端口在所述目标时刻的登录状态仍然表示所述目标端口存在异常的情况下,对所述目标端口所在硬盘进行隔离。
- 根据权利要求9所述的方法,其特征在于,对所述目标端口所在硬盘进行隔离包括:确定所述目标端口所在硬盘为目标硬盘;将所述目标硬盘设置为关闭状态,其中,在所述目标硬盘处于所述关闭状态的情况下,断开所述目标硬盘的全部端口。
- 根据权利要求1所述的方法,其特征在于,在确定本次最后累加的影响因子对应的所述目标端口为故障端口之后,所述方法还包括:获取所述目标链路所在总链路的冗余标记,其中,所述冗余标记包括:指示所述总链路存在异常链路的第一冗余标记,和指示所述总链路不存在所述异常链路的第二冗余标记;在所述总链路存在所述第二冗余标记的情况下,对控制所述故障端口的目标控制器进行故障隔离。
- 一种故障定位设备,其特征在于,包括:硬盘、串行总线和机箱管理器;串行总线,被配置为按照预设链路顺序连接目标链路中至少一个硬盘的目标端口,其中,所述目标链路为存在故障的链路;机箱管理器,被配置为获取所述目标端口的登录状态,在预设集合中查询每个所述登录状态对应的影响因子,按照所述预设链路顺序逐个累加每个所述目标端口对应的影响因子,并将每次累加结果与预设定位阈值进行比较;在所述累加结果大于所述预设定位阈值的情况下,确定最后累加的影响因子对应的所述目标端口为故障端口,其中,所述预设集合用于记录多种登录状态,和每种所述登录状态对应的影响因子。
- 根据权利要求12所述的设备,其特征在于,所述设备还包括:控制器,被配置为将所述串行总线与硬盘的预设端口连接。
- 根据权利要求12所述的设备,其特征在于,所述设备还包括:总线控制芯片,被配置为向至少一个硬盘和所述机箱管理器发布广播事件;所述机箱管理器,还被配置为在收到广播后,发起发现机制获取每个硬盘的位置信息,其中,所述位置信息至少包括:预设机箱信息和预设控制器信息,所述预设机箱信息用于表示每个硬盘所在机箱的唯一标识,所述预设控制器信息用于表示每个机箱中控制器的唯一标识,所述控制器与所述硬盘的预设端口连接。
- 根据权利要求14所述的设备,其特征在于,所述机箱管理器,还被配置为在发起发现机制后,订阅每个所述预设端口的丢帧信息,其中,所述丢帧信息至少包括丢帧数量;在预设区间集合中确定所述丢帧数量对应的预设丢帧区间为目标丢帧区间,其中,所述预设区间集合包括多个所述预设丢帧区间,每个所述预设丢帧区间预先设有对应的预设丢帧状态;确定所述目标丢帧区间对应的目标丢帧状态为所述登录状态。
- 根据权利要求12所述的设备,其特征在于,所述硬盘的预设端口包括:第一预设端口和第二预设端口;所述串行总线包括:第一串行总线和第二串行总线;其中,所述第一串行总线被配置为连接所述第一预设端口;所述第二串行总线被配置为连接所述第二预设端口。
- 根据权利要求16所述的设备,其特征在于,所述机箱管理器,还被配置为获取预设链路中至少一个预设端口的登录状态,其在,所述预设链路为所述第一串行总线的链路或所述第二串行总线的链路;在预设集合中查询每个所述登录状态对应的影响因子,其中,所述预设集合用于记录多种登录状态,和每种所述登录状态对应的影响因子;累加所述预设链路中全部预设端口的影响因子,确定所述预设链路的评价结果;在所述评价结果大于预设冗余阈值的情况下,确定所述预设链路为异常链路;调整所述异常链路所在总链路的冗余标记,其中,所述总链路包括所述第一串行总线和所述第二串行总线,所述冗余标记包括:指示所述总链路存在异常链路的第一冗余标记,和指示所述总链路不存在所述异常链路的第二冗余标记。
- 根据权利要求12所述的设备,其特征在于,所述设备还包括:定时器,被配置为在所述登录状态表示所述目标端口存在异常的情况下启动,设定按照预设时间间隔确定目标时刻;所述机箱管理器,还被配置为获取所述目标端口在目标时刻的登录状态;并在所述目标端口在所述目标时刻的登录状态仍然表示所述目标端口存在异常的情况下,对所述目标端口所在硬盘进行隔离。
- 根据权利要求12所述的设备,其特征在于,所述机箱管理器,还被配置为在确定所述目标链路中的所述故障端口之后,获取所述目标链路所在总链路的冗余标记,其中,所述冗余标记包括:指示所述总链路存在异常链路的第一冗余标记,和指示所述总链路不存在所述异常链路的第二冗余标记;在所述总链路存在所述第二冗余标记的情况下,对控制所述故障端口的目标控制器进行故障隔离。
- 一种故障定位装置,其特征在于,包括:获取模块,被配置为获取存在故障的目标链路中至少一个目标端口的登录状态,其中,所述目标链路用于按照预设链路顺序连接至少一个硬盘的目标端口;查询模块,被配置为在预设集合中查询每个所述登录状态对应的影响因子,其中,所述预设集合用于记录多种登录状态,和每种所述登录状态对应的影响因子;处理模块,被配置为按照所述预设链路顺序逐个累加每个所述目标端口对应的影响因子,得到每次累加后的累加结果,并将每次累加结果与预设定位阈值进行比较;确定模块,被配置为在所述累加结果大于所述预设定位阈值的情况下,确定最后累加的影响因子对应的所述目标端口为故障端口。
- 一种非易失性可读存储介质,其特征在于,所述非易失性可读存储介质被配置为存储程序,其中,在所述程序运行时控制所述非易失性可读存储介质所在设备执行权利要求1至11中任意一项所述故障定位方法。
- 一种电子设备,其特征在于,包括:存储器和处理器,所述处理器被配置为运行存储在所述处理器中的程序,其中,所述程序运行时执行权利要求1至11中任意一项所述故障定位方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202311026712.6A CN116755920B (zh) | 2023-08-15 | 2023-08-15 | 故障定位方法、设备、装置、存储介质及电子设备 |
| CN202311026712.6 | 2023-08-15 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025035892A1 true WO2025035892A1 (zh) | 2025-02-20 |
Family
ID=87953529
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/095818 Pending WO2025035892A1 (zh) | 2023-08-15 | 2024-05-28 | 故障定位方法、设备、装置、存储介质及电子设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116755920B (zh) |
| WO (1) | WO2025035892A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116755920B (zh) * | 2023-08-15 | 2023-11-17 | 苏州浪潮智能科技有限公司 | 故障定位方法、设备、装置、存储介质及电子设备 |
| CN117972685B (zh) * | 2024-03-28 | 2024-06-11 | 广州博今网络技术有限公司 | 一种基于模拟页面与服务器通信方法及系统 |
| CN118626303B (zh) * | 2024-08-13 | 2024-12-10 | 苏州元脑智能科技有限公司 | 存储系统的故障处理方法、装置、产品、存储系统及介质 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20110185226A1 (en) * | 2009-06-02 | 2011-07-28 | Yusuke Douchi | Storage system and control methods for the same |
| CN105227388A (zh) * | 2014-06-20 | 2016-01-06 | 中国电信股份有限公司 | 承载层链路监控的方法和系统 |
| CN105721184A (zh) * | 2014-12-03 | 2016-06-29 | 中国移动通信集团山东有限公司 | 一种网络链路质量的监控方法及装置 |
| CN111858122A (zh) * | 2020-07-29 | 2020-10-30 | 北京浪潮数据技术有限公司 | 一种存储链路的故障检测方法、装置、设备及存储介质 |
| CN113568806A (zh) * | 2021-06-28 | 2021-10-29 | 济南浪潮数据技术有限公司 | 一种sas卡链路状态监控方法、系统、装置及可读存储介质 |
| WO2022155919A1 (zh) * | 2021-01-22 | 2022-07-28 | 华为技术有限公司 | 一种故障处理方法、装置及系统 |
| CN116755920A (zh) * | 2023-08-15 | 2023-09-15 | 苏州浪潮智能科技有限公司 | 故障定位方法、设备、装置、存储介质及电子设备 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101707537B (zh) * | 2009-11-18 | 2012-01-25 | 华为技术有限公司 | 故障链路定位方法、告警根因分析方法及设备、系统 |
| CN105468484B (zh) * | 2014-09-30 | 2020-07-28 | 伊姆西Ip控股有限责任公司 | 用于在存储系统中确定故障位置的方法和装置 |
| CN112783703A (zh) * | 2021-01-15 | 2021-05-11 | 苏州浪潮智能科技有限公司 | 一种sas链路故障定位方法、装置、设备及存储介质 |
| CN116455729A (zh) * | 2023-03-07 | 2023-07-18 | 中国石油大学(华东) | 一种基于链路质量评估模型的故障链路检测与恢复方法 |
| CN116775376A (zh) * | 2023-06-25 | 2023-09-19 | 苏州浪潮智能科技有限公司 | 处理NVMe盘链路故障的方法、系统、设备和存储介质 |
-
2023
- 2023-08-15 CN CN202311026712.6A patent/CN116755920B/zh active Active
-
2024
- 2024-05-28 WO PCT/CN2024/095818 patent/WO2025035892A1/zh active Pending
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20110185226A1 (en) * | 2009-06-02 | 2011-07-28 | Yusuke Douchi | Storage system and control methods for the same |
| CN105227388A (zh) * | 2014-06-20 | 2016-01-06 | 中国电信股份有限公司 | 承载层链路监控的方法和系统 |
| CN105721184A (zh) * | 2014-12-03 | 2016-06-29 | 中国移动通信集团山东有限公司 | 一种网络链路质量的监控方法及装置 |
| CN111858122A (zh) * | 2020-07-29 | 2020-10-30 | 北京浪潮数据技术有限公司 | 一种存储链路的故障检测方法、装置、设备及存储介质 |
| WO2022155919A1 (zh) * | 2021-01-22 | 2022-07-28 | 华为技术有限公司 | 一种故障处理方法、装置及系统 |
| CN113568806A (zh) * | 2021-06-28 | 2021-10-29 | 济南浪潮数据技术有限公司 | 一种sas卡链路状态监控方法、系统、装置及可读存储介质 |
| CN116755920A (zh) * | 2023-08-15 | 2023-09-15 | 苏州浪潮智能科技有限公司 | 故障定位方法、设备、装置、存储介质及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116755920A (zh) | 2023-09-15 |
| CN116755920B (zh) | 2023-11-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2025035892A1 (zh) | 故障定位方法、设备、装置、存储介质及电子设备 | |
| JP7436737B1 (ja) | マルチベンダーを支援するサーバ管理システム | |
| CN112631820A (zh) | 软件系统的故障恢复方法及装置 | |
| CN105607973B (zh) | 一种虚拟机系统中设备故障处理的方法、装置及系统 | |
| CN111628893B (zh) | 分布式存储系统的故障处理方法及装置、电子设备 | |
| JP3957065B2 (ja) | ネットワーク計算機システムおよび管理装置 | |
| JP2007287183A (ja) | ホットスタンバイの構造とそのフォールトトレランス方法 | |
| CN111124722B (zh) | 一种隔离故障内存的方法、设备及介质 | |
| US12340089B2 (en) | Disk array redundancy method and system, computer equipment and storage medium | |
| US12609176B2 (en) | Memory processing method based on a server and apparatus, processor and electronic device | |
| CN117221090A (zh) | 一种多路径软件的路径故障隔离处理方法及装置 | |
| WO2025246553A1 (zh) | Pci设备的故障处理方法及装置、故障处理系统 | |
| CN115065589A (zh) | 数据流量采集灾备处理方法、装置、设备、系统及介质 | |
| CN117221091A (zh) | 存储集群中亚健康节点的隔离方法及其装置、电子设备 | |
| CN105119765B (zh) | 一种智能处理故障体系架构 | |
| CN118626303B (zh) | 存储系统的故障处理方法、装置、产品、存储系统及介质 | |
| CN120892295A (zh) | 服务器硬件故障诊断方法、电子设备和存储介质 | |
| CN119806875A (zh) | 故障处理系统、方法和装置、存储介质及电子设备 | |
| WO2024239569A1 (zh) | 一种集群业务处理方法、服务器及系统 | |
| CN115484267B (zh) | 多集群部署处理方法、装置、电子设备和存储介质 | |
| CN117555711A (zh) | 一种虚拟机管理方法、装置、云计算平台及介质 | |
| CN116909494A (zh) | 服务器的存储切换方法和装置,以及服务器系统 | |
| CN111767163B (zh) | 一种避免应用服务器因交易并发量过大造成宕机的方法及系统 | |
| JP4747909B2 (ja) | ファイバチャネルスイッチにおける障害装置の切り離し方法 | |
| CN114816267A (zh) | 一种存储设备的监控方法及系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24853307 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |