WO2017118152A1 - 一种通信设备机电管理总线故障节点的定位及隔离方法 - Google Patents
一种通信设备机电管理总线故障节点的定位及隔离方法 Download PDFInfo
- Publication number
- WO2017118152A1 WO2017118152A1 PCT/CN2016/102817 CN2016102817W WO2017118152A1 WO 2017118152 A1 WO2017118152 A1 WO 2017118152A1 CN 2016102817 W CN2016102817 W CN 2016102817W WO 2017118152 A1 WO2017118152 A1 WO 2017118152A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- electromechanical management
- communication
- shmc
- management bus
- bus
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/22—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing
- G06F11/2294—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing by remote test
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/06—Management of faults, events, alarms or notifications
- H04L41/0677—Localisation of faults
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/22—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing
- G06F11/2205—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing using arrangements specific to the hardware being tested
- G06F11/221—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing using arrangements specific to the hardware being tested to test buses, lines or interfaces, e.g. stuck-at or open line faults
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/22—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing
- G06F11/2257—Detection or location of defective computer hardware by testing during standby operation or during idle time, e.g. start-up testing using expert systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/06—Management of faults, events, alarms or notifications
- H04L41/0654—Management of faults, events, alarms or notifications using network fault recovery
- H04L41/0659—Management of faults, events, alarms or notifications using network fault recovery by isolating or reconfiguring faulty entities
Definitions
- the invention relates to a communication device, in particular to a method for locating and isolating a fault node of an electromechanical management bus of a communication device.
- the electromechanical management bus based on the electromechanical management system block diagram
- the electromechanical management system is based on two independent hardware serial bus, such as I2C bus and CAN bus
- the program serial bus signal line is small, to achieve Convenient, communication speed can also meet the requirements of electromechanical data transmission.
- serial bus has single node failure and affects the communication of all bus nodes.
- the technical problem to be solved by the present invention is to overcome the problem that the existing electromechanical management system cannot locate and isolate the faulty node after the electromechanical management bus communication is abnormal.
- the technical solution adopted by the present invention is to provide a pass.
- the method for locating and isolating the fault node of the electromechanical management bus of the letter device includes the following steps:
- Step 100 The running SHMC records the communication state of the electromechanical management bus during the communication
- Step 200 The running SHMC performs statistical analysis on the recorded communication state data to determine whether the electromechanical management bus has an unrecoverable abnormality of communication;
- Step 300 When an electromechanical management bus has an unrecoverable abnormality of communication, the running SHMC sends a command to the electromechanical management node attached to the abnormal electromechanical management bus by using the electromechanical management bus of the normal communication, so that the electromechanical management node controls Corresponding bus mechanical switch, coordinate the abnormal bus communication test between each electromechanical management node mounted on the abnormal electromechanical management bus, thereby locate the fault node in the abnormal electromechanical management bus, and return the plate number of the faulty node and Slot number
- Step 400 The running SHMC sends the abnormality alarm of the electromechanical management bus and the board number and the slot number of the faulty node that is abnormal to the remote network management system through the remote network management interface to display the remote alarm positioning indication.
- the communication state is recorded in such a manner that the running SHMC initiates a communication every time using the electromechanical management bus, and the communication state variables are cumulatively operated according to the success or failure of the communication result, and the communication state variable is continuous. The number of communication failures.
- determining that the electromechanical management bus has an unrecoverable abnormality of communication is: determining the recorded communication state data variable, and determining the electromechanical when the communication communication variable communication communication failure number of the electromechanical management bus reaches a predetermined threshold value An unrecoverable exception occurred on the management bus.
- the electromechanical management node hooked up on the abnormal electromechanical management bus includes an IPMC node and a standby SHMC node.
- step 300 specifically includes the following steps:
- step 300 specifically comprises The following steps:
- Step 301 When the running SHMC determines that an unrecoverable abnormality occurs in the one-way electromechanical management bus, the SHMC starts an electromechanical management bus abnormal positioning procedure;
- Step 302 The running SHMC sends a bus disconnection command to all the electromechanical management nodes connected to the abnormal electromechanical management bus through the electromechanical management bus of the normal communication;
- each electromechanical management node drives the mechanical switch to be disconnected, thereby leaving the abnormal electromechanical management bus;
- Step 304 The running SHMC confirms that all the electromechanical management nodes are separated from the abnormal electromechanical management bus, and select two slots and single disks from the electromechanical management node registry;
- Step 305 The running SHMC sends a connection abnormal electromechanical management bus command to the single disk through a normal electromechanical management bus;
- Step 306 the IPMC or SHMC driven mechanical switch of the selected single disk is closed, and is connected to the abnormal electromechanical management bus;
- Step 307 the running SHMC confirms that the selected two single disks are connected to the abnormal electromechanical management bus, and sends a communication test command with the communication address information of one of the single disks IPMC or SHMC to the IPMC or SHMC of the other single disk;
- Step 308 The IPMC or SHMC of the selected single disk receiving the communication test command sends a communication test command receiving response to the running SHMC, and sends test data through the abnormal electromechanical management bus according to the communication address information in the communication test command, and waits The response of the other party;
- Step 309 The running SHMC sends a communication test result acquisition command to the selected single disk IPMC or SHMC that initiates the communication test, and receives response data of the selected single disk IPMC or SHMC to communicate with another selected single disk IPMC or SHMC;
- Step 310 The running SHMC determines the total electromechanical management between the IPMCs or SHMCs of the two single disks accessing the abnormal electromechanical management bus according to the received communication test result response data. Whether the line circuit is abnormal, if the communication is abnormal, step 311 is performed; otherwise, step 312 is performed;
- Step 311 the running SHMC re-select two single disks in the electromechanical management node registry, and then perform step 305;
- Step 312 The running SHMC selects one of the successful communication single disks as a normal node, performs the above communication test on the other slave electromechanical management nodes, and performs communication test on all the electromechanical management nodes on the abnormal electromechanical management bus, and selects and causes The node of the electromechanical management bus is abnormal.
- the invention realizes that the electromechanical management system can complete the mechanical switch through the control bus by serially connecting the bus mechanical switch to each electromechanical management node attached to the electromechanical management bus, so that an irreversible abnormality occurs in one of the two electromechanical management buses. Communication test between each electromechanical management node, so that the electromechanical management system automatically discovers the electromechanical management bus communication abnormality, and locates the electromechanical management node that causes the bus abnormality to be isolated, and can accurately locate the fault point without manual investigation, which is effective not only effective Reduce the labor cost in maintenance, and improve the reliability of the electromechanical management system.
- the remote network management system will promptly report the abnormal electromechanical management node information to the maintenance personnel, so that the maintenance personnel can make timely and effective follow-up maintenance and solve the fault. .
- Figure 1 is a block diagram of an electromechanical management system based on an electromechanical management bus
- FIG. 2 is a block diagram of an electromechanical management node of an electromechanical management system of a communication device according to the present invention
- FIG. 3 is a flowchart of a method for locating and isolating a fault node of an electromechanical management bus of a communication device according to the present invention
- step 300 is a detailed flow chart of step 300 in the present invention.
- FIG. 1 is a block diagram of an electromechanical management node of the electromechanical management system of the communication device of the present invention, and compared with FIG. 1, the electromechanical management bus interface of the system in each electromechanical management node (including the chassis management controller SHMC and the intelligent controller IPMC) The circuit is connected in series with a controlled mechanical switch near the backplane end.
- the mechanical switch is controlled by IPMC or SHMC. The normally closed contact is used.
- the node is connected to the electromechanical management bus; SHMC and IPMC The mechanical switch can be controlled to open, thereby causing the electromechanical management node to detach from the physical management bus from the physical layer.
- the invention provides a method for locating and isolating a fault node of an electromechanical management bus of a communication device, as shown in FIG. 3, comprising the following steps:
- Step 100 The running SHMC records the communication state of the electromechanical management bus during the communication using the electromechanical management bus;
- the communication status is recorded in the following way: the electromechanical management master node (the running SHMC) initiates a communication every time using the electromechanical management bus, and will accumulate the communication state variables (continuous communication failure times) according to the success or failure of the communication result. .
- Step 200 The running SHMC performs statistical analysis on the recorded electromechanical management bus communication state data, and determines whether the electromechanical management bus has an unrecoverable abnormality of communication, that is, the electromechanical management master node cannot access the bus on the bus through the electromechanical management bus. Any from the electromechanical management node, and cannot be recovered;
- the method for judging the abnormality of communication unrecoverable in the electromechanical management bus is: comparing and judging the recorded communication state data variable, and judging when the recorded communication state data shows that the communication communication variable continuous communication failure number value of the electromechanical management bus reaches a predetermined threshold value An unrecoverable exception occurred on the electromechanical management bus.
- Step 300 When an unrecoverable abnormality occurs in the electromechanical management bus, the electromechanical management bus abnormal positioning program on the running SHMC sends a command to the electromechanical management node attached to the abnormal electromechanical management bus by using the electromechanical management bus of the normal communication. (IPMC node and standby SHMC node), so that each electromechanical management node controls its corresponding bus mechanical switch, realizes that each electromechanical management node bus interface circuit is connected to or disconnected from the abnormal electromechanical management bus from the physical layer, and coordinates the abnormal electromechanical management.
- IPMC node and standby SHMC node so that each electromechanical management node controls its corresponding bus mechanical switch, realizes that each electromechanical management node bus interface circuit is connected to or disconnected from the abnormal electromechanical management bus from the physical layer, and coordinates the abnormal electromechanical management.
- Each of the electromechanical management nodes on the bus performs mutual abnormal bus communication test to locate the faulty node in the abnormal electromechanical management bus, and returns the faulty node board number and slot number to realize the faulty node that causes the electromechanical management bus to be abnormal. (Single disc) positioning.
- Step 400 The SHMC sends a bus abnormality alarm and a board number and a slot number of the faulty single disk (fault node) that causes the abnormality to the remote network management system through the remote network management interface, and the bus abnormality sent by the remote network management system receiver power management system is abnormal.
- the alarm and the faulty single-disk positioning information are displayed and displayed to realize the remote alarm positioning indication.
- the single disk that causes the abnormality of the electromechanical management bus is locally illuminated.
- step 300 specifically includes the following steps:
- Step 301 When the running SHMC determines that an unrecoverable abnormality occurs in a certain electromechanical management bus, the SHMC starts an electromechanical management bus abnormal positioning procedure;
- Step 302 The SHMC sends a bus disconnection command to all the electromechanical management nodes (including the IPMC and the SHMC) connected to the abnormal electromechanical management bus through the electromechanical management bus of the normal communication;
- Step 303 After receiving the bus disconnection command, each electromechanical management node drives the machine. The mechanical switch is disconnected, thereby causing all electromechanical management nodes to leave the abnormal electromechanical management bus;
- Step 304 After the running SHMC confirms that all the electromechanical management nodes are separated from the abnormal electromechanical management bus, select two slot single disks from the electromechanical management node registry corresponding to the abnormal electromechanical management bus;
- Step 305 The running SHMC sends a connection abnormal electromechanical management bus command to the selected two single disks through the normal communication electromechanical management bus;
- Step 306 After receiving the bus connection command, the IPMC or SHMC of the selected single disk drives the mechanical switch to close and is connected to the abnormal electromechanical management bus;
- Step 307 After the running SHMC confirms that the selected two single disks are connected to the abnormal electromechanical management bus, send a communication test command to one of the IPMCIPMC or SHMC of the single disk, and the command is accompanied by another single disk IPMC or SHMC communication. Address information;
- Step 308 The IPMC or SHMC of the selected single disk receiving the communication test command sends a communication test command receiving response to the running SHMC, and sends test data through the abnormal electromechanical management bus according to the communication address information in the communication test command, and waits The response of the other party;
- Step 309 The running SHMC sends a communication test result acquisition command to the selected single-disk IPMC or SHMC that initiates the communication test, and receives response data of the single-disk IPMC or SHMC to communicate with another selected single-disk IPMC or SHMC;
- Step 310 The running SHMC determines, according to the received communication test result response data, whether the electromechanical management bus circuit between the IPMCs or SHMCs of the two single disks accessing the abnormal electromechanical management bus is abnormal. If the communication is abnormal, step 311 is performed; otherwise , performing step 312;
- Step 311 The running SHMC reselects two single disks, and then selects step 305 to perform the above communication test until two single disks with successful communication are found;
- Step 312 The running SHMC selects one of the successful communication single disks as a normal node, and performs the above communication test on the other slave electromechanical management nodes until the abnormal electromechanical tube All electromechanical management nodes on the management bus complete the communication test and filter out the nodes that cause the electromechanical management bus to be abnormal.
Landscapes
- Engineering & Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Computer Hardware Design (AREA)
- Quality & Reliability (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Small-Scale Networks (AREA)
Abstract
本发明公开了一种通信设备机电管理总线故障节点的定位及隔离方法,包括:运行的SHMC在通信过程中记录机电管理总线通信状态;并对通信状态数据进行统计分析,判断对应总线是否出现通信不可恢复的异常;当出现通信不可恢复的异常,SHMC用正常机电管理总线向挂接在异常机电管理总线上的机电管理节点发送命令,使其控制对应的总线机械开关,协调异常机电管理总线上的各节点之间进行相互的通信测试,定位故障节点,返回故障节点的定位信息;SHMC通过远程网管系统显示机电管理总线异常告警和导致异常的故障节点的定位信息,实现远程告警定位指示。本发明实现机电管理系统自动发现并定位导致机电管理总线异常故障节点,有效降低维护成本,提高系统可靠性。
Description
本发明涉及通信设备,具体涉及一种通信设备机电管理总线故障节点的定位及隔离方法。
随着通信设备容量的逐渐增加,其功耗也不断增加,致使通信设备的供电及散热日趋复杂,为了更好的实现通信设备供电和散热等机电功能,通信设备开始引入机电管理系统,专用于通信设备机电管理。
如图1所示,基于机电管理总线的机电管理系统框图,该机电管理系统是基于两路独立硬件的串行总线实现的,如I2C总线和CAN总线,该方案串行总线信号线少,实现方便,通信速度也可以满足机电数据传输的要求。但是串行总线存在单节点失效而影响所有总线节点通信的问题,例如单节点总线接口芯片损坏,对地短路,总线控制器和总线防护电路之间也将不能通信;为增加机电管理系统的可靠性,传统的做法是同时启用两组串行总线,两路总线互为主、备,虽然这样做可以提高总线可用性,但是当一条总线异常时,无法定位和隔离损坏节点,需要人工对通信设备所有的单盘进行排查,不仅造成人力资源浪费,而且会影响设备业务。
发明内容
本发明所要解决的技术问题是克服现有机电管理系统在机电管理总线通信异常后,无法定位和隔离故障节点的问题。
为了解决上述技术问题,本发明所采用的技术方案是提供一种通
信设备机电管理总线故障节点的定位及隔离方法,包括以下步骤:
步骤100、运行的SHMC在进行通信的过程中对机电管理总线的通信状态进行记录;
步骤200、运行的SHMC对记录的通信状态数据进行统计分析,判断机电管理总线是否出现通信不可恢复的异常;
步骤300、当一路机电管理总线出现通信不可恢复的异常时,运行的SHMC使用正常通信的机电管理总线向挂接在异常机电管理总线上的机电管理节点发送命令,使所述机电管理节点控制与其对应的总线机械开关,协调挂接在异常机电管理总线上的各机电管理节点之间进行相互的异常总线通信测试,从而定位异常机电管理总线中的故障节点,并返回故障节点的板盘号和槽位号;
步骤400、运行的SHMC通过远程网管接口将机电管理总线异常告警和导致异常的故障节点的板盘号和槽位号发送到远程网管系统进行显示,实现远程告警定位指示。
在上述方法中,所述通信状态进行记录的方式为:运行的SHMC每使用机电管理总线发起一次通信,都将根据通信结果的成败,对通信状态变量进行累加操作,所述通信状态变量为连续通信失败次数。
在上述方法中,判断机电管理总线出现通信不可恢复的异常的方式为:对记录的通信状态数据变量进行判断,当机电管理总线的通信状态变量连续通信失败次数值达到规定阈值时,判断该机电管理总线出现不可恢复的异常。
在上述方法中,所述挂接在异常机电管理总线上的机电管理节点包括IPMC节点和备用SHMC节点。
在上述方法中,步骤300具体包括以下步骤:
5、如权利要求4所述的方法,其特征在于,步骤300具体包括
以下步骤:
步骤301、当运行的SHMC判断一路机电管理总线发生不可恢复的异常时,该SHMC启动机电管理总线异常定位程序;
步骤302、运行的SHMC通过正常通信的机电管理总线向所有连接到异常机电管理总线上的机电管理节点发送总线脱离命令;
步骤303、各机电管理节点驱动机械开关断开,从而脱离异常机电管理总线;
步骤304、运行的SHMC确认所有机电管理节点脱离异常机电管理总线,并从机电管理节点注册表中选择两个槽位单盘;
步骤305、运行的SHMC通过正常机电管理总线向所述单盘发送连接异常机电管理总线命令;
步骤306、被选择单盘的IPMC或SHMC驱动机械开关闭合,与异常机电管理总线连接;
步骤307、运行的SHMC确认被选择的两个单盘连接到异常机电管理总线后,发送附带其中一个单盘IPMC或SHMC的通信地址信息的通信测试命令给另一单盘的IPMC或SHMC;
步骤308、接收到通信测试命令的被选择单盘的IPMC或SHMC向运行的SHMC发送通信测试命令接收应答,并根据通信测试命令中的通信地址信息,通过异常机电管理总线发送测试数据,并等待对方的应答;
步骤309、运行的SHMC向发起通信测试的被选择单盘IPMC或SHMC发送通信测试结果获取命令,并接收该被选择单盘IPMC或SHMC与另一被选择单盘IPMC或SHMC通信的应答数据;
步骤310、运行的SHMC根据收到的通信测试结果应答数据判断接入异常机电管理总线的两个单盘的IPMC或SHMC之间的机电管理总
线电路是否异常,如果通信异常,执行步骤311;否则,执行步骤312;
步骤311、运行的SHMC重新在机电管理节点注册表中选择两个单盘,然后执行步骤305;
步骤312、运行的SHMC选择通信成功单盘中的一个作为正常节点,对其它的从机电管理节点进行上述通信测试,直到对异常机电管理总线上所有机电管理节点都进行完通信测试,筛选出引起机电管理总线异常的节点。
在上述方法中,在机电管理总线异常告警和导致异常的故障节点的板盘号和槽位号在远程网管系统进行显示的同时,导致机电管理总线异常的故障节点在本地进行亮灯告警:
本发明通过为挂接在机电管理总线上的每个机电管理节点串入总线机械开关,使得在两路机电管理总线中的一路总线发生不可恢复异常时,机电管理系统能够通过控制总线机械开关完成每个机电管理节点之间的通信测试,从而实现机电管理系统自动发现机电管理总线通信异常,并定位导致总线异常的机电管理节点进行隔离,不需人工排查,就可以准确定位故障点,不仅有效降低维护中的人力成本,还提高了机电管理系统的可靠性,同时利用远程网络管理系统将导致异常的机电管理节点信息及时反馈给维护人员,方便维护人员做出及时有效的后续维护,解除故障。
图1为基于机电管理总线的机电管理系统框图;
图2为本发明中通信设备的机电管理系统的机电管理节点框图;
图3为本发明提供的一种通信设备机电管理总线故障节点的定位及隔离方法流程图;
图4为本发明中步骤300的具体流程图。
下面结合说明书附图和具体实施例对本发明做出详细的说明。
如图1所示的基于机电管理总线的机电管理系统,该系统的机电管理总线接口电路在靠近背板端,每个机电管理节点没有串入机械开关,当出现总线接口芯片物理层损坏时,无法从总线上脱离出去。而图2为本发明中通信设备的机电管理系统的机电管理节点框图,和图1相比较,该系统在各个机电管理节点(包括机箱管理控制器SHMC和智能控制器IPMC)的机电管理总线接口电路靠近背板端串联接入一个受控的机械开关,机械开关受控于IPMC或SHMC,使用常闭触点,在上电及正常情况下,节点挂接于机电管理总线上;SHMC和IPMC可以控制机械开关打开,从而使机电管理节点从物理层上脱离机电管理总线。
本发明提供的一种通信设备机电管理总线故障节点的定位及隔离方法,如图3所示,包括以下步骤:
步骤100、运行中的SHMC在使用机电管理总线进行通信的过程中对机电管理总线的通信状态进行记录;
通信状态进行记录的方式为:机电管理主节点(正在运行的SHMC)每使用机电管理总线发起一次通信,都将根据本次通信结果的成败,对通信状态变量(连续通信失败次数)进行累加操作。
步骤200、运行中的SHMC对记录的机电管理总线通信状态数据进行统计分析,判断机电管理总线是否出现通信不可恢复的异常,即机电管理主节点无法通过该机电管理总线访问挂接在该总线上的任意从机电管理节点,且不可恢复;
判断机电管理总线出现通信不可恢复的异常的方式为:对记录的通信状态数据变量进行比较判断,当记录的通信状态数据显示机电管理总线的通信状态变量连续通信失败次数值达到规定阈值时,判断该机电管理总线发生不可恢复的异常。
步骤300、当一路机电管理总线出现通信不可恢复的异常时,运行的SHMC上的机电管理总线异常定位程序,使用正常通信的机电管理总线发送命令到挂接在异常机电管理总线上的机电管理节点(IPMC节点和备用SHMC节点),使各个机电管理节点控制与其对应的总线机械开关,实现各机电管理节点总线接口电路从物理层上连入或脱离异常机电管理总线,协调挂接在异常机电管理总线上的各机电管理节点之间进行相互的异常总线通信测试,从而定位异常机电管理总线中的故障节点,并返回故障节点板盘号和槽位号,实现对导致机电管理总线异常的故障节点(单盘)的定位。
步骤400、SHMC通过远程网管接口将总线异常告警和导致异常的故障单盘(故障节点)的板盘号和槽位号发送到远程网管系统,所述远程网管系统接收机电管理系统发送的总线异常告警和故障单盘定位信息(板盘号和槽位号),并进行显示,实现远程告警定位指示;同时导致机电管理总线异常的单盘在本地进行亮灯告警。
在本发明中,如图4所示,步骤300具体包括以下步骤:
步骤301、当运行的SHMC判断某一路机电管理总线发生不可恢复的异常时,该SHMC启动机电管理总线异常定位程序;
步骤302、SHMC通过正常通信的机电管理总线,向所有连接到异常机电管理总线上的机电管理节点(包括IPMC和SHMC)发送总线脱离命令;
步骤303、各机电管理节点在接收到总线脱离命令后,会驱动机
械开关断开,从而使所有机电管理节点脱离异常机电管理总线;
步骤304、运行的SHMC确认所有机电管理节点脱离异常机电管理总线后,从与异常机电管理总线对应的机电管理节点注册表中选择两个槽位单盘;
步骤305、运行的SHMC通过正常通信机电管理总线向被选择的两个单盘发送连接异常机电管理总线命令;
步骤306、被选择单盘的IPMC或SHMC在接收到总线连接命令后,驱动机械开关闭合,与异常机电管理总线连接;
步骤307、运行的SHMC确认被选择的两个单盘连接到异常机电管理总线后,发送通信测试命令给其中的一个单盘的IPMCIPMC或SHMC,命令中附带另一个单盘的IPMC或SHMC的通信地址信息;
步骤308、接收到通信测试命令的被选择单盘的IPMC或SHMC向运行的SHMC发送通信测试命令接收应答,并根据通信测试命令中的通信地址信息,通过异常机电管理总线发送测试数据,并等待对方的应答;
步骤309、运行的SHMC向发起通信测试的被选择单盘IPMC或SHMC发送通信测试结果获取命令,并接收该单盘IPMC或SHMC与另一被选择单盘IPMC或SHMC通信的应答数据;
步骤310、运行的SHMC根据收到的通信测试结果应答数据判断接入异常机电管理总线的两个单盘的IPMC或SHMC之间的机电管理总线电路是否异常,如果通信异常,执行步骤311;否则,执行步骤312;
步骤311、运行的SHMC重新选择两个单盘,然后选择步骤305,进行上述通信测试,直到发现通信成功的两个单盘;
步骤312、运行的SHMC选择通信成功单盘中的一个作为正常节点,对其他的从机电管理节点进行上述通信测试,直到对异常机电管
理总线上所有机电管理节点都进行完通信测试,筛选出引起机电管理总线异常的节点。
显然,本领域的技术人员可以对本发明进行各种改动和变型而不脱离本发明的精神和范围。这样,倘若本发明的这些修改和变型属于本发明权利要求及其等同技术的范围之内,则本发明也意图包含这些改动和变型在内。
Claims (6)
- 一种通信设备机电管理总线故障节点的定位及隔离方法,其特征在于,包括以下步骤:步骤100、运行的SHMC在进行通信的过程中对机电管理总线的通信状态进行记录;步骤200、运行的SHMC对记录的通信状态数据进行统计分析,判断机电管理总线是否出现通信不可恢复的异常;步骤300、当一路机电管理总线出现通信不可恢复的异常时,运行的SHMC使用正常通信的机电管理总线向挂接在异常机电管理总线上的机电管理节点发送命令,使所述机电管理节点控制与其对应的总线机械开关,协调挂接在异常机电管理总线上的各机电管理节点之间进行相互的异常总线通信测试,从而定位异常机电管理总线中的故障节点,并返回故障节点的板盘号和槽位号;步骤400、运行的SHMC通过远程网管接口将机电管理总线异常告警和导致异常的故障节点的板盘号和槽位号发送到远程网管系统进行显示,实现远程告警定位指示。
- 如权利要求1所述的方法,其特征在于,所述通信状态进行记录的方式为:运行的SHMC每使用机电管理总线发起一次通信,都将根据通信结果的成败,对通信状态变量进行累加操作,所述通信状态变量为连续通信失败次数。
- 如权利要求2所述的方法,其特征在于,判断机电管理总线出现通信不可恢复的异常的方式为:对记录的通信状态数据变量进行判断,当机电管理总线的通信状态变量连续通信失败次数值达到规定阈值时,判断该机电管理总线出现不可恢复的异常。
- 如权利要求1所述的方法,其特征在于,所述挂接在异常机 电管理总线上的机电管理节点包括IPMC节点和备用SHMC节点。
- 如权利要求4所述的方法,其特征在于,步骤300具体包括以下步骤:步骤301、当运行的SHMC判断一路机电管理总线发生不可恢复的异常时,该SHMC启动机电管理总线异常定位程序;步骤302、运行的SHMC通过正常通信的机电管理总线向所有连接到异常机电管理总线上的机电管理节点发送总线脱离命令;步骤303、各机电管理节点驱动机械开关断开,从而脱离异常机电管理总线;步骤304、运行的SHMC确认所有机电管理节点脱离异常机电管理总线,并从机电管理节点注册表中选择两个槽位单盘;步骤305、运行的SHMC通过正常机电管理总线向所述单盘发送连接异常机电管理总线命令;步骤306、被选择单盘的IPMC或SHMC驱动机械开关闭合,与异常机电管理总线连接;步骤307、运行的SHMC确认被选择的两个单盘连接到异常机电管理总线后,发送附带其中一个单盘IPMC或SHMC的通信地址信息的通信测试命令给另一单盘的IPMC或SHMC;步骤308、接收到通信测试命令的被选择单盘的IPMC或SHMC向运行的SHMC发送通信测试命令接收应答,并根据通信测试命令中的通信地址信息,通过异常机电管理总线发送测试数据,并等待对方的应答;步骤309、运行的SHMC向发起通信测试的被选择单盘IPMC或SHMC发送通信测试结果获取命令,并接收该被选择单盘IPMC或SHMC与另一被选择单盘IPMC或SHMC通信的应答数据;步骤310、运行的SHMC根据收到的通信测试结果应答数据判断接入异常机电管理总线的两个单盘的IPMC或SHMC之间的机电管理总线电路是否异常,如果通信异常,执行步骤311;否则,执行步骤312;步骤311、运行的SHMC重新在机电管理节点注册表中选择两个单盘,然后执行步骤305;步骤312、运行的SHMC选择通信成功单盘中的一个作为正常节点,对其它的从机电管理节点进行上述通信测试,直到对异常机电管理总线上所有机电管理节点都进行完通信测试,筛选出引起机电管理总线异常的节点。
- 如权利要求1所述的方法,其特征在于,在机电管理总线异常告警和导致异常的故障节点的板盘号和槽位号在远程网管系统进行显示的同时,导致机电管理总线异常的故障节点在本地进行亮灯告警。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US15/744,806 US10725881B2 (en) | 2016-01-07 | 2016-10-21 | Method for locating and isolating failed node of electromechnical management bus in communication device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201610007698.9A CN105577447B (zh) | 2016-01-07 | 2016-01-07 | 一种通信设备机电管理总线故障节点的定位及隔离方法 |
| CN201610007698.9 | 2016-01-07 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2017118152A1 true WO2017118152A1 (zh) | 2017-07-13 |
Family
ID=55887144
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/102817 Ceased WO2017118152A1 (zh) | 2016-01-07 | 2016-10-21 | 一种通信设备机电管理总线故障节点的定位及隔离方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US10725881B2 (zh) |
| CN (1) | CN105577447B (zh) |
| WO (1) | WO2017118152A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112255909A (zh) * | 2020-11-16 | 2021-01-22 | 西安热工研究院有限公司 | 一种冗余通讯总线数据融合方法及系统 |
Families Citing this family (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105577447B (zh) | 2016-01-07 | 2018-10-09 | 烽火通信科技股份有限公司 | 一种通信设备机电管理总线故障节点的定位及隔离方法 |
| CN105978755B (zh) * | 2016-05-16 | 2019-06-11 | 珠海格力电器股份有限公司 | 空调通讯故障检测装置及方法 |
| JP2019057196A (ja) * | 2017-09-22 | 2019-04-11 | 横河電機株式会社 | 情報収集装置、情報収集方法 |
| CN111382018A (zh) * | 2018-12-29 | 2020-07-07 | 上海复控华龙微系统技术有限公司 | 串行通信的故障检测方法及装置、可读存储介质 |
| CN109974196B (zh) * | 2019-04-08 | 2020-08-07 | 珠海格力电器股份有限公司 | 变频机组通信控制方法、装置及变频机组系统 |
| CN110044035A (zh) * | 2019-05-16 | 2019-07-23 | 珠海格力电器股份有限公司 | 变频空调控制器容错电路和控制方法 |
| CN111813589A (zh) * | 2020-06-01 | 2020-10-23 | 北京百卓网络技术有限公司 | 一种分布式集群故障定位方法、装置、设备及存储介质 |
| CN112256628A (zh) * | 2020-10-26 | 2021-01-22 | 山东超越数控电子股份有限公司 | 一种基于国产单片机的多单元服务器故障管理方法 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101291261A (zh) * | 2008-04-28 | 2008-10-22 | 华为技术有限公司 | 一种板内设备测试方法和系统 |
| US20120076003A1 (en) * | 2011-09-27 | 2012-03-29 | Alton Wong | Chassis management modules for advanced telecom computing architecture shelves, and methods for using the same |
| CN102724093A (zh) * | 2012-06-26 | 2012-10-10 | 大唐移动通信设备有限公司 | 一种atca机框及其ipmb连接方法 |
| CN105577447A (zh) * | 2016-01-07 | 2016-05-11 | 烽火通信科技股份有限公司 | 一种通信设备机电管理总线故障节点的定位及隔离方法 |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5421002A (en) * | 1991-08-09 | 1995-05-30 | Westinghouse Electric Corporation | Method for switching between redundant buses in a distributed processing system |
| US6769078B2 (en) * | 2001-02-08 | 2004-07-27 | International Business Machines Corporation | Method for isolating an I2C bus fault using self bus switching device |
| US20070234123A1 (en) * | 2006-03-31 | 2007-10-04 | Inventec Corporation | Method for detecting switching failure |
| CN103592855B (zh) * | 2012-08-14 | 2016-01-27 | 北京航天发射技术研究所 | 一种基于自诊断型的智能配电时序器 |
| CN204302453U (zh) * | 2014-11-28 | 2015-04-29 | 武汉钢铁(集团)公司 | 一种开关故障定位装置 |
| US9835669B2 (en) * | 2014-12-19 | 2017-12-05 | The Boeing Company | Automatic data bus wire integrity verification device |
| US10007629B2 (en) * | 2015-01-16 | 2018-06-26 | Oracle International Corporation | Inter-processor bus link and switch chip failure recovery |
-
2016
- 2016-01-07 CN CN201610007698.9A patent/CN105577447B/zh active Active
- 2016-10-21 US US15/744,806 patent/US10725881B2/en active Active
- 2016-10-21 WO PCT/CN2016/102817 patent/WO2017118152A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101291261A (zh) * | 2008-04-28 | 2008-10-22 | 华为技术有限公司 | 一种板内设备测试方法和系统 |
| US20120076003A1 (en) * | 2011-09-27 | 2012-03-29 | Alton Wong | Chassis management modules for advanced telecom computing architecture shelves, and methods for using the same |
| CN102724093A (zh) * | 2012-06-26 | 2012-10-10 | 大唐移动通信设备有限公司 | 一种atca机框及其ipmb连接方法 |
| CN105577447A (zh) * | 2016-01-07 | 2016-05-11 | 烽火通信科技股份有限公司 | 一种通信设备机电管理总线故障节点的定位及隔离方法 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112255909A (zh) * | 2020-11-16 | 2021-01-22 | 西安热工研究院有限公司 | 一种冗余通讯总线数据融合方法及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN105577447A (zh) | 2016-05-11 |
| US20190012246A1 (en) | 2019-01-10 |
| CN105577447B (zh) | 2018-10-09 |
| US10725881B2 (en) | 2020-07-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2017118152A1 (zh) | 一种通信设备机电管理总线故障节点的定位及隔离方法 | |
| CN106850286B (zh) | 单板上的基板管理控制器及网元管理盘的基板管理控制器 | |
| CN105700969B (zh) | 服务器系统 | |
| TW201348979A (zh) | 串口切換系統、伺服器及串口切換方法 | |
| CN101630298A (zh) | 串行总线从设备地址设置系统 | |
| CN102611600B (zh) | 一种can网络系统的短路位置定位方法及装置 | |
| CN105430327A (zh) | 一种nvr集群备份方法及装置 | |
| TWI530778B (zh) | 具有自動重置功能的機櫃及其自動重置方法 | |
| CN107807630B (zh) | 一种主备设备的切换控制方法、其切换控制系统及装置 | |
| US7912995B1 (en) | Managing SAS topology | |
| CN106209410A (zh) | 使用网关修复通信中断的方法和系统 | |
| CN113992501A (zh) | 一种故障定位系统、方法及计算装置 | |
| CN111399879A (zh) | 一种cpld的固件升级系统和方法 | |
| CN114510378A (zh) | 空调机组的参数备份方法和备份装置、电子设备 | |
| CN108616428A (zh) | 一种远程管理rack机房的移动app实施方法 | |
| CN105068763A (zh) | 一种针对存储故障的虚拟机容错系统和方法 | |
| CN109062184A (zh) | 双机应急救援设备、故障切换方法和救援系统 | |
| CN115543679B (zh) | 漏液检测线检测方法、系统、装置、服务器及电子设备 | |
| CN104010077A (zh) | 一种信息处理方法及电子设备 | |
| CN114064401A (zh) | 定位硬盘故障的方法、装置、电子设备及存储介质 | |
| CN110247833B (zh) | 通信控制方法、装置、子设备和通信系统 | |
| CN109284218A (zh) | 一种检测服务器运行故障的方法及其装置 | |
| CN105119765A (zh) | 一种智能处理故障体系架构 | |
| CN104283712A (zh) | 网络设备及用于网络设备的管理网口配置方法 | |
| CN207869116U (zh) | 一种主备设备的切换控制系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16883280 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16883280 Country of ref document: EP Kind code of ref document: A1 |