WO2023112359A1 - 通信システム、管理装置及び端末 - Google Patents
通信システム、管理装置及び端末 Download PDFInfo
- Publication number
- WO2023112359A1 WO2023112359A1 PCT/JP2022/023197 JP2022023197W WO2023112359A1 WO 2023112359 A1 WO2023112359 A1 WO 2023112359A1 JP 2022023197 W JP2022023197 W JP 2022023197W WO 2023112359 A1 WO2023112359 A1 WO 2023112359A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- terminal
- recovery
- unit
- failure analysis
- management device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/0766—Error or fault reporting or storing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/079—Root cause analysis, i.e. error or fault diagnosis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/0793—Remedial or corrective actions
Definitions
- the present invention relates to a method of analyzing failures occurring in one or more terminals to be managed and automatically executing recovery processing for the failures.
- Patent Document 1 there is a technology that detects failure occurrence by self-analysis of the terminal instead of the management device and implements recovery processing.
- a terminal acquires a list of recovery methods from a management device in advance, and when an abnormality is detected by self-analysis, the terminal refers to the list and executes recovery processing.
- Patent Document 1 failure analysis is entrusted only to the terminal, so it is difficult for the terminal to respond to abnormal events that cannot be self-detected. For example, assume a case where the delivery of firmware from the management device to the terminal fails and the terminal continues to use the old firmware by mistake.
- the distribution source management device knows the version information of the firmware that should be applied, and the terminal cannot self-detect that the firmware applied to itself is outdated. In this way, it is difficult for terminals to self-detect, and there are anomalies that can only be detected by the management device.
- the present invention has been made in view of the above problems.
- the terminal self-detects
- the purpose is to respond to abnormal events that cannot be performed and to enable recovery.
- the present invention is a communication system comprising one or more terminals to be managed and a management apparatus connected to the terminals, wherein the terminals collect log information of their own terminals and transmit the log information to the management apparatus.
- a log information management unit a self-failure analysis unit that analyzes the presence or absence of an abnormal event in the terminal from the log information and determines self-recovery processing when the abnormal event is detected, and a self-recovery processing for the abnormal event.
- a recovery notification transmission unit that notifies the management device; and a self-recovery processing unit that executes self-recovery processing for the abnormal event or recovery processing commanded by the management device.
- the presence or absence of an abnormal event in the terminal is analyzed, and if the abnormal event is detected, a failure analysis unit determines recovery processing, and the terminal responds to the abnormal event detected by the terminal.
- a recovery notification receiving unit that receives a notification of self-recovery processing, and a recovery instruction unit that instructs the terminal to execute recovery processing for an abnormal event detected by the failure analysis unit according to the presence or absence of the notification.
- both the terminal and the management device analyze an abnormal event that has occurred in the terminal based on the log information of the terminal.
- FIG. 1 is a block diagram illustrating the configuration of a communication system and the hardware configuration of a management device and terminals in Embodiment 1;
- FIG. FIG. 5 is a diagram showing an example of a failure analysis condition management table for a management device and a terminal according to the first embodiment;
- FIG. 5 is a diagram showing an example of an event determination condition management table of a management device and a terminal according to the first embodiment;
- FIG. 10 is a diagram illustrating an example of a recovery process management table of a management device and a terminal in Embodiment 1;
- FIG. 4 is a sequence diagram showing the flow of failure analysis and recovery by the management device and the terminal according to the first embodiment, and shows an example of self-recovery by the terminal.
- FIG. 5 is a diagram showing an example of a failure analysis condition management table for a management device and a terminal according to the first embodiment
- FIG. 5 is a diagram showing an example of an event determination condition management table of a management device and a terminal according to the first
- FIG. 4 is a sequence diagram showing the flow of failure analysis and recovery by the management device and terminals in the first embodiment, and shows an example of a recovery command by the management device;
- 5 is a flow chart showing an example of failure analysis processing by a management device and self-failure analysis processing by a terminal in Embodiment 1;
- 7 is a flow chart showing an example of a process for determining whether or not a recovery command is required by the management device according to the first embodiment;
- FIG. 10 is an explanatory diagram showing an example of a screen display for setting various tables in Example 1;
- FIG. 10 is an explanatory diagram showing a screen display example regarding failure detection and restoration information in the first embodiment;
- FIG. 12 is a diagram showing an example of a failure analysis condition management table of the management device in the second embodiment
- FIG. FIG. 10 is a diagram illustrating an example of an event determination condition management table of a management device according to the second embodiment
- FIG. FIG. 11 is a diagram illustrating an example of a recovery process management table of a management device in Embodiment 2
- FIG. 12 is a diagram illustrating an example of a terminal recovery process management table according to the second embodiment
- FIG. 12 is a sequence diagram showing the flow of failure analysis and recovery by the management device and terminals in the third embodiment
- FIG. 11 is an explanatory diagram showing an example of system operation plan information held by a management device in Example 3
- 15 is a flow chart showing an example of restoration instruction permission determination processing by a management device in Embodiment 3.
- FIG. 10 is a diagram illustrating an example of an event determination condition management table of a management device according to the second embodiment
- FIG. 11 is a diagram illustrating an example of a recovery process management table of
- the terminal and management device each manage a list of failure analysis conditions for log information in the failure analysis condition management unit. Further, the event determination condition management section manages a list of event determination conditions based on a combination of matching failure analysis conditions, and a recovery process management section manages a list of recovery processes to be performed for each event.
- the self-failure analysis unit of the terminal refers to the failure analysis condition management unit, analyzes the log information in the terminal for each condition, and extracts failure analysis conditions that match the log information.
- the terminal refers to the event determination condition management unit, and detects an abnormal event that has occurred within the terminal based on the combination of matching failure analysis conditions. Then, the terminal refers to the recovery process management unit, and determines the self-recovery process for the abnormal event and the content of notification to the management device.
- the terminal transmits a notification regarding the execution of the restoration process to the management device with the restoration notification transmission section, and executes the restoration process with the self-restoration processing section, thereby achieving self-recovery from the abnormal state.
- the management device also refers to the failure analysis condition management unit in the failure analysis unit, executes analysis of the log information collected from the terminal, and extracts matching failure analysis conditions. Subsequently, the management device refers to the event determination condition management section, and detects an abnormal event that has occurred in the terminal based on a combination of matching failure analysis conditions.
- the management device refers to the restoration processing management unit and determines restoration processing for the abnormal event. Thereafter, when the recovery notification receiving unit cannot receive the recovery notification from the terminal, the management device transmits a recovery processing execution command to the terminal, thereby realizing recovery from the abnormal state. On the other hand, when the recovery notification is received from the terminal, the command to execute the recovery process is canceled to avoid redundant recovery process execution in the terminal.
- both the terminal and the management device analyze abnormal events that have occurred on the terminal and implement recovery processing.
- Abnormal conditions that can only be detected by the management device such as setting errors, can be detected, and recovery and correction can be automatically taken.
- it is possible to recover from various abnormal states by registering failure analysis conditions, event determination conditions, and recovery processing for abnormal events that can be detected only by either the terminal or the management device as management information.
- by causing the terminal to transmit a notification regarding the recovery process to the management device it is possible to prevent redundant recovery processes.
- the information may be expressed in terms of "AAA table”, but the information may be expressed in any data structure. That is, the "AAA table” can be "AAA information" to indicate that the information is independent of the data structure.
- FIG. Example 1 will be described with reference to FIGS. 1 to 10
- Example 2 will be described with reference to FIG. 11, and Example 3 will be described with reference to FIGS.
- the communication system in FIG. 1 includes one or more terminals 101 (101-a to 101-b) to be managed, a management device 102, and a network 150.
- FIG. 1 when the terminals 101-a to 101-b are not specified individually, the symbol "101" omitting "-" is used.
- the terminal 101 and the management apparatus 102 are connected by a network 150 including either or both of wired communication and wireless communication. to the management device 102 .
- the terminal 101 periodically self-analyzes whether there is an abnormal event that has occurred within itself, determines self-recovery processing when an abnormal event is detected, and notifies about the self-recovery processing. After transmitting to the management device 102, self-recovery processing is executed.
- the management device 102 provides services that utilize the data received from the terminal 101, and also manages failures of the terminal 101. Specifically, the management device 102 periodically analyzes the presence or absence of an abnormal event in the terminal 101 based on log information collected from each terminal 101, and determines restoration processing when an abnormality is detected.
- the management device 102 determines whether or not there is a notification regarding self-restoration processing from the terminal 101 in which the failure occurred, and if the notification is not received, the corresponding terminal 101 is instructed to execute the restoration processing. do.
- FIG. 1 illustrates a configuration in which one management device 102 is responsible for both service provision using received data and fault management of the terminal 101.
- the management device 102 is responsible for only the former service provision.
- the management apparatus 102 that only handles the latter failure management.
- the management device 102 may be installed at the same site as the terminal 101, or may be installed at another base such as on a cloud.
- the terminal 101 is a computer that stores data obtained from the site, log information of the terminal itself, etc. in packets and transmits the packet to the management apparatus 102 , and has a communication function with the management apparatus 102 .
- the terminal 101 has various configurations according to the type of data to be acquired. For example, in addition to collecting and transmitting data from other devices existing in the field, the terminal 101 itself also has a temperature measurement terminal with a communication function when measuring the temperature of the field, and when acquiring an image of the field, it is a communication terminal. It may be a camera terminal or the like having a function.
- the terminal 101 is composed of a communication I/F 111, a CPU 112, an input section 113, an output section 114, and a storage device 115.
- the communication I/F 111 is, for example, a transmission unit that converts a digital signal and a radio signal to each other, converts the generated digital data into a radio signal, and transmits the generated digital data when transmitting/receiving a packet to/from the management apparatus 102 via radio communication. and a receiver for extracting digital data from the received radio signal.
- the terminal 101 transmits/receives packets not only to the management device 102 but also to other devices existing on site, or when the terminal 101 transmits/receives packets to/from the management device 102 via a plurality of communication means, a plurality of communications are performed.
- a form in which the I/F 111 is mounted may be used.
- the CPU 112 executes various computer programs stored in the storage device 115 , thereby realizing various functions of the terminal 101 .
- the input unit 113 is composed of, for example, a keyboard, mouse, or touch panel, and is used by the operator to input various operations and settings.
- the output unit 114 is composed of, for example, a liquid crystal display monitor, and displays setting screens and results of various processes.
- the input unit 113 and the output unit 114 are not essential.
- the storage device 115 includes, for example, a storage device configured by a read-only semiconductor memory or the like and a storage device configured by a rewritable semiconductor memory element or the like, and stores computer programs for realizing various processes, acquired data, and the like. Store.
- the application program 116 manages various settings such as the acquisition method and transmission schedule of the data to be collected, and the CPU 112 connected via the internal bus executes acquisition processing and transmission commands to the communication processing unit 117 .
- the application program 116 manages a method of acquiring log information to be transmitted to the management apparatus 102, a transmission schedule of log information, and the like.
- the application program 116 stores the collected log information in the log buffer 300 and transmits the log information to the management device 102 at a predetermined timing.
- the log information can include acquired information such as the state of the terminal 101 and measured values of sensors connected to the terminal 101 , a time stamp, and an identifier of the terminal 101 .
- the application program 116 can function as a log information management unit that collects log information and transmits the log information to the management device 102 at a predetermined timing.
- the communication processing unit 117 implements transmission/reception processing in communication. Specifically, packet analysis processing including packet assembly processing for transmission and determination of whether or not the packet is addressed to the terminal itself for reception is performed.
- the self-failure analysis unit 118 periodically analyzes whether there is an abnormal event in the terminal by referring to the log information from the log buffer 300 of the terminal itself. Further, when the self-failure analysis unit 118 detects an abnormal event, it determines the self-restoration process to be taken and the contents of notification to the management device 102 . Details of the processing performed by the self-failure analysis unit 118 will be described in detail later with reference to FIG.
- the recovery notification transmission unit 119 transmits notification contents (recovery notification) determined by the self-failure analysis unit 118 to the management device 102 before executing self-recovery processing for the abnormal event detected by the self-failure analysis unit 118 .
- the recovery notification transmission unit 119 outputs notification contents to the communication processing unit 117, and executes transmission processing.
- the self-recovery processing unit 120 executes the self-recovery processing determined by the self-failure analysis unit 118.
- the contents of the self-recovery process are preset processes such as restarting the communication I/F 111, restarting the terminal body, and executing firmware update, and are not limited to specific processes.
- the failure analysis condition management unit 121 manages failure analysis conditions that define analysis rules for log information. Specifically, it manages a failure analysis condition management table 121a, which will be described later with reference to FIG.
- the event determination condition management unit 122 manages event determination conditions that define which events are detected for combinations of failure analysis conditions that match the failure analysis condition management table 121a. Specifically, it manages an event determination condition management table 122a, which will be described later with reference to FIG.
- the recovery process management unit 123 manages the recovery process to be taken for each detected abnormal event and the contents of notification to the management device 102 . Specifically, the recovery processing management table 123a managed in FIG. 4 is managed.
- the terminal 101 may be an independent device or may be a built-in device. Further, as described above, the terminal 101 has various configurations depending on the type of field data to be acquired, and may include, for example, a temperature sensor, a camera module, or an acceleration sensor.
- Communication processing unit 117, self-failure analysis unit 118, recovery notification transmission unit 119, self-recovery processing unit 120, failure analysis condition management unit 121, event determination condition management unit 122, and recovery processing management unit 123 are implemented as programs. It is loaded into the storage device 115 and executed by the CPU 112 .
- the CPU 112 operates as a functional unit that provides a predetermined function by executing processing according to the program of each functional unit.
- CPU 112 functions as self-failure analysis unit 118 by executing processing according to a self-failure analysis program. The same is true for other programs.
- the CPU 112 also operates as a functional unit that provides functions of multiple processes executed by each program.
- Computers and computer systems are devices and systems that include these functional units.
- the management device 102 includes a communication I/F 131, a CPU 132, an input unit 133, an output unit 134, a storage device 135, a communication processing unit 137, a failure analysis condition management unit 141, an event determination condition management unit 142, and a recovery process management unit 143. These components, including, are similar to the components of the terminal 101 described above.
- FIG. 1 similar to the terminal 101 described above, reception of input information from an external device via the communication I/F 131, such as remote login from another external device to the management apparatus 102, and output to the external device When providing information, it is not essential to install the input unit 133 and the output unit 134 .
- the application program 136 is a program that provides users with services that utilize data and log information collected from the terminal 101 .
- the application program 136 is a program that provides an average value per unit time of on-site data (temperature, etc.) received from the terminal 101, data analysis processing such as calculating an average value from the values of the collected data. I do.
- the application program 136 also includes a program for remotely setting and managing the transmission schedule of data and log information in the terminal 101 .
- the application program 136 stores log information collected from each terminal 101 in the log accumulation information 200 .
- the application program 136 also includes a case of managing operation plan information such as a planned stop time of the communication system.
- the application program 136 stores the operation plan information in the system operation plan information 1300 and manages it.
- the storage and management of the operation plan information is not essential in this embodiment. An example of utilizing the operation plan information will be described later in Example 3.
- the failure analysis unit 138 manages the log information of the log accumulation information 200 collected from the terminal 101, analyzes the log information, and analyzes the presence or absence of an abnormal event in each terminal 101. It is similar to the failure analysis unit 118 .
- the failure analysis unit 138 periodically determines the presence or absence of an abnormal event that has occurred in each terminal 101 by analyzing the log information of the collected log accumulation information 200, and determines recovery processing to be taken when an abnormal event is detected. do. Details of the processing performed by the failure analysis unit 138 will be described in detail with reference to FIG. 7, which will be described later.
- the recovery notification receiving unit 139 manages notifications regarding self-recovery processing received from each terminal 101 . Specifically, as a result of analyzing the received packet in the communication processing unit 137, if the packet is a notification related to self-restoration processing, it is notified to the restoration notification reception unit 139, and the restoration notification reception unit 139 receives from which terminal 101, Record what notifications you receive.
- the recovery command unit 140 commands the terminal 101 at the source of the error to execute the recovery process determined by the failure analysis unit 138 .
- the recovery command unit 140 refers to the recovery notification receiving unit 139 and determines whether or not the command is necessary depending on whether or not there is a recovery notification regarding the self-recovery process from the terminal 101 .
- the recovery command unit 140 If the recovery command unit 140 has not received a notification regarding the self-recovery process from the failed terminal 101, it sends a recovery process execution command (hereinafter referred to as a recovery command) to the terminal 101 in question. Specifically, it notifies the communication processing unit 137 of the restoration instruction to transmit the restoration instruction to the terminal 101 .
- a recovery command a recovery process execution command
- the recovery command unit 140 cancels and discards the recovery command, thereby avoiding redundant recovery commands. Details of this process will be described later with reference to FIG.
- the management apparatus 102 may be configured separately into a plurality of management apparatuses 102 including all the configurations in FIG. Therefore, it is not essential to include all of the configuration shown in FIG. 1 in one management apparatus 102 .
- failure analysis condition management tables 121a and 141a managed by the failure analysis condition management unit 121 of the terminal 101 and the failure analysis condition management unit 141 of the management device 102 will be described with reference to FIG.
- the failure analysis condition management tables 121a and 141a manage failure analysis conditions defining analysis rules for various types of log information of the terminal 101.
- FIG. 2 shows a block diagram of the failure analysis condition management table in the first embodiment. .
- the analysis condition ID 201 indicates the identifier of the failure analysis condition, and a unique identifier is set for each failure analysis condition.
- the identifier in this field can be in any format. As illustrated in FIG. 2, the type of log information to be referenced may be expressed as a character string as an identifier, or the identifier may be defined as a serial number such as "No. 1, No. 2, ". .
- the reference information 202 indicates the type of log information to be referred to in the failure analysis condition among the log information of the terminal 101 .
- the log information type is described in the form of a character string that is easy for humans to interpret, such as "LTE connection status", but the format of this description is arbitrary. For example, if individual log information is stored in each register of the terminal 101 and the register value can be substituted for the log information type, the register value may be described in the reference information 202 .
- the comparison method 203 indicates whether the value described in the threshold value 205, which will be described later, is compared as an "absolute value” or as a “relative value” (amount of change) from the previous reference value. For example, in the failure analysis condition management table illustrated in FIG. Determine whether the CPU usage rate of the terminal 101 is 95% (0.95) or more, and if a "relative value" is registered, check whether the current CPU usage rate has increased by 95% or more from the previous reference value. judge.
- FIG. 2 exemplifies two patterns of "absolute value” and “relative value” as options for the comparison method 203.
- Any comparison method such as “average”, “maximum”, “minimum” may be added.
- the threshold value 205 indicates a comparison reference value for the log information specified in the reference information 202.
- the description format of the threshold value 205 is not limited to numerical values, and may be character strings. For example, as exemplified in FIG. Specify the character string "disconnect" in .
- the match count 206 indicates how many times the analysis conditions specified by the comparison method 203, the comparison condition 204, and the threshold 205 for the log information type specified by the reference information 202 are satisfied in succession before the conditions are considered to be met. ing.
- failure analysis condition management tables 121a and 141a are managed by both the terminal 101 and the management device 102, the contents of failure analysis conditions registered in each may be different.
- the row of "failure analysis condition ID: FirmVer” illustrates failure analysis conditions for analyzing whether the firmware version applied to the terminal 101 is less than 14.01.
- the management device 102 manages the latest firmware version information.
- the management device 102 can update the setting value of the threshold value 205, for example, from “14.01” to "15.00". Unless instructed by the management device 102, it is difficult to update the failure analysis condition management table 121a in the same manner as the management device 102. FIG.
- the failure analysis condition related to "failure analysis condition ID: FirmVer” is registered only in the failure analysis condition management unit 141 of the management apparatus 102, and is not registered in the terminal 101, thereby reducing the load of failure analysis processing in the terminal 101. It becomes possible to Also, different threshold values 205 and matching counts 206 may be specified for the same log information for the terminal 101 and the management apparatus 102 .
- failure analysis condition management tables 121a and 141a managed by the terminal 101 and the management device 102 it becomes possible to flexibly detect an abnormality occurring in the terminal 101.
- various failures and anomalies can be detected by registering failure analysis conditions related to anomalies that can be detected only by either the terminal 101 or the management device 102 .
- the event determination condition management tables 122a and 142a managed by the event determination condition management unit 122 of the terminal 101 and the event determination condition management unit 142 of the management device 102 will be described with reference to FIG.
- the event determination condition management tables 122a and 142a manage event determination conditions that define which event should be detected as an abnormal event based on the match determination result for each failure analysis condition.
- FIG. 3 shows a configuration diagram of the event determination condition management tables 122a and 142a in the first embodiment.
- the event ID 301 indicates an identifier for distinguishing the type of abnormal event to be detected, and a unique identifier is registered for each event.
- the naming rule of the identifier is arbitrary, and in the example of FIG. , identifiers may be written.
- a matching condition (1) 302, a matching condition (2) 303, and a matching condition (3) 304 indicate identifiers of failure analysis conditions to be matched in detecting an abnormal event specified by the event ID 301.
- any of the identifiers listed in the analysis condition ID 201 of the failure analysis condition management tables 121a and 141a of FIG. 2 is registered in this field. As illustrated in the row of "Event ID: LackRxCapability (Insufficient reception capability due to increased load)" in FIG. 3, a plurality of failure analysis condition IDs may be specified for a single event.
- the abnormal event is detected only when all of these failure analysis conditions are met.
- failure analysis condition ID CpuRatio (a CPU usage rate of 95% or more is detected three times 300 packets in a row)” is satisfied, the terminal 101 detects “insufficient reception capability due to increased load” as an abnormal event.
- the event determination condition management tables 122a and 142a are managed by both the terminal 101 and the management device 102, the contents of the event determination conditions registered in each may be different.
- the recovery process management tables 123a and 143a managed by the recovery process management unit 123 of the terminal 101 and the recovery process management unit 143 of the management device 102 will be described with reference to FIG.
- the recovery process management tables 123a and 143a manage recovery processes to be performed for each abnormal event, and FIG.
- the event ID 401 indicates an identifier for distinguishing types of abnormal events. It is similar to the event ID 301 of the event determination condition management table in FIG. 3, and the description format conforms to the event ID 301 in FIG.
- the notification message 402 indicates the contents of notification to be sent to the management device 102 when the terminal 101 detects an abnormal event described in the event ID 401.
- registration is made to notify a combination of a character string indicating the content of recovery processing (eg, LteReboot (communication I/F restart related to LTE)) and a detected event ID (eg, LteDisconn (LTE disconnection)).
- the content of this notification may be in any format.
- the management apparatus 102 when the management apparatus 102 receives this notification, the format must be such that it can determine what kind of self-recovery processing the terminal 101 that is the source of the notification executes or what kind of abnormal event is detected. is desirable.
- the field of the notification message 402 may be left blank or omitted in the recovery process management table managed by the management device 102 .
- a recovery measure 403 indicates a recovery process to be taken when the terminal 101 or the management device 102 detects an abnormal event described in the event ID 401.
- the self-recovery processing unit 120 executes self-recovery processing registered in the recovery measure 403 .
- the restoration instruction unit 140 transmits an instruction to execute the restoration process registered in the restoration measure 403 to the target terminal 101 .
- the recovery process is described in a format that prioritizes readability such as "Communication I/F restart (LTE)". It can be in any format.
- the commands necessary for executing the recovery process may be written directly in the recovery measure 403.
- the commands necessary for executing the recovery process may be written directly in the recovery measure 403.
- only a single recovery process is described for each abnormal event. , and information such as the order of execution may be added as necessary.
- the waiting time 404 indicates the waiting time until recovery processing of the recovery measure 403 is taken when an abnormal event described in the event ID 401 is detected.
- the terminal 101 detects an abnormal event described in the event ID 401
- the terminal 101 waits for the time described in the waiting time 404 from the detection of the abnormal event, and then executes the self-recovery process described in the recovery measure 403 .
- the management apparatus 102 when the management apparatus 102 detects an abnormal event, it waits for the time described in the waiting time 404 from the detection of the abnormal event, and then issues an execution command for the recovery processing described in the recovery measure 403 to the recovery command unit. 140 to the terminal 101 .
- recovery processing such as restarting the terminal 101
- data collection from other on-site devices may be interrupted. Inconvenient cases are also assumed.
- the waiting time 404 may be set to "0 minutes" if it is necessary to take immediate recovery processing immediately after detecting an abnormal event. Also, in the example of FIG. 4, the waiting time 404 is specified in minutes, but the unit of time may be changed arbitrarily.
- the recovery process management tables 123a and 143a are managed by both the terminal 101 and the management device 102, the contents of the tables registered in each may be different. For example, it is conceivable to set the value of the waiting time 404 in the management device 102 longer than the waiting time set in the terminal 101 .
- any field may be added in addition to the fields illustrated in FIG.
- a field such as priority may be added in order to clarify recovery processing that should be prioritized in case multiple abnormal events are detected at the same time.
- the failure analysis condition management tables 121a and 141a in FIG. 2, the event determination condition management tables 122a and 142a in FIG. 3, and the recovery process management tables 123a and 143a in FIG. may be registered or changed at any timing.
- a screen display example related to these registrations or changes will be described later with reference to FIG.
- the event determination condition management tables 122a and 142a in FIG. 3 and the recovery process management tables 123a and 143a in FIG. may be in a form in which they are collectively managed in one table.
- FIG. 5 illustrates a case in which recovery from a failure is realized by self-failure analysis and self-restoration processing of the terminal 101, and the details will be described below.
- the type of log information to be transmitted includes at least the log information described in the reference information 202 of the failure analysis condition management table of FIG. However, log information not described in the reference information 202 may be included in the transmission.
- Terminal 101 acquires its own log information in accordance with a method managed by application program 116 , temporarily stores it in log buffer 300 , and then operates communication processing unit 117 according to a transmission schedule also managed by application program 116 . log information is transmitted to the management apparatus 102 via the
- step S501a As illustrated in FIG. 5, for example, if the application program 116 specifies that the log information should be transmitted at a predetermined log transmission period ⁇ T, the application program 116 transmits the log information in step S501a. After that, in step S501b after the log transmission cycle ⁇ T has passed, the log information is transmitted again.
- the self-failure analysis unit 118 of the terminal 101 self-analyzes the presence or absence of an abnormal event in the own terminal based on the log information of the own terminal, These are restoration processing and processing for determining the content of notification to the management apparatus 102 .
- the execution timing of the self-failure analysis processing in steps S502a, S502b, and S502c is managed by the self-failure analysis unit 118 of the terminal 101. For example, as shown in FIG. is executed in
- the self-failure analysis unit 118 executes the self-failure analysis process again after the elapse of the self-failure analysis period ⁇ t1.
- the self-failure analysis unit 118 detects an abnormal event (eg, step S502c)
- the self-failure analysis unit 118 determines the contents of the notification addressed to the management device 102 and the self-restoration process to be executed, and The process of S505 is executed.
- the failure analysis unit 138 of the management device 102 analyzes whether or not an abnormal event has occurred in the terminal 101 based on the log information of the accumulated log information 200 collected from the terminal 101, and detects an abnormal event. In this case, the recovery process to be instructed to the terminal 101 is determined.
- the execution timing of the fault analysis processing in steps S503a to S503c is managed by the fault analysis unit 138 of the management device 102, and is executed at time intervals of the fault analysis cycle ⁇ t2, as illustrated in FIG. 5, for example.
- the failure analysis unit 138 executes the failure analysis process again after the failure analysis period ⁇ t2 has passed, and detects the failure event. If detected (eg, step S503c), the process of step S506, which will be described later, is executed.
- the log transmission cycle ⁇ T by the terminal 101 in step S501, the self-failure analysis cycle ⁇ t1 by the terminal 101 in step S502, and the failure analysis cycle ⁇ t2 by the management device 102 in step S503 are set to different cycles. I don't mind.
- the failure analysis cycle ⁇ t2 must be set longer.
- the self-failure analysis cycle ⁇ t1 by the terminal 101 can be set short regardless of the communication band or communication fee.
- the self-failure analysis cycle ⁇ t1 short it is possible to shorten the time required from failure occurrence to detection, compared to a form in which only the management apparatus 102 executes failure analysis processing.
- steps S501a to S503c are executed periodically
- these may be executed at specified times.
- the management apparatus 102 it is possible to set the management apparatus 102 to execute failure analysis processing at 12:00 and 18:00 every day, and the execution schedule of steps S501a to S503c may be set arbitrarily.
- Step S504 is a process in which the recovery notification transmission unit 119 of the terminal 101 transmits a notification regarding the self-recovery process to the management device 102.
- This notification is the notification content determined when the self-failure analysis unit 118 detects an abnormal event in step S502c. Specifically, in the recovery process management table of FIG. This corresponds to the notification content described in message 402 .
- the recovery notification transmission unit 119 notifies the communication processing unit 117 of the content of the notification, and transmits a packet containing the content of the notification to the management device 102 .
- the communication processing unit 137 analyzes the packet to determine which terminal 101 transmitted the notification, what kind of abnormal event was detected, or what kind of recovery process was performed. It notifies the restoration notification receiving unit 139 of the notification content such as whether to take measures, and records the notification content.
- step S504 is executed after the detection of an abnormal event in step S502c.
- the recovery notification transmitting unit 119 may suspend (delay) the transmission of the notification, and transmit the notification immediately before executing the self-recovery process in step S505, which will be described later.
- Step S505 is processing in which the self-restoration processing unit 120 of the terminal 101 executes the self-restoration processing determined when the abnormal event is detected in step S502c. Specifically, the recovery process described in the recovery measure 403 is executed in the recovery process management tables 123a and 143a of FIG.
- the self-recovery processing unit 120 executes the recovery process. As a result, the terminal 101 can realize self-recovery from the detected abnormal event.
- Step S506 is a process in which the recovery command unit 140 of the management apparatus 102 determines whether or not it is necessary to issue a command to the terminal 101 where the error has occurred, for the recovery process determined in step S503c. be.
- the recovery command process is executed at the timing when the time specified in the waiting time 404 in the recovery process management table of FIG.
- the recovery command unit 140 refers to the recovery notification receiving unit 139 to determine whether the recovery notification described in step S504 has been received from the terminal 101. If it has been received, step S507 will be described later. to cancel the recovery order.
- step S601 If the recovery notification has not been received, the process proceeds to step S601, which will be described later with reference to FIG. . Details of this process will be described later with reference to FIG. In the example of FIG. 5, since the recovery notification has already been received from the terminal 101 in step S504, the process proceeds to step S507.
- step S507 the recovery command unit 140 of the management device 102 cancels the recovery process execution command determined in step S503c. If a recovery notification has already been received from the terminal 101 that caused the error in step S504, the terminal 101 will perform self-recovery processing without an instruction from the management device 102. Therefore, the recovery command should be redundantly transmitted from the management device 102. isn't it. By canceling the recovery command by the recovery command unit 140 in step S507, the terminal 101 can recover from the abnormal event without performing redundant recovery processing.
- FIG. 5 exemplifies a mode in which the terminal 101 transmits a recovery notification in step S504 and then executes self-recovery processing in step S505, the order may be reversed.
- the management device 102 can cancel the redundant recovery command with a higher probability, thereby improving efficiency. can be planned.
- steps S501a to S501c, steps S502a to S502c, steps S503a to S503c, and step S506 are the same as in FIG.
- step S503c after the fault analysis unit 138 detects an abnormal event of the terminal 101 in step S503c, the recovery notification from the terminal 101 is not performed until the recovery instruction necessity determination process in step S506. is not received, the process proceeds to step S601.
- step S601 the recovery command unit 140 of the management device 102 commands the terminal 101 in which the abnormal event occurred to perform the recovery process determined in step S503c. Specifically, the recovery command unit 140 notifies the communication processing unit 137 of the recovery process, and transmits a packet containing the details of the recovery process to be executed to the terminal 101 .
- step S602 the terminal 101, which has received the restoration command transmitted in step S601, executes the restoration processing specified by the management device 102 using the self-restoration processing unit 120. Specifically, when the terminal 101 receives a restoration command from the management device 102, the communication processing unit 117 analyzes the packet, notifies the self-restoration processing unit 120 of restoration processing to be executed, and the self-restoration processing unit 120 executes the restoration process. As a result, recovery from the abnormal event of the terminal 101 is realized.
- the terminal 101 when the terminal 101 does not self-detect an abnormal event and the management apparatus 102 detects an abnormal event, the terminal 101 has not transmitted a recovery notification as described in step S504 of FIG. With this, the management device 102 transmits a recovery command.
- the management device 102 Even if the terminal 101 is capable of self-detecting an abnormal event, if the management device 102 detects it early before the terminal 101 self-detects it, the management device 102 can transmit a recovery command as shown in FIG. , it is possible to realize fault recovery at an early stage, compared with a form in which only the terminal 101 performs self-failure analysis.
- the self-failure analysis processing executed by the self-failure analysis unit 118 of the terminal 101 and the failure analysis processing executed by the failure analysis unit 138 of the management device 102 will be described with reference to FIG. Specifically, it corresponds to the processing executed in steps S502a to S502c and S503a to S503c in FIGS.
- the terminal 101 and the management device 102 refer to the log information of the terminal 101 to analyze the presence or absence of an abnormal event, and if an abnormal event is detected, determine the recovery process to be taken.
- the self-failure analysis processing by the self-failure analysis unit 118 of the terminal 101 the content of the recovery notification to be notified to the management device 102 is also determined in this processing.
- FIG. 7 is a flowchart showing self-failure analysis processing by the terminal 101 and failure analysis processing by the management device 102 in Embodiment 1, and the details will be described below. Note that the management device 102 processes unprocessed log information in the accumulated log information 200 , and the terminal 101 processes unprocessed log information in the log buffer 300 .
- the failure analysis unit 138 (self-failure analysis unit 118) refers to the failure analysis condition management tables 121a and 141a of FIG. This is a process of performing determination of an abnormal event.
- the failure analysis condition management table 121a managed by the failure analysis condition management unit 121 is referenced, and the log information of the own terminal is analyzed.
- the failure analysis condition management table 141a managed by the failure analysis condition management unit 141 is referred to, and the log information collected from each terminal 101 is analyzed. to run.
- the failure analysis unit 138 When analysis processing is completed for all failure analysis conditions registered in the failure analysis condition management tables 121a and 141a and there is a failure analysis condition that matches the conditions, the failure analysis unit 138 (self failure analysis unit 118) performs the analysis.
- a condition ID (analysis condition ID 201 in FIG. 2) is recorded. For example, in the log information of the terminal 101, if the "LTE connection state" is "0 (disconnected)" three times in a row in the past, "analysis condition ID: LteState" is stored in the example of the table in FIG. Record on device 115 , 135 . After the process of step S701 is completed, the process proceeds to step S702.
- Step S702 is a process in which the failure analysis unit 138 (self-failure analysis unit 118) determines whether or not there is an analysis condition ID that matches the failure analysis conditions in the processing of step S701.
- step S703 If there is at least one matching analysis condition ID (YES), the process proceeds to step S703. On the other hand, if there is no matching analysis condition ID (NO), it is clear that there is no abnormal event to be detected, so the process of FIG. 7 is terminated.
- step S703 the failure analysis unit 138 (self-failure analysis unit 118) refers to the event determination condition management tables 122a and 142a of FIG. This is processing for determining an event.
- the event determination condition management table 122a managed by the event determination condition management unit 122 is referenced to execute determination.
- the event determination condition management table 142a managed by the event determination condition management unit 142 is referenced to execute determination.
- the failure analysis unit 138 (self-failure analysis unit 118) records the event ID (event ID 301 in FIG. 3) in the storage devices 115, 135 when a corresponding abnormal event exists.
- step S703 detects an abnormal event of “LTE disconnection” in the terminal 101 and records the event ID.
- step S704 the failure analysis unit 138 (self-failure analysis unit 118) determines whether or not the corresponding event ID 301 exists in the processing of step S703. As in the case of "event ID: LteDisconn (LTE disconnection)" described above, if at least one corresponding event ID 301 exists (YES), the process proceeds to step S705. On the other hand, if the corresponding event ID does not exist (NO), it is determined that there is no abnormal event in the terminal 101, and the processing of FIG. 7 ends.
- step S705 the recovery instruction unit 140 (self-recovery processing unit 120) refers to the recovery processing management tables 123a and 143a of FIG. This is the process of determining the execution timing of the recovery process.
- the recovery instruction unit 140 (self-recovery processing unit 120) refers to the recovery measure 403 and waiting time 404 fields, and executes the corresponding event ID 301 in step S703.
- the recovery process to be executed and the waiting time 404 between which the recovery process should be executed or commanded are determined.
- the recovery command unit 140 self-recovery processing unit 120
- “communication I about LTE” after "5 minutes” in the example of FIG. /F reboot” can be determined to be executed or commanded.
- the self-restoration processing unit 120 refers to the restoration processing management table 123a managed by the restoration processing management unit 123 and executes the processing of step S705. Further, the self-restoration processing unit 120 also refers to the field of the notification message 402 of the restoration processing management table 123a, and determines the contents of the notification to be notified to the management device 102 as well.
- the notification content determined in this process is transmitted to the management device 102 in the recovery notification transmission process in step S504 of FIG. 5, and after the specified waiting time has passed, the recovery process is executed in step S505 of FIG.
- step S705 is executed.
- the failure analysis unit 138 After the specified waiting time has passed, the failure analysis unit 138 performs the recovery command necessity determination processing in step S506 shown in FIGS. , an execution command for the recovery process is transmitted.
- the process of step S705 ends, the process of FIG. 7 ends.
- the terminal 101 and the management device 102 can analyze the presence or absence of an abnormal event in the terminal 101, restore processing that should be performed when an abnormal event is detected, and the execution timing of the restoration processing. can be determined up to
- the restoration instruction necessity determination process executed by the restoration instruction unit 140 of the management device 102 will be described. Specifically, this corresponds to the processing executed in step S506 in FIGS.
- FIG. 8 is a flow chart showing the process of determining whether or not a restoration command is required by the management device 102 according to the first embodiment, and the details will be described below.
- Step S801 is a process in which the restoration instruction unit 140 of the management device 102 refers to the restoration notification reception unit 139 and determines whether or not the restoration notification from the terminal 101 subject to the restoration instruction has been received.
- the recovery command unit 140 determines whether or not a recovery notification has been received from the terminal 101 in which the abnormal event detected in the failure analysis process of FIG. 7 has occurred. If it is related to the abnormal event detected in the process or the recovery process determined in the failure analysis process of FIG. 7, it is determined that it has been received.
- step S801 as a result of the restoration instruction unit 140 referring to the restoration notification reception unit 139, the terminal 101 sends a restoration notification indicating that “LTE disconnection” has been self-detected, or “restarts communication I/F related to LTE”. If a recovery notification to the effect that it will be executed as self-recovery processing has been received, it is determined that it has been received.
- the recovery command unit 140 can determine that the recovery command should be withheld in order to avoid redundant recovery processes. If the management apparatus 102 has already received the restoration notification from the target terminal 101 (YES), the process proceeds to step S802, and if not (NO), the process proceeds to step S803.
- Step S802 is a process in which the recovery command unit 140 determines that the recovery process execution command determined in the failure analysis process of FIG. 7 is unnecessary, and cancels the recovery command. Specifically, the process transitions to the process of canceling the restoration instruction described in step S507 of FIG.
- step S801 it is determined that the recovery notification from the target terminal 101 has been received, so that the recovery command is canceled, thereby avoiding execution of redundant recovery processing in the terminal 101. can be done.
- step S802 ends, the process of FIG. 8 ends.
- Step S803 is a process in which the recovery command unit 140 determines that the recovery process execution command determined in the failure analysis process of FIG. 7 is necessary, and transmits the recovery command to the target terminal 101 . Specifically, the process transitions to the recovery instruction transmission process described in step S601 of FIG.
- step S801 If it is determined in step S801 that the recovery notification has not been received from the target terminal 101, the terminal 101 cannot self-detect an abnormal event. Action can be taken. When the process of step S803 ends, the process of FIG. 8 ends.
- the recovery command unit 140 of the management device 102 can determine whether the recovery process command determined in the failure analysis process of FIG. 7 is necessary. Further, when it is determined that the command is unnecessary based on the determination of the reception of the recovery notification from the terminal 101 to which the recovery order is issued, the recovery command is canceled, thereby preventing the terminal 101 from executing redundant recovery processing. It can be suppressed.
- the output unit 114 of the terminal 101 and the output unit 134 of the management apparatus 102 can display a display screen 900.
- a display area 901 for setting failure analysis conditions and an event determination condition and a display area 903 for setting recovery processing are provided.
- a display area 901 is an area for setting the failure analysis condition management tables 121a and 141a of FIG.
- the input values are set in the failure analysis condition management tables 121a and 141a.
- failure analysis conditions may be added at any time during the operation process of the communication system. Therefore, in the example of FIG. 9, an add button 904 for adding input fields (rows) for failure analysis conditions is provided. making it possible.
- a display area 902 is an area for setting the event determination condition management tables 122a and 142a of FIG.
- the input values are set in the event determination condition management tables 122a and 142a.
- the type of abnormal event to be detected and the number of failure analysis conditions required for match determination may be added at any time during the operation of the system, so an add button 904 is also provided.
- a display area 903 is an area for setting the recovery process management tables 123a and 143a of FIG.
- the display area 903 when the recovery process to be taken for each abnormal event and the waiting time until the recovery process is entered from the input units 113 and 133, the input values are set in the recovery process management table. For the same reason as above, the display area 903 also has an add button 904 .
- a file input button 905 and a file output button 906 are provided.
- a function is provided to read various table values saved in separate external files onto the display screen 900 of FIG. provides a function of outputting table values input on the display screen 900 of FIG. 9 to an external file.
- the output unit 134 of the management device 102 In addition to displaying the screen of FIG. 9 on each of the terminal 101 and the management device 102, for example, the output unit 134 of the management device 102 also displays a screen for setting various table information of the terminal 101, and presses the file output button. It is also conceivable to register various table values of the terminal 101 by sending the external file output by pressing the button 906 to the terminal 101 . Thus, it is not essential that both the terminal 101 and the management device 102 include an input unit and an output unit.
- a screen display example for outputting an abnormal event detected by the self-failure analysis processing by the terminal 101 in FIG. 7 or by the failure analysis processing by the management device 102 and the determined recovery processing will be described with reference to FIG.
- the output unit 114 of the terminal 101 and the output unit 134 of the management apparatus 102 can display a display screen 1000, and a display area 1001 for displaying failure detection and recovery information is set on the display screen 1000. .
- detected abnormal event information analyzed by the self-failure analysis unit 118 of the terminal 101 and the failure analysis unit 138 of the management device 102, and recovery processing information for the abnormal event are displayed.
- the identifier information of the terminal 101 that is the source of the failure, the detected abnormal event information, the detection time of the abnormal event, the detection source information of the abnormal event, and the restoration process to be performed for the abnormal event. is displayed, but any information may be output as necessary.
- the scheduled execution time of the recovery process determined from the waiting time 404 in the recovery process management tables 123a and 143a of FIG. 4 may be added.
- recovery processing such as restarting of the terminal 101
- that may involve interruption of communication system operation is executed, by referring to the display screen 1000, it is possible to take measures such as taking advance preparations for system stoppage. It is also possible to determine whether or not it is necessary.
- a screen display that outputs in text format instead of table format may be used, and the display format of information related to failure detection and recovery is not limited to a specific method. do not have.
- self-analysis and recovery of the terminal 101 enable early failure detection and self-recovery from a communication disconnection state.
- the terminal 101 can also detect an abnormal event that cannot be self-detected, and restore it.
- the terminal 101 takes self-recovery processing, by transmitting a recovery notification to the management device 102, it is possible to suppress redundant recovery orders from the management device 102, and the failure recovery of the terminal 101 can be efficiently performed. can be practically realized. Since these processes are automatically performed, the number of man-hours required for recovery work in the event of a failure can be reduced, and early failure recovery can be realized, which in turn contributes to an improvement in the operating rate of the communication system.
- the management apparatus 102 when the management apparatus 102 performs the failure analysis processing in the failure analysis unit 138, the failure analysis condition management tables 121a and 141a of FIG. It was determined whether or not they matched.
- self-recovery processing by the terminal 101 includes processing that may involve interruption of the operation of the communication system, such as restarting the terminal body. It is desirable to temporarily disable the restoration process by the terminal 101 .
- the management apparatus 102 in the management apparatus 102, the fact that a plurality of terminals 101 simultaneously meet the same failure analysis condition is added as a criterion for detecting an abnormal event, and self-recovery of the terminal 101 according to the type of the abnormal event is also added. A form in which enabling or disabling of processing is also possible will be described.
- the failure analysis condition management unit 141 the event determination condition management unit 142, the recovery processing management unit 143 of the management device 102, and the recovery processing management unit 123 of the terminal 101 manage various Explain the structure of the table.
- Various configurations, processes, and the like according to the second embodiment are the same as those of the first embodiment except for the configuration of the tables shown in FIGS. 11A to 11D.
- FIG. 11A shows a configuration diagram of the failure analysis condition management table 141a managed by the failure analysis condition management unit 141 of the management device 102 in the second embodiment.
- the analysis condition ID 201, reference information 202, comparison method 303, comparison condition 204, threshold value 205, and number of matches 206 are the same as in the failure analysis condition management table 141a in the first embodiment of FIG. 1101 fields are newly added.
- the number of simultaneous detections 1101 indicates how many terminals 101 that meet the analysis condition exist as a result of analyzing the log information collected from each terminal 101 by the management device 102, and whether the conditions are considered to be met. ing.
- the management device 102 determines whether this failure analysis condition is met.
- FIG. 11B shows a configuration diagram of the event determination condition management table 142a managed by the event determination condition management unit 142 of the management device 102 in the second embodiment.
- the event determination condition management table 142a is the same as the event determination condition management table 142a in Example 1 of FIG. In the case of the example of FIG. 11B, based on the matching of "analysis condition ID: Multi-LteState" in FIG. can be detected.
- 11C and 11D are configuration diagrams of the recovery process management table 143a managed by the recovery process management unit 143 of the management device 102 and the recovery process management table 123a managed by the recovery process management unit 123 of the terminal 101 in the second embodiment. is shown.
- a valid/invalid 1102 in FIGS. 11C and 11D is a field for specifying whether to validate or invalidate execution of recovery processing when an abnormal event described in the event ID 401 is detected. If "invalid" is specified in this field, even if an abnormal event described in the event ID 401 is detected, the process described in the restoration measure 403 is not executed.
- step S705 when the management device 102 detects “event ID: NwFailure (carrier failure)”, the failure analysis processing ( In step S705), it is determined to invalidate the recovery process when self-detection of "LTE disconnection” is performed.
- the management device 102 performs the recovery command necessity determination process of FIG. 8 in the same manner as the flow described in FIG. Send to the target terminal. Then, the terminal 101 that has received this restoration command changes the valid/invalid 1102 field corresponding to "event ID: LteDisconn (LTE disconnection)" to "invalid" according to the command, as illustrated in FIG. 11D.
- the terminal 101 does not execute self-recovery processing even if "LTE disconnection" is detected, so that unnecessary self-recovery processing is not required during carrier failure.
- timing to re-enable the disabled recovery process can be set arbitrarily.
- the management device 102 instructs to re-enable the disabled recovery process, and the terminal 101 instructed to disable It may be in the form of automatic reactivation.
- the present embodiment in the failure analysis processing of the management apparatus 102, by adopting the fact that a plurality of terminals 101 simultaneously meet the same failure analysis condition as an abnormal event detection criterion, a single event can be detected. Abnormal events that cannot be detected only by the log information of the terminal 101 can also be detected.
- a field of waiting time 404 is provided in the recovery processing management tables 123a and 143a of FIG. explained.
- the terminal 101 By issuing a restoration command to the terminal 101 at the timing at which the management device 102 permits the execution of restoration processing, it is possible to set an arbitrary waiting time. For example, if the management device 102 holds information related to the operation plan of the communication system and knows the time period during which communication disconnection should not be caused due to the firmware update of the field device, etc., the terminal 101 will restore the communication during that time period. Even if the notification is received, the recovery order is not sent immediately, and the recovery order is withheld until the end of the relevant time period, thereby avoiding a temporary disconnection of communication due to self-recovery processing (such as restarting the terminal 101). be able to.
- the terminal 101 has detected an abnormality, it is actually an intentional abnormal event due to a planned system stop, etc., and there may be cases where self-recovery processing is unnecessary.
- the terminal 101 and another on-site device are connected by wired communication, and the terminal 101 self-detects disconnection of the wired communication when the on-site device is powered off due to a planned shutdown of the communication system.
- the terminal 101 transmits a recovery notification regarding the self-recovery process to the management apparatus 102
- execution of the recovery process is waited until a recovery command is received from the management apparatus 102.
- the terminal 101 executes self-recovery processing after transmitting the recovery notification to the management device 102. , suspend the execution of recovery processing.
- the restoration notification receiving unit 139 when the management device 102 receives a restoration notification from the terminal 101, the restoration notification receiving unit 139 only records the contents of the notification. It also determines permission to execute. Then, only when the management device 102 determines that the restoration is executable, the management device 102 sends back the restoration instruction to the terminal 101 at an appropriate timing. Note that the restoration command transmitted by the management device 102 is a reply to the restoration notification transmitted by the terminal 101 .
- Example 3 Various configurations and processes according to Example 3 are the same as those in Example 1 or Example 2 except for the processes shown in FIGS.
- Steps S501, S502, S504, and S602 of FIG. 12 are the same as those of FIGS. 5 and 6, and are the same as those of the first or second embodiment, so description thereof will be omitted.
- the terminal 101 After transmitting the recovery notification in step S504, the terminal 101 suspends execution of the recovery process until it receives a recovery command from the management device 102, as shown in FIG.
- the management device 102 receives the recovery notification transmitted by the terminal 101 in step S504
- the content of the notification is output to the recovery notification reception unit 139.
- the contents of the notification are also output to the recovery instruction unit 140, and the process proceeds to step S1201.
- step S1201 based on the recovery notification transmitted by the terminal 101 in step S504, the recovery command unit 140 of the management apparatus 102 determines whether the terminal 101 may be permitted to execute self-recovery processing for the notified abnormal event. This is the process of determining with .
- the system operation plan information 1300 (described later in FIG. 13) held by the management device 102 is referenced to determine whether the abnormal event detected by the terminal 101 is due to a planned shutdown of the communication system. When doing so, the transmission of the recovery command is canceled because the self-recovery process is unnecessary.
- step S1202 restores the system in step S1202 at an appropriate timing in terms of the operation plan of the communication system. Transition to command transmission processing. Note that detailed processing of step S1201 will be described later with reference to FIG.
- Step S1202 is a process of transmitting a restoration instruction to the target terminal 101 from the restoration instruction unit 140 of the management device 102 .
- the recovery command transmission process (S601) of FIG. 6 in the first or second embodiment the execution command of the recovery process determined in the failure analysis process (S503) is transmitted, but the recovery command transmission in step S1202 In the process, an execution command for the self-recovery process described in the recovery notification in step S504 is transmitted.
- the recovery process management table managed by the recovery process management unit 143 is referred to, and Sends an instruction to execute the corresponding restoration process.
- the terminal 101 causes the self-restoration processing unit 120 to execute the specified restoration process in step S602.
- the management device 102 determines whether the terminal 101 is permitted to perform the self-recovery process, and controls the transmission cancellation and transmission timing of the recovery command. It is possible to avoid execution of processing and execution of unnecessary recovery processing.
- system operation plan information 1300 held by the management device 102 An example of system operation plan information 1300 held by the management device 102 will be described with reference to FIG.
- the system operation plan information 1300 is managed by the application program 136 of the management device 102, and as described above, when the management device 102 receives a recovery notification, it is used to determine whether the terminal 101 is permitted to execute self-recovery processing. .
- the system operation plan information 1300 is stored in the storage device 135 of the management device 102. FIG.
- the operation plan 1301 in FIG. 13 indicates the operation plan information of the communication system.
- the description format of this field is arbitrary, and as illustrated in FIG. Communication I/F restart, etc.) may be written together.
- a related terminal 1302 is a field for describing the identifier of the terminal 101 targeted for the operation plan described in the operation plan 1301 .
- the identifiers of the terminals 101 are exemplified by subscripts in FIG. It doesn't matter if it is.
- the management apparatus 102 When the management apparatus 102 receives the recovery notification from the terminal 101 and determines whether or not to permit the self-recovery process, the management apparatus 102 refers to the related terminal 1302 to determine which operation plan is used in the recovery process of the terminal 101. It becomes possible to identify which should be considered.
- the target date and time 1303 indicates date and time information for executing the operation plan described in the operation plan 1301 .
- the management device 102 determines that the communication system is scheduled to be shut down during, for example, "2021/9/22 02:00:00 to 03:00:00", and the terminal 101 -a, the abnormal event detected by the terminal 101-b was caused by planning, and the management device 102 can determine that self-recovery is unnecessary.
- system operation plan information 1300 the fields of the operation plan 1301, the related terminal 1302, and the target date and time 1303 are illustrated, but the system operation plan information held by the management device 102 is limited to a specific format. not a thing For example, if various operation plans are applied periodically on a specific day of the week and a specific time period, change the target date and time field 1303 to a format that describes the day of the week and time period, and add a new input field can be added arbitrarily.
- the restoration instruction permission determination process executed by the restoration instruction unit 140 of the management apparatus 102 will be described. Specifically, this corresponds to the processing executed in step S1201 of FIG. This processing is executed when the management device 102 receives a recovery notification from the terminal 101, and determines whether or not the terminal 101 is permitted to execute self-recovery processing based on the system operation plan information 1300.
- FIG. 14 the restoration instruction permission determination process executed by the restoration instruction unit 140 of the management apparatus 102 will be described. Specifically, this corresponds to the processing executed in step S1201 of FIG. This processing is executed when the management device 102 receives a recovery notification from the terminal 101, and determines whether or not the terminal 101 is permitted to execute self-recovery processing based on the system operation plan information 1300.
- FIG. 14 is a flowchart showing an example of restoration order permission determination processing by the management device 102 according to the third embodiment, and the details will be described below.
- step S1401 the management device 102 determines whether the abnormal event detected by the terminal 101 that transmitted the restoration notification was caused by the planned shutdown of the communication system. Specifically, the restoration command unit 140 of the management device 102 refers to the system operation plan information 1300 of FIG. Determine whether it is
- the recovery command unit 140 determines that the event occurred due to a planned shutdown. can.
- step S1402. proceed to
- Step S1402 is a process in which the recovery command unit 140 of the management device 102 determines that a recovery command is unnecessary and cancels the recovery command in response to the recovery notification.
- step S1401 if the recovery command unit 140 can determine that the event is an abnormal event caused by a planned shutdown, self-recovery processing by the terminal 101 is unnecessary. Execution of self-recovery processing can be suppressed.
- the recovery command unit 140 only cancels the recovery command.
- the process of step S1402 ends, the process of FIG. 14 ends.
- Step S1403 is a process for determining whether there is no problem if the recovery command unit 140 immediately executes the recovery process in the terminal 101 that transmitted the recovery notification.
- the recovery command unit 140 of the management device 102 refers to the system operation plan information 1300 of FIG. judge.
- step S1404 If it can be determined that the recovery process can be immediately executed on the terminal 101 (YES), the process proceeds to step S1404, and if it is necessary to suspend the execution as in the above example (NO), the process proceeds to step S1405.
- step S1404 the recovery command unit 140 determines that the terminal 101 that sent the recovery notification may be permitted to immediately execute the recovery process, and transmits a recovery command to the terminal 101 concerned.
- the recovery command unit 140 immediately transitions to the recovery command transmission process described in step S1202 of FIG. 12, and issues a recovery command to the terminal 101.
- the process of step S1404 ends, the process of FIG. 14 ends.
- step S1405 the recovery command unit 140 determines that it is necessary to suspend the execution of the recovery process by the terminal 101 that has transmitted the recovery notification. to the terminal 101 concerned.
- the recovery command unit 140 refers to the system operation plan information 1300 in FIG. Then, a recovery command is issued to the terminal 101 .
- the recovery command unit 140 since it is necessary to wait at least 15 minutes after receiving the recovery notification, the recovery command unit 140 performs recovery command transmission processing after 15 minutes.
- step S1405 By suspending the transmission of the restoration command in step S1405, the terminal 101 can take self-restoration processing at an appropriate timing in terms of the operation plan of the communication system.
- step S1405 of FIG. 14 the transmission of the recovery command is suspended until the recovery process becomes executable. 101 immediately, and the terminal 101 receiving the instruction may suspend the execution of the self-recovery process until a specified timing.
- the management apparatus 102 determines whether or not the terminal 101 can execute the self-recovery process, thereby preventing unnecessary self-recovery from an abnormal event caused by an artificial planned shutdown of the communication system. In addition to suppressing the processing, it becomes possible to execute the restoration processing at a more appropriate timing for the operation plan of the communication system.
- interfering with the operation of the communication system through self-recovery processing is putting the cart before the horse, and by executing recovery processing at a timing that is in line with the operation plan, it is possible to contribute to further improving the operation rate of the communication system.
- the communication systems of the first to third embodiments can be configured as follows.
- a communication system comprising one or more terminals (101) to be managed and a management device (102) connected to the terminals (101), wherein the terminal (101)
- a log information management unit application program 116) that collects the log information of and transmits it to the management device (102); a self-failure analysis unit (118) that determines self-recovery processing when an abnormal event is detected; a self-recovery processing unit (120) that executes self-recovery processing or recovery processing commanded by the management device (102), wherein the management device (102) collects log information from the terminal (101).
- a failure analysis unit (138) that analyzes the presence or absence of an abnormal event in the terminal (101) and determines recovery processing when the abnormal event is detected, and a terminal (101) from the terminal ( a recovery notification receiving unit (139) for receiving a notification of self-recovery processing for the abnormal event detected by the device 101); and a recovery instruction unit (140) that instructs the terminal (101) to
- both the terminal and the management device 102 analyze an abnormal event that has occurred in the terminal 101 based on the log information of the terminal 101, so that self-analysis and restoration of the terminal 101 can detect failures early and prevent communication disconnection.
- the management apparatus 102 can detect abnormal events that cannot be self-detected by the terminal 101 and recover the terminal 101 by means of analysis and recovery instructions.
- the terminal 101 may redundantly execute the recovery processing. For example, assume that the terminal 101 and the management device 102 detect an abnormal event on the terminal 101 at approximately the same time. At this time, the terminal 101 performs self-restoration processing for the detected abnormal event, but the management device 102 instructs the terminal 101 to execute the restoration processing without knowing that self-restoration has been completed. As a result, the terminal 101 will execute redundant recovery processing again, and in the case of recovery by restarting, etc., the continuous operation of the terminal 101 will be hindered.
- the management device 102 when the terminal 101 executes self-recovery processing, the management device 102 explicitly transmits a notification regarding the recovery processing to the management device 102, so that the management device 102 determines that a recovery command is unnecessary, Redundant execution of recovery processing can be avoided.
- a management device (102) having a processor (CPU 132) and a memory (storage device 135) and connected to one or more terminals (101) to be managed, collecting data from the terminals (101)
- a failure analysis unit (138) that analyzes the presence or absence of an abnormal event in the terminal (101) based on the log information obtained and determines recovery processing when the abnormal event is detected, and from the terminal (101) , a recovery notification receiving unit (139) for receiving a notification of self-recovery processing for an abnormal event detected by the terminal (101);
- a management device (102) characterized by comprising a recovery instruction unit (140) that issues a process execution instruction to the terminal (101).
- both the terminal and the management device 102 analyze an abnormal event that has occurred in the terminal 101 based on the log information of the terminal 101, so that self-analysis and restoration of the terminal 101 can detect failures early and prevent communication disconnection.
- the management apparatus 102 can detect abnormal events that cannot be self-detected by the terminal 101 and recover the terminal 101 by means of analysis and recovery instructions.
- the management device (102) comprising: a failure analysis condition management unit (141) for managing failure analysis conditions for the log information; An event determination condition management unit (142) that manages event determination conditions based on combinations, and a recovery processing management unit (143) that manages recovery processing to be performed for each abnormal event, and the failure analysis unit (
- the terminal (138) analyzes the log information according to the failure analysis conditions of the failure analysis condition management unit (141), and analyzes the terminal (101) based on the analysis result and the event determination conditions of the event determination condition management unit (142).
- the recovery process management unit (143) is referred to determine the recovery process for the error event, and the recovery command unit (140) causes the recovery notification reception unit (139) to A management device (102) that commands an execution command of the recovery process to the terminal (101) when a notification of the self-recovery process for the abnormal event is not received from the terminal (101).
- the management apparatus 102 detects an abnormal event from the log information from the terminal 101, and if the notification regarding the restoration process has not been received from the terminal 101, the management apparatus 102 instructs the terminal 101 to execute the restoration process.
- the terminal 101 can be restored by detecting an abnormal event that cannot be self-detected by the terminal 101 .
- the failure analysis condition management unit (141) includes log information (reference information 202) referred to for each failure analysis condition and comparison criteria. a threshold value (205), a comparison condition (204) that defines a magnitude relationship with respect to the threshold value (205), a comparison method (203) that defines whether the threshold value (205) is treated as an absolute value or a relative value; A management device characterized in that it is possible to designate the number of times (matching number 206) that the log information matches the fault analysis condition.
- the failure analysis condition management unit 141 of the management device 102 stores log information (reference information 202) to be referred to for each failure analysis condition, the threshold value 205 as a comparison standard, and the threshold value 205 in the failure analysis condition management table 141a.
- a comparison method 203 that defines whether the threshold value 205 is treated as an absolute value or a relative value, and the number of times 206 that the log information matches the fault analysis condition can be specified. This makes it possible to set failure analysis conditions according to the environment of the terminal 101 .
- recovery process management unit (143) includes a recovery process (recovery measure 403) to be taken for each abnormal event, and A management device characterized by being able to designate a waiting time (404) until execution of a restoration command.
- the recovery process management unit 143 can specify the recovery measure 403 to be taken for each abnormal event and the waiting time 404 from the detection of the abnormal event to the execution of the recovery command in the recovery process management table 143a. It becomes possible to set a recovery process according to an abnormal event that may occur in the terminal 101 .
- the failure analysis condition managed by the failure analysis condition management unit (141) and the event determination managed by the event determination condition management unit (142) an input unit (133) for inputting conditions and recovery processing managed by the recovery processing management unit (143); information on the abnormal event detected by the failure analysis unit (138); and an output unit (134) for outputting the processing.
- the management device 102 can input failure analysis conditions, event determination conditions, and restoration processing through the input unit 133, and can output information on abnormal events and restoration processing for abnormal events through the output unit 114.
- the failure analysis unit (138) detects that two or more of the terminals (101) meet predetermined failure analysis conditions. , a management device that determines an abnormal event that has occurred in the terminal (101).
- the recovery command unit (140) responds to an abnormal event occurring in the terminal (101) to the terminal (101).
- a management device that commands activation or deactivation of predetermined restoration processing.
- the management device 102 enables/disables the self-recovery processing by the terminal 101 according to the detected abnormal event. It is possible to suppress the execution of processing.
- the recovery command unit (140) receives the self-recovery process from the terminal (101) received by the recovery notification receiving unit (139).
- a management device that determines whether or not the recovery process can be executed in response to the notification, and issues an execution command to the terminal (101) at a timing at which the recovery process can be executed.
- the management device 102 determines whether or not the terminal 101 can execute self-recovery processing, thereby suppressing unnecessary self-recovery processing in response to an abnormal event caused by a planned artificial shutdown of the communication system.
- the present invention is not limited to the above-described embodiments, and includes various modifications.
- the above embodiments are described in detail for easy understanding of the present invention, and are not necessarily limited to include all the described configurations.
- any addition, deletion, or replacement of other configurations for a part of the configuration of each embodiment can be applied singly or in combination.
- each of the above configurations, functions, processing units, processing means, etc. may be implemented in hardware, for example, by designing a part or all of them in an integrated circuit. Further, each of the above configurations, functions, etc. may be realized by software by a processor interpreting and executing a program for realizing each function. Information such as programs, tables, and files that implement each function should be recorded in recording devices such as memory, hard disks, SSD (Solid State Drives), or recording media such as IC cards, SD cards, and DVDs. can be done.
- control lines and information lines indicate what is considered necessary for explanation, and not all control lines and information lines are necessarily indicated on the product. In practice, it may be considered that almost all configurations are interconnected.
- a terminal having a processor and a memory and connected to a management device, a log information management unit that collects log information of its own terminal and transmits it to the management device; a self-failure analysis unit that analyzes the presence or absence of an abnormal event in the terminal from the log information and determines self-recovery processing when the abnormal event is detected; a recovery notification transmission unit that notifies the management device of self-recovery processing for the abnormal event; a self-recovery processing unit that executes self-recovery processing for the abnormal event or recovery processing commanded by the management device; The terminal, wherein the restoration processing unit waits execution of the restoration processing until an execution command is received from the management apparatus after notifying the restoration processing via the restoration notification transmission unit.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Quality & Reliability (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Debugging And Monitoring (AREA)
Abstract
Description
以上のように、上記実施例1~3の通信システムは、以下のような構成とすることができる。
特許請求の範囲に記載した以外の本発明の観点の代表的なものとして、次のものがあげられる。
プロセッサとメモリを有して管理装置と接続される端末であって、
自端末のログ情報を収集して、前記管理装置へ送信するログ情報管理部と、
前記ログ情報から当該端末における異常イベントの有無を解析して前記異常イベントを検知した場合には自己復旧処理を決定する自己障害解析部と、
前記異常イベントに対する自己復旧処理を前記管理装置に通知する復旧通知送信部と、
前記異常イベントに対する自己復旧処理又は前記管理装置から命令された復旧処理を実行する自己復旧処理部と、を有し、
前記復旧処理部は、前記復旧通知送信部を介して復旧処理を通知した後、前記管理装置から実行命令を受信するまで、当該復旧処理の実行を待機することを特徴とする端末。
1以上の端末と接続される管理装置が前記端末を管理する管理方法であって、
前記端末が、自端末のログ情報を収集して、前記管理装置へ送信する第1のステップと、
前記端末が、前記ログ情報から当該端末における異常イベントの有無を解析して前記異常イベントを検知した場合には自己復旧処理を決定する第2のステップと、
前記端末が、前記異常イベントに対する自己復旧処理を前記管理装置に通知する第3のステップと、
前記管理装置が、前記端末から収集したログ情報を基に、前記端末における異常イベントの有無を解析して、前記異常イベントを検知した場合には復旧処理を決定する第4のステップと、
前記管理装置が、前記端末から当該端末が検知した異常イベントに対する自己復旧処理の通知を受信する第5のステップと、
前記管理装置が、前記通知の有無に応じて前記障害解析部で検知した異常イベントに対する復旧処理の実行命令を前記端末に指令する第6のステップと、
前記端末が、前記異常イベントに対する自己復旧処理又は前記管理装置から命令された復旧処理を実行する第7のステップと、を含むことを特徴とする管理方法。
Claims (15)
- 管理対象となる1以上の端末と、前記端末と接続される管理装置からなる通信システムであって、
前記端末は、
自端末のログ情報を収集して、前記管理装置へ送信するログ情報管理部と、
前記ログ情報から当該端末における異常イベントの有無を解析して前記異常イベントを検知した場合には自己復旧処理を決定する自己障害解析部と、
前記異常イベントに対する自己復旧処理を前記管理装置に通知する復旧通知送信部と、
前記異常イベントに対する自己復旧処理又は前記管理装置から命令された復旧処理を実行する自己復旧処理部と、を有し、
前記管理装置は、
前記端末から収集したログ情報を基に、前記端末における異常イベントの有無を解析して、前記異常イベントを検知した場合には復旧処理を決定する障害解析部と、
前記端末から当該端末が検知した異常イベントに対する自己復旧処理の通知を受信する復旧通知受信部と、
前記通知の有無に応じて前記障害解析部で検知した異常イベントに対する復旧処理の実行命令を前記端末に指令する復旧命令部と、を有すること特徴とする通信システム。 - プロセッサとメモリを有して、管理対象となる1以上の端末と接続される管理装置であって、
前記端末から収集したログ情報を基に、前記端末における異常イベントの有無を解析して、前記異常イベントを検知した場合には復旧処理を決定する障害解析部と、
前記端末から、当該端末で検知した異常イベントに対する自己復旧処理の通知を受信する復旧通知受信部と、
前記通知の有無に応じて前記障害解析部で検知した異常イベントに対する復旧処理の実行命令を前記端末に指令する復旧命令部と、を有すること特徴とする管理装置。 - 請求項2に記載の管理装置であって、
前記ログ情報に対する障害解析条件を管理する障害解析条件管理部と、
前記ログ情報と合致した1以上の前記障害解析条件の組み合わせに基づくイベント判定条件を管理するイベント判定条件管理部と、
前記異常イベント毎に講じるべき復旧処理を管理する復旧処理管理部と、をさらに有し、
前記障害解析部は、
前記障害解析条件管理部の障害解析条件に従って前記ログ情報を解析して、解析結果と、前記イベント判定条件管理部のイベント判定条件を基に前記端末の異常イベントを検知した後、前記復旧処理管理部を参照して前記異常イベントに対する復旧処理を決定し、
前記復旧命令部は、
前記復旧通知受信部にて前記端末から前記異常イベントに対する自己復旧処理の通知が受信されない場合に、前記復旧処理の実行命令を前記端末に指令することを特徴とする管理装置。 - 請求項3に記載の管理装置であって、
前記障害解析条件管理部は、
前記障害解析条件毎に参照するログ情報と、比較基準となる閾値と、前記閾値に対する大小関係を規定する比較条件と、前記閾値を絶対値と相対値の何れで扱うかを規定する比較方法と、ログ情報が前記障害解析条件に合致した回数を指定可能とすることを特徴とする管理装置。 - 請求項3に記載の管理装置であって、
前記復旧処理管理部は、
前記異常イベント毎に講じるべき復旧処理と、前記異常イベントの検知から復旧命令実行までの待ち時間を指定可能とすることを特徴とする管理装置。 - 請求項3に記載の管理装置であって、
前記障害解析条件管理部で管理する障害解析条件と、前記イベント判定条件管理部で管理するイベント判定条件と、前記復旧処理管理部で管理する復旧処理と、を入力する入力部と、
前記障害解析部で検知された異常イベントの情報と、当該異常イベントに対する復旧処理を出力する出力部と、をさらに有することを特徴とする管理装置。 - 請求項2に記載の管理装置であって、
前記障害解析部は、
2以上の前記端末が所定の障害解析条件に合致したことを以って、前記端末で発生した異常イベントを判定することを特徴とする管理装置。 - 請求項7に記載の管理装置であって、
前記復旧命令部は、
前記端末で発生した異常イベントに応じて、前記端末に対して所定の復旧処理の有効化又は無効化を命令することを特徴とする管理装置。 - 請求項2に記載の管理装置であって、
前記復旧命令部は、
前記復旧通知受信部で受信した前記端末からの自己復旧処理の通知に応じて、当該復旧処理の実行可否を判定し、当該復旧処理を実行可能なタイミングで前記端末へ実行命令を指令することを特徴とする管理装置。 - プロセッサとメモリを有して管理装置と接続される端末であって、
自端末のログ情報を収集して、前記管理装置へ送信するログ情報管理部と、
前記ログ情報から当該端末における異常イベントの有無を解析して前記異常イベントを検知した場合には自己復旧処理を決定する自己障害解析部と、
前記異常イベントに対する自己復旧処理を前記管理装置に通知する復旧通知送信部と、
前記異常イベントに対する自己復旧処理又は前記管理装置から命令された復旧処理を実行する自己復旧処理部と、を有することを特徴とする端末。 - 請求項10に記載の端末であって、
前記ログ情報に対する障害解析条件を管理する障害解析条件管理部と、
前記ログ情報と合致した1以上の前記障害解析条件の組み合わせに基づくイベント判定条件を管理するイベント判定条件管理部と、
前記異常イベント毎に講じるべき復旧処理を管理する復旧処理管理部と、をさらに有し、
前記自己障害解析部は、
前記障害解析条件管理部の障害解析条件に従って前記ログ情報を解析して、解析結果と、前記イベント判定条件管理部のイベント判定条件を基に自端末の異常イベントを検知した後、前記復旧処理管理部を参照して前記異常イベントに対する復旧処理を決定し、
前記自己復旧処理部は、
前記復旧通知送信部を介して当該復旧処理を前記管理装置へ通知した後に前記復旧処理を実行又は前記管理装置からの命令で指定された復旧処理を実行することを特徴とする端末。 - 請求項11に記載の端末であって、
前記障害解析条件管理部は、
前記障害解析条件毎に参照するログ情報と、比較基準となる閾値と、前記閾値に対する大小関係を規定する比較条件と、前記閾値を絶対値と相対値の何れで扱うかを規定する比較方法と、ログ情報が前記障害解析条件に合致した回数を指定可能とすることを特徴とする端末。 - 請求項11に記載の端末であって、
前記復旧処理管理部は、
前記異常イベント毎に講じるべき復旧処理と、前記異常イベントの検知から復旧処理実行までの待ち時間と、前記管理装置に対する通知内容を指定可能とすることを特徴とする端末。 - 請求項11に記載の端末であって、
前記障害解析条件管理部で管理する障害解析条件と、前記イベント判定条件管理部で管理するイベント判定条件と、前記復旧処理管理部で管理する復旧処理を入力する入力部と、
前記自己障害解析部で検知された異常イベントの情報と、当該異常イベントに対する復旧処理を出力する出力部と、をさらに有することを特徴とする端末。 - 請求項10に記載の端末であって、
前記自己復旧処理部は、
前記管理装置からの命令に従って、前記復旧処理の一部を有効化又は無効化することを特徴とする端末。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202280047314.XA CN117642725A (zh) | 2021-12-17 | 2022-06-08 | 通信系统、管理装置和终端 |
| US18/692,132 US12554574B2 (en) | 2021-12-17 | 2022-06-08 | Communication system, management device, and terminal |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2021205356A JP7664154B2 (ja) | 2021-12-17 | 2021-12-17 | 通信システム、管理装置及び端末 |
| JP2021-205356 | 2021-12-17 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023112359A1 true WO2023112359A1 (ja) | 2023-06-22 |
Family
ID=86774211
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2022/023197 Ceased WO2023112359A1 (ja) | 2021-12-17 | 2022-06-08 | 通信システム、管理装置及び端末 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US12554574B2 (ja) |
| JP (1) | JP7664154B2 (ja) |
| CN (1) | CN117642725A (ja) |
| WO (1) | WO2023112359A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7664154B2 (ja) * | 2021-12-17 | 2025-04-17 | 株式会社日立産機システム | 通信システム、管理装置及び端末 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2009087136A (ja) * | 2007-10-01 | 2009-04-23 | Nec Corp | 障害修復システムおよび障害修復方法 |
| JP2018524704A (ja) * | 2015-06-19 | 2018-08-30 | アップテイク テクノロジーズ、インコーポレイテッド | 予測モデルの動的な実行 |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0618377B2 (ja) * | 1983-09-08 | 1994-03-09 | 株式会社日立製作所 | 伝送系 |
| JP3200661B2 (ja) * | 1995-03-30 | 2001-08-20 | 富士通株式会社 | クライアント/サーバシステム |
| JP2006023910A (ja) * | 2004-07-07 | 2006-01-26 | Hitachi Ltd | サーバ障害回復方法およびサーバ障害回復システム |
| US7434102B2 (en) * | 2004-12-29 | 2008-10-07 | Intel Corporation | High density compute center resilient booting |
| JP2006330945A (ja) * | 2005-05-25 | 2006-12-07 | Matsushita Electric Ind Co Ltd | デバイスの遠隔監視・修復システム |
| US8732530B2 (en) * | 2011-09-30 | 2014-05-20 | Yokogawa Electric Corporation | System and method for self-diagnosis and error reporting |
| JP6078984B2 (ja) * | 2012-05-23 | 2017-02-15 | 富士通株式会社 | 処理装置,処理方法,処理プログラム及び管理装置 |
| JP2014182720A (ja) * | 2013-03-21 | 2014-09-29 | Fujitsu Ltd | 情報処理システム、情報処理装置及び障害処理方法 |
| US9729534B2 (en) * | 2015-02-26 | 2017-08-08 | Seagate Technology Llc | In situ device authentication and diagnostic repair in a host environment |
| JP6838568B2 (ja) | 2016-02-05 | 2021-03-03 | コニカミノルタ株式会社 | 情報処理システム及び情報処理方法 |
| US20190303231A1 (en) | 2016-12-27 | 2019-10-03 | Nec Corporation | Log analysis method, system, and program |
| JP7664154B2 (ja) * | 2021-12-17 | 2025-04-17 | 株式会社日立産機システム | 通信システム、管理装置及び端末 |
-
2021
- 2021-12-17 JP JP2021205356A patent/JP7664154B2/ja active Active
-
2022
- 2022-06-08 WO PCT/JP2022/023197 patent/WO2023112359A1/ja not_active Ceased
- 2022-06-08 US US18/692,132 patent/US12554574B2/en active Active
- 2022-06-08 CN CN202280047314.XA patent/CN117642725A/zh active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2009087136A (ja) * | 2007-10-01 | 2009-04-23 | Nec Corp | 障害修復システムおよび障害修復方法 |
| JP2018524704A (ja) * | 2015-06-19 | 2018-08-30 | アップテイク テクノロジーズ、インコーポレイテッド | 予測モデルの動的な実行 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2023090412A (ja) | 2023-06-29 |
| JP7664154B2 (ja) | 2025-04-17 |
| US20240378106A1 (en) | 2024-11-14 |
| CN117642725A (zh) | 2024-03-01 |
| US12554574B2 (en) | 2026-02-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP4995104B2 (ja) | 性能監視条件の設定・管理方法及びその方法を用いた計算機システム | |
| JP2017517060A (ja) | 障害処理方法、関連装置、およびコンピュータ | |
| GB2478625A (en) | Deleting snapshot backups for unstable virtual machine configurations | |
| JPWO2012046293A1 (ja) | 障害監視装置、障害監視方法及びプログラム | |
| US20180173583A1 (en) | Systems and methods for real time computer fault evaluation | |
| CN107526646A (zh) | 监控方法、装置及看门狗系统 | |
| KR20220121008A (ko) | 디바이스 장애 통합 관리 플랫폼 제공 방법 | |
| US20150113320A1 (en) | Processing apparatus, process system, and non-transitory computer-readable recording medium | |
| JP5983102B2 (ja) | 監視プログラム、方法及び装置 | |
| JP7672275B2 (ja) | 対策選定装置、システム及び対策選定方法 | |
| WO2023112359A1 (ja) | 通信システム、管理装置及び端末 | |
| JP6015750B2 (ja) | ログ収集サーバ、ログ収集システム、ログ収集方法 | |
| EP2495660A1 (en) | Information processing device and method for controlling information processing device | |
| JP2014021586A (ja) | プログラムのアップグレードを実施するサーバ、サーバと複数の機器からなるプログラムのアップグレードシステム及びプログラムのアップグレード方法 | |
| JP2003032764A (ja) | 保守装置、メンテナンスシステムおよびメンテナンス方法 | |
| JP2012230451A (ja) | ネットワーク端末故障対応システム、端末装置、サーバ装置、ネットワーク端末故障対応方法及びプログラム | |
| CN119668921A (zh) | 一种容器编排系统故障节点的修复方法、装置及存储介质 | |
| KR101783201B1 (ko) | 서버 통합 관리 시스템 및 방법 | |
| JP5623449B2 (ja) | 報告書作成装置、報告書作成プログラムおよび報告書作成方法 | |
| US11635923B2 (en) | Monitoring system, monitoring method, and monitoring program | |
| JP2011028490A (ja) | システム監視装置、システム監視方法、及びプログラム | |
| JP2015125591A (ja) | プラント監視システムおよびその更新方法 | |
| JP5268820B2 (ja) | 監視装置用プログラムの書き換え方法 | |
| JP2017102834A (ja) | バックアップシステム | |
| WO2025079296A1 (ja) | 管理装置、通信装置、通信システムおよび遠隔更新方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22906896 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202280047314.X Country of ref document: CN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2401000313 Country of ref document: TH |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18692132 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22906896 Country of ref document: EP Kind code of ref document: A1 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 18692132 Country of ref document: US |