WO2007137601A1 - Monitoring data transfer failures in a bus based system - Google Patents
Monitoring data transfer failures in a bus based system Download PDFInfo
- Publication number
- WO2007137601A1 WO2007137601A1 PCT/EP2006/005088 EP2006005088W WO2007137601A1 WO 2007137601 A1 WO2007137601 A1 WO 2007137601A1 EP 2006005088 W EP2006005088 W EP 2006005088W WO 2007137601 A1 WO2007137601 A1 WO 2007137601A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- monitored
- bus
- unit
- monitoring unit
- error
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/06—Management of faults, events, alarms or notifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L12/00—Data switching networks
- H04L12/28—Data switching networks characterised by path configuration, e.g. LAN [Local Area Networks] or WAN [Wide Area Networks]
- H04L12/40—Bus networks
- H04L12/40006—Architecture of a communication node
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0805—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability
- H04L43/0811—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability by checking connectivity
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0805—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability
- H04L43/0817—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability by checking functioning
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L12/00—Data switching networks
- H04L12/28—Data switching networks characterised by path configuration, e.g. LAN [Local Area Networks] or WAN [Wide Area Networks]
- H04L12/40—Bus networks
- H04L12/40169—Flexible bus arrangements
- H04L12/40176—Flexible bus arrangements involving redundancy
- H04L12/40182—Flexible bus arrangements involving redundancy by using a plurality of communication lines
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L69/00—Network arrangements, protocols or services independent of the application payload and not provided for in the other groups of this subclass
- H04L69/40—Network arrangements, protocols or services independent of the application payload and not provided for in the other groups of this subclass for recovering from a failure of a protocol instance or entity, e.g. service redundancy protocols, protocol state redundancy or protocol service redirection
Definitions
- the present invention relates to monitoring data traffic on a bus for detecting failures of the bus and monitored units of a communication system.
- monitoring units are applied for supervising one or more components of a system, for example communications units which are connected to one or more busses for data transmission. These monitoring units determine data transmission failures and may either indicate the failures for example by a signal or shut down the system in case of critical failures. Possible causes of data transmission failures in communication systems may be for example soft- or hardware failures of components, connection failures of plug connections which may be caused for example by mechanical forces such as vibrations or failures occurring in bus lines such as broken lines.
- an improved monitoring unit for supervising a communication system monitors the data traffic from or to monitored units connected by at least one bus and allows to discriminate between a failure of a bus and a unit monitored.
- the improved monitoring unit enables a more specific reaction to the failure, for example to disable a faulty bus or unit or to reboot a faulty unit. This allows to maintain the operation of the communication system even if data transmission failures were monitored and, therefore, offers a more flexible supervision and management of a supervised system than with known monitoring units which simply detect a data transmission failure but do not allow to circumscribe the failure, particularly to determine the cause of the failure.
- the terms "faulty" and "defective" may be used as synonyms.
- a monitoring unit for supervising a communication system comprising at least one bus for connecting the monitoring unit and several monitored units for data communication, wherein each monitored unit is connected to the at least one bus and adapted to transmit data over each of the connected buses, the monitoring unit is connected to the at least one bus and adapted to monitor the data traffic from or to the monitored units on the at least one bus, to analyze the monitored data traffic for data transmission failures, and to discriminate between a failure of a bus and a monitored unit on basis of the analyzed data transmission failures.
- the monitoring unit may be further adapted to generate an error matrix based on the monitoring of the data traffic, wherein the error matrix comprises a line for each of the monitored units and a column for each bus, connecting the monitoring unit and the several monitored units, and each cell of the error matrix contains an error element my of monitored unit i and bus j indicating data transmission failures, and to discriminate between a failure of the at least one bus and one of the monitored units by evaluating the error matrix.
- the error matrix may be stored in the monitoring unit or an external memory.
- the error matrix has the advantage that it allows to quickly detect any failures of busses or monitored units by analyzing the lines and columns.
- the monitoring unit may be further adapted to generate a line i of the error matrix for a certain monitored unit i by performing the following steps: for each bus j connected to the monitored unit i monitor the data traffic on the bus j and generate an error element my depending on failed data transmissions on the bus j, and enter the error element my in said line i of the error matrix.
- This hast the advantage that the error matrix may be generated during a normal operation of the communication system be merely monitoring the data traffic on the different busses from or to a certain monitored unit.
- This process may be performed for each monitored unit in order to generate the error matrix line by line.
- this process may be performed continuously during the normal operation without affecting the normal operation of the system, particularly without reducing the system performance since the monitoring unit only has to select the bus to be monitored and monitor the data transmitted to or received from a certain unit over this bus.
- the monitoring unit may be further adapted to generate a column j of the error matrix for a certain bus j by performing the following steps: for each monitored unit i connected to the bus j monitor the data traffic on the bus j and generate an error element m, j depending on failed data transmissions on the bus j, and enter the error element m, j in said column j of the error matrix.
- the error matrix may be also generated column by column instead of line by line.
- This process has the advantage that the status of a certain monitored unit may be quickly derived from the error matrix which requires only a comparison of the respective error element with the predefined threshold ts1. In principle, this requires only two steps, namely reading out the error element from the error matrix and comparing the read out error element with the predefined threshold ts1. Thus, this process may be implemented either in soft- or in hardware at relatively low costs.
- the monitoring unit may be further adapted to perform the following steps: a) reset a faulty monitored unit i, b) wait for a pre-defined time, particularly for booting the monitored unit i and for some data transmissions of the monitored unit i, c) derive the state of the monitored unit i, d) repeat steps a) to c)for a predefined number if the state of the monitored unit i is determined as faulty in step c). This allows to recover a monitored unit which was detected as faulty.
- the status bj of a bus may be derived quickly from the error matrix.
- this process can also be performed with two steps in principle, namely by reading out the error element from the error matrix and comparing the read out error element with the predefined threshold ts2.
- the monitoring unit may be further adapted to perform the following steps: a) disconnect all monitored units from a faulty bus j, b) connect a monitored unit i with bus j, c) transmit data over the bus j to the monitored unit i and monitor the data traffic, d) if a failed data transmission was monitored in step c), disconnect the monitored unit i from busj, mark monitored unit i connected to bus j as faulty, and disconnect monitored unit I from bus j, e) repeat steps b) to d) for each of the monitored units, f) if all monitored units i are marked as faulty, mark the bus j as defective.
- the bus may for example be isolated by disconnecting all monitored units from the bus. This allows to implement a "self-healing" communication system which does not require the interaction of service personal for repairing the system.
- an error element my may be the number of successive failed data transmission on a certain bus j to or from a certain slave i.
- an error element my may be set to 0 after at least one successful bus transmission on a certain bus j to or from a certain slave i.
- the monitoring unit may comprise at least one of the features: the monitoring unit is capable of disconnecting a certain monitored unit from a bus connected to the monitored unit, the monitoring unit is capable of resetting a monitored unit, and the monitoring unit is adapted to generate frequent data traffic to each monitored unit over each bus.
- a monitored unit may be for example a sub-system of a larger system, and in case of a failure it may not be required to reset the entire system, but only the sub-system.
- the communication system may be an automotive communication system, or a communication system of data processing or telecommunications equipment.
- a further embodiment of the invention provides a software program or product, preferably stored on a data carrier, for implementing the monitoring unit according to any of the above described embodiments, when run on a data processing system such as a computer.
- the monitoring unit may be implemented in software which may be executed by a data processing system for example formed by a processor or controller executing instructions of a program implementing the monitoring unit and stored in a memory.
- a method of supervising a communication system comprising at least one bus for connecting a monitoring unit and several monitored units for data communication, wherein each monitored unit is connected to the at least one bus and transmits data over each of the connected buses, the monitoring unit is connected to the at least one bus and monitors the data traffic from or to the monitored units on the at least one bus, analyzes the monitored data traffic for data transmission failures, and discriminates between a failure of a bus and a monitored unit on basis of the analyzed data transmission failures.
- Embodiments of the invention can be partly or entirely embodied or supported by one or more suitable software programs, which can be stored on or otherwise provided by any kind of data carrier, and which might be executed in or by any suitable data processing unit.
- Software programs or routines can be preferably applied .
- FIG. 1 shows an embodiment of a communication system with a monitoring unit according to the invention for supervising several units of the communication system which are connected by a bus;
- FIG. 2 shows a flowchart of a method for generating a line of an error matrix according to the invention
- FIG. 3 shows a flowchart of a method for generating a column of an error matrix according to the invention
- FIG. 4 shows a flowchart of a method for deriving the state of a certain monitored unit according to the invention
- FIG. 5 shows a flowchart of a method for deriving the state of a certain bus according to the invention
- FIG. 6 shows a flowchart of a method for handling a certain monitored unit which was determined as being faulty according to the invention.
- Fig. 7 shows a flowchart of a method for handling a certain bus which was determined as being faulty according to the invention.
- Fig. 1 shows a communication system with a monitoring unit 10, a bus system comprising three busses 18a, 18b, and 18c, and three monitored units 12, 14, and 16 which are connected to each of the three busses 18a, 18b, and 18c and may communicate via the busses 18a, 18b, 18c.
- the monitored units 12, 14, and 16 may generate data traffic on the busses 18a, 18b, and 18c.
- the data traffic may be monitored by the monitoring unit 10 which is also connected to the bussesi 8a, 18b, 18c.
- the monitored units may be communication modules in a telecommunications equipment, control units in an automobile, or units of data processing equipment such as a server computer.
- the monitoring unit may be implemented by dedicated controller or processor or a computing unit.
- the monitoring unit may be also a part of a monitored unit, for example a remote management unit of a data processing computer which is monitored by the remote management unit.
- the functionality of the monitoring unit which is essential for the invention may be implemented in hard- or software or both.
- the monitoring unit may be a standard microcontroller with a monitoring interface adapted for receiving data traffic on the bus and for controlling an isolation of certain monitored units from the bus and to reset certain monitored units.
- the analysis of the received data traffic for data transmission failures on the bus may be performed by a software stored in a program memory and executed by the microcontroller.
- the monitoring unit is a processor which executes a program adapted for performing the methods as described later in detail for discriminating between failures of busses and monitored units.
- Each monitored unit 12, 14, and 16 is connected to the bussesi ⁇ a, 18b, and 18c via bus connection lines 20a, 20b, 20c, 22a, 22b, 22c, and 24a, 24b, 24c, respectively.
- Each group of bus connection lines 20a, 20b, 20c, 22a, 22b, 22c, and 24a, 24b, 24c comprises switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c, respectively, with which each bus connection line may be separately broken so that each of the monitored units 12, 14, and 16 may be completely or partly electrically isolated from the busses 18a, 18b, and 18c.
- the switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c are controlled by the monitoring unit 10 over a switch control bus 34.
- the monitoring unit 10 has the capability to individually reset the monitored units 12, 14, and 16.
- the monitoring unit 10 thus controls a reset bus 32 connected to a reset input of each monitored unit 12, 14, and 16.
- the reset bus 32 serves to reset the monitored units 12, 14, and 16 in case of an unit failure detected by the monitoring unit 10.
- the monitoring unit 10 addresses the unit over the reset bus 32 and initiates the reset of the addressed unit.
- a monitored unit 12, 14, 16 may comprise three different states: it may be defective, faulty, or non-faulty.
- a failure of a monitored unit comprises the states defective and faulty (thus, these terms may be used herein as synonyms). When a monitored unit is defective or faulty, it may generate data transmission failures on the busses 18a, 18b, 18c.
- the monitoring unit 10 is adapted to generate an error matrix based on the monitoring of the data traffic on the busses 18a, 18b, and 18c.
- the error matrix comprises a line i for each of the monitored units 12, 14, and 16, and a column j for each bus 18a, 18b, and 18c.
- Each cell of the error matrix contains an error element my of monitored unit i and bus j indicating data transmission failures.
- the error matrix is shown for the communication system from Fig. 1 , wherein monitored unit number 1 is 12, 2 is 14, and 3 is 16, and bus number 1 is 18a, 2 is 18b, and 3 is 18c:
- the error matrix may be stored in a memory of the monitoring unit 10, for example in an embedded memory.
- each error element could be formed by a single bit like a flag indicating an error or no error.
- each error element should be an integer value indicating a number of failed transmissions or bit errors of a data transmission failure. This allows are more sophisticated analysis of a failure than a simple flag.
- Data transmission errors may be detected in various ways, for example by checking a checksum of a transmitted data package or by checking handshake or similar data sent and received by various monitored units after receiving or transmitting data over the busses.
- each bus protocol which indicates data transmission failures over a bus is suitable for the purpose of the invention.
- the error element my may indicate the number of data transmission failures of monitored unit i on bus j.
- the error matrix may be analyzed by the monitoring unit 10 for discriminating between failure of a bus j or a monitored unit i as will be described in the following with reference to the flow charts shown in Fig. 2 to 7. First, it is described how lines and columns of the error matrix may be generated with regard to Fig. 2 and 3. Then, the deriving of states of monitored units and busses from the error matrix will be explained with reference to Fig. 4 and 5. Finally, the handling of bus or monitored unit failures will be described with reference to Fig. 6 and 7. It should be noted that the described methods are all implemented in the monitoring unit 10, for example as computer program which is executed by a processor of the monitoring unit. The methods may also be implemented in hardware, for example in a programmable hardware such as a Programmable Gate Array (PGA) or a dedicated hardware such as an Application Specific Integrated Circuit (ASIC).
- PGA Programmable Gate Array
- ASIC Application Specific Integrated Circuit
- a flowchart of a method for generating a line of an error matrix for a communication system containing busses j and monitored units i is shown.
- the parameters j and i may be any number larger than 0 but i should be at least 2, i.e., at least two monitored units are required (otherwise a point-to point connection is monitored which does not make sense in the context of the invention).
- a typical constellation may be for example a communication system with a single bus and several monitored units which are connected to the single bus for data transmission.
- the parameter j is set to 1 in order to start with the first column of the error matrix to be generated.
- step S10 the monitoring unit 10 monitors the data traffic on the bus 1 (reference numeral 18a in Fig. 1 ) connected to a monitored unit i.
- step S12 the monitoring unit 10 generates an error element ITI M depending on failed data transmissions on the bus 1.
- step S14 the monitoring unit 10 enters the error element m ⁇ i in line i and column 1 of the error matrix.
- step S16 the monitoring unit 10 checks whether each bus j connected to the monitored unit i was already monitored. If so, the method stops and a line of the error matrix for the monitored unit i is generated. If not all busses were already checked, the parameter j is increased by 1 in step S18, and the method continues in step S10 with the procedure as described before. This will be repeated until all busses connected to the monitored unit i were checked.
- a flowchart of a method for generating a column of an error matrix for a communication system containing busses j and monitored units i is shown.
- the parameter i is set to 1 in order to start with the first line of the error matrix to be generated.
- the monitoring unit 10 monitors the data traffic on a bus j connected to a monitored unit number 1 (reference numeral 12 in Fig. 1 ).
- the monitoring unit 10 generates an error element m-ij depending on failed data transmissions on the bus j.
- the monitoring unit 10 enters the error element mi j in line 1 and column j of the error matrix.
- step S20 the monitoring unit 10 checks whether each monitored unit i connected to the bus j was already monitored. If so, the method stops and a column of the error matrix for the bus j is generated. If not all monitored units were already checked, the parameter i is increased by 1 in step S22, and the method continues in step S10 with the procedure as described before. This will be repeated until all monitored units connected to bus j were checked.
- the error matrix is calculated by the monitoring unit 10 in observing all data transfers on all busses. Whenever data traffic on a certain bus to or from a certain monitored unit occurs, the corresponding element in the matrix will be updated. If a data transfer on bus j to monitored unit i was successful, my will be set to 0, indicating that there was no error. If a data transfer on bus j to monitored unit i was not successful, my may be for example incremented such as a counter. Thus my provides the number of successive failures on bus j to monitored unit i. In order to ensure that all elements in the error matrix are up to date, the monitoring unit may also generate frequent traffic on all busses to all units. This mechanism may be further improved to reduce traffic on the busses by e.g. time-stamping all elements in the error matrix and to generate traffic only for elements that achieved a certain "age". Thus, certain error elements may be updated.
- Fig. 4 shows a flowchart of a method for deriving the state of a certain monitored unit.
- a first step S23 it is checked whether the state Sj of a monitored unit I is marked as defective. If no, it is assumed in step S24 that a certain monitored unit i is faulty.
- the parameter j is set to 1 in order to check each bus j connected to the monitored unit i which is assumed as being faulty.
- step S28 it is checked whether the error element mn is lower than a predefined threshold ts1.
- the predefined threshold ts1 may be static or dynamic. In the latter case, it may for example depend on the entire operational status of a communication system such as the temperature in order to take care of operation conditions which influence the occurrence of data transmission failures. If in step S28 it is detected that the error element m ⁇ is lowerthan the predefined threshold ts1 , the state Sj of the monitored unit i is determined as non faulty (step S34). This state S 1 may be entered into a monitored unit state vector containing the states Sj of the monitored units i. Since the state Sj of the monitored unit i is determined as non faulty, the method stops.
- step S28 If in step S28 it is detected that the error element m M is higher than the predefined threshold ts1 , the state Sj of the monitored unit i is determined as faulty. This state Sj may be entered into the monitored unit state vector. In this case it is checked in a further step S30 whether all busses j connected to the monitored unit i were already checked. If not, the parameter j is increased by 1 in step S32 and the method continues with step S28. Otherwise, i.e. if all busses j connected to the monitored unit i were already checked, the state Sj of the monitored unit i is determined as faulty, entered in the monitored unit state vector, and the method stops.
- Fig. 5 shows a flowchart of a method for deriving the state of a certain bus.
- a first step S37 it is checked whether the state bj of a bus j is marked as defective. If no, it is assumed in step S38 that a certain bus j is faulty.
- the parameter i is set to 1 in order to check each monitored unit i connected to the bus j which is assumed as being faulty.
- step S42 it is checked whether the error element mi ] is lower than a predefined threshold ts2. If in step S42 it is detected that the error element m ⁇ is lower than the predefined threshold ts2, the state b j of the bus j is determined as non faulty (step S48).
- This state bj may be entered into a bus state vector containing the states bj of the busses j. Since the state bj of the bus j is determined as non faulty, the method stops. If in step S42 it is detected that the error element m ⁇ is higher than the predefined threshold ts2, the state b j of the bus j is determined as faulty. This state b j may be entered into the bus state vector. In this case it is checked in a further step S44 whether all monitored units i connected to the bus j were already checked. If not, the parameter i is increased by 1 in step S46 and the method continues with step S42. Otherwise, i.e. if all monitored units i connected to the bus j were already checked, the state bj of the bus j is determined as faulty, entered in the bus state vector, and the method stops.
- Fig. 6 shows a flowchart of a method for handling a certain monitored unit which was determined as being faulty.
- a faulty monitored unit i is reset by the monitoring unit 10.
- the monitoring unit 10 addresses the unit i over the reset line 32 and initiates the reset of the addressed unit i.
- the monitoring unit 10 waits a pre-defined time for booting the monitored unit i.
- booting as used herein means the process after a reset with which a monitored unit becomes functional and has performed a few data transmissions. For example, this process may contain functional tests of the unit and initialization steps containing data transmissions for making the unit functional.
- step S56 the state Sj of the monitored unit i is derived. This step should be performed after a few data transmissions of the unit in order to be able to generate an error element in the error matrix for the unit.
- step S58 it is checked whether the state Sj of the monitored unit i is faulty. If the state Sj is non faulty, the method stops. If the state Sj is faulty, it is checked in a following step S60 whether a predefined number of resets is already reached. For example, it may be predetermined that a faulty unit may only be reset three times. If the predefined number of resets is already reached, the state Si is set to defective (step S64).
- the error element typically may contain for example the number of failed data transmissions.
- step S52 With the before described method, it is possible to recover a faulty monitored unit if this unit stopped for example due to a software error or anything similar.
- Fig. 7 shows a flowchart of a method for handling a certain bus which was determined as being faulty.
- step S66 all monitored units are disconnected from the faulty bus j and the parameter i is set to 1. Therefore, the monitoring unit 10 opens all switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c, which are connected to bus j, over the switch control bus 34.
- one of bus connection lines 20a, 20b, 20c, 22a, 22b, 22c, and 24a, 24b, 24c or connected port of the monitored unit has issues so that each of the monitored units 12, 14, and 16 is completely electrically isolated from the busses 18a, 18b, and 18c.
- the monitoring unit 10 closes one switch of the switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c in order to connect a monitored unit i to bus j (step S68) by addressing the respective switch over the switch control bus 34 and instruct the switch to close. Thereafter, data is transmitted over the bus j to the monitored unit i and the data traffic is monitored by the monitoring unit 10 (step S70). In the following step S72, it is checked whether failed data transmissions were monitored. If failed data transmissions were monitored, the monitored unit i is permanently disconnected from bus j by the monitoring unit 10 (step S78). Then, the method continues with step S76, in which it is checked whether all monitored units were already checked.
- step S80 the parameter i is increased by 1 in step S80 and the method continues with step S68. If no failed data transmissions were monitored, the method continues with step S74 instead of step S78 and the monitoring unit 10 keeps monitored unit i connected to bus j. If all monitored units are disconnected from the bus j after this algorithm (step S82), then the bus j will be marked defective in the bus state vector b j (step S84). Otherwise, the bus j will be marked as healthy (step S86). With the before described method, all monitored units which communicate over the bus j are one after the other connected with the bus. It may then be detected whether the bus is faulty or any monitored unit, since if each monitored unit shows failed data transmission, the probability of a faulty bus is higher than each of the monitored units is faulty.
- faulty monitored unit should be kept isolated during the above method in order to avoid any influences on the bus "healing".
- the monitoring unit 10 can be further adapted to generate data traffic on the busses to be monitored even if this is not a requirement for the execution of the invention in practice. However, it may accelerate the monitoring and deliver faster and more reliable monitoring results than to wait for data traffic generated by the monitored units.
- the monitoring unit is adapted to generate a frequent data traffic, it should be adapted in that the frequent data traffic is generated similar to a heartbeat, for example data packets should be transmitted periodically from the monitoring unit to the monitored units over the busses.
- an error element of the error matrix may be the number of failed data transmissions, particularly successive failed data transmissions from or to a certain monitored unit over a certain bus.
- the error element may be set to 0 if at least one or more particularly successive successful data transmission from or to the respective monitored unit over the respective bus were supervised. For example, it may be defined that an error element is only set to 0 if three successive data transmissions were monitored from or to a monitored unit over a bus. Then, the likelihood of an error of the bus or the monitored unit is so small that an error element of 0 indicating no error is justified. It should be noted that also other strategies for defining error elements and setting them back to 0 may apply within the scope of the invention.
- the invention allows not only to detect any failures in a communication system applying at least one bus and several monitored units communicating over the at least one bus, but also to discriminate between a failure of a bus and a monitored unit on basis of the analyzed data transmission failures. Furthermore, the invention also provides strategies and methods to "heal" a bus if it was detected as faulty and to recover a monitored unit detected as faulty.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Environmental & Geological Engineering (AREA)
- Debugging And Monitoring (AREA)
- Small-Scale Networks (AREA)
Abstract
The invention relates to a monitoring unit (10) for supervising a communication system, wherein the communication system comprises at least one bus (18a, 18b, 18c) for connecting the monitoring unit (10) and several monitored units (12, 14, 16), each monitored unit (12, 14, 16) being connected to the at least one bus (18a, 18b, 18c) for transmitting data over the connected buses (18a, 18b, 18c), the monitoring unit (10) being connected to the busses (18a, 18b, 18c) for monitoring the data traffic from or to the monitored units (12, 14, 16) over the busses (18a, 18b, 18c), analyzing the monitored data traffic for detecting data transfer failures, thereby discriminating between a failure of a bus (18a, 18b, 18c, 20a, 20b, 20c, 22a, 22b, 22c, 24a, 24b, 24c) and a monitored unit (12, 14, 16).
Description
DESCRIPTION
MONITORING DATA TRANSFER FAILURES IN A BUS BASED
SYSTEM
BACKGROUND ART
[0001] The present invention relates to monitoring data traffic on a bus for detecting failures of the bus and monitored units of a communication system.
[0002] In complex communication systems, monitoring units are applied for supervising one or more components of a system, for example communications units which are connected to one or more busses for data transmission. These monitoring units determine data transmission failures and may either indicate the failures for example by a signal or shut down the system in case of critical failures. Possible causes of data transmission failures in communication systems may be for example soft- or hardware failures of components, connection failures of plug connections which may be caused for example by mechanical forces such as vibrations or failures occurring in bus lines such as broken lines.
DISCLOSURE
[0003] It is an object of the invention to provide an improved monitoring of data traffic. The object is solved by the independent claims. Further embodiments are shown by the dependent claim(s).
[0004] According to embodiments of the present invention, an improved monitoring unit for supervising a communication system monitors the data traffic from or to monitored units connected by at least one bus and allows to discriminate between a failure of a bus and a unit monitored. Thus, the improved monitoring unit enables a more specific reaction to the failure, for example to disable a faulty bus or unit or to reboot a faulty unit. This allows to maintain the operation of the communication system even if data transmission failures were monitored and, therefore, offers a more flexible supervision and management of a supervised system than with known monitoring units which simply detect a data transmission failure but do not allow to circumscribe the failure, particularly to determine the cause of the failure. In the following description,
the terms "faulty" and "defective" may be used as synonyms.
[0005] According to an embodiment of the present invention, a monitoring unit for supervising a communication system is provided, wherein the communication system comprises at least one bus for connecting the monitoring unit and several monitored units for data communication, wherein each monitored unit is connected to the at least one bus and adapted to transmit data over each of the connected buses, the monitoring unit is connected to the at least one bus and adapted to monitor the data traffic from or to the monitored units on the at least one bus, to analyze the monitored data traffic for data transmission failures, and to discriminate between a failure of a bus and a monitored unit on basis of the analyzed data transmission failures.
[0006] According to a further embodiment of the present invention, the monitoring unit may be further adapted to generate an error matrix based on the monitoring of the data traffic, wherein the error matrix comprises a line for each of the monitored units and a column for each bus, connecting the monitoring unit and the several monitored units, and each cell of the error matrix contains an error element my of monitored unit i and bus j indicating data transmission failures, and to discriminate between a failure of the at least one bus and one of the monitored units by evaluating the error matrix. The error matrix may be stored in the monitoring unit or an external memory. The error matrix has the advantage that it allows to quickly detect any failures of busses or monitored units by analyzing the lines and columns. For example, when a certain monitored unit must be analyzed for a failure, only the line corresponding to the monitored unit must be checked for failures, i.e., the error elements of the line must be checked if a failures exists. The same applies for the analysis of a certain bus which requires only the analysis of the column of the error matrix corresponding to this bus.
[0007] According to a further embodiment of the present invention, the monitoring unit may be further adapted to generate a line i of the error matrix for a certain monitored unit i by performing the following steps: for each bus j connected to the monitored unit i monitor the data traffic on the bus j and generate an error element my depending on failed data transmissions on the bus j, and enter the error element my in said line i of the error matrix. This hast the advantage that the error matrix may be generated during a normal operation of the communication system be merely monitoring the data traffic on the different busses from or to a certain monitored unit.
This process may be performed for each monitored unit in order to generate the error matrix line by line. Furthermore, this process may be performed continuously during the normal operation without affecting the normal operation of the system, particularly without reducing the system performance since the monitoring unit only has to select the bus to be monitored and monitor the data transmitted to or received from a certain unit over this bus.
[0008] According to a further embodiment of the present invention, the monitoring unit may be further adapted to generate a column j of the error matrix for a certain bus j by performing the following steps: for each monitored unit i connected to the bus j monitor the data traffic on the bus j and generate an error element m,j depending on failed data transmissions on the bus j, and enter the error element m,j in said column j of the error matrix. The error matrix may be also generated column by column instead of line by line. With this process, a certain unit is determined and then the data traffic on each bus connected to this unit is monitored in order to generate a column of the error matrix.
[0009] According to a further embodiment of the present invention, the monitoring unit may be further adapted to derive the state s, of a certain monitored unit i from the error matrix as follows: assume that the monitored unit i is faulty, for each column j = 1 to n of the error matrix check whether the value of the error element mtl is lower than a predefined threshold ts1 for at least one bus connected to the monitored unit i, and determine the state s, of the monitored unit i as non faulty if the value of at least one error element m,j is lower than the predefined threshold ts1. This process has the advantage that the status of a certain monitored unit may be quickly derived from the error matrix which requires only a comparison of the respective error element with the predefined threshold ts1. In principle, this requires only two steps, namely reading out the error element from the error matrix and comparing the read out error element with the predefined threshold ts1. Thus, this process may be implemented either in soft- or in hardware at relatively low costs.
[0010] According to a further embodiment of the present invention, the monitoring unit may be further adapted to perform the following steps: a) reset a faulty monitored unit i, b) wait for a pre-defined time, particularly for booting the monitored unit i and for some data transmissions of the monitored unit i, c) derive the state of the monitored
unit i, d) repeat steps a) to c)for a predefined number if the state of the monitored unit i is determined as faulty in step c). This allows to recover a monitored unit which was detected as faulty.
[0011] According to a further embodiment of the present invention, the monitoring unit may be further adapted to derive the state bj of a certain bus j from the error matrix as follows: assume that the bus j is faulty, for each line i = 1 to x of the error matrix check whether the value of the error element my is lower than a predefined threshold ts2 for at least one monitored unit connected to the bus j, and determine the state bj of the bus j as non faulty if the value of at least one error element my is lower than the predefined threshold ts2. As described above in connection with the derivation of the status Sj of a monitored unit, also the status bj of a bus may be derived quickly from the error matrix. Furthermore, this process can also be performed with two steps in principle, namely by reading out the error element from the error matrix and comparing the read out error element with the predefined threshold ts2.
[0012] According to a further embodiment of the present invention, the monitoring unit may be further adapted to perform the following steps: a) disconnect all monitored units from a faulty bus j, b) connect a monitored unit i with bus j, c) transmit data over the bus j to the monitored unit i and monitor the data traffic, d) if a failed data transmission was monitored in step c), disconnect the monitored unit i from busj, mark monitored unit i connected to bus j as faulty, and disconnect monitored unit I from bus j, e) repeat steps b) to d) for each of the monitored units, f) if all monitored units i are marked as faulty, mark the bus j as defective. This allows to unambiguously detect a faulty bus in a communication system. After detection of the faulty bus, the bus may for example be isolated by disconnecting all monitored units from the bus. This allows to implement a "self-healing" communication system which does not require the interaction of service personal for repairing the system.
[0013] According to a further embodiment of the present invention, an error element my may be the number of successive failed data transmission on a certain bus j to or from a certain slave i.
[0014] According to a further embodiment of the present invention, an error element my may be set to 0 after at least one successful bus transmission on a certain bus j to
or from a certain slave i.
[0015] According to a further embodiment of the present invention, the monitoring unit may comprise at least one of the features: the monitoring unit is capable of disconnecting a certain monitored unit from a bus connected to the monitored unit, the monitoring unit is capable of resetting a monitored unit, and the monitoring unit is adapted to generate frequent data traffic to each monitored unit over each bus. A monitored unit may be for example a sub-system of a larger system, and in case of a failure it may not be required to reset the entire system, but only the sub-system.
[0016] According to a further embodiment of the present invention, the communication system may be an automotive communication system, or a communication system of data processing or telecommunications equipment.
[0017] A further embodiment of the invention provides a software program or product, preferably stored on a data carrier, for implementing the monitoring unit according to any of the above described embodiments, when run on a data processing system such as a computer. In other words, the monitoring unit may be implemented in software which may be executed by a data processing system for example formed by a processor or controller executing instructions of a program implementing the monitoring unit and stored in a memory.
[0018] According to a further embodiment of the present invention, a method of supervising a communication system is provided, wherein the communication system comprises at least one bus for connecting a monitoring unit and several monitored units for data communication, wherein each monitored unit is connected to the at least one bus and transmits data over each of the connected buses, the monitoring unit is connected to the at least one bus and monitors the data traffic from or to the monitored units on the at least one bus, analyzes the monitored data traffic for data transmission failures, and discriminates between a failure of a bus and a monitored unit on basis of the analyzed data transmission failures.
[0019] Embodiments of the invention can be partly or entirely embodied or supported by one or more suitable software programs, which can be stored on or otherwise provided by any kind of data carrier, and which might be executed in or by any suitable data processing unit. Software programs or routines can be preferably
applied .
BRIEF DESCRIPTION OF DRAWINGS
[0020] Other objects and many of the attendant advantages of embodiments of the present invention will be readily appreciated and become better understood by reference to the following more detailed description of embodiments in connection with the accompanied drawing(s). Features that are substantially or functionally equal or similar will be referred to by the same reference sign(s).
[0021] Fig. 1 shows an embodiment of a communication system with a monitoring unit according to the invention for supervising several units of the communication system which are connected by a bus;
[0022] Fig. 2 shows a flowchart of a method for generating a line of an error matrix according to the invention;
[0023] Fig. 3 shows a flowchart of a method for generating a column of an error matrix according to the invention;
[0024] Fig. 4 shows a flowchart of a method for deriving the state of a certain monitored unit according to the invention;
[0025] Fig. 5 shows a flowchart of a method for deriving the state of a certain bus according to the invention;
[0026] Fig. 6 shows a flowchart of a method for handling a certain monitored unit which was determined as being faulty according to the invention; and
[0027] Fig. 7 shows a flowchart of a method for handling a certain bus which was determined as being faulty according to the invention.
[0028] Fig. 1 shows a communication system with a monitoring unit 10, a bus system comprising three busses 18a, 18b, and 18c, and three monitored units 12, 14, and 16 which are connected to each of the three busses 18a, 18b, and 18c and may communicate via the busses 18a, 18b, 18c. Particularly, the monitored units 12, 14, and 16 may generate data traffic on the busses 18a, 18b, and 18c. The data traffic may be monitored by the monitoring unit 10 which is also connected to the bussesi 8a,
18b, 18c.
[0029] The monitored units may be communication modules in a telecommunications equipment, control units in an automobile, or units of data processing equipment such as a server computer. The monitoring unit may be implemented by dedicated controller or processor or a computing unit. The monitoring unit may be also a part of a monitored unit, for example a remote management unit of a data processing computer which is monitored by the remote management unit. The functionality of the monitoring unit which is essential for the invention may be implemented in hard- or software or both. For example, the monitoring unit may be a standard microcontroller with a monitoring interface adapted for receiving data traffic on the bus and for controlling an isolation of certain monitored units from the bus and to reset certain monitored units. The analysis of the received data traffic for data transmission failures on the bus may be performed by a software stored in a program memory and executed by the microcontroller. Preferably, the monitoring unit is a processor which executes a program adapted for performing the methods as described later in detail for discriminating between failures of busses and monitored units.
[0030] Each monitored unit 12, 14, and 16 is connected to the bussesiδa, 18b, and 18c via bus connection lines 20a, 20b, 20c, 22a, 22b, 22c, and 24a, 24b, 24c, respectively. Each group of bus connection lines 20a, 20b, 20c, 22a, 22b, 22c, and 24a, 24b, 24c comprises switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c, respectively, with which each bus connection line may be separately broken so that each of the monitored units 12, 14, and 16 may be completely or partly electrically isolated from the busses 18a, 18b, and 18c. The switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c are controlled by the monitoring unit 10 over a switch control bus 34.
[0031] Furthermore, the monitoring unit 10 has the capability to individually reset the monitored units 12, 14, and 16. The monitoring unit 10 thus controls a reset bus 32 connected to a reset input of each monitored unit 12, 14, and 16. The reset bus 32 serves to reset the monitored units 12, 14, and 16 in case of an unit failure detected by the monitoring unit 10. For resetting an unit, the monitoring unit 10 addresses the unit over the reset bus 32 and initiates the reset of the addressed unit.
[0032] A monitored unit 12, 14, 16 may comprise three different states: it may be defective, faulty, or non-faulty. A failure of a monitored unit comprises the states defective and faulty (thus, these terms may be used herein as synonyms). When a monitored unit is defective or faulty, it may generate data transmission failures on the busses 18a, 18b, 18c.
[0033] The monitoring unit 10 is adapted to generate an error matrix based on the monitoring of the data traffic on the busses 18a, 18b, and 18c. The error matrix comprises a line i for each of the monitored units 12, 14, and 16, and a column j for each bus 18a, 18b, and 18c. Each cell of the error matrix contains an error element my of monitored unit i and bus j indicating data transmission failures. In the following, the error matrix is shown for the communication system from Fig. 1 , wherein monitored unit number 1 is 12, 2 is 14, and 3 is 16, and bus number 1 is 18a, 2 is 18b, and 3 is 18c:
[0034] The error matrix may be stored in a memory of the monitoring unit 10, for example in an embedded memory. In the most simple case, each error element could be formed by a single bit like a flag indicating an error or no error. However, for a more sophisticated analysis and discrimination, each error element should be an integer value indicating a number of failed transmissions or bit errors of a data transmission failure. This allows are more sophisticated analysis of a failure than a simple flag.
[0035] Data transmission errors may be detected in various ways, for example by checking a checksum of a transmitted data package or by checking handshake or similar data sent and received by various monitored units after receiving or transmitting data over the busses. In principle, each bus protocol which indicates data transmission failures over a bus is suitable for the purpose of the invention.
[0036] The error element my may indicate the number of data transmission failures
of monitored unit i on bus j. The error matrix may be analyzed by the monitoring unit 10 for discriminating between failure of a bus j or a monitored unit i as will be described in the following with reference to the flow charts shown in Fig. 2 to 7. First, it is described how lines and columns of the error matrix may be generated with regard to Fig. 2 and 3. Then, the deriving of states of monitored units and busses from the error matrix will be explained with reference to Fig. 4 and 5. Finally, the handling of bus or monitored unit failures will be described with reference to Fig. 6 and 7. It should be noted that the described methods are all implemented in the monitoring unit 10, for example as computer program which is executed by a processor of the monitoring unit. The methods may also be implemented in hardware, for example in a programmable hardware such as a Programmable Gate Array (PGA) or a dedicated hardware such as an Application Specific Integrated Circuit (ASIC).
[0037] In the following, the generation of lines and columns of the error matrix is described by means of flowcharts showing a kind of serial processing. However, it should be noted that in practice the generation of lines and columns of the error matrix is performed in a kind of parallel process by continuously monitoring the data traffic and updating error elements of the error matrix in parallel.
[0038] In Fig. 2, a flowchart of a method for generating a line of an error matrix for a communication system containing busses j and monitored units i is shown. The parameters j and i may be any number larger than 0 but i should be at least 2, i.e., at least two monitored units are required (otherwise a point-to point connection is monitored which does not make sense in the context of the invention). A typical constellation may be for example a communication system with a single bus and several monitored units which are connected to the single bus for data transmission. In a first step SO, the parameter j is set to 1 in order to start with the first column of the error matrix to be generated. In a following step S10, the monitoring unit 10 monitors the data traffic on the bus 1 (reference numeral 18a in Fig. 1 ) connected to a monitored unit i. In step S12, the monitoring unit 10 generates an error element ITIM depending on failed data transmissions on the bus 1. In step S14, the monitoring unit 10 enters the error element mιi in line i and column 1 of the error matrix. In step S16, the monitoring unit 10 checks whether each bus j connected to the monitored unit i was already monitored. If so, the method stops and a line of the error matrix for the monitored unit i is generated. If not all busses were already checked, the parameter j is increased by 1
in step S18, and the method continues in step S10 with the procedure as described before. This will be repeated until all busses connected to the monitored unit i were checked.
[0039] In Fig. 3, a flowchart of a method for generating a column of an error matrix for a communication system containing busses j and monitored units i is shown. In a first step S2, the parameter i is set to 1 in order to start with the first line of the error matrix to be generated. In a following step S10, the monitoring unit 10 monitors the data traffic on a bus j connected to a monitored unit number 1 (reference numeral 12 in Fig. 1 ). In step S12, the monitoring unit 10 generates an error element m-ij depending on failed data transmissions on the bus j. In step S14, the monitoring unit 10 enters the error element mij in line 1 and column j of the error matrix. In step S20, the monitoring unit 10 checks whether each monitored unit i connected to the bus j was already monitored. If so, the method stops and a column of the error matrix for the bus j is generated. If not all monitored units were already checked, the parameter i is increased by 1 in step S22, and the method continues in step S10 with the procedure as described before. This will be repeated until all monitored units connected to bus j were checked.
[0040] Preferably, the error matrix is calculated by the monitoring unit 10 in observing all data transfers on all busses. Whenever data traffic on a certain bus to or from a certain monitored unit occurs, the corresponding element in the matrix will be updated. If a data transfer on bus j to monitored unit i was successful, my will be set to 0, indicating that there was no error. If a data transfer on bus j to monitored unit i was not successful, my may be for example incremented such as a counter. Thus my provides the number of successive failures on bus j to monitored unit i. In order to ensure that all elements in the error matrix are up to date, the monitoring unit may also generate frequent traffic on all busses to all units. This mechanism may be further improved to reduce traffic on the busses by e.g. time-stamping all elements in the error matrix and to generate traffic only for elements that achieved a certain "age". Thus, certain error elements may be updated.
[0041] Fig. 4 shows a flowchart of a method for deriving the state of a certain monitored unit. In a first step S23, it is checked whether the state Sj of a monitored unit I is marked as defective. If no, it is assumed in step S24 that a certain monitored unit i
is faulty. In the following step S26, the parameter j is set to 1 in order to check each bus j connected to the monitored unit i which is assumed as being faulty. In step S28, it is checked whether the error element mn is lower than a predefined threshold ts1. The threshold ts1 may be an error value which indicates a failure, for example a certain number of failed data transmission from or to the monitored unit i over the busj=1. The predefined threshold ts1 may be static or dynamic. In the latter case, it may for example depend on the entire operational status of a communication system such as the temperature in order to take care of operation conditions which influence the occurrence of data transmission failures. If in step S28 it is detected that the error element m^ is lowerthan the predefined threshold ts1 , the state Sj of the monitored unit i is determined as non faulty (step S34). This state S1 may be entered into a monitored unit state vector containing the states Sj of the monitored units i. Since the state Sj of the monitored unit i is determined as non faulty, the method stops. If in step S28 it is detected that the error element mM is higher than the predefined threshold ts1 , the state Sj of the monitored unit i is determined as faulty. This state Sj may be entered into the monitored unit state vector. In this case it is checked in a further step S30 whether all busses j connected to the monitored unit i were already checked. If not, the parameter j is increased by 1 in step S32 and the method continues with step S28. Otherwise, i.e. if all busses j connected to the monitored unit i were already checked, the state Sj of the monitored unit i is determined as faulty, entered in the monitored unit state vector, and the method stops.
[0042] Fig. 5 shows a flowchart of a method for deriving the state of a certain bus. In a first step S37, it is checked whether the state bj of a bus j is marked as defective. If no, it is assumed in step S38 that a certain bus j is faulty. In the following step S40, the parameter i is set to 1 in order to check each monitored unit i connected to the bus j which is assumed as being faulty. In step S42, it is checked whether the error element mi] is lower than a predefined threshold ts2. If in step S42 it is detected that the error element m^ is lower than the predefined threshold ts2, the state bj of the bus j is determined as non faulty (step S48). This state bj may be entered into a bus state vector containing the states bj of the busses j. Since the state bj of the bus j is determined as non faulty, the method stops. If in step S42 it is detected that the error element m^ is higher than the predefined threshold ts2, the state bj of the bus j is determined as faulty. This state bj may be entered into the bus state vector. In this
case it is checked in a further step S44 whether all monitored units i connected to the bus j were already checked. If not, the parameter i is increased by 1 in step S46 and the method continues with step S42. Otherwise, i.e. if all monitored units i connected to the bus j were already checked, the state bj of the bus j is determined as faulty, entered in the bus state vector, and the method stops.
[0043] Fig. 6 shows a flowchart of a method for handling a certain monitored unit which was determined as being faulty. In step S52, a faulty monitored unit i is reset by the monitoring unit 10. The monitoring unit 10 addresses the unit i over the reset line 32 and initiates the reset of the addressed unit i. In the following step S54, the monitoring unit 10 waits a pre-defined time for booting the monitored unit i. The term "booting" as used herein means the process after a reset with which a monitored unit becomes functional and has performed a few data transmissions. For example, this process may contain functional tests of the unit and initialization steps containing data transmissions for making the unit functional. Then, in a following step S56, the state Sj of the monitored unit i is derived. This step should be performed after a few data transmissions of the unit in order to be able to generate an error element in the error matrix for the unit. In the next step S58, it is checked whether the state Sj of the monitored unit i is faulty. If the state Sj is non faulty, the method stops. If the state Sj is faulty, it is checked in a following step S60 whether a predefined number of resets is already reached. For example, it may be predetermined that a faulty unit may only be reset three times. If the predefined number of resets is already reached, the state Si is set to defective (step S64). The error element typically may contain for example the number of failed data transmissions. Thereafter, the method stops. Otherwise, i.e. if the predefined number of resets is not yet reached, the method continues with step S52. With the before described method, it is possible to recover a faulty monitored unit if this unit stopped for example due to a software error or anything similar.
[0044] Fig. 7 shows a flowchart of a method for handling a certain bus which was determined as being faulty. In step S66, all monitored units are disconnected from the faulty bus j and the parameter i is set to 1. Therefore, the monitoring unit 10 opens all switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c, which are connected to bus j, over the switch control bus 34. Thus, one of bus connection lines 20a, 20b, 20c, 22a, 22b, 22c, and 24a, 24b, 24c or connected port of the monitored unit has issues so that each of the monitored units 12, 14, and 16 is completely electrically isolated from the
busses 18a, 18b, and 18c. Then, the monitoring unit 10 closes one switch of the switches 26a, 26b, 26c, 28a, 28b, 28c, and 30a, 30b, 30c in order to connect a monitored unit i to bus j (step S68) by addressing the respective switch over the switch control bus 34 and instruct the switch to close. Thereafter, data is transmitted over the bus j to the monitored unit i and the data traffic is monitored by the monitoring unit 10 (step S70). In the following step S72, it is checked whether failed data transmissions were monitored. If failed data transmissions were monitored, the monitored unit i is permanently disconnected from bus j by the monitoring unit 10 (step S78). Then, the method continues with step S76, in which it is checked whether all monitored units were already checked. If not, the parameter i is increased by 1 in step S80 and the method continues with step S68. If no failed data transmissions were monitored, the method continues with step S74 instead of step S78 and the monitoring unit 10 keeps monitored unit i connected to bus j. If all monitored units are disconnected from the bus j after this algorithm (step S82), then the bus j will be marked defective in the bus state vector bj (step S84). Otherwise, the bus j will be marked as healthy (step S86). With the before described method, all monitored units which communicate over the bus j are one after the other connected with the bus. It may then be detected whether the bus is faulty or any monitored unit, since if each monitored unit shows failed data transmission, the probability of a faulty bus is higher than each of the monitored units is faulty. On the other hand, if one of the monitored units does not show any data transmission failures over the bus, it is likely that the bus is non faulty and one of the monitored units is faulty on at least this particular bus. It should be noted that faulty monitored unit should be kept isolated during the above method in order to avoid any influences on the bus "healing".
[0045] Furthermore, it should be noted that the concept of bus "healing" as described with reference to Fig. 7 has priority over the concept of the recovery of monitored units as described above with reference to Fig. 6.
[0046] The monitoring unit 10 can be further adapted to generate data traffic on the busses to be monitored even if this is not a requirement for the execution of the invention in practice. However, it may accelerate the monitoring and deliver faster and more reliable monitoring results than to wait for data traffic generated by the monitored units. When the monitoring unit is adapted to generate a frequent data traffic, it should be adapted in that the frequent data traffic is generated similar to a heartbeat, for
example data packets should be transmitted periodically from the monitoring unit to the monitored units over the busses.
[0047] As already mentioned above, an error element of the error matrix may be the number of failed data transmissions, particularly successive failed data transmissions from or to a certain monitored unit over a certain bus. The error element may be set to 0 if at least one or more particularly successive successful data transmission from or to the respective monitored unit over the respective bus were supervised. For example, it may be defined that an error element is only set to 0 if three successive data transmissions were monitored from or to a monitored unit over a bus. Then, the likelihood of an error of the bus or the monitored unit is so small that an error element of 0 indicating no error is justified. It should be noted that also other strategies for defining error elements and setting them back to 0 may apply within the scope of the invention.
[0048] The invention allows not only to detect any failures in a communication system applying at least one bus and several monitored units communicating over the at least one bus, but also to discriminate between a failure of a bus and a monitored unit on basis of the analyzed data transmission failures. Furthermore, the invention also provides strategies and methods to "heal" a bus if it was detected as faulty and to recover a monitored unit detected as faulty.
Claims
1. A monitoring unit (10) for supervising a communication system, wherein the communication system comprises one or a plurality of busses (18a, 18b, 18c) for connecting the monitoring unit (10) and a plurality of monitored units (12, 14, 16), wherein each monitored unit (12, 14, 16) is connected to at least one bus (18a, 18b, 18c) of the one or a plurality of busses (18a, 18b, 18c) and adapted to transmit data thereto the monitoring unit (10) is connected to one or plurality of busses (18a, 18b, 18c) and adapted to monitor the data traffic from or to the monitored units (12,
14, 16), to analyze the monitored data traffic for data transmission failures, and to discriminate between a failure of a bus (18a, 18b, 18c, 20a, 20b, 2Oc1 22a, 22b, 22c, 24a, 24b, 24c) and a monitored unit (12, 14, 16) on basis of the analyzed data transmission failures.
2. The monitoring unit of claim 1 , further adapted to generate an error matrix based on the monitoring the data traffic, wherein the error matrix comprises a row of error elements for each of the monitored units and a column of error elements for each bus or vice versa, wherein each error element indicates a data transmission failure related to a data transfer of the corresponding monitored unit over the corresponding bus, and to discriminate between a failure of a bus and a monitored units by evaluating the error matrix.
3. The monitoring unit of claim 2, wherein each error element represent a count value, wherein the monitoring unit is further adapted to increment an error element depending on subsequent failed data transfers of the corresponding monitored unit over the corresponding bus.
4. The monitoring unit of claim 2 or 3, wherein the monitoring unit is adapted to detect a bus failure, if all error elements of the corresponding row or column indicate data transfer failures.
5. The monitoring unit of claim 2 or any one of the above claims 3 to 4, wherein the monitoring unit is adapted to detect a monitored unit failure, if all error elements of the corresponding column or indicate data transfer failures.
6. The monitoring unit of claim 2 or any one of the above claims 3 to 5, further adapted to derive the state of a certain monitored unit from the error matrix by evaluating for each corresponding column or row of the error matrix, whether the value of the error elements are lower than a predefined threshold for at least one bus connected to the monitored unit (S28, S30, S32), and determine the state of the monitored unit as non faulty if the value of at least one error element is lower than the predefined threshold (S34).
7. The monitoring unit of claim 2 or any one of the above claims 3 to 6, further adapted to derive the state of a certain bus from the error matrix by evaluating for each corresponding line or row of the error matrix whether the value of the error element is lower than a predefined threshold for at least one monitored unit connected to the bus (S42, S44, S46), and determine the state of the bus as non faulty if the value of at least one error element is lower than the predefined threshold (S48).
8. The monitoring unit of claim claims 2 to 7, further adapted to perform the following steps: a) disconnect all monitored units from a faulty bus (S66), b) connect a monitored unit with bus (S68), c) transmit data over the bus to the monitored unit and monitor the data traffic (S70), d) if a failed data transmission was monitored in step c), disconnect the monitored unit from bus and, preliminarily mark the monitored unit connected to bus as faulty (S72, S78), e ) repeat steps b) to e) for different monitored units (S76, S80), and f) if all monitored units are preliminarily marked as faulty, mark the bus as faulty (S84).
9. The monitoring unit of claim 2 or any one of the above claims 3 to 8, wherein an error element is a count value being reset after at least one successful data transfer between the corresponding bus and the corresponding monitored unit
10. The monitoring unit of claim 2 or any one of the above claims 3 to 19, comprising at least one of the features: the monitoring unit is capable of disconnecting a certain monitored unit from a bus connected to the monitored unit, the monitoring unit is capable of resetting a monitored unit, and the monitoring unit is adapted to generate frequent data traffic to each monitored unit over each bus.
11. The monitoring unit of claim 1 or any one of the above claims, wherein the communication system is an automotive communication system, or a communication system of data processing or telecommunications equipment.
12. A method of supervising a communication system, the communication system comprises one or a plurality of busses for connecting a monitoring unit and a plurality of monitored units for data communication, wherein the monitoring unit is connected to the busses and the monitored unit are connected each connected to at least one of the busses, the monitoring unit performing:
monitoring the data traffic from or to the monitored units for detecting transmission failures,
analyzing the detected transmission failures for, and
discriminating between a failure of a bus and a monitored unit.
13. A software program or product, preferably stored on a data carrier, for controlling the execution of all the steps of the preceding method claim, when run on a data processing system such as a computer.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/EP2006/005088 WO2007137601A1 (en) | 2006-05-27 | 2006-05-27 | Monitoring data transfer failures in a bus based system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/EP2006/005088 WO2007137601A1 (en) | 2006-05-27 | 2006-05-27 | Monitoring data transfer failures in a bus based system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2007137601A1 true WO2007137601A1 (en) | 2007-12-06 |
Family
ID=37533467
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2006/005088 Ceased WO2007137601A1 (en) | 2006-05-27 | 2006-05-27 | Monitoring data transfer failures in a bus based system |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2007137601A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DE3506470A1 (en) * | 1985-02-23 | 1986-08-28 | Brown, Boveri & Cie Ag, 6800 Mannheim | Data transmission system |
| EP1225725A2 (en) * | 1999-02-23 | 2002-07-24 | Alcatel Internetworking, Inc. | Multi-service network switch with multiple virtual routers |
| WO2002088972A1 (en) * | 2001-04-26 | 2002-11-07 | The Boeing Company | A system and method for maintaining proper termination and error free communication in a network bus |
-
2006
- 2006-05-27 WO PCT/EP2006/005088 patent/WO2007137601A1/en not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DE3506470A1 (en) * | 1985-02-23 | 1986-08-28 | Brown, Boveri & Cie Ag, 6800 Mannheim | Data transmission system |
| EP1225725A2 (en) * | 1999-02-23 | 2002-07-24 | Alcatel Internetworking, Inc. | Multi-service network switch with multiple virtual routers |
| WO2002088972A1 (en) * | 2001-04-26 | 2002-11-07 | The Boeing Company | A system and method for maintaining proper termination and error free communication in a network bus |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8812913B2 (en) | Method and apparatus for isolating storage devices to facilitate reliable communication | |
| US8510606B2 (en) | Method and apparatus for SAS speed adjustment | |
| EP2137892B1 (en) | Node of a distributed communication system, and corresponding communication system | |
| CN110967641B (en) | Storage battery monitoring system and method | |
| CN113850033B (en) | Redundancy system, redundancy management method and readable storage medium | |
| US10721022B2 (en) | Communication apparatus, communication method, program, and communication system | |
| US20060161714A1 (en) | Method and apparatus for monitoring number of lanes between controller and PCI Express device | |
| CN103810063A (en) | Computer testing system and method | |
| CN105045164A (en) | Degradable triple-redundant synchronous voting computer control system and method | |
| KR20150018993A (en) | Apparatus and Method for implementing fail safe in Battery Management System | |
| EP2784677A1 (en) | Processing apparatus, program and method for logically separating an abnormal device based on abnormality count and a threshold | |
| CN120353741B (en) | Integrated circuit bus system, data processing method and programmable logic unit | |
| CN116680101A (en) | An operating system downtime detection method and device, elimination method and device | |
| US6125454A (en) | Method for reliably transmitting information on a bus | |
| US20240248865A1 (en) | Bus-based communication system, system-on-chip and method therefor | |
| CN101111823A (en) | Information processing device and information processing method | |
| WO2007137601A1 (en) | Monitoring data transfer failures in a bus based system | |
| US20180145869A1 (en) | Debugging method of switches | |
| CN117527653A (en) | Cluster heartbeat management method, system, equipment and medium | |
| CN109885420A (en) | A kind of analysis method, BMC and the storage medium of PCIe link failure | |
| CN114895936A (en) | Server firmware upgrading system and method | |
| CN107942894B (en) | Main input/output submodule, diagnosis method thereof and editable logic controller | |
| CN121000636B (en) | Communication link monitoring method, device, electronic equipment, storage medium and product | |
| US20190354503A1 (en) | Communication apparatus, communication method, program, and communication system | |
| US20070180329A1 (en) | Method of latent fault checking a management network |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 06753935 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A) DATED 31.03.09 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 06753935 Country of ref document: EP Kind code of ref document: A1 |
