WO2025256478A1 - 数据传输方法及装置 - Google Patents
数据传输方法及装置Info
- Publication number
- WO2025256478A1 WO2025256478A1 PCT/CN2025/099684 CN2025099684W WO2025256478A1 WO 2025256478 A1 WO2025256478 A1 WO 2025256478A1 CN 2025099684 W CN2025099684 W CN 2025099684W WO 2025256478 A1 WO2025256478 A1 WO 2025256478A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- node
- message
- fault detection
- sent
- user
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L12/00—Data switching networks
- H04L12/28—Data switching networks characterised by path configuration, e.g. LAN [Local Area Networks] or WAN [Wide Area Networks]
- H04L12/42—Loop networks
- H04L12/437—Ring fault isolation or reconfiguration
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L45/00—Routing or path finding of packets in data switching networks
- H04L45/28—Routing or path finding of packets in data switching networks using route fault recovery
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L49/00—Packet switching elements
- H04L49/90—Buffering arrangements
Definitions
- This application relates to the field of communication technology, and in particular to a data transmission method and apparatus.
- This application provides a data transmission method and apparatus to solve the packet loss problem in ring network switching technology in industrial ring networks.
- the technical solution is as follows:
- a data transmission method is provided, applied to a first node in a ring network.
- the ring network includes multiple nodes connected based on a ring topology, including the first node and a second node.
- the method includes: First, the first node sends a fault detection sequence to the second node, the fault detection sequence including multiple fault detection messages sent at intervals. Then, the first node receives fault notification information sent by the second node, the fault notification information being determined by the second node based on the number of fault detection messages received within a preset window, the fault notification information indicating a link failure between the first node and the second node.
- the first node sends a first buffered message and a second user message based on a failover path, the first buffered message being obtained by buffering the first user message sent to the second node within the preset window, the sequence number of the second user message being greater than the sequence number of the first user message.
- the first node after a link failure occurs between the first node and the second node within a preset window, when the first node retransmits packets via a switchover path, it retransmits the first user packet cached within the preset window that was sent from the first node to the second node, and continues to send subsequent packets of the first user packet based on the switchover path.
- the first node sends the first cached packet within the preset window to the second node based on the switchover path, the first user packet sent by the first node to the second node during the preset window period of the link failure can reach the second node, thereby solving the packet loss problem inherent in ring network switching technology.
- the preset window is determined based on the link bandwidth and the maximum transmission unit (MTU).
- the preset window consists of n minimum detection windows, where n is a positive integer.
- the first node while sending messages to the second node, caches first user messages according to a preset window.
- the first user messages include at least one message sent to the second node within the preset window. In this way, by caching the first user messages, the first node ensures that the cached first messages can be retransmitted after a path switch, avoiding packet loss due to link failures.
- the first node replaces the cached first cached message with the user message sent to the second node within the next preset window. This updates the cached messages and reduces the consumption of cache resources.
- each pair of adjacent user packets in the first user packet includes fault detection information.
- the first node Before sending a packet to the second node based on the switchover path, the first node caches a first count value, which is obtained by the first node counting each packet in the first user packet sent to the second node. Then, the first node sends the first count value to the second node so that the second node can remove redundant packets from the cached first packet based on the first and second count values.
- the second count value is obtained by the second node counting the fault detection information in the received first user packet. Redundant packets are those whose count value in the first user packet is greater than the second count value.
- the second node can identify them based on the first and second count values, thereby avoiding the transmission of redundant packets after the switchover path and improving the bandwidth utilization of packet transmission.
- the transmission interval between every two fault detection messages in the fault detection sequence is a preset interval for sending fault detection messages.
- the link between the first and second nodes is in a non-line-rate message state, at least one fault detection message exists between every two adjacent messages in the user messages sent from the first node to the second node.
- the link between the first and second nodes is in a line-rate message state, one fault detection message exists between every two adjacent messages in the user messages sent from the first node to the second node.
- the fault detection sequence is carried in a reserved field of the Ethernet interface. This reduces the bandwidth resources consumed in message transmission between nodes.
- the fault detection sequence is carried in the message.
- a data transmission method is provided, applied to a second node in a ring network.
- the ring network includes multiple nodes connected based on a ring topology, including a first node and a second node.
- the method includes: First, the second node receives a fault detection sequence sent by the first node, the fault detection sequence including multiple fault detection messages sent at intervals. Then, the second node determines that a link failure has occurred between the first node and the second node based on the number of fault detection messages received within a preset window. Next, the second node sends a fault notification message to the first node, causing the first node to send a first buffered message and a second user message based on a failover path. The first buffered message is obtained by buffering the first user message sent by the first node to the second node within the preset window, and the sequence number of the second user message is greater than the sequence number of the first user message.
- the second node using the preset window as the time granularity for detecting link failures, can determine the period in which the link failure occurred, enabling the first node to determine the cached packets that need to be sent to the second node after subsequent path switching, based on the preset window.
- the second node also caches a second count value, which is obtained by counting the fault detection information in the fault detection sequence received from the first node.
- the second node receives a first count value sent by the first node, which is obtained by the first node counting each message in the first user message sent to the second node.
- the second node removes redundant messages from the first cached messages based on the first and second count values. Redundant messages are those in the first user message whose count value is greater than the second count value.
- the second node can also send a second buffered message and a fourth user message based on the failover path.
- the second buffered message is obtained by buffering the third user message sent to the first node within a preset window, and the sequence number of the fourth user message is greater than that of the third user message.
- the data transmission method of the second aspect may include any implementation of the data transmission method of the first aspect, which will not be elaborated here.
- a data transmission device including a transceiver module and a processing module.
- the transceiver module is used to send a fault detection sequence to a second node; the fault detection sequence includes multiple fault detection messages sent at intervals.
- the transceiver module is also used to receive fault notification information sent by the second node; the fault notification information is determined by the second node based on the number of fault detection messages received within a preset window, and the fault notification information is used to indicate a link failure between the first node and the second node.
- the processing module is used to send a first buffered message and a second user message based on a failover path; the first buffered message is obtained by buffering the first user message sent to the second node within the preset window, and the sequence number of the second user message is greater than the sequence number of the first user message.
- the preset window is determined based on the link bandwidth and the maximum transmission unit.
- the preset window includes n minimum detection windows, where n is a positive integer.
- the minimum detection window is the sum of the minimum interval between two transmission fault detection messages and the serialization time for transmitting a maximum transmission unit based on the link bandwidth.
- the processing module is also used to: cache a first user message according to a preset window; the first user message includes at least one message sent to the second node within the preset window.
- the processing module is specifically used to: replace the cached first cached message with the user message sent to the second node in the next preset window, provided that the link between the first node and the second node is not faulty.
- each pair of adjacent packets in the first user message includes fault detection information.
- the processing module is further configured to: cache a first count value; the first count value is obtained by the first node counting each packet in the first user message sent to the second node.
- the transceiver module is further configured to: send the first count value to the second node, so that the second node removes redundant packets from the cached first message based on the first and second count values; the second count value is obtained by the second node counting the fault detection information in the first user message, and redundant packets are packets in the first user message whose count value is greater than the second count value.
- the transmission interval between every two fault detection messages in the fault detection sequence is a preset interval for transmitting fault detection messages.
- the link between the first and second nodes is in a non-line-rate message transmission state, at least one fault detection message exists in every two adjacent messages in the user messages sent from the second node to the first node.
- the link between the first and second nodes is in a line-rate message transmission state, one fault detection message exists in every two adjacent messages in the user messages sent from the first node to the second node.
- the fault detection sequence is carried in a reserved field of the Ethernet interface.
- the fault detection sequence is carried in the message.
- the aforementioned data transmission apparatus may further include other modules that perform the operational steps of the data transmission method described in the first aspect.
- a data transmission apparatus including a transceiver module.
- the transceiver module is used to receive a fault detection sequence sent by a first node; the fault detection sequence includes multiple fault detection messages sent at intervals.
- a processing module is used to determine that a link failure has occurred between the first node and a second node based on the number of fault detection messages received within a preset window.
- the transceiver module is also used to send a fault notification message to the first node, so that the first node sends a first buffered message and a second user message based on a failover path.
- the first buffered message is obtained by buffering the first user message sent by the first node to the second node within the preset window, and the sequence number of the second user message is greater than the sequence number of the first user message.
- the processing module is specifically used to: determine that the link between the first node and the second node has failed if the number of fault detection messages received within a preset window is less than a preset threshold.
- the processing module is further configured to: cache a second count value; the second count value is obtained by counting the fault detection information in the fault detection sequence.
- the transceiver module is further configured to: receive a first count value sent by the first node; the first count value is obtained by the first node counting each packet in the first user message sent to the second node.
- the processing module is further configured to: remove redundant packets from the first cached packets based on the first and second count values; redundant packets are packets in the first user message whose packet count value is greater than the second count value.
- the transceiver module is also used to: send a second buffered message and a fourth user message based on the switching path.
- the second buffered message is obtained by buffering the third user message sent to the first node within a preset window.
- the sequence number of the fourth user message is greater than the sequence number of the third user message.
- the aforementioned data transmission apparatus may further include other modules that perform the operational steps of the data transmission method described in the first aspect.
- a network device including a memory and a processor, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the data transmission method described in any possible implementation of the first or second aspect.
- a computer program comprising: computer program code, which, when executed by a computer, causes the computer to perform the data transmission method described in any possible implementation of the first or second aspect.
- a chip including a processor for retrieving and executing instructions stored in a memory, causing a communication device on which the chip is mounted to perform the data transmission method described in any possible implementation of the first or second aspect above.
- another chip comprising: an input interface, an output interface, a processor, and a memory, wherein the input interface, the output interface, the processor, and the memory are connected via an internal connection path, and the processor is configured to execute code in the memory, wherein when the code is executed, the processor is configured to execute the data transmission method described in any possible implementation of the first or second aspect above.
- a ninth aspect provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the data transmission method described in any possible implementation of the first or second aspect above.
- Figure 1 is a schematic diagram of a switching path
- FIG. 2 is a schematic diagram of a network architecture provided in this application.
- FIG. 3 is a flowchart illustrating a data transmission method provided in this application.
- Figure 4 is a schematic diagram of the transmission interval of node detection information provided in this application.
- Figure 5 is a timing diagram of the transmission control for a fault detection sequence provided in this application.
- Figure 6 is a timing diagram of the receiving control of a fault detection sequence provided in this application.
- FIG. 7 is a flowchart illustrating a message caching step provided in this application.
- Figure 8 is a schematic diagram of a path switching mechanism provided in this application.
- FIG. 9 is a flowchart illustrating a message deduplication step provided in this application.
- FIG. 10 is a flowchart illustrating another data transmission method provided in this application.
- Figure 11 is a schematic diagram of a buffered message sending method provided in this application.
- Figure 12 is a schematic diagram of another buffered message sending method provided in this application.
- FIG. 13 is a schematic diagram of clock synchronization provided in this application.
- Figure 14 is a schematic diagram of a data transmission device provided in this application.
- FIG. 15 is a schematic diagram of another data transmission device provided in this application.
- Figure 16 is a schematic diagram of the structure of an electronic device provided in this application.
- the terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application.
- the data transmission method provided in the embodiments of this application can be applied to scenarios in the communication field where the network has a ring network topology, i.e., a ring network scenario.
- a ring network scenario i.e., a ring network scenario.
- a ring network is a network consisting of multiple nodes connected in a ring topology.
- an Ethernet ring network is a ring topology composed of a group of IEEE 802.1 compliant Ethernet nodes. Each node is connected to the other two nodes through a ring port based on media access control (MAC).
- the Ethernet MAC can be carried by other service layer technologies (such as Synchronous Digital Hierarchy (SDH), multi-protocol label switching (MPLS) Ethernet pseudowires, etc.), and all nodes can communicate directly or indirectly.
- SDH Synchronous Digital Hierarchy
- MPLS multi-protocol label switching
- a switching path refers to the process where, when a link between two nodes in a ring network fails, the message transmission path between them is changed to the opposite direction of the failed link within the ring network. In other words, it uses the transmission path in the ring network other than the failed link.
- the ring network includes nodes 1, 2, 3, and 4.
- the message transmission path between nodes 1 and 3 is Node 1-Node 4-Node 3. If the link between nodes 3 and 4 fails, the nodes in the ring network execute a switching path scheme, tangenting the message transmission path between nodes 1 and 3 to Node 1-Node 2-Node 3. If the link between nodes 3 and 4 recovers, the message transmission path between nodes 1 and 3 is switched back, becoming Node 1-Node 4-Node 3.
- the Maximum Transmission Unit is used to inform the other party of the maximum size of the data service unit that can be accepted, indicating the payload size that the sender can accept.
- MTU Maximum Transmission Unit
- the maximum length limit for a data frame in Ethernet is 1500 bytes
- the maximum length limit for a data frame in IEEE 802.3 is 1492 bytes. These 1500 bytes or 1492 bytes can be referred to as the Maximum Transmission Unit.
- Ethernet is the most widely used local area network (LAN) communication method and also a protocol.
- An Ethernet interface is the port for network data connection.
- a Medium Independent Interface includes a data interface and a management interface between the MAC and PHY (physical) circuits.
- the data interface includes two independent channels for the transmitter and receiver, each with its own data, clock, and control signals.
- Related types of MII include Reduced Media Independent Interface (RMII), Serial Media Independent Interface (SMII), Serial Gigabit Media Independent Interface (SGMII), and 10Gigabit Media Independent Interface (XGMII), etc.
- Ethernet interfaces exchange information via Ethernet frames.
- Ethernet frames typically include fields such as a preamble, destination MAC address, source MAC address, length, type, and trailer.
- Ethernet frames also include variable parts, such as reserved fields.
- This application provides a data transmission method, particularly a method for retransmitting cached packets within a preset window of link failure after path switching.
- This method is applied to a first node in a ring network, which includes multiple nodes connected based on a ring topology, including a first node and a second node.
- the data transmission method includes: the first node receiving a fault detection sequence sent by the second node, the fault detection sequence including multiple fault detection messages sent at intervals.
- the first node determines that a link failure has occurred between the first node and the second node based on the number of fault detection messages contained in the received fault detection sequence within the preset window.
- the first node sends cached packets and second user packets using a path switching approach.
- the cached packets are obtained by the first node caching first user packets sent to the second node within the preset window, and the sequence number of the second user packets is greater than that of the first user packets.
- the first node sends cached packets within the preset window to the second node based on the path switching, the first user packets sent by the first node to the second node during the preset window of link failure can reach the second node, thereby solving the packet loss problem inherent in ring network switching technology.
- FIG. 2 is a schematic diagram of a network architecture provided in this application.
- the network architecture 200 may include multiple nodes connected in a ring topology.
- the network architecture 200 includes node 201, node 202, node 203, and node 204.
- network architecture 200 could be a data transmission network for an industrial internet park.
- An industrial internet park is an industrial park that uses information and communication technologies to enable network communication between industrial infrastructure within the park, such as control equipment and business terminals.
- Figure 2 illustrates one connection method for the nodes in network architecture 200.
- Node 201 is connected to node 202
- node 202 is also connected to node 203
- node 203 is also connected to node 204
- node 204 is also connected to node 201.
- Nodes 201, 202, 203, and 204 can be network devices.
- Network devices can be switches, routers, gateways, base stations, mobile core networks, optical line terminals (OLTs), wireless access points (APs), or other types of devices, used to forward or process packets from user terminals. This application does not limit the deployment location of the network devices.
- OLTs optical line terminals
- APs wireless access points
- Node 201 includes at least one access interface and two pairs of transceiver ports.
- the at least one access interface is used to connect to at least one terminal device.
- Each pair of transceiver ports includes one transmit port and one receive port.
- One pair of transceiver ports is used to connect to Node 202 to form a bidirectional link, and the other pair of transceiver ports is used to connect to Node 204 to form a bidirectional link.
- Nodes 202, 203, and 204 are similar to Node 201 and will not be described further.
- Each of nodes 201, 202, 203, and 204 is equipped with a memory.
- the memory is a buffer used to store messages sent by node 201's sending port within a preset window.
- node 201 has a buffer on the sending port side connected to node 202 to buffer messages sent by node 201 to node 202 within the preset window.
- node 202 has a buffer on the sending port side connected to node 204 to buffer messages sent by node 201 to node 204 within the preset window.
- node 201 itself has a buffer to buffer messages sent by node 201 to node 202 and to node 204 within the preset window.
- nodes 201, 202, 203, and 204 can be switches, and the network architecture 200 also includes terminal devices to which each node is connected.
- node 201 is connected to terminal device 205 through an access port
- node 202 is connected to terminal device 206 through an access port
- node 203 is connected to terminal device 207 through an access port
- node 204 is connected to terminal device 208 through an access port.
- Terminal devices can also be referred to as terminals, terminal nodes, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc.
- Terminal devices can be access points (APs), mobile phones, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, programmable logic controllers (PLCs) in industrial control, wireless terminals in self-driving, remote medical surgery, smart grids, transportation safety, smart cities, smart homes, and so on.
- the embodiments in this application do not limit the specific technologies or device forms used in the terminal devices.
- Terminal device 101 is used to communicate with other devices through network devices.
- Figure 2 is a simplified schematic diagram for ease of understanding.
- the network architecture 200 may also include other network devices and/or other terminal devices, and the connection relationships between nodes may also vary, which are not shown in Figure 2.
- FIG 3 is a flowchart illustrating a data transmission method provided in this application.
- the data transmission method may include the following steps 301-307.
- Step 301 Node 201 sends a fault detection sequence to Node 202.
- Node 201 sends a fault detection sequence to Node 202 via an Ethernet interface.
- Node 201 can also be referred to as the first node
- Node 202 can also be referred to as the second node.
- node 201 sends each fault detection message in the fault detection sequence to node 202 via the Ethernet interface at intervals.
- fault detection at node 201 After enabling fault detection at node 201, it sends fault detection information from the fault detection sequence to node 202 at intervals. Fault detection enabling can be triggered manually or by node 201 based on preset triggering conditions.
- the interval between each fault detection message in the fault detection sequence is determined based on the link status between node 201 and node 202.
- Figure 4 is a schematic diagram of the transmission interval of node detection information provided in this application. Circles are used to represent fault detection information, and bars are used to represent the interval (interval duration) between fault detection messages.
- the interval between every two fault detection messages in the fault detection sequence is a preset interval for sending fault detection messages.
- the preset interval can be flexibly adjusted according to the link attributes or the immediacy requirements of fault detection.
- node 201 when the link between nodes 201 and 202 is in a non-line-rate state, at least one fault detection message exists in every two adjacent user messages sent from node 201 to node 202.
- Node 201 sends a fault detection message after each message it sends to node 202.
- the interval between two fault detection messages is related to the message length between the fault detection messages and the message transmission rate. Taking the maximum transmission unit of the link between node 201 and node 202 as 1500 bytes and the message transmission rate as equal to the Ethernet interface rate of 1Gbps as an example, the maximum interval between the fault detection messages corresponding to the above non-line speed state is the sum of the preset interval and the transmission time of the maximum transmission unit, which is 12.5 microseconds.
- node 201 when the link between node 201 and node 202 is in the line-rate state, node 201 sends a fault detection message to node 202 every other message.
- the transmission duration of the maximum transmission unit is equal to the quotient of the maximum transmission unit and the message transmission rate (Ethernet interface rate), such as 12 microseconds.
- FIG. 5 is a timing diagram for the transmission control of a fault detection sequence provided in this application.
- Node 201 first initializes T1, T2, t1, and t2.
- T1 is the duration corresponding to the preset window
- T2 is the preset interval
- t1 and t2 are timers.
- the initial values of t1 and t2 are equal to 0.
- the timing value of t1 forms a timing cycle from 0 to the duration corresponding to the preset window
- the timing value of t2 forms a timing cycle from 0 to the preset interval.
- the preset window is determined by the maximum interval of the fault detection information.
- node 201 checks if t1 is less than T1. If t1 is less than T1, it continues timing t1 and t2. Node 201 then checks if t1 is less than T1 and if t2 is equal to T2. If t1 is greater than or equal to T1 or t2 is not equal to T2, it continues timing t1 and t2. If t1 is less than T1 and t2 is equal to T2, it checks if node 201 is sending a message to node 202. If node 201 is not sending a message to node 202, it resets t2 to its initial value and sends a fault detection sequence. It then checks if t1 is less than T1 again and executes subsequent steps.
- the interval of fault detection information in the fault detection sequence is related to the interface bandwidth, maximum transmission unit and other configurations.
- a maximum transmission unit of 1500 bytes, an Ethernet interface rate of 1Gbps and a preset interval of 0.5 microseconds the maximum interval at which node 201 sends fault detection information to node 202 under various link states is the sum of the preset interval and the transmission time of the maximum transmission unit, which is 12.5 microseconds.
- the fault detection sequence can be carried by a reserved sequence of the Ethernet interface (or Ethernet interface), which is compatible with existing Ethernet interfaces and physical layer (PHY).
- Ethernet interface or Ethernet interface
- the Ethernet interface is divided into Class GMII interfaces and Class XGMII interfaces, using different reserved fields to carry fault detection sequences.
- the fault detection sequence is carried by the reserved field of the sequence in the sequence ordered sets specified by the interface protocol.
- the fault detection sequence is carried by the reserved fields of the permissible encodings of TXD ⁇ 7.0>, TX_EN, and TX_ER specified by the interface protocol.
- the fault detection sequence can also be carried by a message.
- Step 302 Node 202 receives the fault detection sequence sent by Node 201.
- Step 303 Node 202 determines that the link between Node 201 and Node 202 has failed based on the number of fault detection messages received within the preset window.
- Node 202 determines that the link between Node 201 and Node 202 has failed based on the comparison between the number of fault detection messages received within the preset window and the preset threshold.
- the number of fault detection messages received by node 202 within a preset window is less than a preset threshold, it determines that the link between node 201 and node 202 has failed.
- the preset window is determined by the maximum interval of fault detection information, i.e., based on the link bandwidth and the maximum transmission unit. For example, with a maximum transmission unit of 1500 bytes, an Ethernet interface rate of 1Gbps, a preset interval of 0.5 microseconds, and a maximum interval of 12.5 microseconds for fault detection information, considering the 0.5 microsecond redundancy for receiving fault detection information at node 202, the minimum detection window equals the sum of the maximum interval of fault detection information and the redundancy, i.e., 13 microseconds.
- the preset window can be the product of a preset threshold and the minimum detection window.
- a preset interval of twice the value is added to the base of 39 microseconds, so the preset window is equal to 39 microseconds + 1 microsecond, which is 40 microseconds.
- Node 202 first initializes T1, T2, t1, and t2.
- T1 is the duration corresponding to the preset window
- T2 is the preset interval
- t1 and t2 are timers.
- the initial values of t1 and t2 are equal to 0.
- the timing value of t1 forms a timing cycle from 0 to the duration corresponding to the preset window
- the timing value of t2 forms a timing cycle from 0 to the preset interval.
- node 202 determines whether t1 is less than T1.
- t1 is greater than or equal to T1
- t1, t2, and cnt are reset to their initial values of 0. If t1 is less than T1, it determines whether fault detection information has been received. If no fault detection information has been received, it re-determines whether t1 is less than T1. If fault detection information has been received, cnt is incremented by 1, and t1 is checked again. If t1 is less than T1, it checks again whether fault detection information has been received and executes subsequent steps. If t1 is greater than or equal to T1, it determines whether the value of cnt is less than a preset threshold.
- cnt If the value of cnt is less than the preset threshold, a fault is determined to have occurred between node 201 and node 202. If the value of cnt is greater than or equal to the preset threshold, t1, t2, and cnt are reset to their initial values and subsequent steps are executed again. If t1 is less than T1, it re-determines whether fault detection information has been received and executes subsequent steps.
- node 202 performs fault detection at the smallest detection window within a preset window.
- Node 202 counts faults based on the received fault detection information; that is, the fault count is 0 when the link is normal.
- the preset interval count determines the minimum interval at which fault detection information is received. When the preset interval count equals the preset interval, fault detection information should be received even if no packets are received. Therefore, node 202 sets the preset interval count to zero when it receives fault detection information.
- the preset window count is used to identify the window boundaries, and the range cycles through the smallest detection window. Assume that the link between node 201 and node 202 fails at time A, which is later in the minimum detection window. Node 202 does not detect the link fault in this minimum detection window. Node 202 detects the fault at time B in the next minimum detection window, meaning the fault count is 1.
- Step 304 Node 202 sends a fault notification message to Node 201.
- the fault notification information is used to indicate that a link failure has occurred between node 201 and node 202.
- node 202 performs fault detection at the smallest detection window within a preset window. Assume that the link between node 201 and node 202 fails at time A, which is later in the smallest detection window. Node 202 does not detect the link fault in this window. In the next smallest detection window, node 202 detects the fault at time B, i.e., the fault count value equals 1, and simultaneously sends a local fault sequence (LFQ) to the peer node 201.
- the local fault sequence (LFQ) can be referred to as fault notification information.
- Step 305 Node 201 receives fault notification information.
- Step 306 Node 201 sends the first buffered message and the second user message based on the inversion path.
- Node 201 sends a first cached message and a second user message within a preset window based on the failover path.
- the first cached message is obtained by Node 201 caching the first user message sent to Node 202 within the preset window of the link failure.
- the second user message is a subsequent message to the first user message; that is, after Node 201 sends and caches the first user message, it sends the second user message to Node 202, and the sequence number of the second user message is greater than that of the first user message.
- the caching method of the first user packet by node 201 is shown in steps 701-705 of Figure 7, and will not be repeated here.
- the specific switching mechanism of the switching path is shown in the steps of Figure 8, and will not be repeated here.
- node 202 when node 201 sends the first cached message based on the switching path, considering that the message receiving side may not have the message deduplication capability, node 202 can perform message deduplication on the first cached message.
- steps 901-908 shown in Figure 9, which will not be repeated here please refer to steps 901-908 shown in Figure 9, which will not be repeated here.
- Step 307 Node 202 receives the first buffered message and the second user message based on the switching path.
- node 202 When node 202 is the destination node of the first cached message and the second user message, it receives the first cached message and the second user message based on the switching path. When the destination node of the first cached message and the second user message is another node, node 202 receives and forwards the first cached message and the second user message based on the switching path.
- node 201 After a link failure occurs between nodes 201 and 202 within a preset window, when node 201 retransmits packets via a switchover path, it retransmits the first user packet cached within the preset window that was sent from node 201 to node 202, and continues to send subsequent packets of the first user packet based on the switchover path.
- node 201 sends the first cached packet within the preset window to node 202 based on the switchover path, the first user packet sent from node 201 to node 202 during the preset window period of the link failure can reach node 202, thereby solving the packet loss problem inherent in ring network switching technology.
- This message caching step may include the following steps 701-705.
- Step 701 Node 201 sends the first user message to node 202 within a preset window.
- the first user message may include at least one message.
- Step 702 Node 201 caches the first user message according to the preset window.
- Node 201 caches the first user message within a preset window.
- Step 703 If the link between node 201 and node 202 does not fail within the preset window, node 201 replaces the cached first cached message with the user message sent by node 201 to node 202 within the next preset window.
- Step 704 If the link between node 201 and node 202 fails within the preset window, node 201 sends the cached first cached message based on the switch path.
- step 306 For the specific steps of node 201 sending the cached first cached message based on the switching path, please refer to step 306 above, which will not be repeated here.
- Step 705 In the next preset window, node 201 replaces the cached first cached message with the user message sent by node 201 to node 202 in the next preset window.
- node 201 caches the sent packets based on a preset window, so that when a link failure occurs, the cached packets are resent based on a switching path to avoid packet loss due to link failure.
- switch 1 When switch 1 receives a packet, if the packet entered from a non-loop port of switch 1, it queries the lower loop table. If no match is found, it defaults to sending the packet from port B. The lower loop table only learns the MAC-port table of its own non-loop ports. If the packet entered from a loop port of switch 1, it queries the lower loop table. If no match is found, it queries the ring network table and sends the packet from the loop port opposite to the entry port. After the switchover, the packet is forwarded to switch 2, where a lower loop table match is found, and the packet is sent from the corresponding non-loop port.
- fault detection sequences are sent between the ring ports of switches 1, 2, and 3 to perform link-level fault detection.
- buffered packets within a preset window are retransmitted.
- packets will loop back after reaching the faulty port.
- the loopback packets will either enter from the faulty port or another ring port.
- the rules for sending packets from the other ring port are determined by the aforementioned failover mechanism, eliminating the need to re-flush tables or change forwarding rules.
- node 201 needs to send the first buffered message within the preset window to node 202. Since the link failure between node 201 and node 202 can occur at any time within the preset window, and node 201's buffered messages are stored based on the preset window, messages that node 201 has already sent to node 202 within the preset window when the link failure occurred can be considered as message redundancy when node 201 sends the first buffered message based on the switching path. Therefore, the message deredundancy steps will be explained in detail below with reference to Figure 9.
- This message deduplication step may include the following steps 901-908.
- Step 901 After enabling fault detection, node 201 sends a message count reset message to node 202.
- Both node 201 and node 202 are configured with message counters.
- the message counter of node 201 is used to count the messages sent to node 202 to obtain a first count value.
- the message counter of node 202 is used to count the messages sent to node 201 to obtain a second count value.
- the message zeroing information is used to instruct the message counter of node 202 to set the second count value to zero.
- the message count zeroing information can be carried by the Ethernet interface reserved field.
- Step 905 Node 202 uses a message counter to count the fault detection information, and obtains and caches the second count value.
- Step 906 When a link failure occurs between node 201 and node 202, node 201 sends a first count value to node 202.
- Step 907 Node 202 receives the first count value.
- Step 908 Node 202 removes redundant packets from the cached packets based on the first count value and the second count value.
- Node 202 compares the second count value with the first count value in the cached message.
- the messages corresponding to the portion where the first count value is greater than the second count value are messages that Node 202 has not received, and the other messages are redundant messages. That is, the messages corresponding to the second count value are redundant messages.
- node 202 can identify it according to the first count value and the second count value, thereby avoiding sending redundant messages after switching paths and improving the bandwidth utilization of message transmission.
- each example illustrates the data transmission method of this application by showing that node 202 receives a fault detection sequence from node 201, and then, based on the fault detection sequence, instructs node 201 to retransmit cached packets according to a preset window and a switching path.
- This method is applicable to scenarios where a one-way link from node 201 to node 202 fails, while a one-way link from node 202 to node 201 remains normal.
- FIG 10 is a flowchart illustrating another data transmission method provided in this application.
- This data transmission method may include the following steps 1001-1006.
- Step 1001 Node 201 sends a fault detection sequence to Node 202.
- Step 1002 Node 202 receives the fault detection sequence.
- Step 1003 Node 202 determines that the link between Node 201 and Node 202 has failed based on the number of fault detection messages received within the preset window.
- Step 1004 Node 202 sends a fault notification message to Node 201.
- steps 301-304 in Figure 3 for steps 1001-1004 above, which will not be repeated here.
- node 201 Because the bidirectional link between node 201 and node 202 has failed, node 201 is unable to receive the fault notification information sent by node 202.
- Step 1005 Node 202 sends the second buffered message and the fourth user message based on the switching path.
- Node 202 sends a second cached message and a fourth user message within a preset window based on the failover path.
- the second cached message is obtained by Node 202 caching the third user message sent to Node 201 within the preset window of the link failure.
- the fourth user message is a subsequent message to the third user message; that is, after Node 202 sends and caches the third user message, it sends the fourth user message to Node 201, and the sequence number of the fourth user message is greater than that of the third user message.
- node 202 considering that node 201 may still be sending messages when node 202 detects that the fault count value is equal to 1, node 202 needs to wait for the maximum transmission unit to finish sending messages before sending the second buffered message and the fourth user message based on the switching path.
- the caching method of the third user's message by node 202 is described in steps 701-705 of Figure 7, and will not be repeated here.
- the specific switching mechanism of the switching path is described in steps 8, and will not be repeated here.
- node 201 when node 202 sends the second cached message based on the switching path, considering that the message receiving side may not have the message deduplication capability, node 201 can perform message deduplication on the second cached message.
- steps 901-908 shown in Figure 9, which will not be repeated here please refer to steps 901-908 shown in Figure 9, which will not be repeated here.
- Step 1006 Node 201 receives the second buffered message and the fourth user message based on the switching path.
- node 201 When node 201 is the destination node of the second buffered message and the fourth user message, it receives the second buffered message and the fourth user message based on the switching path. When the destination node of the second buffered message and the fourth user message is another node, node 201 receives and forwards the second buffered message and the fourth user message based on the switching path.
- node 202 Based on the aforementioned data transmission method, after both bidirectional links between nodes 201 and 202 fail within a preset window, when node 202 detects a failure in the link between node 201 and node 202, it retransmits the third user message cached within the preset window to node 201 during the path switching and retransmission of messages. It also continues to transmit subsequent messages of the third user message based on the switched path.
- nodes 201 and 202 have bidirectional fault detection, when either node 201 or 202 detects a failure in the link from the other end to its own end, it sends a fault notification to the other end and sends the cached message within the preset window to the other end based on the switched path. Even if a bidirectional link failure occurs between nodes 201 and 202, packet loss-free message retransmission during ring network switching can still be achieved.
- Figure 11 is a schematic diagram of a buffered message sending method provided in this application.
- Switch 1 identifies the link fault within two minimum detection windows and sends an LFQ to switch 2.
- Switch 2 receives the LFQ and records the corresponding address and region (e.g., A1/A2/A3) in its cache at that time.
- A3 Record the write address (A3_1) corresponding to the time of receiving LFQ and stop sending messages. Read the data from A1, A2, and A3 to A3_1 in sequence and send it.
- A2 Record the write address (A2_1) corresponding to the time of receiving LFQ and stop sending messages. Read data from A3, A1, and A2 in sequence to send data to A2_1.
- A1 Record the write address (A1_1) corresponding to the time of receiving LFQ and stop sending messages. Read data from A2, A3, and A1 to A1_1 in sequence and send the data.
- Figure 12 is a schematic diagram of another buffered message sending method provided in this application.
- Switch 1 or switch 2 identifies the link failure within two minimum detection windows and sends an LFQ to the peer interface. Due to the bidirectional link failure, although the LFQ is sent bidirectionally, neither switch 1 nor switch 2 receives it. Switch 1 and switch 2 record the moment the link failure is identified (i.e., the moment the fault count equals 1) and the corresponding address and region of the cache at that time.
- the minimum detection window is used as the duration threshold; that is, a total of three minimum detection windows are required from the moment the link failure occurs to triggering the switchover path.
- Switch 1 and Switch 2 determine link faults and switching paths based on the "clock" in units of the minimum detection window. Because the clocks on both sides of the link are asynchronous, the "clocks" of Switch 1 and Switch 2 are not synchronized. The maximum difference is within one minimum detection window. That is, the difference in the minimum detection window is caused by the asynchronous "clocks" at both ends of the link. Considering the minimum detection window of the maximum delay of LFQ, one minimum detection window is used as timeout compensation. A total of 4 minimum detection windows are needed from the moment the link fault occurs (within a certain minimum detection window) to the triggering of folding switching. That is, the sending side stores the packet data of 4 minimum detection windows.
- A1 When A1 receives the fault information, it reads A4, A1, A2, and A3 respectively and performs folding and switching (i.e., switching paths) and sends them.
- folding and switching i.e., switching paths
- A2 receives the fault information and reads A1, A2, A3, and A4 respectively, then folds and retransmits them.
- A3 receives the fault information and reads A2, A3, A4, and A1 respectively, then folds and retransmits them.
- A4 receives the fault information and reads A3, A4, A1, and A2 respectively, then folds and retransmits them.
- switch 1 and switch 2 synchronize the "clock" on both sides of the link by sending synchronization sequences, so that the Ethernet interfaces on both sides of the link are directly triggered by the receiving side of the fault detection sequence for folding and switching.
- the specific steps of clock synchronization can be as follows:
- the Master sends a Sync (synchronization) message (and a follow_up message if configured for two-step mode), and includes the t1 timestamp in the Sync message (or follow_up message).
- the Slave receives the Sync message at time t2, generates the t2 timestamp locally, and extracts the t1 timestamp from the message.
- Master and Slave can be switch 1 and switch 2, respectively.
- Master can be switch 1 and Slave can be switch 2.
- Master can be switch 2 and Slave can be switch 1.
- the slave sends a delay_req message at time t3 and generates a t3 timestamp locally.
- the Master receives the delay_req message at time t4, generates a t4 timestamp locally, and then carries the t4 timestamp in the delay_resp message and sends it back to the Slave.
- the slave node receives the delay_resp message and extracts the t4 timestamp from it. Finally, the slave node obtains a set of timestamps (t1, t2, t3, t4).
- Offset [(t2-t1)-(t4-t3)-(t-ms-t-sm)]/2
- the Slave can calculate the time offset between itself and the Master based on the four timestamps t1, t2, t3, and t4, and then adjust its local time accordingly, thus achieving time synchronization between the Slave and the Master.
- switch 1 or switch 2 After clock synchronization, for a bidirectional fault in the link between switch 1 and switch 2, with a link length not exceeding 200 meters, switch 1 or switch 2 identifies the fault and sends an LFQ to the other end within two minimum detection windows. Due to the bidirectional link fault, although the LFQ is sent bidirectionally, neither side receives it. Switch 1 or switch 2 records the time of fault identification (i.e., the fault count value equals 1) and the corresponding buffer position. Since the maximum time from sending the LFQ to receiving the LFQ will not exceed the minimum detection window, the minimum detection window is used as the duration threshold. That is, from the moment the link fault occurs (within a certain minimum detection window) to triggering the folding switch, a total of three minimum detection windows, i.e., the preset window, are required.
- the time of fault identification i.e., the fault count value equals 1
- the minimum detection window is used as the duration threshold. That is, from the moment the link fault occurs (within a certain minimum
- the fault detection and failover determination of switches 1 and 2 on both sides of the link are based on the "clock” in units of the minimum detection window.
- the "clocks" on both sides of the link are synchronized, so the “clocks” of switches 1 and 2 are synchronized.
- the fault detection result on the receiving side can be used to judge and process the sending side buffer of the same interface.
- A1 receives the fault information and reads A3, A1, and A2 respectively, then folds and retransmits them.
- A2 receives the fault information and reads A1, A2, and A3 respectively, then folds and retransmits them.
- A3 receives the fault information and reads A2, A3, and A1 respectively, then folds and retransmits them.
- this application also provides a data transmission device 1400, which is used to execute the data transmission method described above. As shown in FIG14, the device includes:
- the transceiver module 1410 is used to send a fault detection sequence to the second node; the fault detection sequence includes multiple fault detection messages sent at intervals.
- the transceiver module 1410 is also used to receive fault notification information sent by the second node; the fault notification information is determined by the second node based on the number of fault detection messages received within a preset window, and the fault notification information is used to indicate that a link failure has occurred between the first node and the second node.
- the processing module 1420 is used to send a first cached message and a second user message based on the switching path.
- the first cached message is obtained by caching the first user message sent to the second node within a preset window.
- the sequence number of the second user message is greater than the sequence number of the first user message.
- the preset window is determined based on the link bandwidth and the maximum transmission unit.
- the preset window includes n minimum detection windows, where n is a positive integer.
- the minimum detection window is the sum of the minimum interval between two transmission fault detection messages and the serialization time for transmitting a maximum transmission unit based on the link bandwidth.
- the processing module 1420 is also configured to: cache a first user message according to a preset window; the first user message includes at least one message sent to the second node within the preset window.
- the processing module 1420 is specifically used to: replace the cached first cached message with a user message sent to the second node in the next preset window, provided that the link between the first node and the second node is not faulty.
- each pair of adjacent packets in the first user message includes fault detection information.
- Processing module 1420 is further configured to: cache a first count value; the first count value is obtained by the first node counting each packet in the first user message sent to the second node.
- Transceiver module 1410 is further configured to: send the first count value to the second node, so that the second node removes redundant packets from the cached first message based on the first count value and the second count value; the second count value is obtained by the second node counting the fault detection information in the first user message, and redundant packets are packets in the first user message whose count value is greater than the second count value.
- the transmission interval between every two fault detection messages in the fault detection sequence is a preset interval for transmitting fault detection messages.
- the link between the first and second nodes is in a non-line-rate message transmission state, at least one fault detection message exists in every two adjacent messages in the user messages sent from the second node to the first node.
- the link between the first and second nodes is in a line-rate message transmission state, one fault detection message exists in every two adjacent messages in the user messages sent from the first node to the second node.
- the fault detection sequence is carried in a reserved field of the Ethernet interface.
- the fault detection sequence is carried in the message.
- the device shown in Figure 14 above is only illustrated by the division of the above-described functional modules.
- the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
- the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
- this application also provides a data transmission device 1500, which is used to execute the data transmission method described above. As shown in FIG15, the device includes:
- the transceiver module 1510 is used to receive the fault detection sequence sent by the first node; the fault detection sequence includes multiple fault detection messages sent at intervals.
- the processing module 1520 is used to determine that a link failure has occurred between the first node and the second node based on the number of fault detection messages received within a preset window.
- the transceiver module 1510 is also used to send fault notification information to the first node, so that the first node can send a first cached message and a second user message based on the switching path.
- the first cached message is obtained by caching the first user message sent by the first node to the second node within a preset window.
- the sequence number of the second user message is greater than the sequence number of the first user message.
- the processing module 1520 is specifically used to: determine that the link between the first node and the second node has failed when the number of fault detection messages received within a preset window is less than a preset threshold.
- processing module 1520 is further configured to: cache a second count value; the second count value is obtained by counting the fault detection information in the fault detection sequence.
- Transceiver module 1510 is further configured to: receive a first count value sent by the first node; the first count value is obtained by the first node counting each message in the first user message sent to the second node.
- Processing module 1520 is further configured to: remove redundant messages from the first cached messages based on the first count value and the second count value; redundant messages are messages in the first user message whose message count value is greater than the second count value.
- the transceiver module 1510 is also used to: send a second buffered message and a fourth user message based on the switching path.
- the second buffered message is obtained by buffering the third user message sent to the first node within a preset window, and the sequence number of the fourth user message is greater than the sequence number of the third user message.
- the device shown in Figure 15 is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, which will not be repeated here.
- Figure 16 is a schematic diagram of the structure of an electronic device provided in this embodiment.
- the electronic device 1600 includes a processor 1610, a bus 1620, a memory 1630, a communication interface 1640, and a memory unit 1650 (also referred to as a main memory unit).
- the processor 1610, the memory 1630, the memory unit 1650, and the communication interface 1640 are connected through the bus 1620.
- the processor 1610 may be a CPU, but it may also be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
- DSPs digital signal processors
- a general-purpose processor may be a microprocessor or any conventional processor.
- the processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.
- GPU graphics processing unit
- NPU neural network processing unit
- ASIC application-specific integrated circuit
- the communication interface 1640 is used to enable communication between the electronic device 1600 and external devices or components. In this embodiment, when the electronic device 1600 is used to implement the function of any node in Figure 2, the communication interface 1640 is used as a physical port for sending and receiving data packets.
- Bus 1620 may include a pathway for transferring information between the aforementioned components (such as processor 1610, memory unit 1650, and memory 1630). In addition to a data bus, bus 1620 may also include a power bus, control bus, and status signal bus. However, for clarity, all buses are labeled as bus 1620 in the figure.
- Bus 1620 may be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc.
- PCIe Peripheral Component Interconnect Express
- EISA Extended Industry Standard Architecture
- Ubus or UB Unified Bus
- CXL Compute Express Link
- CLIX Cache Coherent Interconnect for Accelerators
- electronic device 1600 may include multiple processors.
- a processor may be a multi-core (multi-CPU) processor.
- a processor may refer to one or more devices, circuits, and/or computing units used to process data (e.g., computer program instructions).
- Figure 16 only shows an example of an electronic device 1600 including a processor 1610 and a memory 1630.
- the processor 1610 and the memory 1630 are used to indicate a type of device or equipment.
- the number of each type of device or equipment can be determined according to business needs.
- Memory unit 1650 can correspond to the storage medium used to store cached messages in the above method embodiments.
- Memory unit 1650 can be volatile memory or non-volatile memory, or it can include both volatile and non-volatile memory.
- the non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.
- the volatile memory can be random access memory (RAM), which is used as an external cache.
- RAM dynamic random access memory
- DRAM dynamic random access memory
- SDRAM synchronous dynamic random access memory
- DDR SDRAM double data rate synchronous dynamic random access memory
- ESDRAM enhanced synchronous dynamic random access memory
- SLDRAM synchronous linked dynamic random access memory
- DR RAM direct rambus RAM
- the memory 1630 can correspond to the storage medium used to store computer instructions and other information in the above method embodiments, such as a disk, like a mechanical hard disk or a solid-state hard disk.
- the aforementioned electronic device 1600 can be a general-purpose device or a special-purpose device.
- electronic device 1600 can be an edge device (e.g., a box carrying a chip with processing capabilities).
- electronic device 1600 can also be a network device, a server, or other device with computing capabilities.
- the electronic device 1600 may correspond to the data transmission device 1400 or the data transmission device 1500 in this embodiment, and may correspond to the corresponding subject executing the method according to FIG3.
- the above and other operations and/or functions of each module in the data transmission device 1400 or the data transmission device 1500 are respectively for implementing the corresponding process of the method in FIG3. For the sake of brevity, they will not be described in detail here.
- the method steps in this embodiment can be implemented in hardware or by a processor executing software instructions.
- the software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art.
- An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium.
- the storage medium can also be a component of the processor.
- the processor and storage medium can reside in an ASIC.
- the ASIC can reside in an electronic device.
- the processor and storage medium can also exist as discrete components in an electronic device.
- implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof.
- software When implemented using software, it can be implemented entirely or partially in the form of a computer program product.
- the computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially.
- the computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device.
- the computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
- the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means.
- the computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media.
- the available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
- SSD solid-state drive
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Small-Scale Networks (AREA)
Abstract
提供一种数据传输方法及装置,涉及通信技术领域。该方法包括:第一节点向第二节点发送故障探测序列,序列包括间隔的多个故障探测信息。第一节点接收第二节点发送的故障通知信息,故障通知信息是第二节点根据预设窗口内接收到的故障探测信息的个数确定的。第一节点基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。如此,由于第一节点基于倒换路径向第二节点发送预设窗口内的缓存报文,链路发生故障的预设窗口期间内第一节点向第二节点发送的第一用户报文能够到达第二节点,从而解决了环网倒换技术存在的丢包问题。
Description
本申请要求于2024年06月11日提交国家知识产权局、申请号为202410749823.8、申请名称为“数据传输方法及装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及通信技术领域,尤其涉及一种数据传输方法及装置。
为保持工业以太网络的可靠性,现有的工业以太网络的拓扑大量采用环形组网。考虑到环形组网的工业以太网络可能出现链路故障,通常采用环网倒换技术将原本通过故障链路传输的报文切换至环形组网中倒换路径进行传输。但现有的环网倒换技术在故障检测和故障通告期间,原本通过故障链路传输的报文仍然在发送,在故障检测和故障通告期间发送的报文无法通过故障链路传输至目的地址,因此存在丢包问题。
本申请实施例提供一种数据传输方法及装置,以解决工业环网中环网倒换技术存在的丢包问题,技术方案如下:
第一方面,提供一种数据传输方法,应用于环形网络中的第一节点,环形网络包括基于环形拓扑结构连接的多个节点,多个节点包括第一节点和第二节点。该方法包括:首先,第一节点向第二节点发送故障探测序列,故障探测序列包括间隔发送的多个故障探测信息。然后,第一节点接收第二节点发送的故障通知信息,故障通知信息是第二节点根据预设窗口内接收到的故障探测信息的个数确定的,故障通知信息用于指示第一节点和第二节点之间的链路发生故障。接下来,第一节点基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。
基于上述数据传输方法,在第一节点和第二节点的链路在预设窗口内发生故障后,第一节点在倒换路径重新发送报文时,将预设窗口内缓存的第一节点向第二节点发送的第一用户报文再次发送,并基于倒换路径继续发送第一用户报文的后续报文。如此,由于第一节点基于倒换路径向第二节点发送预设窗口内的第一缓存报文,链路发生故障的预设窗口期间内第一节点向第二节点发送的第一用户报文能够到达第二节点,从而解决了环网倒换技术存在的丢包问题。
作为一种可能的实现方式,预设窗口是基于链路带宽和最大传输单元(maximum transmission unit,MTU)确定的。预设窗口包括n个最小检测窗口,n为正整数。最小检测窗口为两个发送故障探测信息的最小间隔与基于链路带宽发送一个最大传输单元的串行化时长的和。例如,n=2、3、4等。
作为一种可能的实现方式,第一节点在向第二节点发送报文的同时,根据预设窗口缓存第一用户报文。第一用户报文包括在预设窗口内向第二节点发送的至少一个报文。如此,第一节点通过对第一用户报文的缓存,保证倒换路径后能够对第一缓存报文进行重新发送,避免由于链路故障导致的丢包。
可选地,第一节点在第一节点和第二节点之间的链路未发生故障的情况下,采用下个预设窗口内向第二节点发送的用户报文替换已缓存的第一缓存报文。如此,实现了缓存报文的更新,减少了对缓存资源的占用。
作为一种可能的实现方式,第一用户报文中的每相邻两个报文之间包括一个故障探测信息。第一节点在基于倒换路径向第二节点发送报文之前,缓存第一计数值,第一计数值是第一节点对向第二节点发送第一用户报文中的每个报文进行计数得到的。然后,第一节点向第二节点发送第一计数值,以使第二节点根据第一计数值和第二计数值去除第一缓存报文中的冗余报文。其中,第二计数值是第二节点对接收到的第一用户报文中的故障探测信息进行计数得到的。冗余报文是第一用户报文中计数值大于第二计数值的报文。如此,针对缓存报文中第一节点在预设窗口内已向第二节点发送的部分报文,第二节点能够根据第一计数值和第二计数值将其识别出来,从而在倒换路径后避免发送冗余报文,提高了报文传输的带宽利用率。
作为一种可能的实现方式,在第一节点和第二节点之间的链路处于空闲状态的情况下,故障探测序列的每两个故障探测信息的发送间隔为发送故障探测信息的预设间隔。在第一节点和第二节点之间的链路处于报文非线速状态的情况下,第一节点向第二节点发送的用户报文中每相邻两个报文中存在至少一个故障探测信息。在第一节点和第二节点之间的链路处于报文线速状态的情况下,第一节点向第二节点发送的用户报文中每相邻两个报文之间存在一个故障探测信息。
作为一种可能的实现方式,故障探测序列承载于以太接口保留字段。如此,减少对节点之间报文传输的带宽资源的占用。
作为一种可能的实现方式,故障探测序列承载于报文。
第二方面,提供一种数据传输方法,应用于环形网络中的第二节点,环形网络包括基于环形拓扑结构连接的多个节点,多个节点包括第一节点和第二节点。该方法包括:首先,第二节点接收第一节点发送的故障探测序列,故障探测序列包括间隔发送的多个故障探测信息。然后,第二节点根据预设窗口内接收到的故障探测信息的个数,确定第一节点和第二节点之间的链路发生故障。接下来,第二节点向第一节点发送故障通知信息,以使第一节点基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内第一节点向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。
作为一种可能的实现方式,第二节点在预设窗口内接收到的故障探测信息的个数小于预设阈值的情况下,确定第一节点和第二节点之间的链路发生故障。如此,第二节点根据预设窗口作为检测链路故障的时间粒度,能够确定链路发生故障的时段,使第一节点在后续倒换路径后根据预设窗口确定已缓存的需要向第二节点发送的缓存报文。
作为一种可能的实现方式,第二节点还缓存第二计数值,第二计数值是对从第一节点接收到的故障探测序列的故障探测信息进行计数得到的。第二节点接收第一节点发送的第一计数值,第一计数值是第一节点对向第二节点发送第一用户报文中的每个报文进行计数得到的。第二节点根据第一计数值和第二计数值去除第一缓存报文中的冗余报文。冗余报文是第一用户报文中报文计数值大于第二计数值的报文。
作为一种可能的实现方式,第二节点还可以基于倒换路径发送第二缓存报文和第四用户报文。第二缓存报文是对预设窗口内向第一节点发送的第三用户报文进行缓存得到的,第四用户报文的报文序号大于第三用户报文的报文序号。如此,在第一节点和第二节点之间的链路存在双向故障的情况下,第一节点和第二节点均无法接收到故障通知信息,第二节点也能够根据故障探测信息,基于倒换路径发送第二缓存报文和第四用户报文,从而解决了链路双向故障场景下的丢包问题。
作为一种可能的实现方式,第二方面的数据传输方法可以包括第一方面的数据传输方法的任一种实施方式,在此不再赘述。
关于第二方面的技术原理和有益效果,可以参考前述第一方面的相关描述,在此不再赘述。
第三方面,提供一种数据传输装置,包括收发模块和处理模块。收发模块,用于向第二节点发送故障探测序列;故障探测序列包括间隔发送的多个故障探测信息。收发模块,还用于接收第二节点发送的故障通知信息;故障通知信息是第二节点根据预设窗口内接收到的故障探测信息的个数确定的,故障通知信息用于指示第一节点和第二节点之间的链路发生故障。处理模块,用于基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。
作为一种可能的实现方式,预设窗口是基于链路带宽和最大传输单元确定的。
作为一种可能的实现方式,预设窗口包括n个最小检测窗口,n为正整数,最小检测窗口为两个发送故障探测信息的最小间隔与基于链路带宽发送一个最大传输单元的串行化时长的和。
作为一种可能的实现方式,处理模块还用于:根据预设窗口缓存第一用户报文;第一用户报文包括在预设窗口内向第二节点发送的至少一个报文。
作为一种可能的实现方式,处理模块具体用于:在第一节点和第二节点之间的链路未发生故障的情况下,采用下个预设窗口内向第二节点发送的用户报文替换已缓存的第一缓存报文。
作为一种可能的实现方式,第一用户报文中的每相邻两个报文之间包括一个故障探测信息。处理模块还用于:缓存第一计数值;第一计数值是第一节点对向第二节点发送第一用户报文中的每个报文进行计数得到的。收发模块还用于:向第二节点发送第一计数值,以使第二节点根据第一计数值和第二计数值去除第一缓存报文中的冗余报文;第二计数值是第二节点对第一用户报文中的故障探测信息进行计数得到的,冗余报文是第一用户报文中报文计数值大于第二计数值的报文。
作为一种可能的实现方式,在第一节点和第二节点之间的链路处于空闲状态的情况下,故障探测序列的每两个故障探测信息的发送间隔为发送故障探测信息的预设间隔。在第一节点和第二节点之间的链路处于报文非线速状态的情况下,第二节点向第一节点发送的用户报文中每相邻两个报文中存在至少一个故障探测信息。在第一节点和第二节点之间的链路处于报文线速状态的情况下,第一节点向第二节点发送的用户报文中每相邻两个报文中存在一个故障探测信息。
作为一种可能的实现方式,故障探测序列承载于以太接口保留字段。
作为一种可能的实现方式,故障探测序列承载于报文。
作为一种可能的实现方式,上述数据传输装置还可以包括执行第一方面所述的数据传输方法的操作步骤的其他模块。
关于第三方面的技术原理和有益效果,可以参考前述第一方面的相关描述,在此不再赘述。
第四方面,提供一种数据传输装置,包括收发模块。收发模块,用于接收第一节点发送的故障探测序列;故障探测序列包括间隔发送的多个故障探测信息。处理模块,用于根据预设窗口内接收到的故障探测信息的个数,确定第一节点和所二节点之间的链路发生故障。收发模块,还用于向第一节点发送故障通知信息,以使第一节点基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内第一节点向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。
作为一种可能的实现方式,处理模块具体用于:在预设窗口内接收到的故障探测信息的个数小于预设阈值的情况下,确定第一节点和第二节点之间的链路发生故障。
作为一种可能的实现方式,处理模块还用于:缓存第二计数值;第二计数值是对故障探测序列的故障探测信息进行计数得到的。收发模块还用于:接收第一节点发送的第一计数值;第一计数值是第一节点对向第二节点发送第一用户报文中的每个报文进行计数得到的。处理模块还用于:根据第一计数值和第二计数值去除第一缓存报文中的冗余报文;冗余报文是第一用户报文中报文计数值大于第二计数值的报文。
作为一种可能的实现方式,收发模块还用于:基于倒换路径发送第二缓存报文和第四用户报文,第二缓存报文是对预设窗口内向第一节点发送的第三用户报文进行缓存得到的,第四用户报文的报文序号大于第三用户报文的报文序号。
作为一种可能的实现方式,上述数据传输装置还可以包括执行第一方面所述的数据传输方法的操作步骤的其他模块。
关于第四方面的技术原理和有益效果,可以参考前述第一方面的相关描述,在此不再赘述。
第五方面,提供一种网络设备,包括存储器及处理器,所述存储器中存储有至少一条指令,所述至少一条指令由所述处理器加载并执行,以实现上述第一方面或第二方面中任一可能的实现方式所述的数据传输方法。
第六方面,提供一种计算机程序(产品),所述计算机程序(产品)包括:计算机程序代码,当所述计算机程序代码被计算机运行时,使得所述计算机执行上述第一方面或第二方面中任一可能的实现方式所述的数据传输方法。
第七方面,提供一种芯片,包括处理器,用于从存储器中调用并运行所述存储器中存储的指令,使得安装有所述芯片的通信设备执行上述第一方面或第二方面中任一可能的实现方式所述的数据传输方法。
第八方面,提供另一种芯片,包括:输入接口、输出接口、处理器和存储器,所述输入接口、输出接口、所述处理器以及所述存储器之间通过内部连接通路相连,所述处理器用于执行所述存储器中的代码,当所述代码被执行时,所述处理器用于执行上述第一方面或第二方面中任一可能的实现方式所述的数据传输方法。
第九方面,提供一种计算机可读存储介质,所述存储介质中存储有至少一条指令,所述指令由处理器加载并执行以实现上述第一方面或第二方面中任一可能的实现方式所述的数据传输方法。
图1为一种倒换路径的示意图;
图2为本申请提供的一种网络架构的示意图;
图3为本申请提供的一种数据传输方法的流程示意图;
图4为本申请提供的一种节点探测信息的发送间隔示意图;
图5为本申请提供的一种故障探测序列的发送控制时序图;
图6为本申请提供的一种故障探测序列的接收控制时序图;
图7为本申请提供的一种报文缓存步骤的流程示意图;
图8为本申请提供的一种倒换路径的机制的示意图;
图9为本申请提供的一种报文去冗步骤的流程示意图;
图10为本申请提供的另一种数据传输方法的流程示意图;
图11为本申请提供的一种缓存报文发送方式的示意图;
图12为本申请提供的另一种缓存报文发送方式的示意图;
图13为本申请提供的时钟同步的示意图;
图14为本申请提供的一种数据传输装置的示意图;
图15为本申请提供的另一种数据传输装置的示意图;
图16为本申请提供的一种电子设备的结构示意图。
本申请的实施方式部分使用的术语仅用于对本申请的具体实施例进行解释,而非旨在限定本申请。例如,本申请实施例提供的数据传输方法能够应用于通信领域中网络为环网拓扑结构的场景,即环网场景中。下面对本申请可能涉及的技术进行简单介绍。
(1)环网
环网是指环形网络,即基于环形拓扑结构连接的多个节点组成的网络。例如以太环网,是由一组IEEE 802.1兼容的以太网节点组成的环形拓扑,每个节点通过基于媒体访问控制(media access control,MAC)的环端口与其他两个节点相连,而以太网MAC可以由其他服务层技术承载(如同步数字体系(Synchronous Digital Hierarchy,SDH)、多协议标签交换(multi-protocol label switching,MPLS)的以太网伪线等),所有节点间能够直接或者间接通信。
(2)倒换路径
倒换路径是指环网中的两个节点之间的链路发生故障时,两个节点之间的报文传输路径从故障链路的相反环网方向,即环网中除故障链路之外的传输路径进行报文传输。如图1所示,环网包括节点1、节点2、节点3和节点4,节点1和节点3之间的报文传输路径为节点1-节点4-节点3。若节点3和节点4之间的链路发生故障,则环网中的节点执行倒换路径的方案,将节点1和节点3之间的报文传输路径正切为节点1-节点2-节点3。若节点3和节点4之间的链路恢复,则将节点1和节点3之间的报文传输路径进行回切,即节点1-节点4-节点3。
(3)最大传输单元
最大传输单元用来通知对方所能接受数据服务单元的最大尺寸,说明发送方能够接受的有效载荷大小。例如,以太网数对数据帧的长度限制的最大值为1500字节,IEEE 802.3对数据帧的长度限制的最大值为1492字节,上述1500字节或1492字节可以被称为最大传输单元。
(4)以太网接口
以太网(ethernet)是应用最广泛的局域网通讯方式,同时也是一种协议。而以太网接口就是网络数据连接的端口。例如媒体独立接口(Medium Independent Interface,MII),包括MAC和PHY(physical)之间的一个数据接口和一个管理接口。数据接口包括分别用于发送器和接收器的两条独立信道,每条信道都有自己的数据、时钟和控制信号。媒体独立接口的相关类型包括简化媒体独立接口(Reduced Media Independant Interface,RMII)、串行媒体独立接口(Serial Media Independent Interface,SMII)、串行千兆媒体独立接口(Serial Gigabit Media Independent Interface,SGMII)GMII、10Gb独立媒体接口(10Gigabit Media Independent Interface,XGMII)等。
以太网接口之间通过以太网帧进行信息传递,以太网帧通常包括前导码、目的MAC地址、源MAC地址、长度、类型、帧尾等字段。此外,以太网帧还包括可变部分,例如保留(reserved)字段等。
目前针对环网的环网倒换技术,例如ERPS等协议,主要由故障探测、故障通告、清表、表项重建等步骤组成,只有在清表以后报文在传输路径中由单播转为广播才能避免后续的报文丢包,在此之前的流程中由于报文的发送侧仍然在基于已经出现链路故障的传输路径发送报文,而报文的接收侧无法收到报文,不可避免地存在丢包。
本申请实施例提供了一种数据传输方法,尤其是一种在倒换路径后对链路出现故障的预设窗口内的缓存报文进行重发的数据传输方法。该方法应用于环形网络中的第一节点,环形网络包括基于环形拓扑结构连接的多个节点,多个节点包括第一节点和第二节点。该数据传输方法包括:第一节点接收第二节点发送的故障探测序列,故障探测序列包括间隔发送的多个故障探测信息。第一节点根据预设窗口内接收到的故障探测序列包含的故障探测信息的个数,确定第一节点和第二节点之间的链路发生故障。第一节点采用倒换路径的方式发送缓存报文和第二用户报文,缓存报文是第一节点对预设窗口内向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。如此,由于第一节点基于倒换路径向第二节点发送预设窗口内的缓存报文,链路发生故障的预设窗口期间内第一节点向第二节点发送的第一用户报文能够到达第二节点,从而解决了环网倒换技术存在的丢包问题。
下面将结合附图对本申请实施例的实施方式进行详细描述。
图2为本申请提供的一种网络架构的示意图。网络架构200可以包括基于环形拓扑结构连接的多个节点。网络架构200包括节点201、节点202、节点203和节点204。
例如,网络架构200可以是一个工业互联网园区的数据传输网络。工业互联网园区是通过信息通信技术实现控制设备、业务终端等工业园区内的工业基础设施的联网通信的工业园区。
图2示出了网络架构200中各个节点的一种连接方式。节点201与节点202连接,节点202还与节点203连接,节点203还与节点204连接,节点204还与节点201连接。
节点201、节点202、节点203和节点204可以是网络设备。
网络设备可以是交换机、路由器、网关、基站、移动核心网、光线路终端(optical line terminal,OLT)、无线访问点(access coint,AP)或其他类型的设备,用于对用户终端的报文进行转发或处理。本申请实施例对网络设备的部署位置不予限定。
节点201包括至少一个接入接口以及两对收发端口。至少一个接入接口用于与至少一个终端设备连接。每对收发端口包括一个发送端口和一个接收端口,一对收发端口用于与节点202连接构成双向链路,另一对收发端口用于与节点204连接构成双向链路。节点202、节点203、节点204与节点201类似,在此不再赘述。
节点201、节点202、节点203和节点204中的每个节点都设置有存储器。以节点201为例,存储器为缓冲器(buffer),用于存储节点201的发送端口在预设窗口内发送的报文。例如,节点201在与节点202连接的发送端口侧设置有一个缓冲器,用于缓存节点201在预设窗口内向节点202发送的报文。又如,节点202在与节点204连接的发送端口侧设置有一个缓冲器,用于缓存节点201在预设窗口内向节点204发送的报文。还如,节点201内设置有一个缓冲器,用于缓存节点201在预设窗口内向节点202发送的报文,以及节点201在预设窗口内向节点204发送的报文。
作为一种可能的实现方式,节点201、节点202、节点203和节点204可以是交换机,网络架构200还包括终端设备每个节点各自连接的终端设备。例如,节点201通过一个接入端口与终端设备205连接,节点202通过一个接入端口与终端设备206连接,节点203通过一个接入端口与终端设备207连接,节点204通过一个接入端口与终端设备208连接。
终端设备也可以称为终端(Terminal)、终端节点、用户设备(user equipment,UE)、移动台(mobile station,MS)、移动终端(mobile terminal,MT)等。终端设备可以是访问接入点(access point,AP)、手机(mobile phone)、平板电脑(pad)、带无线收发功能的电脑、虚拟现实(virtual reality,VR)终端设备、增强现实(augmented reality,AR)终端设备、工业控制(industrial control)中的可编程逻辑控制器(rrogrammable logic controller,PLC)等线终端、无人驾驶(self driving)中的无线终端、远程手术(remote medical surgery)中的无线终端、智能电网(smart grid)中的无线终端、运输安全(ttransportation safety)中的无线终端、智慧城市(smart city)中的无线终端、智慧家庭(smart home)中的无线终端等等。本申请的实施例对终端设备所采用的具体技术和具体设备形态不做限定。终端设备101用于通过网络设备与其他设备进行通信。
应理解,图2仅为便于理解而示例的简化示意图,该网络架构200中还可以包括其它网络设备,和/或,其它终端设备,各节点间的连接关系也可以有其他变化,图2中未予以画出。
应当指出的是,本申请实施例中的方案还可以应用于其它网络中,例如其他类型的园区网络、数据中心网络、移动承载网络等,相应的名称也可以用其它网络架构中的对应功能的名称进行替代。
接下来,将结合附图对本申请实施例提供的数据传输方法进行具体描述。在这里以图2中的网络架构200中的各节点执行数据传输方法为例,对数据传输方法的具体步骤进行说明。
图3为本申请提供的一种数据传输方法的流程示意图。请参考图3,数据传输方法可以包括如下步骤301-步骤307。
步骤301、节点201向节点202发送故障探测序列。
节点201通过以太网接口向节点202发送故障探测序列。其中,本实施例的节点201也可称为第一节点,节点202也可称为第二节点。
作为一种可能的实现方式,节点201根据间隔,通过以太网接口向节点202发送故障探测序列中的各个故障探测信息。
在节点201执行故障探测使能后,向节点202间隔发送故障探测序列中各个故障探测信息。其中,故障探测使能可以是由人工配置触发,也可以是由节点201根据预设触发条件触发。
可选地,故障探测序列中各个故障探测信息的间隔是基于节点201和节点202之间的链路状态确定的。请参考图4,图4为本申请提供的一种节点探测信息的发送间隔示意图,圆形用于表示故障探测信息,条形用于表示故障探测信息之间的间隔(间隔时长)。
例如,节点201和节点202之间的链路处于空闲(idie)状态的情况下,故障探测序列的每两个故障探测信息的发送间隔为发送故障探测信息的预设间隔。
其中,预设间隔可以根据链路属性或故障探测的即时性需求进行灵活调整,如预设间隔T等于0.1微秒、0.5微秒、1微秒、2.6微秒等,本实施例中以T=0.5微秒为例进行后续说明。
又如,节点201和节点202之间的链路处于报文非线速状态的情况下,节点201向节点202发送的用户报文中每相邻两个报文中存在至少一个故障探测信息。节点201每向节点202发送一个报文,在报文发送结束后发送一个故障探测信息。
其中,每两个故障探测信息的间隔与故障探测信息之间的报文长度,以及报文传输速率相关。以节点201和节点202之间的链路的最大传输单元为1500字节、报文传输速率等于以太接口速率1Gbps为例,上述非线速状态对应的故障探测信息之间的最大间隔为预设间隔与最大传输单元的传输时长之和,即12.5微秒。
还如,节点201和节点202之间的链路处于报文线速状态的情况下,节点201每间隔一个报文向节点202发送一个故障探测信息。
其中,最大传输单元的传输时长等于最大传输单元与报文传输速率(以太接口速率)的商,如12微秒。
基于上述三种链路状态,节点201发送故障探测序列的总体判定方式可以参考如图5,图5为本申请提供的一种故障探测序列的发送控制时序图。节点201首先初始化T1、T2、t1和t2,T1为预设窗口对应的时长,T2为预设间隔,t1和t2为计时器,t1和t2的初始化的初值等于0,t1的计时值以0至预设窗口对应的时长构成计时循环,t2的计时值以0至预设间隔构成计时循环。其中预设窗口是由故障探测信息的最大间隔确定的。然后,节点201判断t1是否小于T1,若t1小于T1,继续执行t1和t2的计时,节点201再判断t1是否小于T1,t2是否等于T2,t1大于等于T1或t2不等于T2,继续执行t1和t2的计时,若t1小于T1,t2等于T2,判断节点201是否正在向节点202发送报文,若节点201未向节点202发送报文,将t2重置为初始值,并发送故障探测序列,再次判断t1是否小于T1并执行后续步骤。若节点201正在向节点202发送报文,节点201判断t1是否小于T1,t2是否等于T2,若是,重新判断节点201是否正在向节点202发送报文并执行后续步骤,若否,在t1等于T1时重置t1和t2为初始值。
如此,故障探测序列中故障探测信息的间隔与接口带宽、最大传输单元等配置相关联,在最大传输单元为1500字节、以太接口速率1Gbps、预设间隔为0.5微秒的情况下,节点201在各种链路状态下向节点202发送故障探测信息的最大间隔为预设间隔与最大传输单元的传输时长之和,即12.5微秒。
作为一种可能的实现方式,由于故障探测序列由以太网接口传输,故障探测序列可以由以太网接口(或称以太接口)的保留序列承载,兼容现有的以太接口和物理层(PHY)。
可选地,根据以太网接口的MAC与PHY接口的类型,分为类GMII接口和类XGMII接口,采用不同保留字段承载故障探测序列。例如,针对类XGMII接口,故障探测序列由该类接口协议规定的序列命令配置(sequence ordered sets)中sequence的保留(reserved)字段承载。又如,针对类GMII接口,故障探测序列由该类接口协议规定的permissible encodings ofTXD<7.0>,TX_EN,and TX_ER中的保留字段承载。
作为一种可能的实现方式,故障探测序列还可以由报文承载。
步骤302、节点202接收节点201发送的故障探测序列。
节点202通过以太网接口接收节点201发送的故障探测序列。
步骤303、节点202根据预设窗口内接收到的故障探测信息的个数,确定节点201和节点202之间的链路发生故障。
节点202根据预设窗口内接收到的故障探测信息的个数与预设阈值的比较结果,确定节点201和节点202之间的链路发生故障。
作为一种可能的实现方式,节点202在预设窗口内接收到的故障探测信息的个数小于预设阈值的情况下,确定节点201和节点202之间的链路发生故障。
可选地,预设窗口是由故障探测信息的最大间隔确定的,即是基于链路带宽和最大传输单元确定的。例如,在最大传输单元为1500字节、以太接口速率1Gbps、预设间隔为0.5微秒、故障探测信息的最大间隔为12.5微秒,则考虑到节点202接收故障探测信息的设置收发冗余为0.5微秒,最小检测窗口等于故障探测信息的最大间隔与收发冗余之和即13微秒。预设窗口可以是预设阈值与最小检测窗口的乘积。其中,预设阈值是考虑到收发延迟等,在一个最小检测窗口的时长的基础上添加n-1个最小检测窗口后得到的,例如n=2、3、4、5等。如此,预设窗口等于预设阈值与最小检测窗口的乘积,例如n=3,最小检测窗口等于13us,则预设窗口等于3*13微秒=39微秒。在本实施例中,考虑到收发延迟等,在39微秒的基础上添加两倍预设间隔,则预设窗口等于39微秒+1微秒即40微秒。
接下来结合图6对节点202的故障判定逻辑进行说明,图6为本申请提供的一种故障探测序列的接收控制时序图。节点202首先初始化T1、T2、t1、t2,T1为预设窗口对应的时长,T2为预设间隔,t1和t2为计时器,t1和t2的初始化的初值等于0,t1的计时值以0至预设窗口对应的时长构成计时循环,t2的计时值以0至预设间隔构成计时循环。然后,节点202判断t1是否小于T1,若t1大于或等于T1,将t1、t2和cnt重置为初始值0,若t1小于T1,判断是否接收到故障探测信息,若未接收到故障探测信息,重新判断t1是否小于T1,若接收到故障探测信息,cnt增加1,再判断t1是否小于T1,若t1小于T1,再次判断是否接收到故障探测信息并执行后续步骤,若t1大于或等于T1,判断cnt的值是否小于预设阈值(门限值),若cnt的值小于预设阈值,确定节点201和节点202之间发生故障,若cnt的值大于或等于预设阈值,重置t1、t2、cnt为初始值并再次执行后续步骤,若t1小于T1,重新判断是否接收到故障探测信息并执行后续步骤。
作为一种可能的实现方式,节点202以预设窗口中的最小检测窗口为颗粒度进行故障探测。节点202根据接收到的故障探测信息进行故障计数,即链路正常时故障计数值等于0,预设间隔计数决定接收到故障探测信息的最小间隔,预设间隔计数值等于预设间隔时在报文无接收的情况下应接收到故障探测信息,则节点202在接收到故障探测信息时将预设间隔计数值置零。预设窗口计数用于标识窗口边界,范围以最小检测窗口进行循环。假设节点201与节点202之间的链路在A时刻发生故障,A处于最小检测窗口的靠后位置,该最小检测窗口中节点202未识别出链路故障,节点202在下一个最小检测窗口识别B时刻发生故障,即故障计数值等于1。
步骤304、节点202向节点201发送故障通知信息。
故障通知信息用于指示节点201和节点202之间的链路发生故障。
作为一种可能的实现方式,节点202以预设窗口中的最小检测窗口为颗粒度进行故障探测。假设节点201与节点202之间的链路在A时刻发生故障,A处于最小检测窗口的靠后位置,该最小检测窗口中节点202未识别出链路故障,节点202在下一个最小检测窗口识别B时刻发生故障,即故障计数值等于1并同时向对端即节点201发送本地故障序列(local fault sequence,LFQ)。本地故障序列(local fault sequence,LFQ)可以被称为故障通知信息。
步骤305、节点201接收故障通知信息。
步骤306、节点201基于倒换路径发送第一缓存报文和第二用户报文。
节点201基于倒换路径发送预设窗口内缓存的第一缓存报文,以及第二用户报文。其中,第一缓存报文是节点201在链路故障的预设窗口内对向节点202发送的第一用户报文进行缓存得到的。第二用户报文是第一用户报文的后续报文,即节点201在发送第一用户报文并缓存后,向节点202发送的第二用户报文,第二用户报文的报文序号大于第一用户报文的报文序号。
在本实施例中,上述节点201对第一用户报文的缓存方式请参考图7所示的步骤701-步骤705,在此不再赘述。上述倒换路径的具体倒换机制请参考图8所示的步骤,在此不再赘述。
在本实施例中,节点201在基于倒换路径发送第一缓存报文时,考虑到报文接收侧可能不具备报文去冗能力,节点202可以对第一缓存报文进行报文去冗,报文去冗的具体步骤请参考图9所示的步骤901-步骤908,在此不再赘述。
步骤307、节点202基于倒换路径接收第一缓存报文和第二用户报文。
节点202在第一缓存报文和第二用户报文的目的节点为节点202时,基于倒换路径接收第一缓存报文和第二用户报文。节点202在第一缓存报文和第二用户报文的目的节点为其他节点时,基于倒换路径接收并转发第一缓存报文和第二用户报文。
基于上述数据传输方法,在节点201和节点202的链路在预设窗口内发生故障后,节点201在倒换路径重新发送报文时,将预设窗口内缓存的节点201向节点202发送的第一用户报文再次发送,并基于倒换路径继续发送第一用户报文的后续报文。如此,由于节点201基于倒换路径向节点202发送预设窗口内的第一缓存报文,链路发生故障的预设窗口期间内节点201向节点202发送的第一用户报文能够到达节点202,从而解决了环网倒换技术存在的丢包问题。
上文结合图3-图6对数据传输方法的整体流程进行了说明,接下来结合图7对报文缓存的具体步骤进行详细说明。
请参考图7,图7为本申请提供的一种报文缓存步骤的流程示意图。该报文缓存步骤可以包括如下步骤701-步骤705。
步骤701、节点201在预设窗口内向节点202发送第一用户报文。
第一用户报文可以包括至少一个报文。
步骤702、节点201根据预设窗口缓存第一用户报文。
节点201在预设窗口内缓存第一用户报文。
步骤703、节点201在预设窗口内节点201与节点202之间的链路未发生故障的情况下,在下个预设窗口采用下个预设窗口内节点201向节点202发送的用户报文替换已缓存的第一缓存报文。
步骤704、节点201在预设窗口内节点201和节点202之间的链路发生故障的情况下,基于倒换路径发送已缓存的第一缓存报文。
节点201基于倒换路径发送已缓存的第一缓存报文的具体步骤请参考上述步骤306,在此不再赘述。
步骤705、节点201在下个预设窗口采用下个预设窗口内节点201向节点202发送的用户报文替换已缓存的第一缓存报文。
基于上述步骤701-步骤705,节点201基于预设窗口对已发送的报文进行缓存,从而在链路发生故障时基于倒换路径重新发送缓存报文,避免由于链路故障导致的丢包。
上文结合图7对报文缓存的具体步骤进行了说明,节点201涉及基于倒换路径重发缓存报文和第二用户报文,接下来结合图8对倒换路径的机制进行详细说明。
请参考图8,图8为本申请提供的一种倒换路径的机制的示意图。以交换机1、交换机2和交换机3为例,每个交换机包括端口A至端口F,其中,端口A和端口B是与环形拓扑结构中其他交换机连接的环端口,端口C至端口F是与终端设备连接的非环端口。
交换机1接收报文,若该报文是从交换机1的非环端口进入,查询下环表,未命中则默认从端口B发出。其中,下环表仅学习自身非环端口的MAC-端口(port)表。若报文是从交换机1的环端口进入,查询下环表,未命中时查询环网表,从进入环端口的对端环口发出。倒换后报文转发至交换机2,查询下环表命中,从对应的非环端口发出。
结合图3所示的数据传输方法,交换机1、交换机2和交换机3之间的环端口之间发送故障探测序列,已进行链路级的故障检测,并根据上述倒换路径的机制将预设窗口内的缓存报文进行重发。如此,假设交换机2和交换机3之间的链路发生故障,报文到达故障端口后进行环回,环回后的报文是从故障端口进入或另一个环端口进入,另一个环端口发出报文的规则,均是由上述倒换路径的机制确定的,不需要重新刷表或更改转发规则。
上文结合图8对倒换路径的机制进行了说明,而倒换路径后节点201需要向节点202发送的预设窗口内的第一缓存报文,由于节点201和节点202之间的链路发生故障的时刻可以是预设窗口内的任意时刻,而节点201的缓存报文是基于预设窗口存储的,针对节点201在链路发生故障的预设窗口内已向节点202发送的报文,在节点201基于倒换路径发送第一缓存报文时,可被视为报文冗余。因此,接下来结合图9对报文去冗步骤进行详细说明。
请参考图9,图9为本申请提供的一种报文去冗步骤的流程示意图。该报文去冗步骤可以包括如下步骤901-步骤908。
步骤901、节点201在故障探测使能后向节点202发送报文计数置零信息。
节点201和节点202均配置有报文计数器,节点201的报文计数器用于对向节点202发送的报文进行计数,得到第一计数值,节点202的报文计数器用于对节点201发送来的报文进行计数,得到第二计数值,报文置零信息用于指示节点202的报文计数器将第二计数值置零。
作为一种可能的实现方式,报文计数置零信息可以由以太网接口保留字段承载。
步骤902、节点202根据报文计数置零信息将报文计数器置零。
步骤903、节点201向节点202发送报文及故障探测序列。
节点201每向节点202发送一个报文之后,向节点202发送一个故障探测信息,多个报文包含的多个故障探测信息组成故障探测序列。其中,多个报文可以是第一用户报文包含的报文。
步骤904、节点201采用报文计数器对故障探测信息进行计数,得到并缓存第一计数值。
节点201每发送一个故障探测信息,报文计数器的第一计数值加一,并将第一计数值和报文一起进行缓存。
步骤905、节点202采用报文计数器对故障探测信息进行计数,得到并缓存第二计数值。
节点202每接收一个故障探测信息,报文计数器的第二计数值加一,并将第二计数值和报文一起进行缓存。
步骤906、在节点201和节点202之间的链路出现故障时,节点201向节点202发送第一计数值。
步骤907、节点202接收第一计数值。
步骤908、节点202根据第一计数值和第二计数值去除缓存报文中的冗余报文。
节点202将缓存报文中第二计数值与第一计数值进行比较,第一计数值大于第二计数值的部分对应的报文是节点202未接收的报文,其他报文为冗余报文,即第二计数值对应的报文是冗余报文。
如此,针对缓存报文中节点201在预设窗口内已向节点202发送的部分报文,节点202能够根据第一计数值和第二计数值将其识别出来,从而在倒换路径后避免发送冗余报文,提高了报文传输的带宽利用率。
上文各个实施例中均以节点202接收节点201的故障探测序列,节点202再根据故障探测序列判断链路故障时指示节点201根据预设窗口基于倒换路径进行缓存报文的重发为示例,对本申请的数据传输方法进行说明,该方式适用于节点201至节点202的单向链路发生故障,节点202至节点201的单向链路正常的场景。本申请提供的数据传输方法中,为了解决双向链路故障的场景下的报文重发的丢包问题,除了某个节点自身确定链路故障并通知对端节点进行倒换路径的缓存报文重发,还可以是一个节点根据接收到的故障探测序列确定链路故障,自身据预设窗口基于倒换路径进行缓存报文的重发。
请参考图10,图10为本申请提供的另一种数据传输方法的流程示意图。该数据传输方法可以包括如下步骤1001-步骤1006。
步骤1001、节点201向节点202发送故障探测序列。
步骤1002、节点202接收故障探测序列。
步骤1003、节点202根据预设窗口内接收到的故障探测信息的个数,确定节点201和节点202之间的链路发生故障。
步骤1004、节点202向节点201发送故障通知信息。
作为一种可能的实现方式,上述步骤1001-步骤1004请参考图3所示的步骤301-步骤304,在此不再赘述。
由于节点201和节点202之间的双向链路发生故障,节点201无法接收到节点202发送的故障通知信息。
步骤1005、节点202基于倒换路径发送第二缓存报文和第四用户报文。
节点202基于倒换路径发送预设窗口内缓存的第二缓存报文,以及第四用户报文。其中,第二缓存报文是节点202在链路故障的预设窗口内对向节点201发送的第三用户报文进行缓存得到的。第四用户报文是第三用户报文的后续报文,即节点202在发送第三用户报文并缓存后,向节点201发送的第四用户报文,第四用户报文的报文序号大于第三用户报文的报文序号。
作为一种可能的实现方式,考虑到节点202检测到故障计数值等于1的时刻节点201可能仍有报文发送,节点202需要等待最大传输单元的报文发送完毕后基于倒换路径发送第二缓存报文和第四用户报文。
在本实施例中,上述节点202对第三用户报文的缓存方式请参考图7所示的步骤701-步骤705,在此不再赘述。上述倒换路径的具体倒换机制请参考图8所示的步骤,在此不再赘述。
在本实施例中,节点202在基于倒换路径发送第二缓存报文时,考虑到报文接收侧可能不具备报文去冗能力,节点201可以对第二缓存报文进行报文去冗,报文去冗的具体步骤请参考图9所示的步骤901-步骤908,在此不再赘述。
步骤1006、节点201基于倒换路径接收第二缓存报文和第四用户报文。
节点201在第二缓存报文和第四用户报文的目的节点为节点201时,基于倒换路径接收第二缓存报文和第四用户报文。节点201在第二缓存报文和第四用户报文的目的节点为其他节点时,基于倒换路径接收并转发第二缓存报文和第四用户报文。
基于上述数据传输方法,在节点201和节点202之间的双向链路在预设窗口内均发生故障后,节点202在检测到节点201至节点202的链路发生故障时,就倒换路径重新发送报文时,将预设窗口内缓存的节点202向节点201发送的第三用户报文再次发送,并基于倒换路径继续发送第三用户报文的后续报文。如此,由于节点201和节点202存在双向故障探测,节点201或节点202在检测到对端至本端的链路存在故障时,向对端发送故障通知信息,并在本端基于倒换路径向对端发送预设窗口内的缓存报文,即使节点201与节点202之间发生双向链路故障,也能够实现环网倒换的无丢包报文重传。
接下来结合附图对缓存报文的发送方式进行说明。
请参考图11,图11为本申请提供的一种缓存报文发送方式的示意图。
如图11所示,假设交换机2(例如节点202)至交换机1(例如节点203)的链路故障而交换机1至交换机2的链路正常,链路长度不超过200米(传输时长小于1微秒)。交换机1在两个最小检测窗口内识别链路故障并发送LFQ至交换机2。交换机2接收到LFQ并记录此时对应的缓存对应的地址及区域(如A1/A2/A3)。
A3:记录接收LFQ时刻对应写地址(A3_1)并停止发送报文,依次读取A1、A2、A3中到A3_1数据发送。
A2:记录接收LFQ时刻对应写地址(A2_1)并停止发送报文,依次读取A3,A1,A2中到A2_1数据发送。
A1:记录接收LFQ时刻对应写地址(A1_1)并停止发送报文,依次读取A2,A3,A1中到A1_1数据发送。
请参考图12,图12为本申请提供的另一种缓存报文发送方式的示意图。
如图12所示,假设交换机2(例如节点202)与交换机1(例如节点203)的链路双向故障。交换机1或交换机2在两个最小检测窗口内识别链路故障并发送LFQ到对端接口。由于链路双向故障,LFQ虽然双向发送但交换机1和交换机2均无法收到。交换机1和交换机2记录下识别到链路故障的时刻即故障计数值等于1的时刻以及此时对应的缓存对应的地址及区域。因为LFQ被发送到对端接口至对端接口接收到LFQ最大时间不会超过最小检测窗口,因此将最小检测窗口作为时长阈值,即链路发生故障时刻到触发倒换路径一共需要3个最小检测窗口。
交换机1和交换机2判断链路故障和倒换路径基于最小检测窗口为单位的“时钟”,因链路两侧异步时钟,所以交换机1和交换机2的“时钟”非同步,最大差异在一个最小检测窗口内,即因为链路两端的“时钟”异步而造成最多最小检测窗口差异,考虑到LFQ最大时延最小检测窗口,采用一个最小检测窗口作为超时补偿,综合链路发生故障时刻(处于某个最小检测窗口内)到触发折叠倒换一共需要4个最小检测窗口,即发送侧存储4个最小检测窗口的报文数据。
A1接收到故障信息,分别读出A4、A1、A2、A3进行折叠倒换(即倒换路径)发送。
A2接收到故障信息,分别读出A1、A2、A3、A4进行折叠倒换发送。
A3接收到故障信息,分别读出A2、A3、A4、A1进行折叠倒换发送。
A4接收到故障信息,分别读出A3、A4、A1、A2进行折叠倒换发送。
作为一种可能的实现方式,交换机1和交换机2通过发送同步序列实现链路两侧的“时钟”同步,从而使链路两侧的以太网接口直接由故障探测序列的接收侧触发折叠倒换,其时钟同步的具体步骤可以如下:
Master在t1时刻发送Sync(synchronization,同步)报文(如果配置为two-step方式,还会发送follow_up报文),并将t1时间戳携带在Sync报文(或follow_up报文)中。
Slave在t2时刻接收到Sync报文,在本地产生t2时间戳,并从报文中提取t1时间戳。
其中,Master和Slave可以是交换机1和交换机2。例如,Master为交换机1,Slave为交换机2。又如,Master为交换机2,Slave为交换机1。
Slave在t3时刻发送delay_req报文,并在本地产生t3时间戳。
Master在t4时刻接收到delay_req报文,并在本地产生t4时间戳,然后将t4时间戳携带在delay_resp报文中,回传给Slave。
Slave接收到delay_resp报文,从报文中提取t4时间戳。最后Slave节点得到了一组时间戳(t1、t2、t3、t4)。
假设Master到Slave的发送链路延迟是t-ms,Slave到Master的发送链路延迟是t-sm,Slave和Master之间的时间偏差为Offset,则:
t2-t1=t-ms+Offset
t4-t3=t-sm-Offset
(t2-t1)-(t4-t3)=(t-ms+Offset)-(t-sm-Offset)
因此,Offset=[(t2-t1)-(t4-t3)-(t-ms-t-sm)]/2
如果t-ms=t-sm,即Master和Slave之间的收发链路延迟对称,那么:
Offset=[(t2-t1)-(t4-t3)]/2
这样Slave就可以根据t1、t2、t3、t4四个时间戳计算出自己和Master之间的时间偏差Offset,再对本地时间进行偏差调整,就实现了Slave与Master的时间同步。
如图13所示,在采用“时钟”同步后,针对交换机1和交换机2链路双向故障,链路长度不超过200米。交换机1或交换机2在2个最小检测窗口内识别故障并发送LFQ到对端。由于链路双向故障,LFQ虽然双向发送,但是双方都收不到。交换机1或交换机2记录下识别故障时刻(即故障计数值等于1)以及对应的缓存位置,因为发送LFQ到对端至接收到LFQ最大时间不会超过最小检测窗口,所以将最小检测窗口作为时长阈值,即链路发生故障时刻(处于某个最小检测窗口内)到触发折叠倒换一共需要3个最小检测窗口即预设窗口。
链路两侧的交换机1和交换机2判断故障和倒换路径判断基于最小检测窗口为单位的“时钟”,链路两侧“时钟”同步,所以交换机1和交换机2的“时钟”已同步,可以采用接收侧故障判断结果对同一接口的发送侧缓存进行判断处理:
A1接收到故障信息,分别读出A3、A1、A2进行折叠倒换发送。
A2接收到故障信息,分别读出A1、A2、A3进行折叠倒换发送。
A3接收到故障信息,分别读出A2、A3、A1进行折叠倒换发送。
为了配合本申请实施例提供的上述数据传输方法,本申请实施例还提供了一种数据传输装置1400,该装置用于执行上述数据传输方法。如图14所示,该装置包括:
收发模块1410,用于向第二节点发送故障探测序列;故障探测序列包括间隔发送的多个故障探测信息。
收发模块1410,还用于接收第二节点发送的故障通知信息;故障通知信息是第二节点根据预设窗口内接收到的故障探测信息的个数确定的,故障通知信息用于指示第一节点和第二节点之间的链路发生故障。
处理模块1420,用于基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。
作为一种可能的实现方式,预设窗口是基于链路带宽和最大传输单元确定的。
作为一种可能的实现方式,预设窗口包括n个最小检测窗口,n为正整数,最小检测窗口为两个发送故障探测信息的最小间隔与基于链路带宽发送一个最大传输单元的串行化时长的和。
作为一种可能的实现方式,处理模块1420还用于:根据预设窗口缓存第一用户报文;第一用户报文包括在预设窗口内向第二节点发送的至少一个报文。
作为一种可能的实现方式,处理模块1420具体用于:在第一节点和第二节点之间的链路未发生故障的情况下,采用下个预设窗口内向第二节点发送的用户报文替换已缓存的第一缓存报文。
作为一种可能的实现方式,第一用户报文中的每相邻两个报文之间包括一个故障探测信息。处理模块1420还用于:缓存第一计数值;第一计数值是第一节点对向第二节点发送第一用户报文中的每个报文进行计数得到的。收发模块1410还用于:向第二节点发送第一计数值,以使第二节点根据第一计数值和第二计数值去除第一缓存报文中的冗余报文;第二计数值是第二节点对第一用户报文中的故障探测信息进行计数得到的,冗余报文是第一用户报文中报文计数值大于第二计数值的报文。
作为一种可能的实现方式,在第一节点和第二节点之间的链路处于空闲状态的情况下,故障探测序列的每两个故障探测信息的发送间隔为发送故障探测信息的预设间隔。在第一节点和第二节点之间的链路处于报文非线速状态的情况下,第二节点向第一节点发送的用户报文中每相邻两个报文中存在至少一个故障探测信息。在第一节点和第二节点之间的链路处于报文线速状态的情况下,第一节点向第二节点发送的用户报文中每相邻两个报文中存在一个故障探测信息。
作为一种可能的实现方式,故障探测序列承载于以太接口保留字段。
作为一种可能的实现方式,故障探测序列承载于报文。
应理解的是,上述图14提供的装置在实现其功能时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将设备的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的装置与方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
为了配合本申请实施例提供的上述数据传输方法,本申请实施例还提供了一种数据传输装置1500,该装置用于执行上述数据传输方法。如图15所示,该装置包括:
收发模块1510,用于接收第一节点发送的故障探测序列;故障探测序列包括间隔发送的多个故障探测信息。
处理模块1520,用于根据预设窗口内接收到的故障探测信息的个数,确定第一节点和所二节点之间的链路发生故障。
收发模块1510,还用于向第一节点发送故障通知信息,以使第一节点基于倒换路径发送第一缓存报文和第二用户报文,第一缓存报文是对预设窗口内第一节点向第二节点发送的第一用户报文进行缓存得到的,第二用户报文的报文序号大于第一用户报文的报文序号。
作为一种可能的实现方式,处理模块1520具体用于:在预设窗口内接收到的故障探测信息的个数小于预设阈值的情况下,确定第一节点和第二节点之间的链路发生故障。
作为一种可能的实现方式,处理模块1520还用于:缓存第二计数值;第二计数值是对故障探测序列的故障探测信息进行计数得到的。收发模块1510还用于:接收第一节点发送的第一计数值;第一计数值是第一节点对向第二节点发送第一用户报文中的每个报文进行计数得到的。处理模块1520还用于:根据第一计数值和第二计数值去除第一缓存报文中的冗余报文;冗余报文是第一用户报文中报文计数值大于第二计数值的报文。
作为一种可能的实现方式,收发模块1510还用于:基于倒换路径发送第二缓存报文和第四用户报文,第二缓存报文是对预设窗口内向第一节点发送的第三用户报文进行缓存得到的,第四用户报文的报文序号大于第三用户报文的报文序号。
应理解的是,上述图15提供的装置在实现其功能时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将设备的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的装置与方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
图16为本实施例提供的一种电子设备的结构示意图。如图16所示,电子设备1600包括处理器1610、总线1620、存储器1630、通信接口1640和内存单元1650(也可以称为主存(Main Memory)单元)。处理器1610、存储器1630、内存单元1650和通信接口1640通过总线1620相连。
应理解,在本实施例中,处理器1610可以是CPU,该处理器1610还可以是其他通用处理器、数字信号处理器(Digital Signal Processing,DSP)、ASIC、FPGA或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者是任何常规的处理器等。
处理器还可以是图形处理器(Graphics Processing Unit,GPU)、神经网络处理器(Neural Network Processing Unit,NPU)、微处理器、专用集成电路(Application Specific Integrated Circuit,ASIC)、或一个或多个用于控制本申请方案程序执行的集成电路。
通信接口1640用于实现电子设备1600与外部设备或器件的通信。在本实施例中,电子设备1600用于实现图2中任一节点的功能时,通信接口1640用于作为收发数据包的物理端口。
总线1620可以包括一通路,用于在上述组件(如处理器1610、内存单元1650和存储器1630)之间传送信息。总线1620除包括数据总线之外,还可以包括电源总线、控制总线和状态信号总线等。但是为了清楚说明起见,在图中将各种总线都标为总线1620。总线1620可以是快捷外围部件互连标准(Peripheral Component Interconnect Express,PCIe)总线,或扩展工业标准结构(Extended Industry Standard Architecture,EISA)总线、统一总线(Unified Bus,Ubus或UB)、计算机快速链接(Compute Express Link,CXL)、缓存一致互联协议(Cache Coherent Interconnect for Accelerators,CCIX)等。总线1620可以分为地址总线、数据总线、控制总线等。
作为一个示例,电子设备1600可以包括多个处理器。处理器可以是一个多核(multi-CPU)处理器。这里的处理器可以指一个或多个设备、电路、和/或用于处理数据(例如计算机程序指令)的计算单元。
值得说明的是,图16中仅以电子设备1600包括1个处理器1610和1个存储器1630为例,此处,处理器1610和存储器1630分别用于指示一类器件或设备,具体实施例中,可以根据业务需求确定每种类型的器件或设备的数量。
内存单元1650可以对应上述方法实施例中用于存储缓存报文的存储介质。内存单元1650可以是易失性存储器或非易失性存储器,或可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(read-only memory,ROM)、可编程只读存储器(programmable ROM,PROM)、可擦除可编程只读存储器(erasable PROM,EPROM)、电可擦除可编程只读存储器(electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(random access memory,RAM),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(static RAM,SRAM)、动态随机存取存储器(DRAM)、同步动态随机存取存储器(synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(double data date SDRAM,DDR SDRAM)、增强型同步动态随机存取存储器(enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(synchlink DRAM,SLDRAM)和直接内存总线随机存取存储器(direct rambus RAM,DR RAM)。
存储器1630可以对应上述方法实施例中用于存储计算机指令等信息的存储介质,例如,磁盘,如机械硬盘或固态硬盘。
上述电子设备1600可以是一个通用设备或者是一个专用设备。例如,电子设备1600可以是边缘设备(例如,携带具有处理能力芯片的盒子)等。可选地,电子设备1600也可以是网络设备、服务器或其他具有计算能力的设备。
应理解,根据本实施例的电子设备1600可对应于本实施例中的数据传输装置1400或数据传输装置1500,并可以对应于执行根据图3中方法中的相应主体,并且数据传输装置1400或数据传输装置1500中的各个模块的上述和其它操作和/或功能分别为了实现图3中方法的相应流程,为了简洁,在此不再赘述。
本实施例中的方法步骤可以通过硬件的方式来实现,也可以由处理器执行软件指令的方式来实现。软件指令可以由相应的软件模块组成,软件模块可以被存放于随机存取存储器、闪存、只读存储器、可编程只读存储器、可擦除可编程只读存储器、电可擦除可编程只读存储器、寄存器、硬盘、移动硬盘、CD-ROM或者本领域熟知的任何其它形式的存储介质中。一种示例性的存储介质耦合至处理器,从而使处理器能够从该存储介质读取信息,且可向该存储介质写入信息。当然,存储介质也可以是处理器的组成部分。处理器和存储介质可以位于ASIC中。另外,该ASIC可以位于电子设备中。当然,处理器和存储介质也可以作为分立组件存在于电子设备中。
在上述实施例中,可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机程序或指令。在计算机上加载和执行所述计算机程序或指令时,全部或部分地执行本申请实施例所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络、网络设备、用户设备或者其它可编程装置。所述计算机程序或指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机程序或指令可以从一个网站站点、计算机、服务器或数据中心通过有线或无线方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存取的任何可用介质或者是集成一个或多个可用介质的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质,例如,软盘、硬盘、磁带;也可以是光介质,例如,数字视频光盘(digital video disc,DVD);还可以是半导体介质,例如,固态硬盘(solid state drive,SSD)。以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到各种等效的修改或替换,这些修改或替换都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以权利要求的保护范围为准。
Claims (17)
- 一种数据传输方法,其特征在于,应用于环形网络中的第一节点,所述环形网络包括基于环形拓扑结构连接的多个节点,所述多个节点包括所述第一节点和第二节点,所述方法包括:向所述第二节点发送故障探测序列;所述故障探测序列包括间隔发送的多个故障探测信息;接收所述第二节点发送的故障通知信息;所述故障通知信息是所述第二节点根据预设窗口内接收到的故障探测信息的个数确定的,所述故障通知信息用于指示所述第一节点和所述第二节点之间的链路发生故障;基于倒换路径发送第一缓存报文和第二用户报文,所述第一缓存报文是对所述预设窗口内向所述第二节点发送的第一用户报文进行缓存得到的,所述第二用户报文的报文序号大于所述第一用户报文的报文序号。
- 根据权利要求1所述的方法,其特征在于,所述预设窗口是基于链路带宽和最大传输单元确定的。
- 根据权利要求2所述的方法,其特征在于,所述预设窗口包括n个最小检测窗口,n为正整数,所述最小检测窗口为两个发送故障探测信息的最小间隔与基于所述链路带宽发送一个最大传输单元的串行化时长的和。
- 根据权利要求1-3中任一项所述的方法,其特征在于,所述方法还包括:根据所述预设窗口缓存所述第一用户报文;所述第一用户报文包括在所述预设窗口内向所述第二节点发送的至少一个报文。
- 根据权利要求4所述的方法,其特征在于,所述方法还包括:在所述第一节点和所述第二节点之间的链路未发生故障的情况下,采用下个所述预设窗口内向所述第二节点发送的用户报文替换已缓存的所述第一缓存报文。
- 根据权利要求1-5中任一项所述的方法,其特征在于,所述第一用户报文中的每相邻两个报文之间包括一个故障探测信息,在所述基于倒换路径发送第一缓存报文和第二用户报文之前,还包括:缓存第一计数值;所述第一计数值是所述第一节点对向所述第二节点发送所述第一用户报文中的每个报文进行计数得到的;向所述第二节点发送所述第一计数值,以使所述第二节点根据所述第一计数值和第二计数值去除所述第一缓存报文中的冗余报文;所述第二计数值是所述第二节点对所述第一用户报文中的故障探测信息进行计数得到的,所述冗余报文是所述第一用户报文中报文计数值大于所述第二计数值的报文。
- 根据权利要求1-6中任一项所述的方法,其特征在于,在所述第一节点和所述第二节点之间的链路处于空闲状态的情况下,所述故障探测序列的每两个故障探测信息的发送间隔为发送故障探测信息的预设间隔;在所述第一节点和所述第二节点之间的链路处于报文非线速状态的情况下,所述第一节点向所述第二节点发送的用户报文中每相邻两个报文中存在至少一个故障探测信息;在所述第一节点和所述第二节点之间的链路处于报文线速状态的情况下,所述第一节点向所述第二节点发送的用户报文中每相邻两个报文中存在一个故障探测信息。
- 根据权利要求1-7中任一项所述的方法,其特征在于,所述故障探测序列承载于以太接口保留字段,或所述故障探测序列承载于报文。
- 一种数据传输方法,其特征在于,应用于环形网络中的第二节点,所述环形网络包括基于环形拓扑结构连接的多个节点,所述多个节点包括所述第一节点和第二节点,所述方法包括:接收所述第一节点发送的故障探测序列;所述故障探测序列包括间隔发送的多个故障探测信息;根据预设窗口内接收到的故障探测信息的个数,确定所述第一节点和所述第二节点之间的链路发生故障;向所述第一节点发送故障通知信息,以使所述第一节点基于倒换路径发送第一缓存报文和第二用户报文,所述第一缓存报文是对所述预设窗口内所述第一节点向所述第二节点发送的第一用户报文进行缓存得到的,所述第二用户报文的报文序号大于所述第一用户报文的报文序号。
- 根据权利要求9所述的方法,其特征在于,所述根据预设窗口内接收到的故障探测信息的个数,确定所述第一节点和所述第二节点之间的链路发生故障,包括:在所述预设窗口内接收到的故障探测信息的个数小于预设阈值的情况下,确定所述第一节点和所述第二节点之间的链路发生故障。
- 根据权利要求9或10所述的方法,其特征在于,所述方法还包括:缓存第二计数值;所述第二计数值是对所述故障探测序列的故障探测信息进行计数得到的;接收所述第一节点发送的第一计数值;所述第一计数值是所述第一节点对向所述第二节点发送所述第一用户报文中的每个报文进行计数得到的;根据所述第一计数值和所述第二计数值去除所述第一缓存报文中的冗余报文;所述冗余报文是所述第一用户报文中报文计数值大于所述第二计数值的报文。
- 根据权利要求9-11中任一项所述的方法,其特征在于,所述方法还包括:基于倒换路径发送第二缓存报文和第四用户报文,所述第二缓存报文是对所述预设窗口内向所述第一节点发送的第三用户报文进行缓存得到的,所述第四用户报文的报文序号大于所述第三用户报文的报文序号。
- 一种数据传输装置,其特征在于,包括:收发模块,用于向所述第二节点发送故障探测序列;所述故障探测序列包括间隔发送的多个故障探测信息;所述收发模块,还用于接收所述第二节点发送的故障通知信息;所述故障通知信息是所述第二节点根据预设窗口内接收到的故障探测信息的个数确定的,所述故障通知信息用于指示所述第一节点和所述第二节点之间的链路发生故障;处理模块,用于基于倒换路径发送第一缓存报文和第二用户报文,所述第一缓存报文是对所述预设窗口内向所述第二节点发送的第一用户报文进行缓存得到的,所述第二用户报文的报文序号大于所述第一用户报文的报文序号。
- 一种数据传输装置,其特征在于,包括:收发模块,用于接收第一节点发送的故障探测序列;所述故障探测序列包括间隔发送的多个故障探测信息;处理模块,用于根据预设窗口内接收到的故障探测信息的个数,确定所述第一节点和所二节点之间的链路发生故障;所述收发模块,还用于向所述第一节点发送故障通知信息,以使所述第一节点基于倒换路径发送第一缓存报文和第二用户报文,所述第一缓存报文是对所述预设窗口内所述第一节点向所述第二节点发送的第一用户报文进行缓存得到的,所述第二用户报文的报文序号大于所述第一用户报文的报文序号。
- 一种网络设备,其特征在于,包括处理器和存储器,所述处理器用于执行所述存储器中存储的指令,使得所述网络设备执行如权利要求1-12中任一项所述的方法。
- 一种包含指令的计算机程序产品,其特征在于,当所述指令被网络设备运行时,使得所述网络设备执行如权利要求1-12中任一项所述的方法。
- 一种计算机可读存储介质,其特征在于,包括计算机程序指令,当所述计算机程序指令由网络设备执行时,所述网络设备执行如权利要求1-12中任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410749823.8 | 2024-06-11 | ||
| CN202410749823.8A CN121125394A (zh) | 2024-06-11 | 2024-06-11 | 数据传输方法及装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025256478A1 true WO2025256478A1 (zh) | 2025-12-18 |
Family
ID=97942216
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/099684 Pending WO2025256478A1 (zh) | 2024-06-11 | 2025-06-06 | 数据传输方法及装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121125394A (zh) |
| WO (1) | WO2025256478A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN121418451A (zh) * | 2025-12-24 | 2026-01-27 | 中交路桥科技有限公司 | 一种桥隧结构健康监测系统及环网故障自愈通信方法 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6101166A (en) * | 1998-05-01 | 2000-08-08 | Emulex Corporation | Automatic loop segment failure isolation |
| CN101340346A (zh) * | 2008-08-11 | 2009-01-07 | 中兴通讯股份有限公司 | 一种以太环网系统中环控制的方法及装置 |
| CN102238069A (zh) * | 2010-04-29 | 2011-11-09 | 杭州华三通信技术有限公司 | 一种链路切换过程中的数据处理方法和装置 |
| CN105141493A (zh) * | 2015-07-27 | 2015-12-09 | 浙江宇视科技有限公司 | 环网故障时的业务帧处理方法及系统 |
| CN107733980A (zh) * | 2017-09-11 | 2018-02-23 | 深圳市盛路物联通讯技术有限公司 | 基于环形网络的物联网数据传输系统及第一中继器 |
| CN111130893A (zh) * | 2019-12-27 | 2020-05-08 | 中国联合网络通信集团有限公司 | 一种报文传输方法和装置 |
-
2024
- 2024-06-11 CN CN202410749823.8A patent/CN121125394A/zh active Pending
-
2025
- 2025-06-06 WO PCT/CN2025/099684 patent/WO2025256478A1/zh active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6101166A (en) * | 1998-05-01 | 2000-08-08 | Emulex Corporation | Automatic loop segment failure isolation |
| CN101340346A (zh) * | 2008-08-11 | 2009-01-07 | 中兴通讯股份有限公司 | 一种以太环网系统中环控制的方法及装置 |
| CN102238069A (zh) * | 2010-04-29 | 2011-11-09 | 杭州华三通信技术有限公司 | 一种链路切换过程中的数据处理方法和装置 |
| CN105141493A (zh) * | 2015-07-27 | 2015-12-09 | 浙江宇视科技有限公司 | 环网故障时的业务帧处理方法及系统 |
| CN107733980A (zh) * | 2017-09-11 | 2018-02-23 | 深圳市盛路物联通讯技术有限公司 | 基于环形网络的物联网数据传输系统及第一中继器 |
| CN111130893A (zh) * | 2019-12-27 | 2020-05-08 | 中国联合网络通信集团有限公司 | 一种报文传输方法和装置 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN121418451A (zh) * | 2025-12-24 | 2026-01-27 | 中交路桥科技有限公司 | 一种桥隧结构健康监测系统及环网故障自愈通信方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121125394A (zh) | 2025-12-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN100534048C (zh) | 分布式以太网系统及基于该系统的故障检测方法 | |
| CN100550715C (zh) | 在以太网上支持同步数字系列/同步光纤网自动保护交换的方法 | |
| CN104022906B (zh) | 用于弹性无线分组通信的系统和方法 | |
| JP5801175B2 (ja) | パケット通信装置および方法 | |
| US20230018911A1 (en) | Troubleshooting method, device, and readable storage medium | |
| CN105871674B (zh) | 环保护链路故障保护方法、设备及系统 | |
| CN101094157A (zh) | 利用链路聚合实现网络互连的方法 | |
| US8184650B2 (en) | Filtering of redundant frames in a network node | |
| WO2010060250A1 (zh) | 以太环网的地址刷新方法及装置 | |
| CN108206759A (zh) | 一种转发报文的方法、设备及系统 | |
| CN103873336A (zh) | 分布式弹性网络互连的业务承载方法及装置 | |
| WO2025256478A1 (zh) | 数据传输方法及装置 | |
| WO2022063207A1 (zh) | 处理时间同步故障的方法、装置及系统 | |
| CN108243114A (zh) | 一种转发报文的方法、设备及系统 | |
| CN111885554A (zh) | 基于双无线蓝牙通信的链路切换方法及相关设备 | |
| CN101436975A (zh) | 一种在环网中实现快速收敛的方法、装置及系统 | |
| CN101217445B (zh) | 防止环路产生的方法和以太环网系统 | |
| WO2012062097A1 (zh) | 多环以太网及其保护方法 | |
| WO2012000374A1 (zh) | 接口业务集中处理方法和系统 | |
| CN100454880C (zh) | 一种实现环网保护的方法及系统 | |
| CN104202184A (zh) | 一种快速回切业务的方法和装置 | |
| JP5672836B2 (ja) | 通信装置、通信方法、および通信プログラム | |
| CN101237319B (zh) | 以太环网中的主节点、时间同步方法和以太环网系统 | |
| JP5176623B2 (ja) | イーサネットの伝送方法、伝送装置およびシステム | |
| CN101291258B (zh) | 用于通讯平台多框互连时的以太网环路处理方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25821190 Country of ref document: EP Kind code of ref document: A1 |