WO2010051737A1 - 一种负载均衡分组交换结构及其构造方法 - Google Patents
一种负载均衡分组交换结构及其构造方法 Download PDFInfo
- Publication number
- WO2010051737A1 WO2010051737A1 PCT/CN2009/074739 CN2009074739W WO2010051737A1 WO 2010051737 A1 WO2010051737 A1 WO 2010051737A1 CN 2009074739 W CN2009074739 W CN 2009074739W WO 2010051737 A1 WO2010051737 A1 WO 2010051737A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- self
- routing
- switching module
- data
- level
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L49/00—Packet switching elements
- H04L49/10—Packet switching elements characterised by the switching fabric construction
- H04L49/101—Packet switching elements characterised by the switching fabric construction using crossbar or matrix
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L49/00—Packet switching elements
- H04L49/55—Prevention, detection or correction of errors
- H04L49/552—Prevention, detection or correction of errors by ensuring the integrity of packets received through redundant connections
Definitions
- the present invention relates to the field of communications technologies, and in particular, to a load balancing packet switching structure and a method for constructing the same.
- the so-called switch fabric is a network device that implements the path selection of data units and sends the data units to the next destination address.
- a load balancing switch structure is generally used to balance the arriving traffic. This structure allows the distribution of traffic to be in an equalized state within the switching fabric, i.e., the ports of the switching fabric are the same as the utilization of the respective internal circuitry. This maximizes the throughput of the switch fabric and reduces congestion within the switch fabric.
- the load-balancing Birkhoff-von Neumann switch fabric consists of two levels of cross-load-loading (Load-balancing) switching, four Birkhoff-von Neumann switch exchanges, and one in two VOQ (virtual output queue) between levels.
- the first level of switching completes the load balancing
- the second level of switching completes the data packet switching. Since the two-stage connection of the switch fabric is deterministic and periodic, no scheduling between the input and output ports is required.
- the connection mode selection must be such that each successive W time gap is required, and each input is connected to each output exactly once. It can be seen that the load balancing switch structure described above solves the problem of data blocking of the switch fabric.
- TCP Transmission Control Protocol
- the technical problem to be solved by the present invention is to provide a load balancing packet switching structure and a construction method thereof, so that the load balancing Birkhoff-von Neumann switching structure can solve the problem of packet out-of-order and improve end-to-end throughput.
- the present invention provides a method for constructing a load balancing packet switching structure, the method comprising: dividing a load balancing packet switching structure based on a self-routing hub into a first-level switching module having a load balancing function and having a self-routing for completing a packet The second level switching module of the forwarding function;
- the packets stored in the virtual output group queue are sent to the first-level exchange, combined into data blocks of preset length, and then divided into data pieces of equal length, and a self-routing for self-routing is added to the data slice.
- the data piece having the self-routing tag is transmitted to the reordering buffer of the destination output port after being transmitted by the first-level switching module and the second-level switching module, and the data pieces are reassembled into the virtual output group according to the self-routing tag carried by the data piece.
- the data blocks into which the group queues are combined.
- the middle line group is set between the first level switching module and the second level switching module.
- the self-routing hub-based load balancing packet switching structure adopts distributed self-routing.
- the present invention also provides a load balancing packet switching structure, including a first-level switching module based on a self-routing hub for performing load balancing functions, and a function for completing packet data self-routing and forwarding functions.
- a secondary switching module wherein a virtual output group queue is set in front of the input end of the first level switching module, and a reordering buffer is set after the second level switching module output end, and the virtual output group queue is used
- the reordering buffer is configured to arrange packet data blocks belonging to the same input group according to self-routing address information for subsequent processing, the first level switching and
- the second level of exchange is a middle line group connection.
- the virtual output group queue is set in front of the input end of the first-level switching module, and the switching module output is in the second-level switching module.
- the reordering buffer is set, and before the packets stored in the virtual output group queue are sent to the first level of switching, the packets are combined into data blocks of preset length, and then divided into pieces of data of equal length, and at the same time A self-routing label for implementing self-routing is added to the data slice.
- the load balancing packet switching structure provided by the present invention cancels the intermediate stage of the virtual output queue VOQ between the first level switching and the second level switching described in the background art, so that the present invention
- the load balancing packet switching structure described does not have a queuing delay problem, thereby avoiding packet out of order. Therefore, the present invention solves load balancing
- the Birkhoff-von Neumann switch fabric can group out-of-order problems and improve end-to-end throughput.
- FIG. 1 is a schematic diagram of a load balancing Birkhoff-von Neumann switching structure in the prior art
- FIG. 2a is a schematic flow chart of a general construction method of a multipath self-routing switching structure according to an embodiment of the present invention
- FIG. 3 is a schematic diagram of a load balancing packet switching structure model according to an embodiment of the present invention; An example of an algorithm in the embodiment of the present invention;
- FIG. 5 is a schematic diagram of an algorithm 2 according to an embodiment of the present invention.
- the embodiment of the present invention adopts a packet switching structure based on a self-routing hub, and the switching structure is mainly constructed by using a hub and a line group technology on the basis of a routable multi-level interconnection network.
- a routable multi-level interconnection network constitutes a packet switching structure based on a self-routing hub.
- First construct a sly routable network (usually choose the partition network with the best layout complexity). Then replace the 2x2 routing units in the network with 2G-to-G self-routing group hubs, and replace the connections between the various stages in the network with G parallel lines, thus establishing one with M outputs (input ) Group, each group contains a network of G output (input) ports.
- the 2G-to-G hub has two input ports and two sets of output ports.
- the two addresses in the two output groups are called 0-output group and the address is large, called 1-output group.
- two input groups Called 0-input group and 1-input group. Ports in the same output group are indistinguishable because the effect of switching to any port in the same group is equivalent for a single signal.
- a 2G-to-G hub is equivalent to a 2x2 basic routing unit because the addresses of the G ports in each input (output) group are the same.
- a 2G-to-G hub refers to a 2Gx2G sorting switch module that exchanges the G addresses with the largest address among the 2G input signals to the G output ports with the largest output address, and routes the remaining G signals to G output ports with the smallest output address. As shown in FIG.
- a re-sequencing buffer (RB: Re-sequencing Buffer) can be set after the second-level switching module to construct a load-balanced packet switching structure.
- the first-level switching module plays the role of load balancing. It is responsible for homogenizing the input network traffic and forwarding it to the input of the second-level switching module. Then, through the self-routing label carried by the data, the second-level switching module can use the self-routing feature to send the data to the final destination port.
- the (output) ports form an input (output) group, thus forming M groups at the input and output of the switch fabric.
- the G root internal links that are commonly connected to different hubs within the switch fabric also form a line group.
- IGi ( OGi ) represent a specific input (output) group
- VOGQ Virtual Output Group Queue VOGQ is logically equivalent to the virtual output queue VOQ, except that each VOGQ queue is responsible for storing data from G input ports.
- the load-balanced packet-switched architecture is scheduled in units of time slots.
- the processing of packets per time slot can be roughly divided into the following consecutive phases, and should be run as pipelined as possible to speed up the processing:
- Encapsulation phase Packets stored in VOGQ are first combined into data blocks of a certain maximum length.
- Algorithm 1 Algorithm 1
- the data blocks stored in VOGQ (i,j) are first uniformly sliced into M equal-length pieces of data, which is the load marked in Figure 3. Thereafter, the self-routing labels MG, IG, and OG are inserted before the corresponding data slice for use in routing through the two-level structure.
- the MGs are inserted in order from the smallest to the largest in front of the M pieces of data corresponding to the specific VOGQ.
- M is the total number of groups */
- DataBlock[M][M]; /*DataBlock(i,j) is the data block stored in VOGQ(i) */
- SlicePayload[M][M][M]; /*SlicePayload(i,j,k) represents the data slice load after DataBlock(i) is split */
- DataSlice[M][M][M]; /*DataSlice(i,j,k) represents the tagged data slice */
- IG[M] ⁇ 0, 1 , 2 , 3 , ⁇ , M-1 ⁇ ;
- MG[M] ⁇ 0, 1 , 2 , 3 , ⁇ , M-1 ⁇ ;
- the pieces of data with the same IG tag are reassembled and the MG tag will be used to restore the blocks in order.
- the data block is then re-cut into packets by packet size.
- the processed packet can be sent off the output port.
- M is the total number of groups */
- DataBlock[M][M]; /*DataBlock(i,j) is the data block stored in VOGQ(i) */
- SlicePayload[M][M][M]; /*SlicePayload(i,j,k) represents the data slice load after DataBlock(i) is split */
- MG[M] ⁇ 0, 1 , 2 , 3 , ⁇ , M-1 ⁇ ;
- Data block recovery is completed, re-segmented by packet size can leave the switch fabric output * / Figure 5, according to Algorithm 2, on each output group OG, we first collect data pieces of the same input group IG 5 Then, according to the MG tag of the data piece, the pieces of data are combined in order from small to large. After removing the extra self-routing labels, we restored the data blocks. After cutting these data blocks by packet size, they can be sent out to the output port.
- the two input groups of the switch fabric are connected to a 2G-to-G self-channel
- the size of the 2G-to-G self-routing hub is 2Gx2G.
- the M data slices undergo the same transmission delay in the switch fabric, and the same time slot arrives at the reorder buffer of the output port, thereby reassembling into uncut data blocks according to the self-routing labels. And then sent to the line card on the output. Since the G input data packets are respectively stored in the corresponding queue of the VOGQ according to the output port address, and the output ports have M, then the pre-position VOGQ of one input group actually has M virtual output queues. Thus, each time the M time slots are passed, each virtual output queue of the VOGQ can send a data packet once.
- each virtual output queue of VOGQ can send a data packet once.
- the size of the load balancing switch fabric is not limited.
- the switch fabric is a fully distributed self-routing, which also provides a technical and physical basis for the large-scale implementation of the load balancing switch fabric.
- the reordering buffer is set, and before the packets stored in the virtual output group queue are sent to the first level of switching, the packets are combined into a preset length of data blocks, and then divided into equal length data pieces. Adding self-routing address information for self-routing on the data slice, and reassembling the data pieces into the virtual output group queue according to the self-routing address information carried by the data piece in the reordering buffer. data block.
- the load balancing packet switching structure provided by the embodiment of the present invention cancels the intermediate level of the virtual output queue VOQ between the first level switching and the second level switching described in the background art, so that The load balancing packet switching structure described in the embodiment of the invention does not have a queuing delay problem, thereby avoiding packet out of order. Therefore, the embodiment of the invention solves the problem that the load balancing Birkhoff-von Neumann switch fabric can be out of order and improves the end-to-end throughput.
- a virtual output group queue is set in front of the input end of the first-level switching module, and after the output terminal of the second-level switching module Setting a reordering buffer, and before the packets stored in the virtual output group queue are sent to the first level exchange, grouping the packets into data blocks of preset length, and then dividing into data pieces of equal length, and simultaneously on the data slice Adding a self-routing label for implementing self-routing, in the reordering cache, re-combining the data pieces into data blocks combined by the virtual output group queue according to the self-routing label carried by the data piece, thereby solving load balancing
- the Birkhoff-von Neumann switch fabric can group out-of-order problems and improve end-to-end throughput.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Computer Security & Cryptography (AREA)
- Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- Data Exchanges In Wide-Area Networks (AREA)
Abstract
本发明提供了一种负载均衡分组交换结构及其构造方法, 其中, 所述构造方法包括:将基于自路由集线器的负载均衡分组交换结构分成第一级交换模块和第二级交换模块; 在第一级交换模块输入端前设置虚拟输出群组队列, 在第二级交换模块输出端后设置重排序缓存,在虚拟输出群组队列存储的分组发送到第一级交换前, 将分组组合成预置长度的数据块, 然后再分割成等长的数据片, 同时添加自路由标签, 传输后到达重排序缓存后, 再把数据片重新组合成所述数据块。 本发明所述种负载均衡分组交换结构及其构造方法, 解决了负载均衡 Birkhoff-von Neumann交换结构的分组乱序的问题,提高了端到端的吞吐量。
Description
一种负载均衡分组交换结构及其构造方法 技术领域
本发明涉及通信技术领域,尤其涉及一种负载均衡分组交换结构及其构造 方法。
背景技术
电信应用中, 所谓的交换结构是一种网络设备, 该设备实现数据单元的路 径选择, 并将数据单元发送到下一个目标地址。
由于交换结构的内部容量是有限的, 因此, 当到达交换结构的流量不均衡 时,会出现有些端口或者内部线路已经处于饱和状态, 而有些端口或者内部线 路却仍处于空闲状态的情况。 为了避免上述不均衡状态的出现, 一般采用负载 均衡交换结构来均衡到达的流量。该结构使流量的分布在交换结构内部处于均 衡状态, 即交换结构的端口和各个内部线路的利用率相同。这样便可以最大限 度的提高交换结构的吞吐量, 降低交换结构内部的阻塞。
负载均衡 Birkhoff-von Neumann交换结构恰好能够解决交换结构内部阻塞 的问题。
如图 1所示,负载均衡 Birkhoff-von Neumann交换结构包含两级纵横式负载 均衡 ( Load-balancing ) 交换、 4白克霍夫—冯诺伊曼 ( Birkhoff-von Neumann switch ) 交换和一个处于两级之间的虚拟输出队列 VOQ(virtual output queue)。 第一级交换完成负载均衡, 第二级交换完成数据分组交换。 由于交换结构两级 的连接方式是确定和周期性的, 所以不需要任何输入输出端口间的调度。连接 模式的选择必须要求在每一个连续的 W个时间间隙,每个输入端都要和每个输 出端恰好连接一次。可见, 上述的负载均衡交换结构解决了交换结构数据阻塞 的问题。
但是, 在每个输入端口, 由于业务流量是不相同和不均衡的, 于是不同流 所含有的数据分组的数量也是不同的, 这就使得中间级虚拟输出队列 VOQ的
长度不同。 又由于队列服务是独立于各自长度的, 因此, 上述负载均衡交换机 构又出现了队列排队延迟, 数据分组乱序的问题。 而分组乱序传输可能导致
TCP ( Transmission Control Protocol传输控制协议快速恢复, 使得 TCP滑动窗 口减半, 端到端的吞吐量也会减半。
发明内容
为此, 本发明所要解决的技术问题是: 提供一种负载均衡分组交换结构及 其构造方法,使得负载均衡 Birkhoff-von Neumann交换结构能够解决分组乱序 的问题, 提高端到端的吞吐量。
于是,本发明提供了一种负载均衡分组交换结构的构造方法,该方法包括: 将基于自路由集线器的负载均衡分组交换结构分成具有完成负载均衡功 能的第一级交换模块和具有完成分组自路由转发功能的第二级交换模块;
在所述第一级交换模块输入端前设置虚拟输出群组队列,在所述第二级交 换模块输出端后设置重排序緩存, 所述虚拟输出群组队列, 用于存储带有自路 由地址信息的分组数据块, 所述重排序緩存, 用于将属于同一个输入群组的分 组数据块按自路由地址信息排列, 以便后续处理;
所述虚拟输出群组队列存储的分组发送到第一级交换前,组合成预置长度 的数据块, 然后被分割成等长的数据片, 同时在数据片上添加用于实现自路由 的自路由标签;
拥有自路由标签的数据片经第一级交换模块和第二级交换模块传输后到 达目的输出端口的重排序緩存,根据数据片携带的自路由标签,把数据片重新 组合成所述虚拟输出群组队列组合成的数据块。
其中, 在所述第一级交换模块和第二级交换模块之间设置中间线群组。 其中, 所述基于自路由集线器的负载均衡分组交换结构采用分布式自路 由。
本发明还提供一种负载均衡分组交换结构,包括基于自路由集线器用于完 成负载均衡功能的第一级交换模块和用于完成分组数据自路由转发功能的第
二级交换模块,其中,在所述第一级交换模块输入端前设置虚拟输出群组队列, 在所述第二级交换模块输出端后设置重排序緩存, 所述虚拟输出群组队列, 用 于存储带有自路由地址信息的分组数据块, 所述重排序緩存, 用于将属于同一 个输入群组的分组数据块按自路由地址信息排列, 以便后续处理, 所述第一级 交换和第二级交换之间为中间线群组连接。 可见,通过将基于自路由集线器的负载均衡分组交换结构分成第一级交换 模块和第二级交换模块, 在第一级交换模块输入端前设置虚拟输出群组队列, 在第二级交换模块输出端后设置重排序緩存,并在所述虚拟输出群组队列存储 的分组发送到第一级交换前,将分组组合成预置长度的数据块, 然后被分割成 等长的数据片, 同时在数据片上添加用于实现自路由的自路由标签,在重排序 緩存中,根据数据片携带的自路由标签,再把数据片重新组合成所述虚拟输出 群组队列组合成的数据块。 通过上述结构的变化, 可见, 本发明提供的负载均 衡分组交换结构取消了背景技术中所述的第一级交换和第二级交换之间的虚 拟输出队列 VOQ这一中间级, 使得本发明所述的负载均衡分组交换结构不存 在排队延迟问题, 进而避免了分组乱序。 所以, 本发明解决了负载均衡
Birkhoff-von Neumann交换结构能够分组乱序的问题, 提高端到端的吞吐量。 附图说明
图 1为现有技术中负载均衡 Birkhoff-von Neumann交换结构示意图; 图 2a为本发明实施例所述多路径自路由交换结构一般构造方法流程示意 图;
图 2b为图 2a所述多路径 N=128 G=8 M=16的多路径自路由交换结构构造 方法流程示意图; 图 3为本发明实施例所述负载均衡分组交换结构模型示意图; 图 4为本发明实施例所述算法一示例图; 图 5为本发明实施例所述算法二示例图。
具体实施方式
下面, 结合附图对本发明进行详细描述。 本发明实施例采用基于自路由集线器的分组交换结构,而该交换结构主要 是利用集线器和线组技术, 在可路由多级互连网络的基础上来构造。
如图 2a所示, 一个 ΜχΜ可路由的多级互连网络构成一个基于自路由集线 器的分组交换结构, 一般的, 设 N=2n , N = MxG, M=2m,G=2g, 先构造一个 ΜχΜ的可路由网络(通常选择版图复杂性最优的分治网络)。 然后将网络中 各级 2x2路由单元替换为 2G-to-G自路由群组集线器,把网络中各级间的连线替 换成 G条平行的线束, 这样就建立了一个拥有 M个输出 (输入)群组, 每群组 包含 G个输出 (输入)端口的 ΝχΝ网络。 2G-to-G集线器具有两组输入端口和 两组输出端口的, 两个输出组中地址小的称为 0-输出组和地址大的称为 1 -输出 组; 同理, 两个输入组称为 0-输入组和 1-输入组。 同一个输出组中的端口是不 用区分的, 这是因为对于一个信号而言, 交换到同一组中任何一个端口的效果 都是等价的。
如图 2b所示, 当线束大小 G为 8时,将线组和 16-to-8集线器应用于图 2a所示 的 16x 16网络, 就得到了一个 128x 128网络。
逻辑上, 2G-to-G集线器等同于 2x2基本路由单元, 因为它每个输入 (输出) 组中的 G个端口的地址是相同的。 一个 2G-to-G集线器是指一个 2Gx2G的排序 交换模块,它将 2G个输入信号中地址最大的 G个信号交换到具有最大输出地址 的 G个输出端口, 并将其余的 G个信号路由到具有最小输出地址的 G个输出端 口。 如图 3所示, 基于上述构造的自路由集线器的分组交换结构, 通过叠加 2 个基于自路由集线器的分组交换结构以及在第一级交换模块前添置虚拟输出 群组队列( VOGQ: Virtual Output Group Queuing ), 在第二级交换模块后面设置 重排序緩存 ( RB:Re-sequencing Buffer )便可构造负载均衡的分组交换结构。 实际上, 第一级交换模块起到了负载均衡的作用, 它负责将输入的网络流 量均匀化后转送到第二级交换模块输入端。之后,通过数据携带的自路由标签, 第二级交换模块就可利用自路由特性将数据送到最终目的端口。 每 G个输入
(输出)端口组成一个输入(输出)群组, 这样在交换结构的输入输出端各形 成了 M个群组。 交换结构内部共同连接不同集线器的 G根内部链路也相应组成 一个线群组。 为了便于表达, 设 IGi ( OGi )代表一个特定的输入(输出 )群组, MGi代表前后两级交换模块之间的线群组(i=0,l, ...M-1 )。
虚拟输出群队列 VOGQ逻辑上等同于虚拟输出队列 VOQ,不同的是每个 VOGQ队列负责存储来自 G个输入端口的数据, 实质上 VOGQ由 M个虚拟输出 队列组成。 假定 VOGQ ( i,j )代表存储来自输入群组 IGi,目的地为输出群组 OGj 数据的队列, (i, j=0,l, ...M-1 ), 同时设当前 VOGQ ( i,j ) 队列长度为 L ,即有 L 个分组在緩存中等待传送。
一般而言, 负载均衡的分组交换结构按时隙为单位进行调度,每时隙对分 组的处理可大致分为以下几个连续阶段,并且应尽可能以流水线方式运行来加 快处理速度:
1)到达阶段: 新的分组在此阶段到达输入端 IGs. 其中到达输入群组 IGi去往输 出群组 OGj的分组被存储于 VOGQ ( i,j ) 队列中
2)封装阶段: 存储于 VOGQ中的分组首先被组合成最大长度一定的数据块
( data block )。 然后根据算法 1 , 这些数据块将在分割和打标签后成为小的 数据片 (data slice ), 并等待进一步传输。 参见图 3中的数据片格式 )。
3) 均衡阶段: 通过使用 MG自路由标签, 所有输入群组 IG同时将封装后的数 据片送到两级间的线群组。 当数据片到达中间线群组后, MG地址将作为 MG签重新插入到 IG标签与数据负载之间。 参见图 3中的数据片格式(β )。
4)转发阶段: 数据片将进一步使用 OG标签自路由地穿越第 2级转发模块, 并 最终到达其预期的输出群组。当数据片到达输出端 OGs的重排序緩存 RB时, OG标签将被丟弃。 参见图 3中的数据片格式(γ )。
5) 离开阶段: 在重排序緩存 RB中, 根据算法 1被分割的数据块将使用算法 2进 行重组, 并离开交换结构的输出端。
下面, 针对所述算法 1和算法 2进行详细说明。
算法 1:
对于每一个输入群组 IG, 在封装阶段, 存储于 VOGQ ( i,j ) 的数据块首先 被均匀地切割成 M个等长的数据片, 即为图 3中标记出的负载。 之后, 自路由 标签 MG、 IG和 OG被插入到相应数据片之前, 以供穿越两级结构路由时使用。 其中 MG将按从小到大的顺序依次插入到特定 VOGQ所对应的 M份数据片前。
/*算法 1伪代码, M为总的群组数 */
DataBlock[M][M]; /*DataBlock(i,j)为 VOGQ(i )内存储的数据块 */
SlicePayload[M][M][M]; /*SlicePayload(i,j,k)代表 DataBlock(i )分割后的 数据片负载 */
DataSlice[M][M][M]; /*DataSlice(i,j,k)代表加标记后的数据片 */
IG[M]={0, 1 , 2 , 3 , ···, M-1};
OG[M]={ 0, 1 , 2 , 3 , ···, M-1};
MG[M]={ 0, 1 , 2 , 3 , ···, M-1};
/* IG,OG,MG数组存储了自路由标签 */
for (i=0;i<M;i++) /*对每个输入群组进行处理, 实现时输入群组间并行运 行 */
for(j=0;j<M;j++){
Segment(DataBlock[i][j]);/*将数据块均匀分割为 M份负载,生成未 加自路由标签的数据片 SlicePayload */
for (k=0;k<M;k++) {
AddTag(SlicePayload[i] [j] [k],IG[i],OG[j],MG[k]);
/*为切割后的负载依次加标签, 生成数据片 DataSlice(i,j,k)*/
}
}
图 4中, 对于输入群组 IG5, 在封装阶段, 存储于 VOGQ (5,j)的数据块首 先被均匀地切割成 M=8个等长数据片。 之后, 自路由标签 MG, IG和 OG被插入 到数据片之前。 对于 VOGQ (5,j) , 8个数据片的 IG=5, OG=j, 而 MG则按 0到 7 依次编号。 这样便将输入端去往某输出群组的流量均匀地导出到所有输出群 组。
算法 2:
具有相同 IG标签的数据片被重组在一起, MG标签将用于按顺序将数据块 复原。 然后将数据块按分组大小重新切割为分组。处理后的分组就可被送离输 出端口。
如图 5所示, 算法 2的 C语言伪代码如下:
/*算法 2伪代码, M为总的群组数 */
DataBlock[M][M]; /*DataBlock(i,j)为 VOGQ(i )内存储的数据块 */
SlicePayload[M][M][M]; /*SlicePayload(i,j,k)代表 DataBlock(i )分割后的 数据片负载 */
DataSlice[M][M][M]; /*DataSlice(i,j,k)代表加标签后的数据片 */ IG[M]={0, 1 , 2 , 3 , ···, M-1};
OG[M]={ 0, 1 , 2 , 3 , ···, M-1};
MG[M]={ 0, 1 , 2 , 3 , ···, M-1};
/* IG,OG,MG数组存储了自路由标签 */
for G=0;j<M;j++) /*对每个输出群组进行处理, 实现时输出群组间并行运
行 */
for (i=0; i<M;j++){
for (k=0; k<M;k++)
{
DeleteTag(DataSlice(i,j,k), IG[j], MG[k]);
/*去除 DataSlice的自路由标签 IG,MG, 还原成未加标签的 SlicePayload; OG 标志已在自路由过程中去除,见图 3格式 γ*/
Recover(DataBlock(i,j), SlicePayload(i,j,k));
/*将数据片负载 SlicePayload(i,j,k)按 k从小到大依次组合,最终恢复
DataBlock(ij)*/
}
}
/*数据块恢复完成, 按分组大小重新分割后可离开交换结构输出端 */ 图 5中, 根据算法 2, 在每个输出群组 OG上, 我们首先搜集同属输入群组 IG5的数据片, 然后根据数据片的 MG标签,从小到大依次将这些数据片组合起 来。 去除多余的自路由标签后, 我们就复原了数据块。 再将这些数据块按分组 大小切割后, 就可将它们送出输出端口了。
我们在负载均衡的分组交换结构输入端前置的 VOGQ中对去往每个输出 端口的分组进行封装切割,在输出的后置的重排序緩存中对这些经过切割的数 据片进行重排序。 由于交换结构的输出端口为 M, 即需要把分组组合成数据块 然后平均切割成 M个数据片, 而一个 2G-to-G自路由集线器的组大小为 G, 于是 M和 G的大小关系影响着分组组合封装和输出的方法。 本实施例给出 M和 G的 三种关系的封装和传输方法。
1 ) M=G: 这种情况最筒单。 交换结构的两个输入群组连接到一个 2G-to-G自路
由群组集线器, 而 2G-to-G自路由集线器的规模是 2Gx2G。 在封装的时候, 把 VOGQ的某一个虚拟输出队列的一个数据块切割成了 M个数据片, 于是一个 2G-to-G自路由群组集线器的每个输入端有 M个数据片。 由于 M=G, 所以对每 个 VOGQ的某一个虚拟输出队列进行封装切割而成的 M个数据片可以在一个 时隙全部送到输入端。 由于交换结构中间没有緩存, 于是这 M个数据片在交换 结构中经过相同的传输延迟, 同一个时隙到达输出端口后置的重排序緩存,从 而根据自路由标签重新组合成没有切割的数据块, 然后送到输出端的线卡上。 由于 G个输入的数据分组根据输出端口地址分别存放到 VOGQ相应的队列中, 而输出端口有 M个, 于是一个输入群组的前置 VOGQ实际上有 M个虚拟输出队 列组成。 这样每经过 M个时隙, VOGQ的每一个虚拟输出队列就可以发送一次 数据分组。
2 ) M<G:由于 M=2m,G=2g, 于是 G是 M的 2X倍( x为正整数)。 由于 VOGQ的每个 虚拟输出队列的一个数据块被封装切割成了 M个数据片, 而一个 2G-to-G的规 模是 2Gx2G, 如果每次只传输 VOGQ的某一个虚拟输出队列, 则只使用了自路 由集线器的 2M个输入(或者输出)端口, 而自路由集线器共有 2G个输入(或 者输出)端口。为了充分利用自路由集线器的路由交换能力,我们对每个 VOGQ 的 2X个虚拟输出队列进行封装切割,这样一个自路由集线器的两个输入端共输 入 2x2xxM=2G个数据片。这样每经过 M/2X个时隙, VOGQ的每一个虚拟输出队 列就可以发送一次数据分组。
3 ) M>G:由于 M=2m,G=2g, 于是 M是 G的 2X倍( x为正整数)。 由于 VOGQ的每个 虚拟输出队列的一个数据块被封装切割成了 M个数据片, 如果每次封装传输 VOGQ的一个虚拟输出队列, 则一个自路由集线器的两个输入端共产生 2M个 数据片, 而一个 2G-to-G的规模是 2Gx2G, 这样便超过了自路由集线器的路由 交换能力。 为了解决这个问题, 我们把 M个数据片分成 2X个部分, 这样每个部 分有 G个数据片。 同时为了防止负载均衡模块内部的阻塞, 我们把负载均衡交 换结构的输入组也分成 2X个部分, 这样每部分有 G个输入组。 在一个时隙中, 一个输入组部分的 G个输入组分别向 0至(G-1 ), G至(2G-1 ), ( M-1-G ) 至(M-1 )输出组发送数据, 这样 2X个时隙便完成了一个轮转, 即 VOGQ的一
个虚拟输出队列完成了一次数据发送。 于是每经过 2χχΜ个时隙, VOGQ中的 每一个虚拟输出队列完成一次数据的传输。 由于本发明实施例中的交换结构采用基于自路由集线器的分组交换结构, 而这种结构可以递归构造, 于是这个负载均衡交换结构的规模不受限制。 同时 该交换结构是完全分布式的自路由,也为该负载均衡交换结构的大规模实现提 供了技术和物理上的基础。 综上所述,通过将基于自路由集线器的负载均衡分组交换结构分成第一级 交换模块和第二级交换模块,在第一级交换模块输入端前设置虚拟输出群组队 列,在第二级交换模块输出端后设置重排序緩存, 并在所述虚拟输出群组队列 存储的分组发送到第一级交换前,将分组组合成预置长度的数据块, 然后被分 割成等长的数据片, 同时在数据片上添加用于实现自路由的自路由地址信息, 在重排序緩存中,根据数据片携带的自路由地址信息,再把数据片重新组合成 所述虚拟输出群组队列组合成的数据块。 通过上述结构的变化, 可见, 本发明 实施例提供的负载均衡分组交换结构取消了背景技术中所述的第一级交换和 第二级交换之间的虚拟输出队列 VOQ这一中间级, 使得本发明实施例所述的 负载均衡分组交换结构不存在排队延迟问题, 进而避免了分组乱序。 所以, 本 发明实施例解决了负载均衡 Birkhoff-von Neumann交换结构能够分组乱序的问 题, 提高端到端的吞吐量。 通过将基于自路由集线器的负载均衡分组交换结构分成第一级交换模块 和第二级交换模块,在第一级交换模块输入端前设置虚拟输出群组队列,在第 二级交换模块输出端后设置重排序緩存,并在所述虚拟输出群组队列存储的分 组发送到第一级交换前,将分组组合成预置长度的数据块, 然后被分割成等长 的数据片, 同时在数据片上添加用于实现自路由的自路由标签,在重排序緩存 中,根据数据片携带的自路由标签,再把数据片重新组合成所述虚拟输出群组 队列组合成的数据块, 解决了负载均衡 Birkhoff-von Neumann交换结构能够分 组乱序的问题, 提高端到端的吞吐量。
以上所述仅为本发明的较佳实施例而已, 并不用以限制本发明, 凡在本发 明的精神和原则之内, 所作的任何修改、 等同替换、 改进等, 均应包含在本发 明的保护范围之内。
Claims
1、 一种负载均衡分组交换结构的构造方法, 其特征在于, 包括: 将基于自路由集线器的负载均衡分组交换结构分成具有完成负载均衡功 能的第一级交换模块和具有完成分组自路由转发功能的第二级交换模块; 在所述第一级交换模块输入端前设置虚拟输出群组队列,在所述第二级交 换模块输出端后设置重排序緩存, 所述虚拟输出群组队列, 用于存储带有自路 由地址信息的分组数据块, 所述重排序緩存, 用于将属于同一个输入群组的分 组数据块按自路由地址信息排列, 以便后续处理; 所述虚拟输出群组队列存储的分组发送到第一级交换前,组合成预置长度 的数据块, 然后被分割成等长的数据片, 同时在数据片上添加用于实现自路由 的自路由标签; 拥有自路由标签的数据片经第一级交换模块和第二级交换模块传输后到 达目的输出端口的重排序緩存,根据数据片携带的自路由标签,把数据片重新 组合成所述虚拟输出群组队列组合成的数据块。
2、 根据权利要求 1所述的方法, 其特征在于, 在所述第一级交换模块和第 二级交换模块之间设置中间线群组。
3、 根据权利要求 1或 2所述的方法, 其特征在于, 所述基于自路由集线器 的负载均衡分组交换结构采用分布式自路由。
4、 一种负载均衡分组交换结构, 包括基于自路由集线器用于完成负载均 衡功能的第一级交换模块和用于完成分组数据自路由转发功能的第二级交换 模块, 其特征在于, 在所述第一级交换模块输入端前设置虚拟输出群组队列, 在所述第二级交换模块输出端后设置重排序緩存, 所述虚拟输出群组队列, 用 于存储带有自路由地址信息的分组数据块, 所述重排序緩存, 用于将属于同一 个输入群组的分组数据块按自路由地址信息排列, 以便后续处理, 所述第一级 交换和第二级交换之间为中间线群组连接。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US12/739,729 US8902887B2 (en) | 2008-11-04 | 2009-10-31 | Load-balancing structure for packet switches and its constructing method |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN 200810217261 CN101404616A (zh) | 2008-11-04 | 2008-11-04 | 一种负载均衡分组交换结构及其构造方法 |
| CN200810217261.3 | 2008-11-04 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2010051737A1 true WO2010051737A1 (zh) | 2010-05-14 |
Family
ID=40538491
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2009/074739 Ceased WO2010051737A1 (zh) | 2008-11-04 | 2009-10-31 | 一种负载均衡分组交换结构及其构造方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US8902887B2 (zh) |
| CN (1) | CN101404616A (zh) |
| WO (1) | WO2010051737A1 (zh) |
Families Citing this family (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101404616A (zh) * | 2008-11-04 | 2009-04-08 | 北京大学深圳研究生院 | 一种负载均衡分组交换结构及其构造方法 |
| WO2011050541A1 (zh) * | 2009-10-31 | 2011-05-05 | 北京大学深圳研究生院 | 最小缓存复杂度的负载均衡分组交换结构及其构造方法 |
| US9042383B2 (en) * | 2011-06-30 | 2015-05-26 | Broadcom Corporation | Universal network interface controller |
| CN107483574B (zh) | 2012-10-17 | 2021-05-28 | 阿里巴巴集团控股有限公司 | 一种负载均衡下的数据交互系统、方法及装置 |
| US9253121B2 (en) | 2012-12-31 | 2016-02-02 | Broadcom Corporation | Universal network interface controller |
| CN103152281B (zh) * | 2013-03-05 | 2014-09-17 | 中国人民解放军国防科学技术大学 | 基于两级交换的负载均衡调度方法 |
| CN103595658B (zh) * | 2013-11-18 | 2016-09-21 | 清华大学 | 无需闭环流控的可扩展定长多路径交换系统 |
| CN103596168A (zh) * | 2013-11-18 | 2014-02-19 | 无锡赛思汇智科技有限公司 | 一种无线通讯中自适应抗干扰的消息发送与接收方法及装置 |
| US9560124B2 (en) * | 2014-05-13 | 2017-01-31 | Google Inc. | Method and system for load balancing anycast data traffic |
| CN105357320A (zh) * | 2015-12-09 | 2016-02-24 | 浪潮电子信息产业股份有限公司 | 一种多Web服务器负载均衡系统 |
| CN107196868B (zh) * | 2017-05-19 | 2019-10-18 | 合肥工业大学 | 一种应用于片上网络的负载均衡系统 |
| CN110324265B (zh) * | 2018-03-29 | 2021-09-07 | 阿里巴巴集团控股有限公司 | 流量分发方法、路由方法、设备及网络系统 |
| US11146491B1 (en) | 2020-04-09 | 2021-10-12 | International Business Machines Corporation | Dynamically balancing inbound traffic in a multi-network interface-enabled processing system |
| EP4395364A4 (en) * | 2021-10-15 | 2024-07-10 | Huawei Technologies Co., Ltd. | EXCHANGE APPARATUS, EXCHANGE METHOD AND EXCHANGE DEVICE |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH01190196A (ja) * | 1988-01-26 | 1989-07-31 | Fujitsu Ltd | 自己ルーティング通話路制御方式 |
| CN101141660A (zh) * | 2006-09-05 | 2008-03-12 | 北京大学深圳研究生院 | 自路由交换集线器及其方法 |
| CN101141374A (zh) * | 2006-09-05 | 2008-03-12 | 北京大学深圳研究生院 | 自路由集线器以分治网络构成交换结构的方法 |
| CN101350779A (zh) * | 2008-08-26 | 2009-01-21 | 北京大学深圳研究生院 | 基于自路由集线器的电路式分组交换方法 |
| CN101388847A (zh) * | 2008-10-17 | 2009-03-18 | 北京大学深圳研究生院 | 一种负载均衡电路式分组交换结构及其构建方法 |
| CN101404616A (zh) * | 2008-11-04 | 2009-04-08 | 北京大学深圳研究生院 | 一种负载均衡分组交换结构及其构造方法 |
Family Cites Families (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CA1297567C (en) * | 1987-02-06 | 1992-03-17 | Kazuo Hajikano | Self routing-switching system |
| US5157654A (en) * | 1990-12-18 | 1992-10-20 | Bell Communications Research, Inc. | Technique for resolving output port contention in a high speed packet switch |
| US5253251A (en) * | 1991-01-08 | 1993-10-12 | Nec Corporation | Switching system with time-stamped packet distribution input stage and packet sequencing output stage |
| US5341369A (en) * | 1992-02-11 | 1994-08-23 | Vitesse Semiconductor Corp. | Multichannel self-routing packet switching network architecture |
| KR100321784B1 (ko) * | 2000-03-20 | 2002-02-01 | 오길록 | 중재 지연 내성의 분산형 입력 버퍼 스위치 시스템 및그를 이용한 입력 데이터 처리 방법 |
| JP2002077238A (ja) * | 2000-08-31 | 2002-03-15 | Fujitsu Ltd | パケットスイッチ装置 |
| KR100459036B1 (ko) * | 2001-12-18 | 2004-12-03 | 엘지전자 주식회사 | 에이티엠 스위치 시스템의 트레인 패킷 구성 방법 |
| GB0208797D0 (en) * | 2002-04-17 | 2002-05-29 | Univ Cambridge Tech | IP-Capable switch |
| US7310333B1 (en) * | 2002-06-28 | 2007-12-18 | Ciena Corporation | Switching control mechanism for supporting reconfiguaration without invoking a rearrangement algorithm |
| US7590102B2 (en) * | 2005-01-27 | 2009-09-15 | Intel Corporation | Multi-stage packet switching system |
-
2008
- 2008-11-04 CN CN 200810217261 patent/CN101404616A/zh active Pending
-
2009
- 2009-10-31 WO PCT/CN2009/074739 patent/WO2010051737A1/zh not_active Ceased
- 2009-10-31 US US12/739,729 patent/US8902887B2/en not_active Expired - Fee Related
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH01190196A (ja) * | 1988-01-26 | 1989-07-31 | Fujitsu Ltd | 自己ルーティング通話路制御方式 |
| CN101141660A (zh) * | 2006-09-05 | 2008-03-12 | 北京大学深圳研究生院 | 自路由交换集线器及其方法 |
| CN101141374A (zh) * | 2006-09-05 | 2008-03-12 | 北京大学深圳研究生院 | 自路由集线器以分治网络构成交换结构的方法 |
| CN101350779A (zh) * | 2008-08-26 | 2009-01-21 | 北京大学深圳研究生院 | 基于自路由集线器的电路式分组交换方法 |
| CN101388847A (zh) * | 2008-10-17 | 2009-03-18 | 北京大学深圳研究生院 | 一种负载均衡电路式分组交换结构及其构建方法 |
| CN101404616A (zh) * | 2008-11-04 | 2009-04-08 | 北京大学深圳研究生院 | 一种负载均衡分组交换结构及其构造方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN101404616A (zh) | 2009-04-08 |
| US8902887B2 (en) | 2014-12-02 |
| US20110176425A1 (en) | 2011-07-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2010051737A1 (zh) | 一种负载均衡分组交换结构及其构造方法 | |
| CN100405344C (zh) | 用于在交换结构中分发缓冲区状态信息的装置和方法 | |
| EP1856921B1 (en) | Multi-stage packet switching system with alternate traffic routing | |
| EP2317702B1 (en) | Methods and apparatus related to a distributed switch fabric | |
| US20060165111A1 (en) | Replication of multicast data packets in a multi-stage switching system | |
| US20110010474A1 (en) | Low latency request dispatcher | |
| US7590102B2 (en) | Multi-stage packet switching system | |
| US11496398B2 (en) | Switch fabric packet flow reordering | |
| US20020126669A1 (en) | Apparatus and methods for efficient multicasting of data packets | |
| CN101853211B (zh) | 涉及可变大小信元的共享存储器缓冲区的方法和设备 | |
| CN101388847A (zh) | 一种负载均衡电路式分组交换结构及其构建方法 | |
| WO2011050541A1 (zh) | 最小缓存复杂度的负载均衡分组交换结构及其构造方法 | |
| US8233496B2 (en) | Systems and methods for efficient multicast handling | |
| CN111953618B (zh) | 一种多级并行交换架构下的解乱序方法、装置及系统 | |
| US20080031262A1 (en) | Load-balanced switch architecture for reducing cell delay time | |
| JP4588259B2 (ja) | 通信システム | |
| WO2002065145A1 (en) | Method and system for sorting packets in a network | |
| US10164906B1 (en) | Scalable switch fabric cell reordering | |
| He et al. | Load-balanced multipath self-routing switching structure by concentrators | |
| CN108540398A (zh) | 反馈型负载均衡交叉缓冲调度算法 | |
| JP2786246B2 (ja) | 自己ルーチング通話路 | |
| WO2007074423A2 (en) | Method and system for byte slice processing data packets at a packet switch | |
| JP5338404B2 (ja) | スイッチ装置におけるReordering処理方法 | |
| CN121367680A (zh) | 内部包产生的可扩缩方法 | |
| Finochietto et al. | Hardware primitives for packet flow processing architectures |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| WWE | Wipo information: entry into national phase |
Ref document number: 12739729 Country of ref document: US |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 09824392 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 09824392 Country of ref document: EP Kind code of ref document: A1 |