WO2017019106A1 - Traffic on defined network having aggregations of network devices - Google Patents

Traffic on defined network having aggregations of network devices Download PDF

Info

Publication number
WO2017019106A1
WO2017019106A1 PCT/US2015/043020 US2015043020W WO2017019106A1 WO 2017019106 A1 WO2017019106 A1 WO 2017019106A1 US 2015043020 W US2015043020 W US 2015043020W WO 2017019106 A1 WO2017019106 A1 WO 2017019106A1
Authority
WO
WIPO (PCT)
Prior art keywords
network
aggregation
network devices
observation matrix
corresponding observation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2015/043020
Other languages
French (fr)
Inventor
Puneet Sharma
Mehdi Malboubi
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hewlett Packard Enterprise Development LP
Original Assignee
Hewlett Packard Enterprise Development LP
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hewlett Packard Enterprise Development LP filed Critical Hewlett Packard Enterprise Development LP
Priority to PCT/US2015/043020 priority Critical patent/WO2017019106A1/en
Publication of WO2017019106A1 publication Critical patent/WO2017019106A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/12Discovery or management of network topologies
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/14Network analysis or design
    • H04L41/142Network analysis or design using statistical or mathematical methods
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L43/00Arrangements for monitoring or testing data switching networks
    • H04L43/02Capturing of monitoring data
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L43/00Arrangements for monitoring or testing data switching networks
    • H04L43/04Processing captured monitoring data, e.g. for logfile generation

Definitions

  • network flow measurements are used for a variety of applications, including managing utilization of the network and optimizing traffic throughput.
  • network measurements are primarily obtained from Simple Network Management Protocol ("SNMP") resources such as link counters.
  • SNMP Simple Network Management Protocol
  • Many data centers utilize dedicated network monitoring links between servers and or racks of servers and network monitor system in order to communicate SNMP link load data, representing direct network measurements at an endpoint of the link (e.g., server, rack).
  • SNMP link load data representing direct network measurements at an endpoint of the link (e.g., server, rack).
  • the volume of monitoring traffic resulting from the number of data flows, as well as end-to-end network paths such flows utilize, typically far exceeds the bandwidth allocated for monitoring in the network.
  • FIG. 1 illustrates an example system for estimating traffic on a defined network in which multiple aggregations of network devices are provided.
  • FIG. 2 determines an example method for determining optimized compressed network measurements from each aggregation of a defined network.
  • FIG. 3 illustrates an example of a data center for implementing one or more examples.
  • Examples described herein estimate traffic in a defined network using compressed network measurements which are reflective of data flows through a given aggregation of network devices.
  • an observation matrix is determined for each aggregation of network devices (e.g., servers or rack with servers) using topology and routing information.
  • the corresponding observation matrix can be
  • each of multiple aggregations of network devices communicated to each of multiple aggregations of network devices in order to obtain, from each aggregation, a compressed representation of the network traffic through individual networked devices such as servers of that aggregation.
  • optical means an outcome configured to induce or augment an objective or characteristic at the expense of another characteristic or objective.
  • aggregation in context of network devices is intended to including a physical and/or logical aggregation (e.g., defined group) of network devices. Examples of
  • aggregations include physically clustered network devices (e.g., see racks as described with FIG. 1 and 3), as well as logical groupings of physically clustered devices (e.g., two or more racks) and/or multiple devices which are logically defined and/or physically interconnected into individual
  • examples recognize that the actual measurements are made by servers or groups of servers, but such measurements typically cannot be shared (e.g., with a network controller) because of bandwidth constraints.
  • network elements often contain local resources for making actual measurements, but the local measurements cannot be communicated directly to a central location where all flows within the network are known, as such
  • examples such as described provide for the communication of information that represents a compressed and optimal profile of the traffic being exchanged through individual devices of the node.
  • a network is identified in terms of aggregations of network devices (e.g., edge devices), from which data flows of the network originate.
  • Each aggregation of edge devices is provided an observation matrix which (i) accounts for routing and traffic flow across the entire network, (ii) is specific to the particular aggregation of devices, and (iii) when applied to actual measurements of each edge device of the aggregation, generate an aggregate set of measurements which are representative of the traffic exchanged through the aggregation.
  • Examples described herein provide the methods, techniques, and actions performed by a computing device that are performed
  • Examples may be implemented as hardware, or a combination of hardware (e.g., a
  • processor(s) and executable instructions (e.g., stored on a machine- readable storage medium). These instructions can be stored in one or more memory resources of the computing device.
  • a programmatically performed step may or may not be automatic.
  • the programmatic modules or components may be any combination of hardware (e.g., processor(s)) and programming to implement the functionalities of the modules or components described herein.
  • the programming for the components may be processor executable instructions stored on at least one non-transitory machine- readable storage medium and the hardware for the components may include at least one processing resource to execute those instructions.
  • the at least one machine-readable storage medium may storage instructions that, when executed by the at least one processing resource, implement the components.
  • Some examples described herein can generally involve the use of computing devices, including processing and memory resources.
  • examples described herein may be implemented, in whole or in part, on computing devices such as desktop computers, cellular or smart phones, personal digital assistants (PDAs), laptop computers, printers, digital picture frames, and tablet devices.
  • PDAs personal digital assistants
  • Memory, processing, and network resources may all be used in connection with the establishment, use, or performance of any example described herein (including with the
  • processors These instructions may be carried on a computer-readable medium.
  • Machines shown or described with figures below provide examples of processing resources and computer-readable mediums on which
  • Examples of computer-readable mediums include permanent memory storage devices, such as hard drives on personal computers or servers.
  • Other examples of computer storage mediums include portable storage units, such as CD or DVD units, flash memory (such as carried on smart phones, multifunctional devices or tablets), and magnetic memory.
  • Computers, terminals, network enabled devices e.g., mobile devices, such as cell phones
  • examples may be implemented in the form of computer- programs, or a computer usable carrier medium capable of carrying such a program.
  • FIG. 1 illustrates an example system for estimating traffic on a defined network in which multiple aggregations of network devices are provided.
  • a defined network 10 includes a network monitoring system 110 which communicates with multiple aggregations 120 of network devices 122 using network channels of the defined network 10.
  • the network monitoring system 110 can estimate traffic flow throughout the defined network 10, using actual measurements made by sensors on edge devices and other traffic measuring resources that are distributed on the network 10.
  • the network monitoring system 110 can utilize routing information and network topology to determine parameters for enabling network traffic measurements at each aggregation to be compressed in a manner that is meaningful and
  • defined network 10 can represent, for example, a data center network in which the aggregations of network devices include racks of servers. While some examples are described in the context of data center networks, examples as described can be applicable other kinds of networks in which groups or clusters of network devices are utilized, including other kinds of software defined networks and Ethernet type networks.
  • each aggregation 120 includes a cluster, group or stack of networked or edge devices 122 and a set of aggregation components 124.
  • the network devices 122 can be physically and/or logically aggregated to share a set of physical and/or network resources for operating on the network 10.
  • each aggregation 120 can correspond to a group of servers which share cabling or physical channels for communicating on the network 10, as well as data switches for enabling individual servers of the aggregation to access the network channels of the defined network 10 to communicate with other servers on other aggregations.
  • the aggregations 120 can each also include one or more aggregation resources 124, which can include switches which are shared amongst the network devices 122 to provide access to data channels of the network 10.
  • the aggregation resources 124 can also include local resources to make network measurements at each aggregation 120. Numerous other kinds of resources can also be shared at each aggregation 120, including ports, physical housing structures, and logical resources.
  • Each aggregation 120 can include a set of links 121 for communicating network measurements to the network monitoring system 110.
  • the network monitoring system 110 includes functionality for predicting and/or analyzing various aspects of the defined network 10.
  • network monitoring system 110 incudes a network inference engine 112, a statistical component for performing path analysis 114, and a compression parameter determination component 116.
  • the network inference engine 112 can develop, for example, a network inference model or data structure to analyze or predict traffic patterns and
  • the network inference engine 112 can be used to develop network traffic matrices or models which can in turn, be used to identify information which affects the health, performance and/or efficiency of the defined network 10 as a whole.
  • an output of the network inference engine 112 can be used to determine (i) a cause of a network anomaly, (ii) predict future network anomalies and traffic patterns, (iii) identify "heavy hitters" which utilize a disproportionate amount of bandwidth on the defined network 10, and/or (iv) determine traffic volume and throughput for various data channels.
  • the relevance or accuracy of the network inference engine 112 can be based on the quantity, depth and accuracy of network measurements which the network inference engine 112 uses as input.
  • network measurements which are made at each aggregation 120 and communicated to the network monitoring system 110 via links 121.
  • the defined network 10 may typically use the links 121, which are dedicated for network measurement
  • network traffic measurement data can be obtained through standard Simple Network Management Protocol (SNMP) which is resident on individual racks of servers. While such network traffic measurement data may be readily available locally at the rack, the number of data flows which have an end point at the rack are exponentially greater than the number of servers present within each rack. More generally, given a first aggregation 120 with n network devices, and a second aggregation 120 with m network devices, the total number of data flows just between the first and second aggregations is nxm. Thus, as shown by an example of FIG.
  • SNMP Simple Network Management Protocol
  • the network monitoring system 110 can operate to obtain a compressed or reduced, but representative, form of network measurements from the individual aggregations 120, for use with the network inference engine 112.
  • the directly obtained network measurement data 123 can correspond to SNMP link load data, as communicated by network measuring resources (e.g. aggregate resources 124) of each aggregation 120.
  • examples determine and utilize a set (or matrix) of compression parameters ("compression parameter data set 115") to optimize a compression of network
  • Each set of compression parameter data set 115 can also be determined to be specific to each aggregation 120, so that the resulting compression parameter data set 115 reflects the measured network data at that aggregation 120.
  • the optimization provided through the use of compression parameter data set 115 can be provided at expense of bandwidth.
  • the compressed network measurements 125 can
  • the optimization provided through the use of compression parameter data set 115 can be provided at expense of bandwidth.
  • examples such as provided with FIG. 1 are distinct from other approaches, some of which, for example, seek to obtain supplemental network measurement through sampling in order to determine more significant data flows.
  • an example of FIG 1 generates the compressed network measurement 125 to be reflective of all of the data flows of the aggregation 120.
  • the network monitoring system 110 implements operations to optimize the relevance and accuracy of the compression operator implemented on the network measurements which are made at a given location (e.g., aggregation 120), with available bandwidth serving as the cost for optimizing the network measurements.
  • the compressed network measurements are representative of all the data flows of the corresponding aggregation 120.
  • the compressed network measurements 125 can also take into account multipath routing of data flows.
  • the network monitoring system 110 implements processes for determining the optimal observation matrix (i.e. compression operator or compression parameter data set) 115 based in part on topology information 117 and the routing information (as determined from the routing matrix 119). Multiple compression parameter data sets 115 can be determined, with each compression parameter data set 115 being determined for optimization of network measurements made at a
  • the compression parameter data set 115 can be determined as an observation matrix which provides coefficients to each aggregation 120.
  • the aggregation resource 124 can include network measurement functionality to receive and apply the compression parameter data set 115 for that aggregation 120.
  • the values of the compression parameter data set 115 enable each aggregation 122 to return a set of network measurements 125 which are compressed by coefficients of the compression parameter data set 115, but the compressed set of network measurements 125 are also highly representative of the traffic profile or state of the corresponding aggregation 120.
  • examples provide that the calculations used to obtain the compression parameter data sets 115 and corresponding sets of compressed network measurements 125 are relatively light computationally, while the
  • network monitoring system 110 generates compression parameter data set 115 which account for multipath routing of individual data flows.
  • network monitoring system 110 can include components and functionality such as shown with statistical path analysis logic 114.
  • the statistical path analysis logic 114 can determine statistical expectation of data traffic from individual flows on separate paths of the defined network 10.
  • compression parameter determination logic 116 can use the statistical expectations in determining the compression parameter data sets 115 of each aggregation 120. In determining the statistical expectations the statistical path analysis logic 114 can utilize the probability distribution function of the routing matrix or different realization of the routing matrix can be used.
  • examples such as described provide a set of linear combination of per-flow measurements which are supplementary to directly measured network information (such as provided by SNMP link loads). Accordingly, an example of FIG. 1 can utilize compressed network measurements 125 to increase the estimation accuracy of network data flows, which are recognized as being highly variable over time/space, using compressed sensing techniques. As described with some other examples, the compression parameter data sets 115 for each aggregation 120 can be designed as observation matrices, which can be distributed to the
  • FIG. 2 determines an example method for determining optimized compressed network measurements from each aggregation of a defined network.
  • network topology and routing information can be obtained for the defined network 10 (210).
  • an example method of FIG. 2 can be performed at network monitor system 110 or similar network element, where a routing table for the network is located (212).
  • an observation matrix is determined for each aggregation 120 (220).
  • Each observation matrix may include or correspond to a compression parameter data set 115 for a particular aggregation 120.
  • a traffic matrix optimization problem can be posed as:
  • ⁇ " is a vector presentation of the traffic matrix which is unknown and must be estimated, is compression operator, also called the observation matrix, 1 " 'is the compressed network measurements (or linear combinations of unknowns), and is the routing matrix.
  • the set of compression operator 115 can be termed as an Optimal Local Observation Matrices (OLOM), which can be determined and applied independently of the OLOM of other aggregations.
  • OLOM Optimal Local Observation Matrices
  • a smaller traffic matrix estimation problem can be formulated by considerin local observation matrix as:
  • L represents the number of racks (or aggregations)
  • r a sub-set of link-loads (or rows of H)
  • the local observation matrices can be pre-distributed among racks/servers in an offline manner.
  • the optimal observation matrix - 3 ⁇ 4 3 ⁇ 4 is thus designed for each aggregation 120, and aggregation-specific observation matrices are distributed among servers in each aggregation 120, where measurement compression/aggregation modules are available and can be utilized to generate a desirable set of linear combinations of X (shown with 3 ⁇ 4 tf ). These aggregated measurements are then communicated to the network monitoring system where actual traffic matrix estimation is performed using a network inference technique of choice.
  • a good local observation matrix can be designed by minimizing the sum of all off- diagonal elements of the corresponding Gram matrix (G) defined as below where T denotes the Trans ose operation.
  • the routing matrix H includes coefficients or parametric values to reflect multi-path routing of individual data flows, and the coefficients can be used with the observation matrix. Specifically, to consider the randomness of some entries of H (due to multi-path routing in many data center network), a statistical expectation over H can be considered in the objective function. Note, for notation simplicity, .- ' -Ti is considered as A:
  • a computational process such as provided by the Newton method, can be used to solve the resulting optimization problem where the gradient of objective function is:
  • the network monitoring system 110 may receive a compressed representation of network measurements made by each aggregation 120 (232).
  • the compressed information can serve as supplemental information to directly measured network information, which for example can be
  • the network monitoring system 110 can apply a selected network inference technique based on the received measured network information, including the compressed representation of network
  • FIG. 3 illustrates an example of a data center for implementing one or more examples.
  • a data center 300 includes a plurality of racks, represented by racks 320, 330 and a network management system 310.
  • Each rack 320, 330 can include an aggregation of servers 322, 332, a set of top of rack switches 324, 334, and one or more aggregation switches 326, 336 and core switches 328, 338.
  • the aggregation switches 326, 336 can obtain and communicate network measurements to the network monitoring system 310 through a set of links 311. While FIG. 3 illustrates an example in which an aggregation is provided by a rack, in other examples, an aggregation can be defined to extent to multiple physically connected and/or logically defined racks.
  • the network monitoring system 310 can include an optimization component 312 for determining a rack specific observation matrix 315, and a network inference engine 314. As described with other examples, the network monitoring system 310 can generate the observation matrix 315 for the particular rack 320. In some variations, the observation matrix 315 can be communicated offline to the respective rack 320. The aggregation switch 326 of the receiving rack can implement the observation matrix 315 to return a compressed set of network measurements 317 on the link 311. The compressed set of measurements 317 can supplement directly measured values 313 which can also be communicated using the links 311. The network measuring system 310 can use the directly measured values 313 and the compressed set of measurements 317 as input for the inference engine 314.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Algebra (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Mathematical Physics (AREA)
  • Probability & Statistics with Applications (AREA)
  • Pure & Applied Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • Data Exchanges In Wide-Area Networks (AREA)

Abstract

An observation matrix is determined for each aggregation of network devices (e.g., rack with servers) using topology and routing information. The corresponding observation matrix can be communicated to each of multiple aggregations of network devices in order to obtain, from each aggregation, a compressed representation of the network traffic through individual network devices of that aggregation.

Description

TRAFFIC ON DEFINED NETWORK
HAVING AGGREGATIONS OF NETWORK DEVICES
BACKGROUND
[0001] For many types of networks, network flow measurements are used for a variety of applications, including managing utilization of the network and optimizing traffic throughput. In many applications, network measurements are primarily obtained from Simple Network Management Protocol ("SNMP") resources such as link counters. Many data centers utilize dedicated network monitoring links between servers and or racks of servers and network monitor system in order to communicate SNMP link load data, representing direct network measurements at an endpoint of the link (e.g., server, rack). However, the volume of monitoring traffic resulting from the number of data flows, as well as end-to-end network paths such flows utilize, typically far exceeds the bandwidth allocated for monitoring in the network.
BRIEF DESCRIPTION OF THE DRAWINGS
[0002] FIG. 1 illustrates an example system for estimating traffic on a defined network in which multiple aggregations of network devices are provided.
[0003] FIG. 2 determines an example method for determining optimized compressed network measurements from each aggregation of a defined network.
[0004] FIG. 3 illustrates an example of a data center for implementing one or more examples.
DETAILED DESCRIPTION
[0005] Examples described herein estimate traffic in a defined network using compressed network measurements which are reflective of data flows through a given aggregation of network devices. According to examples described, an observation matrix is determined for each aggregation of network devices (e.g., servers or rack with servers) using topology and routing information. The corresponding observation matrix can be
communicated to each of multiple aggregations of network devices in order to obtain, from each aggregation, a compressed representation of the network traffic through individual networked devices such as servers of that aggregation.
[0006] As used herein, the term "optimal" or variations thereof (e.g., "optimized") means an outcome configured to induce or augment an objective or characteristic at the expense of another characteristic or objective.
[0007] Additionally, the term "aggregation" (and variants thereof) in context of network devices is intended to including a physical and/or logical aggregation (e.g., defined group) of network devices. Examples of
aggregations include physically clustered network devices (e.g., see racks as described with FIG. 1 and 3), as well as logical groupings of physically clustered devices (e.g., two or more racks) and/or multiple devices which are logically defined and/or physically interconnected into individual
aggregations.
[0008] In estimation theory in general, Y=AX represents a fundamental relationship in which Y is an actual measurement, X is that is to be estimated, and A is an observation matrix which maps the actual and estimated measurements. With regard to network traffic monitoring, examples recognize that the actual measurements are made by servers or groups of servers, but such measurements typically cannot be shared (e.g., with a network controller) because of bandwidth constraints. Thus, network elements often contain local resources for making actual measurements, but the local measurements cannot be communicated directly to a central location where all flows within the network are known, as such
communication would require far too much bandwidth than allocated for monitoring purpose. Rather, examples such as described provide for the communication of information that represents a compressed and optimal profile of the traffic being exchanged through individual devices of the node.
[0009] According to examples, a network is identified in terms of aggregations of network devices (e.g., edge devices), from which data flows of the network originate. Each aggregation of edge devices is provided an observation matrix which (i) accounts for routing and traffic flow across the entire network, (ii) is specific to the particular aggregation of devices, and (iii) when applied to actual measurements of each edge device of the aggregation, generate an aggregate set of measurements which are representative of the traffic exchanged through the aggregation.
[0010] Examples described herein provide the methods, techniques, and actions performed by a computing device that are performed
programmatically, or as a computer-implemented method. Examples may be implemented as hardware, or a combination of hardware (e.g., a
processor(s)) and executable instructions (e.g., stored on a machine- readable storage medium). These instructions can be stored in one or more memory resources of the computing device. A programmatically performed step may or may not be automatic.
[0011] Examples described herein can be implemented using
programmatic modules or components. The programmatic modules or components may be any combination of hardware (e.g., processor(s)) and programming to implement the functionalities of the modules or components described herein. In examples described herein, such combinations of hardware and programming may be implemented in a number of different ways. For example, the programming for the components may be processor executable instructions stored on at least one non-transitory machine- readable storage medium and the hardware for the components may include at least one processing resource to execute those instructions. In such examples, the at least one machine-readable storage medium may storage instructions that, when executed by the at least one processing resource, implement the components.
[00012] Some examples described herein can generally involve the use of computing devices, including processing and memory resources. For example, examples described herein may be implemented, in whole or in part, on computing devices such as desktop computers, cellular or smart phones, personal digital assistants (PDAs), laptop computers, printers, digital picture frames, and tablet devices. Memory, processing, and network resources may all be used in connection with the establishment, use, or performance of any example described herein (including with the
performance of any method or with the implementation of any system).
[0013] Furthermore, examples described herein may be implemented through the use of instructions that are executable by one or more
processors. These instructions may be carried on a computer-readable medium. Machines shown or described with figures below provide examples of processing resources and computer-readable mediums on which
instructions for implementing examples described herein can be carried and/or executed. In particular, the numerous machines shown with examples include processor(s) and various forms of memory for holding data and instructions. Examples of computer-readable mediums include permanent memory storage devices, such as hard drives on personal computers or servers. Other examples of computer storage mediums include portable storage units, such as CD or DVD units, flash memory (such as carried on smart phones, multifunctional devices or tablets), and magnetic memory. Computers, terminals, network enabled devices (e.g., mobile devices, such as cell phones) are all examples of machines and devices that utilize processors, memory, and instructions stored on computer-readable mediums. Additionally, examples may be implemented in the form of computer- programs, or a computer usable carrier medium capable of carrying such a program.
[0014] SYSTEM DESCRIPTION
[0015] FIG. 1 illustrates an example system for estimating traffic on a defined network in which multiple aggregations of network devices are provided. As shown, a defined network 10 includes a network monitoring system 110 which communicates with multiple aggregations 120 of network devices 122 using network channels of the defined network 10. Among other functions, the network monitoring system 110 can estimate traffic flow throughout the defined network 10, using actual measurements made by sensors on edge devices and other traffic measuring resources that are distributed on the network 10. As described in greater detail, the network monitoring system 110 can utilize routing information and network topology to determine parameters for enabling network traffic measurements at each aggregation to be compressed in a manner that is meaningful and
representative of the network traffic incident on or passing through all of the network devices 122 of the aggregation 120, but the compressed monitoring data has significantly less bandwidth requirements than direct network measurement data which can be obtained over the same set of data channels. [0016] In an example of FIG. 1, defined network 10 can represent, for example, a data center network in which the aggregations of network devices include racks of servers. While some examples are described in the context of data center networks, examples as described can be applicable other kinds of networks in which groups or clusters of network devices are utilized, including other kinds of software defined networks and Ethernet type networks.
[0017] Depending on the type of network, each aggregation 120 includes a cluster, group or stack of networked or edge devices 122 and a set of aggregation components 124. The network devices 122 can be physically and/or logically aggregated to share a set of physical and/or network resources for operating on the network 10. By way of example, each aggregation 120 can correspond to a group of servers which share cabling or physical channels for communicating on the network 10, as well as data switches for enabling individual servers of the aggregation to access the network channels of the defined network 10 to communicate with other servers on other aggregations. The aggregations 120 can each also include one or more aggregation resources 124, which can include switches which are shared amongst the network devices 122 to provide access to data channels of the network 10. The aggregation resources 124 can also include local resources to make network measurements at each aggregation 120. Numerous other kinds of resources can also be shared at each aggregation 120, including ports, physical housing structures, and logical resources. Each aggregation 120 can include a set of links 121 for communicating network measurements to the network monitoring system 110.
[0018] The network monitoring system 110 includes functionality for predicting and/or analyzing various aspects of the defined network 10. In an example of FIG. 1, network monitoring system 110 incudes a network inference engine 112, a statistical component for performing path analysis 114, and a compression parameter determination component 116. The network inference engine 112 can develop, for example, a network inference model or data structure to analyze or predict traffic patterns and
characterizations of data flows, as well as other aspects of the defined network 10 (e.g., network traffic on data channels, network ingress/egress of aggregations 120, and/or network devices 122). For example, the network inference engine 112 can be used to develop network traffic matrices or models which can in turn, be used to identify information which affects the health, performance and/or efficiency of the defined network 10 as a whole. For example, an output of the network inference engine 112 can be used to determine (i) a cause of a network anomaly, (ii) predict future network anomalies and traffic patterns, (iii) identify "heavy hitters" which utilize a disproportionate amount of bandwidth on the defined network 10, and/or (iv) determine traffic volume and throughput for various data channels.
[0019] The relevance or accuracy of the network inference engine 112 can be based on the quantity, depth and accuracy of network measurements which the network inference engine 112 uses as input. In many applications, network measurements which are made at each aggregation 120 and communicated to the network monitoring system 110 via links 121. However, in many kinds of defined networks 10, there is insufficient bandwidth for communicating full sets of network measurements (or other network measurement sources) which would reflect traffic or state of individual data paths within the defined network 10. The defined network 10 may typically use the links 121, which are dedicated for network measurement
communications, but the links 121 are far less in quantity and bandwidth than data paths of the defined network 10. For example, in data centers, network traffic measurement data can be obtained through standard Simple Network Management Protocol (SNMP) which is resident on individual racks of servers. While such network traffic measurement data may be readily available locally at the rack, the number of data flows which have an end point at the rack are exponentially greater than the number of servers present within each rack. More generally, given a first aggregation 120 with n network devices, and a second aggregation 120 with m network devices, the total number of data flows just between the first and second aggregations is nxm. Thus, as shown by an example of FIG. 1, when links 121 between network measurement sources and the network monitor system 110 are provided, the bandwidth allocated for monitoring on these links is limited and generally is not enough to carry monitoring data for the number of data flows, or paths for such data flows. As a result, locally obtained network measurement data, such as show by directly measured data ("DMD") 123 (such as SNMP link data) from a particular aggregation 120 at a moment has a limited range of use, as the link bandwidth constrains the amount of direct measurement (e.g., SNMP link load data) which the network monitoring system 110 can ultimately receive, as compared to the data flows and end- to-end network paths.
[0020] Accordingly, the network monitoring system 110 can operate to obtain a compressed or reduced, but representative, form of network measurements from the individual aggregations 120, for use with the network inference engine 112. Numerous techniques exist for obtaining compressed network measurements as a mechanism to augment or supplement directly obtained network measurement data 123. By way of example, the directly obtained network measurement data 123 can correspond to SNMP link load data, as communicated by network measuring resources (e.g. aggregate resources 124) of each aggregation 120. In order to obtain supplementary network measurement data, examples determine and utilize a set (or matrix) of compression parameters ("compression parameter data set 115") to optimize a compression of network
measurements ("compressed network measurements 125" or "CNM 125") made at each aggregation 120 to accurately reflect the network
measurements which affect all of the data flows of the respective aggregation 120. Each set of compression parameter data set 115 can also be determined to be specific to each aggregation 120, so that the resulting compression parameter data set 115 reflects the measured network data at that aggregation 120. The optimization provided through the use of compression parameter data set 115 can be provided at expense of bandwidth. In some implementations, the compressed network measurements 125 can
supplement SNMP link loads, provided over the links between aggregations and the network monitor system.
[0021] The optimization provided through the use of compression parameter data set 115 can be provided at expense of bandwidth. In this respect, examples such as provided with FIG. 1 are distinct from other approaches, some of which, for example, seek to obtain supplemental network measurement through sampling in order to determine more significant data flows. In contrast to such approaches, an example of FIG 1 generates the compressed network measurement 125 to be reflective of all of the data flows of the aggregation 120. [0022] In an example of FIG. 1, the network monitoring system 110 implements operations to optimize the relevance and accuracy of the compression operator implemented on the network measurements which are made at a given location (e.g., aggregation 120), with available bandwidth serving as the cost for optimizing the network measurements. Moreover, in some implementations, the compressed network measurements are representative of all the data flows of the corresponding aggregation 120. Still further, the compressed network measurements 125 can also take into account multipath routing of data flows.
[0023] In an example of FIG. 1, the network monitoring system 110 implements processes for determining the optimal observation matrix (i.e. compression operator or compression parameter data set) 115 based in part on topology information 117 and the routing information (as determined from the routing matrix 119). Multiple compression parameter data sets 115 can be determined, with each compression parameter data set 115 being determined for optimization of network measurements made at a
corresponding aggregation 120. As described with other examples, the compression parameter data set 115 can be determined as an observation matrix which provides coefficients to each aggregation 120. The aggregation resource 124 can include network measurement functionality to receive and apply the compression parameter data set 115 for that aggregation 120. The values of the compression parameter data set 115 enable each aggregation 122 to return a set of network measurements 125 which are compressed by coefficients of the compression parameter data set 115, but the compressed set of network measurements 125 are also highly representative of the traffic profile or state of the corresponding aggregation 120. As further described, examples provide that the calculations used to obtain the compression parameter data sets 115 and corresponding sets of compressed network measurements 125 are relatively light computationally, while the
representation of the individual aggregations 120 provided by the
corresponding compressed sets of network measurements 125 have characteristics of being both relatively fine in granularity and representative of network traffic characteristics of all of the data flows that are formed through that aggregation 120. [0024] In some implementations, network monitoring system 110 generates compression parameter data set 115 which account for multipath routing of individual data flows. In order to count for multipath routing, network monitoring system 110 can include components and functionality such as shown with statistical path analysis logic 114. The statistical path analysis logic 114 can determine statistical expectation of data traffic from individual flows on separate paths of the defined network 10. The
compression parameter determination logic 116 can use the statistical expectations in determining the compression parameter data sets 115 of each aggregation 120. In determining the statistical expectations the statistical path analysis logic 114 can utilize the probability distribution function of the routing matrix or different realization of the routing matrix can be used.
[0025] Among other benefits, examples such as described provide a set of linear combination of per-flow measurements which are supplementary to directly measured network information (such as provided by SNMP link loads). Accordingly, an example of FIG. 1 can utilize compressed network measurements 125 to increase the estimation accuracy of network data flows, which are recognized as being highly variable over time/space, using compressed sensing techniques. As described with some other examples, the compression parameter data sets 115 for each aggregation 120 can be designed as observation matrices, which can be distributed to the
aggregations 120 in an offline manner. Examples as described can improve the estimation accuracy resulting from the compressed network
measurements 125, while also achieving high Compression Ratio (CR) under limited link bandwidths.
[0026] METHODOLOGY
[0027] FIG. 2 determines an example method for determining optimized compressed network measurements from each aggregation of a defined network. In describing an example of FIG. 2, reference may be made to elements of FIG. 1 for purpose of illustrating suitable components for performing a step or sub-step being described.
[0028] With reference to an example of FIG. 2, network topology and routing information can be obtained for the defined network 10 (210). In some examples, an example method of FIG. 2 can be performed at network monitor system 110 or similar network element, where a routing table for the network is located (212).
[0029] With the routing and topology information, an observation matrix is determined for each aggregation 120 (220). Each observation matrix may include or correspond to a compression parameter data set 115 for a particular aggregation 120. In many networks, including data centers, an assumption can be made that the network exhibits sparse behavior, with large fluctuations of data flows present. Given sparseness and fluctuations, a traffic matrix optimization problem can be posed as:
X = mmx \\Y° - A°X\\2 + λ\\Χ\\ where
Y = HX
Figure imgf000011_0003
Figure imgf000011_0002
[0030] In the equations provided, Λ" is a vector presentation of the traffic matrix which is unknown and must be estimated, is compression operator, also called the observation matrix, 1 "'is the compressed network measurements (or linear combinations of unknowns), and is the routing matrix. For any given aggregation 120, the set of compression operator 115 can be termed as an Optimal Local Observation Matrices (OLOM), which can be determined and applied independently of the OLOM of other aggregations.
[0031] A smaller traffic matrix estimation problem can be formulated by considerin local observation matrix as:
Figure imgf000011_0001
based on the assumption that local observation matrices Ai?i can be designed, independently. In designing each OLOM for each grouping 120, L represents the number of racks (or aggregations), r a sub-set of link-loads (or rows of H), can be assumed to have contributions from flows observed by ■ΐ'ί . In some examples, the local observation matrices can be pre-distributed among racks/servers in an offline manner. The optimal observation matrix - ¾¾ is thus designed for each aggregation 120, and aggregation-specific observation matrices are distributed among servers in each aggregation 120, where measurement compression/aggregation modules are available and can be utilized to generate a desirable set of linear combinations of X (shown with ¾ tf). These aggregated measurements are then communicated to the network monitoring system where actual traffic matrix estimation is performed using a network inference technique of choice.
[0032] Assuming the compressed sensing network inference, a good local observation matrix can be designed by minimizing the sum of all off- diagonal elements of the corresponding Gram matrix (G) defined as below where T denotes the Trans ose operation.
G = Αυ for i = 1,
Figure imgf000012_0001
[0033] In some examples, the routing matrix H includes coefficients or parametric values to reflect multi-path routing of individual data flows, and the coefficients can be used with the observation matrix. Specifically, to consider the randomness of some entries of H (due to multi-path routing in many data center network), a statistical expectation over H can be considered in the objective function. Note, for notation simplicity, .-'-Ti is considered as A:
Figure imgf000012_0002
[0034] A computational process, such as provided by the Newton method, can be used to solve the resulting optimization problem where the gradient of objective function is:
VF - 4A(EH[HTH] + ATA I) E denotes the statistical expectation operator, and the iterations of Newton algorithm(for ith column of is continued until convergence as:
A(: , i)M A( , ¾™~ n^F for i = 1, , N
[0035] The network monitoring system 110 may receive a compressed representation of network measurements made by each aggregation 120 (232). The compressed information can serve as supplemental information to directly measured network information, which for example can be
communicated over dedicated data links.
[0036] The network monitoring system 110 can apply a selected network inference technique based on the received measured network information, including the compressed representation of network
measurements (234).
[0037] DATA CENTER EXAMPLE
[0038] FIG. 3 illustrates an example of a data center for implementing one or more examples. In an example of FIG. 3, a data center 300 includes a plurality of racks, represented by racks 320, 330 and a network management system 310. Each rack 320, 330 can include an aggregation of servers 322, 332, a set of top of rack switches 324, 334, and one or more aggregation switches 326, 336 and core switches 328, 338. In an example shown, the aggregation switches 326, 336 can obtain and communicate network measurements to the network monitoring system 310 through a set of links 311. While FIG. 3 illustrates an example in which an aggregation is provided by a rack, in other examples, an aggregation can be defined to extent to multiple physically connected and/or logically defined racks.
[0039] The network monitoring system 310 can include an optimization component 312 for determining a rack specific observation matrix 315, and a network inference engine 314. As described with other examples, the network monitoring system 310 can generate the observation matrix 315 for the particular rack 320. In some variations, the observation matrix 315 can be communicated offline to the respective rack 320. The aggregation switch 326 of the receiving rack can implement the observation matrix 315 to return a compressed set of network measurements 317 on the link 311. The compressed set of measurements 317 can supplement directly measured values 313 which can also be communicated using the links 311. The network measuring system 310 can use the directly measured values 313 and the compressed set of measurements 317 as input for the inference engine 314.
[0040] Although illustrative embodiments have been described in detail herein with reference to the accompanying drawings, variations to specific embodiments and details are encompassed by this disclosure. It is intended that the scope of embodiments described herein be defined by claims and their equivalents. Furthermore, it is contemplated that a particular feature described, either individually or as part of an embodiment, can be combined with other individually described features, or parts of other embodiments. Thus, absence of describing combinations should not preclude the inventor(s) from claiming rights to such combinations.

Claims

What is claimed is:
1. A method for estimating traffic on a defined network, the method being implemented by one or more processors and
comprising :
obtaining information which is indicative of (i) a topology of the defined network, the topology including multiple aggregations of network devices, and (ii) routing information that identifies a plurality of data flows amongst network devices of the defined network;
determining, using at least the topology and the routing information, a corresponding observation matrix for each of the multiple aggregations of network devices; and
communicating the corresponding observation matrix to each of the multiple aggregations of network devices in order to obtain, from each aggregation, a compressed representation of the network traffic through individual network devices of that aggregation.
2. The method of claim 1, further comprising generating a representation of the network traffic through the defined network using the compressed representation of each of the multiple aggregations of network devices.
3. The method of claim 1, wherein communicating the
corresponding observation matrix to each of the multiple aggregations includes distributing the corresponding observation matrix to each of the multiple aggregations of network devices.
4. The method of claim 1, wherein determining the corresponding observation matrix for the multiple aggregations of network devices includes determining a statistical expectation, using at least the routing information, in order to account for multi-path routing of data provided in part through the network devices of each aggregation.
5. The method of claim 4, wherein determining the corresponding observation matrix for the multiple aggregations of network devices includes determining, from the statistical expectation, a set of non- binary coefficients for the corresponding observation matrix of each aggregation.
6. The method of claim 4, wherein determining the statistical expectation includes using at least one of a probability distribution function of the routing matrix or a different set of realizations of the routing matrix.
7. The method of claim 1, wherein the corresponding observation matrix of each aggregation is optimized to improve estimation accuracy.
8. The method of claim 1, wherein the defined network corresponds to a data center, each aggregation of network devices corresponds to a rack of servers, and the method is performed at a network monitoring center.
9. A computer system comprising :
a set of memory resources to store a set of instructions to monitor a network of the computer system;
one or more processors to execute the set of instructions to: obtain information which is indicative of (i) a topology of the defined network, the topology including multiple aggregations of network devices, and (ii) routing information that identifies a plurality of data flows amongst network devices of the defined network;
determine, using at least the topology and the routing information, a corresponding observation matrix for each of the multiple aggregations of network devices; and
communicate the corresponding observation matrix to each of the multiple aggregations of network devices in order to obtain, from each aggregation, a compressed representation of the network traffic through individual network devices of that aggregation.
10. The system of claim 9, wherein the one or more processors execute the set of instructions to generate a representation of the network traffic through the defined network using the compressed representation of each of the multiple aggregations of network devices.
11. The system of claim 9, wherein the one or more processors execute the set of instructions to communicate the corresponding observation matrix to each of the multiple aggregations by
distributing the corresponding observation matrix to each of the multiple aggregations of network devices.
12. The system of claim 9, wherein the one or more processors execute the set of instructions to (i) determine a statistical
expectation, using at least the routing information, in order to account for multi-path routing of data provided in part through the network devices of each aggregation, and (ii) from the statistical expectation, determine a set of non-binary coefficients for the corresponding observation matrix of each aggregation, the set of non-binary coefficients representing multi-path routing of data for individual flows.
13. The system of claim 9, wherein the one or more processors optimize the observation matrix of each aggregation for reduction of sparseness.
14. The system of claim 9, wherein the one or more processors are implemented on a network monitor computer, and wherein each of the multiple aggregations corresponds to a rack of servers in a data center.
15. A data center system comprising :
a network monitor computer; a plurality of racks, each rack including an aggregation of multiple servers;
wherein the network monitor computer uses a network connection to each rack to :
determine, using at least a topology and routing information, a corresponding observation matrix for each of the multiple aggregations of network devices; and communicate the corresponding observation matrix to each of the multiple aggregations of network devices in order to obtain, from each aggregation, a compressed representation of the network traffic through individual network devices of that aggregation; and wherein each rack of the plurality of racks operates to :
receive the corresponding observation matrix from the network monitor computer;
apply the observation matrix to network traffic measurements made on the rack for the aggregation of servers, and
send a determination from applying the observation matrix to the traffic measurements to the network monitor computer.
PCT/US2015/043020 2015-07-30 2015-07-30 Traffic on defined network having aggregations of network devices Ceased WO2017019106A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/US2015/043020 WO2017019106A1 (en) 2015-07-30 2015-07-30 Traffic on defined network having aggregations of network devices

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2015/043020 WO2017019106A1 (en) 2015-07-30 2015-07-30 Traffic on defined network having aggregations of network devices

Publications (1)

Publication Number Publication Date
WO2017019106A1 true WO2017019106A1 (en) 2017-02-02

Family

ID=57884985

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2015/043020 Ceased WO2017019106A1 (en) 2015-07-30 2015-07-30 Traffic on defined network having aggregations of network devices

Country Status (1)

Country Link
WO (1) WO2017019106A1 (en)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050190695A1 (en) * 1999-11-12 2005-09-01 Inmon Corporation Intelligent collaboration across network systems
KR20070120737A (en) * 2006-06-20 2007-12-26 경희대학교 산학협력단 Flow-based traffic measurement method and apparatus
JP2008085812A (en) * 2006-09-28 2008-04-10 Oki Electric Ind Co Ltd Network monitoring system, network monitoring method, and network monitoring program
US20100281388A1 (en) * 2009-02-02 2010-11-04 John Kane Analysis of network traffic
JP2012244302A (en) * 2011-05-17 2012-12-10 Nippon Telegr & Teleph Corp <Ntt> Network monitoring apparatus and network monitoring method

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050190695A1 (en) * 1999-11-12 2005-09-01 Inmon Corporation Intelligent collaboration across network systems
KR20070120737A (en) * 2006-06-20 2007-12-26 경희대학교 산학협력단 Flow-based traffic measurement method and apparatus
JP2008085812A (en) * 2006-09-28 2008-04-10 Oki Electric Ind Co Ltd Network monitoring system, network monitoring method, and network monitoring program
US20100281388A1 (en) * 2009-02-02 2010-11-04 John Kane Analysis of network traffic
JP2012244302A (en) * 2011-05-17 2012-12-10 Nippon Telegr & Teleph Corp <Ntt> Network monitoring apparatus and network monitoring method

Similar Documents

Publication Publication Date Title
CN107710238B (en) Deep Neural Network Processing on Hardware Accelerator with Stacked Memory
US10528682B2 (en) Automatic performance characterization of a network-on-chip (NOC) interconnect
Zhang et al. Gemini: Practical reconfigurable datacenter networks with topology and traffic engineering
Zhong et al. An efficient SDN load balancing scheme based on variance analysis for massive mobile users
Huu et al. Modeling and experimenting combined smart sleep and power scaling algorithms in energy-aware data center networks
Cui et al. Cross-platform machine learning characterization for task allocation in IoT ecosystems
US20230300074A1 (en) Optimal Control of Network Traffic Visibility Resources and Distributed Traffic Processing Resource Control System
Meng et al. QoE-driven big data management in pervasive edge computing environment
US9667499B2 (en) Sparsification of pairwise cost information
Gao et al. JCSP: Joint caching and service placement for edge computing systems
Tootaghaj et al. Evaluating the combined impact of node architecture and cloud workload characteristics on network traffic and performance/cost
Mahmoudi et al. MBL-DSDN: a novel load balancing algorithm in distributed software-defined networks based on micro-clustering and B-LSTM methods
Sharma et al. Gpu cluster scheduling for network-sensitive deep learning
Mehta et al. Distributed cost-optimized placement for latency-critical applications in heterogeneous environments
Huo et al. A software‐defined networks‐based measurement method of network traffic for 6G technologies
Zhang et al. Prophet: Toward fast, error-tolerant model-based throughput prediction for reactive flows in dc networks
WO2017019106A1 (en) Traffic on defined network having aggregations of network devices
Tang et al. Modeling and performance analysis of energy harvesting wireless communication systems with reliable energy backup
Sharma et al. A Network Calculus Model for SFC Realization and Traffic Bounds Estimation in Data Centers
Lu et al. High-elasticity virtual cluster placement in multi-tenant cloud data centers
US20200252290A1 (en) Network Bandwidth Configuration
Josephraj et al. TAFLE: Task‐Aware Flow Scheduling in Spine‐Leaf Network via Hierarchical Auto‐Associative Polynomial Reg Net
US20260039557A1 (en) Management of large-scale networks
Ennaceur et al. Engineering edge-cloud offloading of big data for channel modelling in thz-range communications
Żal et al. An energy-efficient control algorithms for switching fabrics

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15899896

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15899896

Country of ref document: EP

Kind code of ref document: A1