WO2025130623A1 - 分布式模型训练的网络拥塞控制方法及相关设备 - Google Patents

分布式模型训练的网络拥塞控制方法及相关设备 Download PDF

Info

Publication number
WO2025130623A1
WO2025130623A1 PCT/CN2024/136899 CN2024136899W WO2025130623A1 WO 2025130623 A1 WO2025130623 A1 WO 2025130623A1 CN 2024136899 W CN2024136899 W CN 2024136899W WO 2025130623 A1 WO2025130623 A1 WO 2025130623A1
Authority
WO
WIPO (PCT)
Prior art keywords
model training
congestion control
task
network
time period
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/136899
Other languages
English (en)
French (fr)
Inventor
汪洋
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
China Telecom Corp Ltd Technology Innovation Center
China Telecom Corp Ltd
Original Assignee
China Telecom Corp Ltd Technology Innovation Center
China Telecom Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by China Telecom Corp Ltd Technology Innovation Center, China Telecom Corp Ltd filed Critical China Telecom Corp Ltd Technology Innovation Center
Publication of WO2025130623A1 publication Critical patent/WO2025130623A1/zh
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L47/00Traffic control in data switching networks
    • H04L47/10Flow control; Congestion control
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/16Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks using machine learning or artificial intelligence

Definitions

  • the present disclosure relates to the field of artificial intelligence technology, and in particular to a network congestion control method, device, equipment, medium and computer program product for distributed model training.
  • RDMA remote direct memory access
  • the model training tasks do not have a high traffic occupation demand for the network at all times. These tasks may generate periodic high bandwidth pressure on the system network.
  • the data synchronization phase of the same model training task may overlap periodically, resulting in periodic network congestion on the shared link, thereby extending the overall training time of each model training task.
  • Network congestion refers to the situation where the network transmission performance decreases due to the limited resources of the storage and forwarding nodes when the number of packets transmitted in the packet switching network is too large, resulting in a continuous overloaded network state, which may cause additional network overhead and reduce the efficiency of model training.
  • a network congestion control method for distributed model training comprising: monitoring a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during the execution of the model training task when the network bandwidth occupancy is higher than a preset threshold; based on a pre-constructed network traffic model of the distributed system and the target model training task, calling corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of existing model training tasks in the distributed system.
  • the monitoring of a target model training task to be executed in a distributed system includes: monitoring whether the distributed system receives the model training task to be executed; if the monitored distributed system receives the model training task to be executed, determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period; if the model training task to be executed has a periodic peak network bandwidth occupancy time period, determining the model training task to be executed as the target model training task.
  • determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period includes: determining the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task according to a preset network congestion control algorithm; and determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period according to the peak network bandwidth occupancy time period of the model training task to be executed.
  • the preset network congestion control algorithm is a Data Center Quantized Congestion Notification (DCQCN) algorithm.
  • DCQCN Data Center Quantized Congestion Notification
  • corresponding network congestion control parameters are called to perform network congestion control on the distributed system, including: obtaining a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; according to the first execution priority and the second execution priority, determining the execution order of the target model training and the existing model training task, and calling corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the model training task with a higher task execution priority is executed first, and the model training task with a higher task execution priority is postponed.
  • the calling of corresponding network congestion control parameters to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system including: the calling of corresponding network congestion control parameters to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.
  • a network congestion control device for distributed model training including: a training task monitoring module, configured to monitor a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during which the network bandwidth occupancy is higher than a preset threshold during the execution of the model training task; a network congestion control module, configured to call corresponding network congestion control parameters to perform network congestion control on the distributed system based on a pre-built network traffic model of the distributed system and the target model training task, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.
  • an electronic device which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the network congestion control method for distributed model training described in any one of the above-mentioned instructions by executing the executable instructions.
  • a computer-readable storage medium on which a computer program is stored.
  • the computer program is executed by a processor, the network congestion control method for distributed model training described in any one of the above is implemented.
  • a computer program product including a computer program, which, when executed by a processor, implements any of the above-mentioned network congestion control methods for distributed model training.
  • the network congestion control method, device, equipment, medium and computer program product for distributed model training provided in the embodiments of the present disclosure, after monitoring the existence of a target model training task to be executed with a periodic peak network bandwidth occupancy time period in the distributed system, combined with a pre-built network traffic model, calls the corresponding network congestion control parameters to control the network congestion of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from that of the existing model training tasks in the distributed system.
  • the embodiments of the present disclosure can reduce the probability of network congestion and improve network utilization by automatically adjusting the network congestion parameters of the model training task according to user needs.
  • FIG1 is a schematic diagram showing an exemplary application system architecture of a network congestion control method for distributed model training in an embodiment of the present disclosure
  • FIG2 shows a schematic diagram of a network perception system in an embodiment of the present disclosure
  • FIG3 shows a flow chart of a network congestion control method for distributed model training in an embodiment of the present disclosure
  • FIG4 shows a flow chart of another network congestion control method for distributed model training in an embodiment of the present disclosure
  • FIG5 shows a schematic diagram of a network congestion control device for distributed model training in an embodiment of the present disclosure
  • FIG6 shows a block diagram of an electronic device in an embodiment of the present disclosure
  • FIG. 7 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure.
  • RDMA Remote Direct Memory Access, remote direct memory access
  • PFC Priority Flow Control, is a flow control technology used to manage the priority of data flows in Ethernet networks
  • DCQCN Data Center Quality of Service Congestion Notification, data center service quality congestion notification, is a technology used for congestion notification and quality of service in data center networks.
  • Figure 1 shows a schematic diagram of an exemplary application system architecture of a network congestion control method for distributed model training in an embodiment of the present disclosure.
  • the system architecture may include a spine device 101, a leaf device 102, a server 103, and a client 104.
  • the spine device 101 can be connected to the first leaf device 1021 and the second leaf device 1022 respectively, and the first leaf device can be connected to the first server 1031, the second server 1032, the third server 1033, and the fourth server 1034 respectively.
  • the first server 1031 is taken as an example, and the first server 1031 is connected to the client 104.
  • the current cluster-based distributed model training system performs end-to-end model training by supporting different frameworks (such as TensorFlow, etc.).
  • frameworks such as TensorFlow, etc.
  • the model training task is carried on multiple nodes of the cluster.
  • the distributed system divides the model training task into smaller batches, performs calculations on each node, and then updates the unified model weights through a synchronization mechanism.
  • a combination of multiple graphics processors (GPUs) on multiple servers can be selected for model training tasks.
  • the model weights are synchronized between multiple servers through the network every time the model undergoes an iteration or a fixed period. Multiple servers usually encounter situations where weights are updated using a shared network link.
  • both the PFC-based port congestion control and the flow-based DCQCN congestion control mechanism have adjustment mechanisms that can set the network bandwidth for specific tasks (such as large model training tasks).
  • a network perception system 105 is added to the network architecture of the related technology.
  • FIG2 shows a schematic diagram of a network perception system in an embodiment of the present disclosure.
  • the system 105 includes: a traffic model building module 1051, a network congestion perception module 1052 and a parameter memory module 1053.
  • the traffic model construction module 1051 has two user interfaces, one interface is oriented to the current distributed system, mainly for system topology perception and network perception of the model training task, and recording the bandwidth sensitivity of each node of the model training task; the other interface is oriented to the model training task, and searches and adapts a part of the hyperparameters of the model training task within a preset area.
  • the traffic model table of the initial model training task in the system is constructed. The user can refer to the traffic model table to intervene in the actual model training plan and select the most appropriate parameter configuration under the current system topology to complete the training of the initial traffic model.
  • the traffic model table can show the traffic conditions of the model training tasks under different parameters, as shown in Table 1:
  • W1 can be a model parameter group consisting of multiple hyperparameters set by default in the system
  • TH1 and LT1 can be the default network throughput and default task latency calculated based on the current default model parameter group
  • W2 can be a model parameter group consisting of multiple hyperparameters of the initial model training task
  • TH2 and LT2 can be the network throughput and task latency calculated based on the model parameter group of the current initial model training task. Compare the effects of the two model parameter groups and select the better data to complete the initial traffic model training.
  • the above-mentioned preset area is determined by selecting the search range based on the hyperparameter customization of the training initial traffic model, and can be adjusted according to actual conditions.
  • the embodiments of the present disclosure do not make specific limitations on this.
  • the network congestion perception module 1052 is used to perceive network communication traffic during the model training task execution phase, including an analysis perception phase of building a traffic model (i.e., an initial traffic model) for existing simulation training tasks in the system at startup, a sampling perception phase of newly connected network traffic (i.e., new model training tasks) during the execution of the model training tasks, and an inherent perception phase of adjusting the network congestion control parameters of the model training tasks.
  • a traffic model i.e., an initial traffic model
  • sampling perception phase of newly connected network traffic i.e., new model training tasks
  • an inherent perception phase of adjusting the network congestion control parameters of the model training tasks.
  • the analysis and perception stage refers to the scenario in which there is a model training task in the system, in which the network usage rate of the model training task is perceived, and a network traffic model of the current model training task is constructed as the initial traffic model of the current system.
  • a preset network congestion control algorithm can be used to perceive the periodic data flow and ignore the bursty data flow.
  • the above-mentioned preset network congestion control algorithm may refer to the DCQCN algorithm, that is, the traffic in the network link is evaluated by analyzing the number of explicit congestion notification ECN marks in the algorithm. It should be noted that any network congestion control algorithm that can perceive the periodic data flow of the model training task can be used, and the present disclosure does not make specific limitations on this.
  • the sampling perception stage begins.
  • a preset network congestion control algorithm is used to perceive the abnormal congestion of network traffic to determine whether the current model training task has periodic data traffic (i.e., periodic peak network bandwidth occupancy time period). If so, the system returns to the inherent perception stage and re-perceives whether there is any new model training task access in the system.
  • the inherent perception stage is used to start the congestion parameter search function. For example, there are two model training tasks in the current system. By adjusting the RDMA network congestion control parameters of a certain node, the network occupancy rate is adjusted to the priority first training task according to a certain ratio, so that the system's network resources are more inclined to complete the first training task when congestion occurs. After adjustment, after several rounds of training iterations, the network congestion perception module determines whether the system actively postpones the time point when the second training task enters the network to occupy high bandwidth (that is, whether it finds a concurrent periodicity suitable for the two model training tasks).
  • the network congestion perception stage recognizes that the system's congestion control reaches the best state for the current situation, the system reaches stability, and the completion time of the model training task is expected to be the shortest.
  • the network occupancy rate of the two tasks on the shared link will be improved in the high occupancy demand stage, similar to shifting the phase of two periodic sine waves to avoid the area where the peaks overlap, reduce the pressure on the overall network bandwidth, and increase the network bandwidth occupied by each.
  • the priority of the model training task can be the task execution priority pre-set by the user. It should be noted that the embodiment of the present disclosure does not make specific limitations on this.
  • the parameter memory module 1053 has two interfaces, one interface for the model training task, and the other interface for the network congestion control parameters during the execution of the model training task.
  • the module is used to register and store the hyperparameters of the model training task, the network congestion control parameters and the network traffic model in the current system, and is also used to provide and configure the hyperparameters of the model training task and the initialization values of the network congestion control parameters.
  • the module can also be opened to users, which is convenient for users to set parameter information in a personalized system, thereby improving the flexibility of the system.
  • the system can provide users with the function of analyzing the network traffic of application services and systems, identify network traffic information for model training tasks, and not specifically specify the system topology of model training jobs, which can help users optimize network configuration in different systems.
  • it also combines business operations and network performance parameters, comprehensively considers the overall training accuracy and training efficiency, provides users with recommended parameters under dual considerations of training operations and network configuration when using the system, and provides a two-dimensional training operation parameter model.
  • FIG1 the number of spine devices, leaf devices, servers, and clients in FIG1 is merely illustrative, and any number of spine devices, leaf devices, servers, and clients may be provided according to actual needs, and the embodiments of the present disclosure are not limited thereto.
  • an embodiment of the present disclosure provides a network congestion control method for distributed model training, which can be executed by any electronic device with computing and processing capabilities.
  • FIG3 shows a flow chart of a network congestion control method for distributed model training in an embodiment of the present disclosure. As shown in FIG3 , the method includes the following steps:
  • the target model training task is a model training task with a periodic peak network bandwidth occupancy time period
  • the peak network bandwidth occupancy time period is a task time period during which the network bandwidth occupancy is higher than a preset threshold during the execution of the model training task.
  • a model training task with a periodic peak network bandwidth occupancy time period is determined as a target model training task, wherein the peak network bandwidth occupancy time period may refer to a task time period during which the network bandwidth occupancy of the model training task is higher than a preset threshold during the execution of the model training task.
  • the preset threshold can be set by the user according to actual conditions, and the present embodiment does not specifically limit this.
  • S304 based on the pre-built network traffic model of the distributed system and the target model training task, call the corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.
  • a pre-built network traffic model of a distributed system may be a network traffic model constructed based on one or more existing model training tasks of periodic peak network bandwidth occupancy time periods in the distributed system; the peak network bandwidth occupancy time period of the model training task is adjusted by calling the corresponding network congestion control parameters in the target model training task to distinguish it from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.
  • the embodiment of the present disclosure detects that there is a target model training task to be executed in the distributed system with a periodic peak network bandwidth occupancy time period, it combines the pre-built network traffic model and calls the corresponding network congestion control parameters to control the network congestion of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from that of the existing model training tasks in the distributed system.
  • the embodiment of the present disclosure can reduce the probability of network congestion and improve network utilization by automatically adjusting the network congestion parameters of the model training task according to user needs.
  • the above S302 includes monitoring whether the distributed system receives the model training task to be executed; if the distributed system is monitored to receive the model training task to be executed, determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period; if the model training task to be executed has a periodic peak network bandwidth occupancy time period, determining the model training task to be executed as the target model training task.
  • the DCQCN algorithm can be used to determine whether there is a periodic peak network bandwidth occupancy time period in the model training task to be executed. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether the model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.
  • a model training task to be executed has a periodic peak network bandwidth occupancy time period, including: determining the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task according to a preset network congestion control algorithm; and determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period according to the peak network bandwidth occupancy time period of the model training task to be executed.
  • the preset network congestion control algorithm is a Data Center Quantized Congestion Notification (DCQCN) algorithm.
  • DCQCN Data Center Quantized Congestion Notification
  • the above S304 includes obtaining a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; according to the first execution priority and the second execution priority, determining the execution order of the target model training and the existing model training tasks, and calling corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the model training tasks with higher task execution priorities are executed first, and the model training tasks with higher task execution priorities are postponed.
  • first execution priority and the second execution priority can be set in advance by the user, and the embodiments of the present disclosure do not specifically limit this.
  • the pre-constructed network traffic model of the distributed system is a network traffic model constructed based on one or more existing model training tasks of periodic peak network bandwidth occupancy time periods in the distributed system.
  • DCQCN can be used to determine whether an existing model training task has a periodic peak network bandwidth occupancy time period. It should be noted that the present embodiment can also adopt any other network congestion control algorithm that can determine whether a model training task has a periodic peak network bandwidth occupancy time period, and the present embodiment does not make specific limitations on this.
  • the above S304 includes: calling corresponding network congestion control parameters to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.
  • the business virtualization that can maximize network performance can be combined to similar hardware implementations to provide users with effective network configuration and system topology recommendations.
  • the cloud service provider under the condition that the cloud service provider provides the same resources, more efficient services are provided to users, and at the same time, it can help and guide users to have more control over the use of resources, thereby improving the user's cognitive experience. It can also help cloud service providers improve the efficiency of user resource allocation. Cloud service providers can provide complementary resources to multiple users according to user needs, thereby saving resource costs.
  • FIG4 shows a flow chart of another network congestion control method for distributed model training in an embodiment of the present disclosure. As shown in FIG4 , the method includes the following steps:
  • an initial model training task received in a distributed system is judged, and an initial network traffic model is trained based on a model training task with a periodic peak network bandwidth occupancy time period, that is, the model training task with a periodic peak network bandwidth occupancy time period is registered in a user-facing interface, while the model training task with only a bursty peak network bandwidth occupancy is not registered.
  • the DCQCN algorithm can be used to determine whether an existing model training task has a periodic peak network bandwidth occupancy time period. It should be noted that the present embodiment can also adopt any other network congestion control algorithm that can determine whether a model training task has a periodic peak network bandwidth occupancy time period, and the present embodiment does not make specific limitations on this.
  • DCQCN can be used to determine whether there is a change in the network traffic execution cycle in the current distributed system to determine whether a new model training task is connected to the system. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether the model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.
  • S405 determine whether there is a periodic peak network bandwidth occupancy time period in the received new model training task. If yes, execute S406; if no, mark the task, do not process it later, and execute S407.
  • DCQCN can be used to determine whether the model training task to be executed has a periodic peak network bandwidth occupancy time period. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether the model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.
  • S406 Search and configure parameters, and optimize the distributed system.
  • the network congestion control parameters of the model training task are adjusted so that the peak network bandwidth occupation time period of the model training task is different from the peak network bandwidth occupation time period of the existing model training tasks in the distributed system.
  • the current network congestion control parameters are kept unchanged, and the hyperparameters in the model training task are searched and configured so that the distributed system reaches the optimal state.
  • the system when multiple model training tasks are executed simultaneously in the system, the system enters a stable periodic operation.
  • the estimated completion time of the model training task is the shortest, it proves that the current distributed system has achieved optimal performance.
  • the module completes the congestion perception task of this stage.
  • the model training it organizes and analyzes the search for hyperparameters, the adjustment of network congestion parameters, and the network traffic model obtained by the final training. The above data is sent to the user as a model training report and reasonable suggestions are made.
  • the monitoring task of the new model training task in the system can be triggered at the same time. At this time, the system continues to monitor whether there is a new model training task.
  • Rt represents the network congestion parameter group during congestion perception
  • Ro represents the congestion parameter group before congestion perception
  • g represents the change rate of congestion control bandwidth, g ⁇ [-1,1].
  • the g value should also take into account the robustness of changes in the entire system;
  • the above-mentioned network congestion parameter group is a general term for a series of network congestion parameters used to adjust the peak network bandwidth occupancy time period of the model training task, such as network bandwidth, etc.
  • the network congestion parameters involved in the actual model training can be added to the network congestion parameter group.
  • the embodiment of the present disclosure does not specifically limit the parameter types contained in the network congestion parameter group.
  • the time difference between tasks before and after parameter search can be reflected by the following formula:
  • ⁇ T represents the time difference between the completion time of the model training task after parameter search configuration and the completion time of the original model training task
  • T ⁇ ( ⁇ ) represents the completion time of the model training task under the current parameter configuration
  • P train_n_nonsense represents the comprehensive value of the network congestion parameters of the model training task after parameter search
  • P train_n_nonsense represents the comprehensive value of the hyperparameters of the model training before parameter search
  • P net_n_nonsense represents the comprehensive value of the network congestion parameters of the model training before parameter search.
  • the above-mentioned network congestion parameter can be the comprehensive value of the network congestion parameters composed of R t , Ro and g at the current data collection moment, which is a constant;
  • the hyperparameter is used to optimize the output effect of the current model after training, and can be, for example, batch size, working nodes, etc.
  • the above-mentioned hyperparameter can be the comprehensive value of a hyperparameter group composed of one or more hyperparameters obtained by adjusting the current network congestion parameters at the current data collection moment, which is a constant.
  • the present disclosure also provides a network congestion control device for distributed model training, as described in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
  • FIG5 shows a schematic diagram of a network congestion control device for distributed model training in an embodiment of the present disclosure.
  • the device includes: a training task monitoring module 501 and a network congestion control module 502 .
  • the training task monitoring module 501 is configured to monitor the target model training tasks to be executed in the distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during the execution of the model training task when the network bandwidth occupancy is higher than a preset threshold;
  • the network congestion control module 502 is configured to call the corresponding network congestion control parameters to perform network congestion control on the distributed system based on a pre-built network traffic model and target model training task of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.
  • the embodiment of the present disclosure detects that there is a target model training task to be executed in the distributed system with a periodic peak network bandwidth occupancy time period, it combines the pre-built network traffic model and calls the corresponding network congestion control parameters to control the network congestion of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from that of the existing model training tasks in the distributed system.
  • the embodiment of the present disclosure can reduce the probability of network congestion and improve network utilization by automatically adjusting the network congestion parameters of the model training task according to user needs.
  • the above-mentioned training task monitoring module 501 is also configured to monitor whether the distributed system receives the model training task to be executed; if the distributed system is monitored to receive the model training task to be executed, it is determined whether the model training task to be executed has a periodic peak network bandwidth occupancy time period; if the model training task to be executed has a periodic peak network bandwidth occupancy time period, the model training task to be executed is determined as the target model training task.
  • the above-mentioned training task monitoring module 501 is also configured to determine the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task according to a preset network congestion control algorithm; and determine whether there is a periodic peak network bandwidth occupancy time period for the model training task to be executed based on the peak network bandwidth occupancy time period of the model training task to be executed.
  • the preset network congestion control algorithm is a Data Center Quantized Congestion Notification (DCQCN) algorithm.
  • DCQCN Data Center Quantized Congestion Notification
  • the above-mentioned network congestion control module 502 is also configured to obtain a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; according to the first execution priority and the second execution priority, the execution order of the target model training and the existing model training tasks is determined, and the corresponding network congestion control parameters are called to perform network congestion control on the distributed system, so that the model training tasks with higher task execution priorities are executed first, and the model training tasks with higher task execution priorities are postponed.
  • the pre-constructed network traffic model of the distributed system is a network traffic model constructed based on one or more existing model training tasks of periodic peak network bandwidth occupancy time periods in the distributed system.
  • the above-mentioned network congestion control module 502 is also configured to call corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.
  • the electronic device 600 according to this embodiment of the present disclosure is described below with reference to Fig. 6.
  • the electronic device 600 shown in Fig. 6 is only an example and should not bring any limitation to the functions and scope of use of the embodiment of the present disclosure.
  • the electronic device 600 is in the form of a general computing device.
  • the components of the electronic device 600 may include but are not limited to: at least one processing unit 610, at least one storage unit 620, and a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610).
  • the processing unit 610 can execute the following steps of the above method embodiment: monitoring a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during the execution of the model training task when the network bandwidth occupancy is higher than a preset threshold; based on a pre-constructed distributed system network traffic model and a target model training task, calling corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.
  • the storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and/or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
  • RAM random access memory unit
  • ROM read-only memory unit
  • the storage unit 620 may also include a program/utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
  • program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
  • Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
  • the electronic device 600 may also communicate with one or more external devices 640 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and/or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input/output (I/O) interface 650.
  • the electronic device 600 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and/or public networks, such as the Internet) via a network adapter 660.
  • networks e.g., local area networks (LANs), wide area networks (WANs), and/or public networks, such as the Internet
  • the network adapter 660 communicates with other modules of the electronic device 600 via a bus 630. It should be understood that, although not shown in the figure, other hardware and/or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
  • the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
  • a non-volatile storage medium which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.
  • a computing device which can be a personal computer, a server, a terminal device, or a network device, etc.
  • the process described in the above reference flowchart can be implemented as a computer program product, which includes: a computer program, which implements the above-mentioned network congestion control method of distributed model training when executed by a processor.
  • a computer-readable storage medium is also provided, and the computer-readable storage medium may be a readable signal medium or a readable storage medium.
  • FIG. 7 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure.
  • a program product 700 capable of implementing the above method of the present disclosure is stored on the computer-readable storage medium.
  • various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above “Exemplary Method” section of this specification according to various exemplary implementations of the present disclosure.
  • a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
  • a readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
  • the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
  • the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages.
  • the program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
  • the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
  • LAN local area network
  • WAN wide area network
  • the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
  • a non-volatile storage medium which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.
  • a computing device which can be a personal computer, a server, a mobile terminal, or a network device, etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Signal Processing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Computing Systems (AREA)
  • Physics & Mathematics (AREA)
  • Data Exchanges In Wide-Area Networks (AREA)

Abstract

本公开提供了一种分布式模型训练的网络拥塞控制方法、装置、设备、介质及计算机程序产品,涉及人工智能技术领域。该方法包括:监测到分布式系统中待执行的目标模型训练任务,其中,目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;基于预先构建的分布式系统的网络流量模型和目标模型训练任务,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。本公开能够降低网络拥塞情况发生的概率,提高网络利用率。

Description

分布式模型训练的网络拥塞控制方法及相关设备
相关申请的交叉引用
本公开要求于2023年12月22日提交的申请号为202311789769.1、名称为“分布式模型训练的网络拥塞控制方法、装置、设备和介质”的中国专利申请的优先权,该中国专利申请的全部内容通过引用全部并入本文。
技术领域
本公开涉及人工智能技术领域,尤其涉及一种分布式模型训练的网络拥塞控制方法、装置、设备、介质及计算机程序产品。
背景技术
由于超大模型的机器学习应用往往需要大量的数据计算,因此,实际应用中为实现超大模型的机器学习应用通常会采用分布式模型进行训练作为工程解决方案。分布式机器学习将模型训练过程分散到多个计算节点上,每个节点都处理部分数据,以此来加速模型训练过程,实现大规模数据集的处理以及训练更复杂的模型。
在基于远程直接内存访问RDMA的分布式系统中,对于网络的拥塞控制通常采用公平策略,虽然RDMA技术提升了分布式系统的网络通信性能,但是RDMA的通用性适配并没有针对机器学习中的应用进行优化,可例如,存在两个模型训练任务的数据流量同时对网络产生压力时,根据RDMA拥塞控制的策略,会公平分配网络资源给不同的训练任务,而这种公平分配也使得两个模型训练任务在执行期间周期性的进入争抢网络资源和释放网络资源状态。特别的,当这两个模型训练任务是同类型任务时,该情况较为明显,系统中的网络拥塞将会使两个任务的执行时间都相对延迟。
相关技术中,实际的机器学习模型训练过程,模型训练任务并不是全时段的对网络有高流量的占用需求,这些任务对系统网络有可能产生周期性的高带宽压力。系统中出现多个模型训练任务时,相同的模型训练任务的数据同步阶段可能周期性的重叠,从而导致共享链路上周期性的网络拥塞,进而延长每个模型训练任务各自的整体训练时长。网络拥塞即在分组交换网络中传送分组的数目太多时,由于存储转发节点的资源有限而造成网络传输性能下降的情况,达成一种持续过载的网络状态,由此可能造成额外的网络开销,降低模型训练的效率。
需要说明的是,在上述背景技术部分公开的信息仅用于加强对本公开的背景的理解,因此可以包括不构成对本领域普通技术人员已知的现有技术的信息。
发明内容
本公开提供一种分布式模型训练的网络拥塞控制方法、装置、设备、介质及计算机程序产品,至少在一定程度上克服相关技术中模型训练时发生网络拥塞情况的问题。
本公开的其他特性和优点将通过下面的详细描述变得显然,或部分地通过本公开的实践而习得。
根据本公开的一个方面,提供了一种分布式模型训练的网络拥塞控制方法,包括:监测到分布式系统中待执行的目标模型训练任务,其中,所述目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,所述峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;基于预先构建的所述分布式系统的网络流量模型和所述目标模型训练任务,调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
在一些实施例中,所述监测到分布式系统中待执行的目标模型训练任务,包括:监测分布式系统是否接收到待执行的模型训练任务;若所述监测分布式系统接收到所述待执行的模型训练任务,则判断所述待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段;若所述待执行的模型训练任务存在周期性峰值网络带宽占用时间段,则将所述待执行的模型训练任务确定为目标模型训练任务。
在一些实施例中,判断所述待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段,包括:根据预设网络拥塞控制算法,确定待执行的模型训练任务在模型训练任务执行过程中峰值网络带宽占用时间段;根据所述待执行的模型训练任务的峰值网络带宽占用时间段,判断所述待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段。
在一些实施例中,所述预设网络拥塞控制算法为数据中心量化拥塞通知DCQCN算法。
在一些实施例中,基于预先构建的所述分布式系统的网络流量模型和所述目标模型训练任务,调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,包括:获取第一执行优先级和第二执行优先级,其中,所述第一执行优先级为目标模型训练任务的任务执行优先级,所述第二执行优先级为已有模型训练任务的任务执行优先级;根据第一执行优先级和第二执行优先级,确定所述目标模型训练和所述已有模型训练任务的执行顺序,并调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述任务执行优先级较高的模型训练任务优先执行,所述任务执行优先级较高的模型训练任务推迟执行。
在一些实施例中,所述预先构建的所述分布式系统的网络流量模型为根据分布式系统中一个或多个已有的周期性峰值网络带宽占用时间段的模型训练任务构建的网络流量模型。
在一些实施例中,所述调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同,包括:所述调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内任意一个已有模型训练任务的峰值网络带宽占用时间段均不同。
根据本公开的另一个方面,还提供了一种分布式模型训练的网络拥塞控制装置,包括:训练任务监测模块,设置为监测到分布式系统中待执行的目标模型训练任务,其中,所述目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,所述峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;网络拥塞控制模块,设置为基于预先构建的所述分布式系统的网络流量模型和所述目标模型训练任务,调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
根据本公开的另一个方面,还提供了一种电子设备,该电子设备包括:处理器;以及存储器,用于存储所述处理器的可执行指令;其中,所述处理器配置为经由执行所述可执行指令来执行上述任意一项所述的分布式模型训练的网络拥塞控制方法。
根据本公开的另一个方面,还提供了一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述任意一项所述的分布式模型训练的网络拥塞控制方法。
根据本公开的另一个方面,还提供了一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现上述任意一项的分布式模型训练的网络拥塞控制方法。
本公开的实施例中提供的分布式模型训练的网络拥塞控制方法、装置、设备、介质及计算机程序产品,在监测到分布式系统中存在周期性峰值网络带宽占用时间段的待执行的目标模型训练任务后,结合预先构建的网络流量模型,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。本公开实施例能够通过将模型训练任务的网络拥塞参数根据用户需求进行自动调整,降低网络拥塞情况发生的概率,提高网络利用率。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,并不能限制本公开。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并与说明书一起用于解释本公开的原理。显而易见地,下面描述中的附图仅仅是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1示出了本公开实施例中分布式模型训练的网络拥塞控制方法的示例性应用系统架构示意图;
图2示出了本公开实施例中一种网络感知系统的示意图;
图3示出本公开实施例中一种分布式模型训练的网络拥塞控制方法流程图;
图4示出本公开实施例中另一种分布式模型训练的网络拥塞控制方法流程图;
图5示出本公开实施例中一种分布式模型训练的网络拥塞控制装置示意图;
图6示出本公开实施例中一种电子设备的框图;
图7示出本公开实施例中一种计算机可读存储介质示意图。
具体实施方式
现在将参考附图更全面地描述示例实施方式。然而,示例实施方式能够以多种形式实施,且不应被理解为限于在此阐述的范例;相反,提供这些实施方式使得本公开将更加全面和完整,并将示例实施方式的构思全面地传达给本领域的技术人员。所描述的特征、结构或特性可以以任何合适的方式结合在一个或更多实施方式中。
此外,附图仅为本公开的示意性图解,并非一定是按比例绘制。图中相同的附图标记表示相同或类似的部分,因而将省略对它们的重复描述。附图中所示的一些方框图是功能实体,不一定必须与物理或逻辑上独立的实体相对应。可以采用软件形式来实现这些功能实体,或在一个或多个硬件模块或集成电路中实现这些功能实体,或在不同网络和/或处理器装置和/或微控制器装置中实现这些功能实体。
为便于理解,在介绍本公开实施例之前,首先对本公开实施例中涉及到的几个名词进行解释如下:
RDMA:Remote Direct Memory Access,远程直接内存访问;
PFC:Priority Flow Control,优先级流量控制,是一种流控制技术,它在以太网网络中用于管理数据流的优先级;
DCQCN:Data Center Quality of Service Congestion Notification,数据中心服务质量拥塞通知,是一种在数据中心网络中用于拥塞通知和质量服务的技术。
下面结合附图,对本公开实施例的具体实施方式进行详细说明。
图1示出了本公开实施例中分布式模型训练的网络拥塞控制方法的示例性应用系统架构示意图。如图1所示,相关技术中,该系统架构可以包括脊设备101、叶设备102、服务器103和客户端104。其中,脊设备101可分别与第一叶设备1021和第二叶设备1022相连,第一叶设备可分别与第一服务器1031、第二服务器1032、第三服务器1033和第四服务器1034相连。此处以第一服务器1031举例,第一服务器1031与客户端104相连。
相关技术中,当前基于集群的分布式模型训练系统通过支持不同的框架(例如张量流TensorFlow等)进行端到端的模型训练,此时,模型训练任务承载在集群的多个节点上,分布式系统将模型训练任务切分成更小的批次,在各自节点进行计算,再通过同步机制更新统一的模型权重。在人脸识别的场景下,模型训练任务作业时可以选择多个服务器上的多个图形处理器GPU组合进行。在常用的分布式框架训练期间,模型每经历一个迭代或固定周期,多个服务器之间通过网络进行模型权重的同步,多个服务器通常会遇到共享网络链路进行权重更新的情况。
在基于RDMA的分布式系统中,对于网络的拥塞控制使用公平策略,但在特殊场景(可例如大模型训练场景)的需求下,基于PFC的端口拥塞控制和基于流的DCQCN拥塞控制机制都有调节机制,能完成对特定任务(可例如大模型训练任务)网络带宽进行设定。
因此,在本公开的一个实施例中,在相关技术的网络架构下,增加了网络感知系统105。具体地,图2示出了本公开实施例中一种网络感知系统的示意图。该系统105包括:流量模型构建模块1051、网络拥塞感知模块1052和参数记忆模块1053。
在本公开的一个实施例中,流量模型构建模块1051拥有两个用户接口,一个接口面向当前的分布式系统,主要是进行模型训练任务的系统拓扑感知和网络感知,记录模型训练任务各节点的带宽敏感度;另一个接口面向模型训练任务,对模型训练任务的一部分超参数进行预设区域内的搜索适配,例如,数据并行的架构中,对模型训练任务的批量大小、工作节点等影像数据流量的超参数进行离散值配置。经过适配不同的超参数组合,构建系统中初始模型训练任务的流量模型表,用户可以参考流量模型表,来干预实际的模型训练计划,选择适合在当前系统拓扑下最合适的参数配置,来完成初始流量模型的训练。
具体地,流量模型表可以展示出不同参数下的模型训练任务的流量情况,具体的,可如表1所示:
表1
其中,W1可以为系统中默认设置的多个超参数构成的模型参数组,TH1和LT1可以为基于当前默认的模型参数组计算得到的默认网络吞吐量和默认任务时延;W2可以为初始模型训练任务的多个超参数构成的模型参数组,TH2和LT2可以为基于当前初始模型训练任务的模型参数组计算得到的网络吞吐量和任务时延。比较两个模型参数组带来的效果,选择更优的数据来完成初始流量模型训练。
需要说明的是,上述预设区域是根据训练初始流量模型的超参数定制化选择搜索范围确定的,可根据实际情况进行调整,本公开实施例对此不做具体限定。
在本公开的一个实施例中,网络拥塞感知模块1052用来在模型训练任务执行阶段进行网络通信流量的感知,其中包括在启动时对系统中已有的模拟训练任务进行流量模型构建(即初始流量模型)的分析感知阶段,在模型训练任务执行期间对新接入的网络流量(即新模型训练任务)的抽样感知阶段,以及对模型训练任务的网络拥塞控制参数进行调整的固有感知阶段。
在本公开的一个实施例中,分析感知阶段指在系统中存在一项模型训练任务的场景下,对该模型训练任务的网络使用率进行感知,构建当前模型训练任务的网络流量模型,作为当前系统的初始流量模型,具体地,可通过一预设网络拥塞控制算法,感知周期性的数据流动,并忽略突发性的数据流动。上述预设网络拥塞控制算法可以指DCQCN算法,即通过分析算法中显式拥塞通知ECN标记数量来评估网络链路中的流量,需要说明的是,可以采用任意可感知到模型训练任务周期性的数据流动的网络拥塞控制算法,本公开对此不做具体限定。
在本公开的一个实施例中,获取初始流量模型后,开始进入抽样感知阶段。当存在新的模型训练任务在系统中运行时,通过一预设网络拥塞控制算法对网络流量的异常拥塞感知,确定当前模型训练任务是否存在周期性的数据流量(即周期性峰值网络带宽占用时间段)。若存在,则系统返回固有感知阶段,同时重新感知系统中是否还存在新模型训练任务接入。
在本公开的一个实施例中,固有感知阶段用于启动拥塞参数搜索功能,例如,当前系统中存在两个模型训练任务,通过调整一个某个节点的RDMA网络拥塞控制参数,将网络的占用率按一定比例调整为优先级第一项训练任务,使得系统的网络资源在发生拥塞时,更偏向于完成第一个训练任务。调整后,经过几轮训练迭代,网络拥塞感知模块判断系统是否主动推迟第二项训练任务进入占用网络高带宽的时间点(即是否找到适合两个模型训练任务的并发周期性)。当网络拥塞感知阶段识别到系统的拥塞控制达到当前情况的最佳状态时,该系统达到稳定,预计模型训练任务的完成时间最短,两项任务对共享链路上的网络占用率,在高占用需求阶段性能都会提升,类似将两个周期性的正弦波通过相位移动,规避波峰叠加的区域,降低整体网络带宽的压力,同时提高各自占用的网络带宽。
需要说明的是,模型训练任务的优先级可为用户预先设定的任务执行优先级,需要说明的是,本公开实施例对此不做具体限定。
在本公开的一个实施例中,参数记忆模块1053拥有两个接口,一个接口面向模型训练任务,另一个接口面向该模型训练任务执行期间的网络拥塞控制参数。该模块用于注册和存储模型训练任务的超参数、网络拥塞控制参数和当前系统中的网络流量模型,还用于提供和配置模型训练任务的超参数、网络拥塞控制参数的初始化值。该模块还可以开放给用户,方便用户在个性化的系统中设定参数信息,提高了系统的灵活性。
需要说明的是,上述初始化值在实际应用中通常由系统提供默认值,也可以由用户进行自主配置,本公开实施例对此不做具体限定。
在本公开的一个实施例中,该系统可为用户提供分析应用业务和系统的网络流量的功能,对模型训练任务进行的网络流量信息识别,并不特别的指定模型训练作业的系统拓扑结构,可以帮助用户在不同系统中进行网络的最优化配置。同时还结合业务作业和网络性能参数,综合考虑整体的训练精度和训练效率,为用户在使用系统时提供训练作业和网络配置双重考虑下的推荐参数,提供两个维度的训练作业参数模型。
本领域技术人员可以知晓,图1中的脊设备、叶设备、服务器和客户端的数量仅仅是示意性的,根据实际需要,可以具有任意数目的脊设备、叶设备、服务器和客户端。本公开实施例对此不作限定。
在上述系统架构下,本公开实施例中提供了一种分布式模型训练的网络拥塞控制方法,该方法可以由任意具备计算处理能力的电子设备执行。
图3示出本公开实施例中一种分布式模型训练的网络拥塞控制方法流程图,如图3所示,该方法包括如下步骤:
S302,监测到分布式系统中待执行的目标模型训练任务,其中,目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段。
在本公开的一个实施例中,一个或多个新接入分布式系统的模型训练任务中,存在周期性峰值网络带宽占用时间段的模型训练任务被确定为目标模型训练任务,其中,峰值网络带宽占用时间段可以指模型训练任务在任务执行期间,该模型训练任务的网络带宽占用高于预设阈值的任务时间段。需要说明的是,预设阈值可由用户根据实际情况进行设定,本公开实施例对此不做具体限定。
S304,基于预先构建的分布式系统的网络流量模型和目标模型训练任务,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
在本公开的一个实施例中,预先构建的分布式系统的网络流量模型可以是根据分布式系统中一个或多个已有的周期性峰值网络带宽占用时间段的模型训练任务构建的网络流量模型;通过调用目标模型训练任务中相应的网络拥塞控制参数来调整该模型训练任务的峰值网络带宽占用时间段,使其与分布式系统内已有模型训练任务的峰值网络带宽占用时间段区别开。
由上述可知,本公开实施例在监测到分布式系统中存在周期性峰值网络带宽占用时间段的待执行的目标模型训练任务后,结合预先构建的网络流量模型,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。本公开实施例能够通过将模型训练任务的网络拥塞参数根据用户需求进行自动调整,降低网络拥塞情况发生的概率,提高网络利用率。
在本公开的一个实施例中,上述S302,包括监测分布式系统是否接收到待执行的模型训练任务;若监测分布式系统接收到待执行的模型训练任务,则判断待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段;若待执行的模型训练任务存在周期性峰值网络带宽占用时间段,则将待执行的模型训练任务确定为目标模型训练任务。
在本公开的一个实施例中,可通过DCQCN算法判断待执行的模型训练任务中是否存在周期性峰值网络带宽占用时间段,需要说明的是,本公开实施例还可采用其他任意可判断模型训练任务是否具有周期性峰值网络带宽占用时间段的网络拥塞控制算法,本公开实施例对此不做具体限定。
在本公开的一个实施例中,判断待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段,包括:根据预设网络拥塞控制算法,确定待执行的模型训练任务在模型训练任务执行过程中峰值网络带宽占用时间段;根据待执行的模型训练任务的峰值网络带宽占用时间段,判断待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段。
在本公开的一个实施例中,预设网络拥塞控制算法为数据中心量化拥塞通知DCQCN算法。
在本公开的一个实施例中,上述S304,包括获取第一执行优先级和第二执行优先级,其中,第一执行优先级为目标模型训练任务的任务执行优先级,第二执行优先级为已有模型训练任务的任务执行优先级;根据第一执行优先级和第二执行优先级,确定目标模型训练和已有模型训练任务的执行顺序,并调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使任务执行优先级较高的模型训练任务优先执行,任务执行优先级较高的模型训练任务推迟执行。
需要说明的是,第一执行优先级和第二执行优先级均可由用户预先进行设定,本公开实施例对此不做具体限定。
在本公开的一个实施例中,预先构建的分布式系统的网络流量模型为根据分布式系统中一个或多个已有的周期性峰值网络带宽占用时间段的模型训练任务构建的网络流量模型。
在本公开的一个实施例中,可通过DCQCN判断已有的模型训练任务是否存在周期性峰值网络带宽占用时间段,需要说明的是,本公开实施例还可采用其他任意可判断模型训练任务是否具有周期性峰值网络带宽占用时间段的网络拥塞控制算法,本公开实施例对此不做具体限定。
在本公开的一个实施例中,上述S304,包括:调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与分布式系统内任意一个已有模型训练任务的峰值网络带宽占用时间段均不同。
在本公开的一个实施例中,还可在多用户的云计算业务场景中,基于不同业务,将可以结合最大化网络性能的业务虚拟化到相近的硬件实现上,给用户提供有效的网络配置和系统拓扑建议。
在本公开的一个实施例中,在云服务商提供相同资源的条件下,为用户提供更高效的服务,同时也能帮助和指导用户对资源的利用有更多的可控性,提升用户认知体验。还可以帮助云服务商提升用户资源的配置效率,云服务商可以根据用户的需求,将可以互补的资源提供给多用户,节省资源成本。
图4示出本公开实施例中另一种分布式模型训练的网络拥塞控制方法流程图,如图4所示,该方法包括如下步骤:
S401,在系统启动时,在流量模型构建模块注册训练任务,对初始网络流量模型进行构建。
在本公开的一个实施例中,对分布式系统中接收到的初始模型训练任务进行判断,基于存在周期性峰值网络带宽占用时间段的模型训练任务训练初始网络流量模型,即对存在周期性峰值网络带宽占用时间段的模型训练任务,在面向用户的接口进行注册,而对仅存在突发性峰值网络带宽占用的模型训练任务不进行注册。
在本公开的一个实施例中,可通过DCQCN算法判断已有的模型训练任务是否存在周期性峰值网络带宽占用时间段,需要说明的是,本公开实施例还可采用其他任意可判断模型训练任务是否具有周期性峰值网络带宽占用时间段的网络拥塞控制算法,本公开实施例对此不做具体限定。
S402,在初始网络流量模型构建完成后,开始进入抽样感知阶段。
S403,判断是否存在新模型训练任务接入当前分布式系统。若是,则执行S404;若否,则执行S410。
在本公开的一个实施例中,可通过DCQCN判断当前分布式系统中网络流量执行周期是否出现变化,来判断系统中是否接入新模型训练任务,需要说明的是,本公开实施例还可采用其他任意可判断模型训练任务是否具有周期性峰值网络带宽占用时间段的网络拥塞控制算法,本公开实施例对此不做具体限定。
S404,进行网络拥塞感知,启动拥塞感知功能。
S405,判断接收到的新模型训练任务中是否存在周期性峰值网络带宽占用时间段。若是,则执行S406;若否,则标记该任务,后续不做处理,并执行S407。
在本公开的一个实施例中,可通过DCQCN判断待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段,需要说明的是,本公开实施例还可采用其他任意可判断模型训练任务是否具有周期性峰值网络带宽占用时间段的网络拥塞控制算法,本公开实施例对此不做具体限定。
S406,进行参数搜索及配置,并对分布式系统进行调优。
在本公开的一个实施例中,按照用户预先设置的任务执行优先级,调整模型训练任务的网络拥塞控制参数,以使模型训练任务的峰值网络带宽占用时间段与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。保持当前网络拥塞控制参数不变,对模型训练任务中的超参数进行搜索并配置,使得分布式系统达到最优状态。
S407,对当前参数进行存储。
S408,判断当前系统是否达到最优。若是,则执行S409;若否,则执行S406。
在本公开的一个实施例中,当多个模型训练任务在系统中同时执行时,系统进入稳定周期性的运行,预计模型训练任务的完成时间最短时,则证明当前分布式系统已达到性能最优。
S409,保持当前参数配置。
S410,判断当前系统是否同时满足系统结束以及当前系统中注册的新任务小于1。若是,则执行S411;若否,则执行S402。
S411,向用户输出模型训练报告。
当系统结束且系统中注册的模型训练任务小于1时,模块完成本阶段的拥塞感知任务,在模型训练期间对超参数的搜索、对网络拥塞参数的调整和最终训练得到的网络流量模型进行整理分析,将上述数据作为模型训练报告发送给用户并提出合理建议。
由上述可知,本公开实施例在监测到分布式系统中存在周期性峰值网络带宽占用时间段的待执行的目标模型训练任务后,结合预先构建的网络流量模型,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。可以进行对模型训练任务的标记和网络监控,在同时执行多个训练任务时,有侧重性的调节网络优先级,在不影响训练业务准确率的情况下,对并发训练任务进行网络流量模型的识别,通过定制化的拥塞控制,即根据系统的推荐隐性使系统进入高利用率模式,达成网络利用率的最大化。
在本公开的一个实施例中,当存在新模型训练任务接入当前分布式系统时,在开始进行网络拥塞情况感知的同时,可同时触发系统中新模型训练任务的监测任务,此时,仍继续监测系统中是否存在新模型训练任务。
在本公开的一个实施例中,可通过如下公式进行网络拥塞感知:
Rt=Ro×(1+g)                 (1)
其中,Rt表示拥塞感知时网络拥塞参数组,Ro表示拥塞感知前的拥塞参数组,g表示拥塞控制带宽的变化率,g∈[-1,1]。
其中,g值越远离0,变化率越高,回到高速带宽的速度越快,需要说明的是,g值同时应考虑到整个系统变化的鲁棒性;上述网络拥塞参数组为用于调整模型训练任务峰值网络带宽占用时间段的一系列网络拥塞参数的总称,可例如网络带宽等,在实际模型训练中涉及的网络拥塞参数均可加入到该网络拥塞参数组中,本公开实施例对网络拥塞参数组中包含的参数类型不做具体限定。
在本公开的一个实施例中,可通过如下公式体现参数搜索前后任务之间的时间差:
其中,ΔT表示参数搜索配置后模型训练任务完成时间与原始模型训练任务完成时间的时间差,T×(·)表示当前参数配置下的模型训练任务的完成时间,表示参数搜索后模型训练任务的超参数综合值,表示参数搜索后模型训练任务的网络拥塞参数综合值,Ptrain_n_nonsense表示参数搜索前模型训练的超参数综合值,Pnet_n_nonsense表示参数搜索前模型训练的网络拥塞参数综合值。
需要说明的是,上述ΔT、T、Ptrain_n_nonsense和Ptrain_n_nonsense均为一个常数;上述网络拥塞参数可以是当前数据采集时刻下由Rt、Ro和g构成的网络拥塞参数的综合值,为一个常数;超参数用于使当前模型训练后输出结果的效果最优,可例如批量大小、工作节点等,上述超参数可以是当前数据采集时刻下基于当前网络拥塞参数进行调节得到的一个或多个超参数构成的超参数组的综合值,为一个常数。
基于同一发明构思,本公开实施例中还提供了一种分布式模型训练的网络拥塞控制装置,如下面的实施例所述。由于该装置实施例解决问题的原理与上述方法实施例相似,因此该装置实施例的实施可以参见上述方法实施例的实施,重复之处不再赘述。
图5示出本公开实施例中一种分布式模型训练的网络拥塞控制装置示意图,如图5所示,该装置包括:训练任务监测模块501和网络拥塞控制模块502。
其中,训练任务监测模块501,设置为监测到分布式系统中待执行的目标模型训练任务,其中,目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;网络拥塞控制模块502,设置为基于预先构建的分布式系统的网络流量模型和目标模型训练任务,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
由上述可知,本公开实施例在监测到分布式系统中存在周期性峰值网络带宽占用时间段的待执行的目标模型训练任务后,结合预先构建的网络流量模型,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。本公开实施例能够通过将模型训练任务的网络拥塞参数根据用户需求进行自动调整,降低网络拥塞情况发生的概率,提高网络利用率。
在本公开的一个实施例中,上述训练任务监测模块501,还设置为监测分布式系统是否接收到待执行的模型训练任务;若监测分布式系统接收到待执行的模型训练任务,则判断待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段;若待执行的模型训练任务存在周期性峰值网络带宽占用时间段,则将待执行的模型训练任务确定为目标模型训练任务。
在本公开的一个实施例中,上述训练任务监测模块501,还设置为根据预设网络拥塞控制算法,确定待执行的模型训练任务在模型训练任务执行过程中峰值网络带宽占用时间段;根据待执行的模型训练任务的峰值网络带宽占用时间段,判断待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段。
在本公开的一个实施例中,预设网络拥塞控制算法为数据中心量化拥塞通知DCQCN算法。
在本公开的一个实施例中,上述网络拥塞控制模块502,还设置为获取第一执行优先级和第二执行优先级,其中,第一执行优先级为目标模型训练任务的任务执行优先级,第二执行优先级为已有模型训练任务的任务执行优先级;根据第一执行优先级和第二执行优先级,确定目标模型训练和已有模型训练任务的执行顺序,并调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使任务执行优先级较高的模型训练任务优先执行,任务执行优先级较高的模型训练任务推迟执行。
在本公开的一个实施例中,预先构建的分布式系统的网络流量模型为根据分布式系统中一个或多个已有的周期性峰值网络带宽占用时间段的模型训练任务构建的网络流量模型。
在本公开的一个实施例中,上述网络拥塞控制模块502,还设置为调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与分布式系统内任意一个已有模型训练任务的峰值网络带宽占用时间段均不同。
所属技术领域的技术人员能够理解,本公开的各个方面可以实现为系统、方法或程序产品。因此,本公开的各个方面可以具体实现为以下形式,即:完全的硬件实施方式、完全的软件实施方式(包括固件、微代码等),或硬件和软件方面结合的实施方式,这里可以统称为“电路”、“模块”或“系统”。
下面参照图6来描述根据本公开的这种实施方式的电子设备600。图6显示的电子设备600仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
图6示出本公开实施例中一种电子设备的框图。下面参照图6来描述根据本公开的这种实施方式的电子设备600。图6显示的电子设备600仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图6所示,电子设备600以通用计算设备的形式表现。电子设备600的组件可以包括但不限于:上述至少一个处理单元610、上述至少一个存储单元620、连接不同系统组件(包括存储单元620和处理单元610)的总线630。
其中,所述存储单元存储有程序代码,所述程序代码可以被所述处理单元610执行,使得所述处理单元610执行本说明书上述“示例性方法”部分中描述的根据本公开各种示例性实施方式的步骤。例如,所述处理单元610可以执行上述方法实施例的如下步骤:监测到分布式系统中待执行的目标模型训练任务,其中,目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;基于预先构建的分布式系统的网络流量模型和目标模型训练任务,调用相应的网络拥塞控制参数对分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
存储单元620可以包括易失性存储单元形式的可读介质,例如随机存取存储单元(RAM)6201和/或高速缓存存储单元6202,还可以进一步包括只读存储单元(ROM)6203。
存储单元620还可以包括具有一组(至少一个)程序模块6205的程序/实用工具6204,这样的程序模块6205包括但不限于:操作系统、一个或者多个应用程序、其它程序模块以及程序数据,这些示例中的每一个或某种组合中可能包括网络环境的实现。
总线630可以为表示几类总线结构中的一种或多种,包括存储单元总线或者存储单元控制器、外围总线、图形加速端口、处理单元或者使用多种总线结构中的任意总线结构的局域总线。
电子设备600也可以与一个或多个外部设备640(例如键盘、指向设备、蓝牙设备等)通信,还可与一个或者多个使得用户能与该电子设备600交互的设备通信,和/或与使得该电子设备600能与一个或多个其它计算设备进行通信的任何设备(例如路由器、调制解调器等等)通信。这种通信可以通过输入/输出(I/O)接口650进行。并且,电子设备600还可以通过网络适配器660与一个或者多个网络(例如局域网(LAN),广域网(WAN)和/或公共网络,例如因特网)通信。如图所示,网络适配器660通过总线630与电子设备600的其它模块通信。应当明白,尽管图中未示出,可以结合电子设备600使用其它硬件和/或软件模块,包括但不限于:微代码、设备驱动器、冗余处理单元、外部磁盘驱动阵列、RAID系统、磁带驱动器以及数据备份存储系统等。
通过以上的实施方式的描述,本领域的技术人员易于理解,这里描述的示例实施方式可以通过软件实现,也可以通过软件结合必要的硬件的方式来实现。因此,根据本公开实施方式的技术方案可以以软件产品的形式体现出来,该软件产品可以存储在一个非易失性存储介质(可以是CD-ROM,U盘,移动硬盘等)中或网络上,包括若干指令以使得一台计算设备(可以是个人计算机、服务器、终端装置、或者网络设备等)执行根据本公开实施方式的方法。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机程序产品,该计算机程序产品包括:计算机程序,所述计算机程序被处理器执行时实现上述分布式模型训练的网络拥塞控制方法。
在本公开的示例性实施例中,还提供了一种计算机可读存储介质,该计算机可读存储介质可以是可读信号介质或者可读存储介质。图7示出本公开实施例中一种计算机可读存储介质示意图,如图7所示,该计算机可读存储介质上存储有能够实现本公开上述方法的程序产品700。在一些可能的实施方式中,本公开的各个方面还可以实现为一种程序产品的形式,其包括程序代码,当所述程序产品在终端设备上运行时,所述程序代码用于使所述终端设备执行本说明书上述“示例性方法”部分中描述的根据本公开各种示例性实施方式的步骤。
本公开中的计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。
在本公开中,计算机可读存储介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了可读程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。可读信号介质还可以是可读存储介质以外的任何可读介质,该可读介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。
在一些实施例中,计算机可读存储介质上包含的程序代码可以用任何适当的介质传输,包括但不限于无线、有线、光缆、RF等等,或者上述的任意合适的组合。
在具体实施时,可以以一种或多种程序设计语言的任意组合来编写用于执行本公开操作的程序代码,所述程序设计语言包括面向对象的程序设计语言—诸如Java、C++等,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算设备上执行、部分地在用户设备上执行、作为一个独立的软件包执行、部分在用户计算设备上部分在远程计算设备上执行、或者完全在远程计算设备或服务器上执行。在涉及远程计算设备的情形中,远程计算设备可以通过任意种类的网络,包括局域网(LAN)或广域网(WAN),连接到用户计算设备,或者,可以连接到外部计算设备(例如利用因特网服务提供商来通过因特网连接)。
应当注意,尽管在上文详细描述中提及了用于动作执行的设备的若干模块或者单元,但是这种划分并非强制性的。实际上,根据本公开的实施方式,上文描述的两个或更多模块或者单元的特征和功能可以在一个模块或者单元中具体化。反之,上文描述的一个模块或者单元的特征和功能可以进一步划分为由多个模块或者单元来具体化。
此外,尽管在附图中以特定顺序描述了本公开中方法的各个步骤,但是,这并非要求或者暗示必须按照该特定顺序来执行这些步骤,或是必须执行全部所示的步骤才能实现期望的结果。附加的或备选的,可以省略某些步骤,将多个步骤合并为一个步骤执行,以及/或者将一个步骤分解为多个步骤执行等。
通过以上实施方式的描述,本领域的技术人员易于理解,这里描述的示例实施方式可以通过软件实现,也可以通过软件结合必要的硬件的方式来实现。因此,根据本公开实施方式的技术方案可以以软件产品的形式体现出来,该软件产品可以存储在一个非易失性存储介质(可以是CD-ROM,U盘,移动硬盘等)中或网络上,包括若干指令以使得一台计算设备(可以是个人计算机、服务器、移动终端、或者网络设备等)执行根据本公开实施方式的方法。
本领域技术人员在考虑说明书及实践这里公开的发明后,将容易想到本公开的其它实施方案。本公开旨在涵盖本公开的任何变型、用途或者适应性变化,这些变型、用途或者适应性变化遵循本公开的一般性原理并包括本公开未公开的本技术领域中的公知常识或惯用技术手段。说明书和实施例仅被视为示例性的,本公开的真正范围和精神由所附的权利要求指出。

Claims (11)

  1. 一种分布式模型训练的网络拥塞控制方法,包括:
    监测到分布式系统中待执行的目标模型训练任务,其中,所述目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,所述峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;
    基于预先构建的所述分布式系统的网络流量模型和所述目标模型训练任务,调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
  2. 根据权利要求1所述的分布式模型训练的网络拥塞控制方法,其中,所述监测到分布式系统中待执行的目标模型训练任务,包括:
    监测分布式系统是否接收到待执行的模型训练任务;
    若所述监测分布式系统接收到所述待执行的模型训练任务,则判断所述待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段;
    若所述待执行的模型训练任务存在周期性峰值网络带宽占用时间段,则将所述待执行的模型训练任务确定为目标模型训练任务。
  3. 根据权利要求2所述的分布式模型训练的网络拥塞控制方法,其中,判断所述待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段,包括:
    根据预设网络拥塞控制算法,确定待执行的模型训练任务在模型训练任务执行过程中峰值网络带宽占用时间段;
    根据所述待执行的模型训练任务的峰值网络带宽占用时间段,判断所述待执行的模型训练任务是否存在周期性峰值网络带宽占用时间段。
  4. 根据权利要求3所述的分布式模型训练的网络拥塞控制方法,其中,所述预设网络拥塞控制算法为数据中心量化拥塞通知DCQCN算法。
  5. 根据权利要求1所述的分布式模型训练的网络拥塞控制方法,其中,基于预先构建的所述分布式系统的网络流量模型和所述目标模型训练任务,调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,包括:
    获取第一执行优先级和第二执行优先级,其中,所述第一执行优先级为目标模型训练任务的任务执行优先级,所述第二执行优先级为已有模型训练任务的任务执行优先级;
    根据第一执行优先级和第二执行优先级,确定所述目标模型训练和所述已有模型训练任务的执行顺序,并调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述任务执行优先级较高的模型训练任务优先执行,所述任务执行优先级较高的模型训练任务推迟执行。
  6. 根据权利要求1所述的分布式模型训练的网络拥塞控制方法,其中,所述预先构建的所述分布式系统的网络流量模型为根据分布式系统中一个或多个已有的周期性峰值网络带宽占用时间段的模型训练任务构建的网络流量模型。
  7. 根据权利要求6所述的分布式模型训练的网络拥塞控制方法,其中,所述调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同,包括:
    调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内任意一个已有模型训练任务的峰值网络带宽占用时间段均不同。
  8. 一种分布式模型训练的网络拥塞控制装置,包括:
    训练任务监测模块,设置为监测到分布式系统中待执行的目标模型训练任务,其中,所述目标模型训练任务为存在周期性峰值网络带宽占用时间段的模型训练任务,所述峰值网络带宽占用时间段为模型训练任务执行过程中网络带宽占用高于预设阈值的任务时间段;
    网络拥塞控制模块,设置为基于预先构建的所述分布式系统的网络流量模型和所述目标模型训练任务,调用相应的网络拥塞控制参数对所述分布式系统进行网络拥塞控制,以使所述目标模型训练任务的峰值网络带宽占用时间段与所述分布式系统内已有模型训练任务的峰值网络带宽占用时间段不同。
  9. 一种电子设备,包括:
    处理器;以及
    存储器,用于存储所述处理器的可执行指令;
    其中,所述处理器配置为经由执行所述可执行指令来执行权利要求1~7中任意一项所述的分布式模型训练的网络拥塞控制方法。
  10. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现权利要求1~7中任意一项所述的分布式模型训练的网络拥塞控制方法。
  11. 一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现权利要求1~7中任意一项所述的分布式模型训练的网络拥塞控制方法。
PCT/CN2024/136899 2023-12-22 2024-12-04 分布式模型训练的网络拥塞控制方法及相关设备 Pending WO2025130623A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202311789769.1 2023-12-22
CN202311789769.1A CN117544565A (zh) 2023-12-22 2023-12-22 分布式模型训练的网络拥塞控制方法、装置、设备和介质

Publications (1)

Publication Number Publication Date
WO2025130623A1 true WO2025130623A1 (zh) 2025-06-26

Family

ID=89784396

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/136899 Pending WO2025130623A1 (zh) 2023-12-22 2024-12-04 分布式模型训练的网络拥塞控制方法及相关设备

Country Status (2)

Country Link
CN (1) CN117544565A (zh)
WO (1) WO2025130623A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120639639A (zh) * 2025-07-11 2025-09-12 浙江平慧信息技术有限公司 一种自构建网络搭建方法及系统

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117544565A (zh) * 2023-12-22 2024-02-09 中国电信股份有限公司技术创新中心 分布式模型训练的网络拥塞控制方法、装置、设备和介质
CN119149254A (zh) * 2024-11-19 2024-12-17 山东海量信息技术研究院 分布式计算系统的训练方法、装置、程序产品及介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111131080A (zh) * 2019-12-26 2020-05-08 电子科技大学 分布式深度学习流调度方法、系统、设备
US20220012642A1 (en) * 2020-07-09 2022-01-13 International Business Machines Corporation Dynamic network bandwidth in distributed deep learning training
CN114089889A (zh) * 2021-02-09 2022-02-25 京东科技控股股份有限公司 模型训练方法、装置以及存储介质
CN115220899A (zh) * 2022-08-20 2022-10-21 抖音视界有限公司 模型训练任务的调度方法、装置及电子设备
CN117544565A (zh) * 2023-12-22 2024-02-09 中国电信股份有限公司技术创新中心 分布式模型训练的网络拥塞控制方法、装置、设备和介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111131080A (zh) * 2019-12-26 2020-05-08 电子科技大学 分布式深度学习流调度方法、系统、设备
US20220012642A1 (en) * 2020-07-09 2022-01-13 International Business Machines Corporation Dynamic network bandwidth in distributed deep learning training
CN114089889A (zh) * 2021-02-09 2022-02-25 京东科技控股股份有限公司 模型训练方法、装置以及存储介质
CN115220899A (zh) * 2022-08-20 2022-10-21 抖音视界有限公司 模型训练任务的调度方法、装置及电子设备
CN117544565A (zh) * 2023-12-22 2024-02-09 中国电信股份有限公司技术创新中心 分布式模型训练的网络拥塞控制方法、装置、设备和介质

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120639639A (zh) * 2025-07-11 2025-09-12 浙江平慧信息技术有限公司 一种自构建网络搭建方法及系统

Also Published As

Publication number Publication date
CN117544565A (zh) 2024-02-09

Similar Documents

Publication Publication Date Title
CN110505099B (zh) 一种基于迁移a-c学习的服务功能链部署方法
WO2025130623A1 (zh) 分布式模型训练的网络拥塞控制方法及相关设备
WO2020038175A1 (en) Data-stream allocation method for link aggregation and related devices
CN112737823A (zh) 一种资源切片分配方法、装置及计算机设备
CN107645407B (zh) 一种适配QoS的方法和装置
EP3371938B1 (en) Adaptive subscriber-driven resource allocation for push-based monitoring
CN112887217B (zh) 控制数据包发送方法、模型训练方法、装置及系统
CN113315806B (zh) 一种面向云网融合的多接入边缘计算架构
JP7661629B2 (ja) データ伝送制御方法、装置、コンピュータ機器、及びコンピュータプログラム
JP6137657B2 (ja) 帯域調整方法および帯域調整コントローラ
CN107846371B (zh) 一种多媒体业务QoE资源分配方法
CN112887156B (zh) 一种基于深度强化学习的动态虚拟网络功能编排方法
US11252078B2 (en) Data transmission method and apparatus
WO2020026018A1 (zh) 文件的下载方法、装置、设备/终端/服务器及存储介质
CN117499403A (zh) 算力网络的计算任务调度方法及装置
CN112714081A (zh) 一种数据处理方法及其装置
US20230156520A1 (en) Coordinated load balancing in mobile edge computing network
CN116896511B (zh) 一种专线入云的服务限速方法、装置、设备及存储介质
WO2025246804A1 (zh) 拥塞控制方法、装置、电子设备、计算机可读介质及计算机程序产品
CN120186086A (zh) 数据传输方法、系统、设备、存储介质及程序产品
CN108965025A (zh) 云计算系统中流量的管理方法和装置
WO2023198174A1 (en) Methods and systems for predicting sudden changes in datacenter networks
CN110572851A (zh) 一种数据上传方法、系统、装置及计算机可读存储介质
CN118828959A (zh) 接入业务的atsss策略确定方法、装置、设备及介质
CN116471506A (zh) 光网络带宽资源分配方法、装置、电子设备及介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24906107

Country of ref document: EP

Kind code of ref document: A1