WO2017177567A1 - 云调度器中应对不确定需求的多资源调度方法 - Google Patents

云调度器中应对不确定需求的多资源调度方法 Download PDF

Info

Publication number
WO2017177567A1
WO2017177567A1 PCT/CN2016/089898 CN2016089898W WO2017177567A1 WO 2017177567 A1 WO2017177567 A1 WO 2017177567A1 CN 2016089898 W CN2016089898 W CN 2016089898W WO 2017177567 A1 WO2017177567 A1 WO 2017177567A1
Authority
WO
WIPO (PCT)
Prior art keywords
resource
value
cloud
indicates
task
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/089898
Other languages
English (en)
French (fr)
Inventor
姚建国
马汝辉
徐鑫
管海兵
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Shanghai Jiao Tong University
Original Assignee
Shanghai Jiao Tong University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Shanghai Jiao Tong University filed Critical Shanghai Jiao Tong University
Priority to US16/093,585 priority Critical patent/US11157327B2/en
Publication of WO2017177567A1 publication Critical patent/WO2017177567A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/50Allocation of resources, e.g. of the central processing unit [CPU]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/50Allocation of resources, e.g. of the central processing unit [CPU]
    • G06F9/5083Techniques for rebalancing the load in a distributed system
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/48Program initiating; Program switching, e.g. by interrupt
    • G06F9/4806Task transfer initiation or dispatching
    • G06F9/4843Task transfer initiation or dispatching by program, e.g. task dispatcher, supervisor, operating system
    • G06F9/4881Scheduling strategies for dispatcher, e.g. round robin, multi-level priority queues
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/04Network management architectures or arrangements
    • H04L41/042Network management architectures or arrangements comprising distributed management centres cooperatively managing the network
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/08Configuration management of networks or network elements
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L67/00Network arrangements or protocols for supporting network services or applications
    • H04L67/01Protocols
    • H04L67/10Protocols in which an application is distributed across nodes in the network
    • H04L67/1001Protocols in which an application is distributed across nodes in the network for accessing one among a plurality of replicated servers
    • H04L67/1004Server selection for load balancing
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L67/00Network arrangements or protocols for supporting network services or applications
    • H04L67/01Protocols
    • H04L67/10Protocols in which an application is distributed across nodes in the network
    • H04L67/1097Protocols in which an application is distributed across nodes in the network for distributed storage of data in networks, e.g. transport arrangements for network file system [NFS], storage area networks [SAN] or network attached storage [NAS]

Definitions

  • the present invention relates to the field of cloud resource scheduling, and in particular to a multi-resource scheduling method for responding to uncertain requirements in a cloud scheduler.
  • Cloud computing uses a new computing model and network service model to handle computing tasks in the data center, so that a large number of users can remotely access a variety of computing resources, such as CPU, GPU, memory, storage space and network bandwidth. Through virtualization, isolation, and other technologies, these computing resources can be distributed to cloud applications running by users in accordance with on-demand policies.
  • Multi-resource allocation technology is a key technology in cloud computing, because the efficiency and fairness of resource allocation directly affect the performance and economic benefits of the entire cloud computing platform.
  • Resource allocation is the process of efficiently and fairly distributing available physical resources (including computing, storage, and network) to remote cloud users through the network. Due to the uncertainty of resource demand and supply, designing an optimal resource allocation strategy faces great challenges. From the perspective of cloud service providers, in order to maximize resource utilization and economic profit, cloud computing resources cannot be fully provided to cloud users. From the perspective of cloud users, in order to complete the tasks in the cloud application on time, the estimation of cloud resources may exceed the actual demand.
  • DPF Dominant Resource Fairness
  • the DRF focuses on the high fairness of resource allocation, but this may lead to a decrease in efficiency, so a unified trade-off efficiency and fair multi-resource allocation framework is proposed (C.Joe-Wong, S.Sen, T.Lan , and M. Chiang, "Multi-resource allocation: Fairness-efficiency tradeoffs in a unifying framework," in Proc. of IEEE INFOCOM 2012. IEEE, March 2012, pp. 1206-1214.).
  • This framework contains two sets of formulas for modeling fairness, namely Fairnesson Dominant Shares (FDS) and Generalized Fairness on Jobs (GFJ).
  • FDS The primary resource is the resource that allocates the most resources to the cloud user, which is the most important resource we should pay attention to.
  • the j-th cloud user For the i-th resource, the j-th cloud user needs the i-th resource of the number R ij to process a task. make ⁇ ij is expressed as a task to be processed, and the jth cloud user needs the sharing ratio of the i-th resource. Then, the maximum share ratio ⁇ j required for each task of the jth cloud user can be expressed as: Then ⁇ j x j is the main resource sharing ratio of the jth cloud user.
  • ⁇ and ⁇ are the selected parameters.
  • the parameter ⁇ gives the type of fairness
  • the parameter ⁇ represents the degree of attention to efficiency
  • the larger absolute value of ⁇ means more emphasis on efficiency than fairness.
  • a calculation formula indicating the fairness of the main resource sgn( ⁇ ) represents a symbol function
  • represents a real number
  • ⁇ j represents a maximum sharing ratio of resources required for each task of the jth cloud user
  • x j represents an allocation to the jth
  • the number of tasks of cloud users ⁇ k represents the maximum sharing ratio of resources required for each task of the kth cloud user
  • x k represents the number of tasks assigned to the kth cloud user
  • n represents the number of cloud users.
  • GFJ GFJ
  • the existing traditional resource allocation strategy does not take into account the cloud users' demand for resources is uncertain.
  • the cloud application's demand for resources is dynamically changed due to workload changes.
  • the current multi-resource allocation strategy is based on the assumption that the cloud user's task requires resources. It is static and unchanged. Therefore, the existing strategies lack robustness. In a real application environment, these strategies calculate an impractical allocation scheme, resulting in a large performance loss.
  • this strategy calculate an impractical allocation scheme, resulting in a large performance loss.
  • the cloud platform can provide a total of 9 vCPUs and 18 vGPUs for two cloud users.
  • Each task of cloud user 1 requires 1 vCPU and 4 vGPUs.
  • Each task of cloud user 2 requires 3 vCPUs and 1 vGPU. .
  • each task of cloud user 1 needs to use a total of 1/9 of vCPU and 2/9 of vGPU, so the cloud user 1's primary resource is vGPU; likewise, the primary resource corresponding to cloud user 2 is vCPU, because each of his tasks requires the use of a total of 1/3 of the vCPU and 1/18 of the vGPU. Therefore, DRF will allocate 3 tasks to cloud user 1, including 3 vCPUs and 12 vGPUs, and allocate 2 tasks to cloud user 2, including 6 vCPUs and 2 vGPUs, so that a total of 5 tasks can be processed (see figure). 1 (a)), the main resource sharing ratio of the two cloud users here is 2/3.
  • an object of the present invention is to provide a multi-resource scheduling method for responding to uncertain demands in a cloud scheduler.
  • a multi-resource scheduling method for responding to uncertain requirements in a cloud scheduler includes:
  • Step 1 Set the following parameters:
  • the initial value of the obstacle factor ⁇ is ⁇ 0 ;
  • f(X) represents a fairness function corresponding to the optimization target
  • X represents the number of task assignments
  • ⁇ k represents a value corresponding to the obstacle factor ⁇ at the kth outer loop
  • h i (X) represents a slack variable corresponding to the i-th resource
  • N represents a natural number set
  • P i represents a demand change vector corresponding to the i-th resource
  • x j represents the number of tasks assigned to the jth cloud user
  • the initial search point X 0 * of f(X) is set to X 0 ;
  • the initial search point x 0 * is set to X 0 ;
  • the outer loop iteration number k is assigned a value of 0;
  • Step 3 Assign the flag newtonFlag to fail, where the flag newtonFlag is a flag indicating whether the Newton iteration method can calculate the result, and fail indicates no;
  • Step 4 If it is a positive definite matrix, go to step 5, otherwise go to step 8; Indicates secondary derivation, Representing the k-th outer loop, corresponding to the obstacle function whose task allocation quantity is X k , and X k represents the task allocation quantity calculated by the k-th iteration;
  • Step 5 Assign the inner loop iteration number s to 0;
  • Step 6 Assign the search point x s+1 * of the s+1th inner loop of f(X) Increase the value of s by one; Indicates a derivation;
  • Step 7 If Or s>MaxIter, go to step 8, otherwise go to step 6; Indicates the number of task assignments calculated in the sth inner loop. Indicates the number of task assignments calculated in the s-1th inner loop. Indicates that in the kth outer loop, the number of tasks assigned is Barrier function, Indicates that in the kth outer loop, the number of tasks assigned is The obstacle function, TolFun represents the function tolerance value of the iterative termination, TolX represents the number of task assignments for the iteration termination X tolerance value, and MaxIter represents the maximum number of iterations;
  • Step 8 If s ⁇ MaxIter, assign newtonFlag to success, go to step 9; if s>MaxIter, then go to step 9; where success is YES;
  • Step 9 If newtonFlag is equal to fail, go to step 10, otherwise go to step 15;
  • Step 10 Assign the inner loop iteration number s to 0, and assign the search direction d k corresponding to the kth outer loop
  • Step 11 If s>0, assign d s to the value Go to step 12; otherwise, go to step 12; where d s-1 represents the search direction corresponding to the s-1th inner loop;
  • Step 12 Calculate the search step size ⁇ s using the exact linear search method and assign x s+1 * Increase the value of s by one; Indicates the number of task assignments calculated in the s+1th inner loop; where d s represents the search direction corresponding to the sth inner loop; for example, the search step size ⁇ s can be calculated by an accurate linear search method, where
  • a linear search method refer to "Newton's method with exact line search for solving the algebraic Riccati equation," Fakult ⁇ at f ⁇ ur Mathematik, TU Chemnitz Zwickau, 09107 Chemnitz, FRG, Tech. Rep. SPC 95-24, 1995; available as SPC95_24 .ps by anonymous ftp from ftp.tu-chemnitz.de, directory/pub/Local/mathematik/Benner;
  • Step 13 If Or s>MaxIter, go to step 14, otherwise go to step 11;
  • Step 14 If s>MaxIter, assign k to 0, go to step 17; if s ⁇ MaxIter, continue to step 15;
  • Step 15 Will Assignment ⁇ k+1 is assigned to r* ⁇ k , which increases the value of k by 1; Indicates the number of task assignments calculated by the k+1th outer loop, ⁇ k+1 represents the obstacle factor in the k+1th outer loop, and ⁇ k represents the obstacle factor in the kth outer loop;
  • Step 16 If Or k>MaxIter, go to step 17, otherwise go to step 3; Indicates the number of task assignments calculated for the kth outer loop. Indicates the number of task assignments calculated for the k-1th outer loop. Representation corresponds to Fair value, Representation corresponds to Fair value
  • Step 17 Output the optimal fair value Optimal number of task assignments for cloud users
  • the multi-resource scheduling method is based on an ellipsoid uncertainty model
  • the ellipsoid uncertainty model namely:
  • U i represents a demand set corresponding to the i-th resource
  • P i represents a demand change vector corresponding to the i-th resource
  • the method further comprises the steps of:
  • Step 18 follow the optimal fair value Optimal number of task assignments for cloud users Perform resource scheduling.
  • the decreasing coefficient r takes a value of 0.1
  • the function tolerance value TolFun of the iterative termination is equal to 1e-6;
  • the iterative termination task assignment quantity tolerance value TolX is equal to 1e-10;
  • MaxIter The maximum number of iterations, is equal to 1000.
  • the present invention has significant beneficial effects. Specifically, the present invention should have the following three outstanding contributions to the multi-resource allocation strategy for dynamically changing demands:
  • the present invention proposes an ellipsoidal uncertainty model that can be used to capture characteristics of uncertain resource requirements.
  • the present invention lists representation formulas for non-linear optimization problems such as multi-resource allocation in the cloud scheduler that take into account dynamic changes in demand.
  • the resource allocation scheme calculated by the present invention can still obtain optimal fairness and efficiency under the condition of dynamic change of demand, and has good robustness.
  • Figure 1 and Figure 2 show the number of task assignments calculated using the DRF in the comparison scenario before and after the resource demand change.
  • Figure 1 corresponds to scenario 1 with constant resource requirements
  • Figure 2 corresponds to scenario 2 with variable resource requirements.
  • 3 is a multi-resource allocation management architecture illustrating the robustness of the present invention.
  • FIG. 4 and 5 are three-dimensional diagrams of the fair function, FIG. 4 is an FDS, and FIG. 5 is a GFJ.
  • Figures 6 and 7 show the constraint equation and the minimum point.
  • Figure 6 shows the change in the ellipsoid without considering the change in demand.
  • Figure 8 is a graph of fair value versus P in an ellipsoid uncertainty model.
  • Figure 9 and Figure 10 are graphs showing the number of tasks assigned to two cloud users as a function of P in the ellipsoid uncertainty model.
  • Figure 9 shows the FDS and
  • Figure 10 shows the GFJ.
  • Figure 11 and Figure 12 are graphs of the remaining numbers of vCPU and vGPU as a function of P in the ellipsoid uncertainty model, Figure 11 is the FDS, and Figure 12 is the GFJ.
  • the present invention provides a new multi-resource allocation strategy for the cloud scheduler capable of handling dynamic change requirements.
  • the strategy uses two calculation formulas for fairness efficiency, namely FDS and GFJ, as cost functions in the optimization problem.
  • FDS and GFJ calculation formulas for fairness efficiency
  • the robustness equivalence of the original nonlinear optimization problem is easy to calculate, so the present invention models the features of these sets of resource uncertainties, namely the ellipsoid uncertain model.
  • the model places each coefficient vector in a super-ellipsoidal space and serves as a measure of the magnitude of the measurement uncertainty.
  • the invention mainly relates to three main inventions:
  • the robust multi-resource allocation management framework is improved by the traditional resource allocation architecture.
  • the robust multi-resource allocation module is added to cope with the dynamic changes of cloud users' resource requirements.
  • the cost function adopts FDS and GFJ, taking the fairness and efficiency of resource allocation into consideration, and the quantity of resources as a constraint avoids excessive or too little allocation of resources.
  • the ellipsoid uncertainty model can capture the characteristics of demand changes, making the robust equivalence of nonlinear optimization problems easy to calculate, and then calculate the resource allocation scheme that can cope with the dynamic changes of demand.
  • the resource allocation scheme calculated by FDS and GFJ can cope with the dynamic changes of cloud users' demand for resources, and can achieve high efficiency and good performance. Fairness.
  • each task corresponds to a monitor that monitors the operational parameters of the corresponding task and communicates with the scheduler.
  • the scheduler serves two purposes: (1) accepting operational parameters from the monitor, and using these operational parameters as input to the scheduling calculation; (2) automatically scheduling the tasks in the composite task, and using the scheduling of the tasks as Output.
  • the goal of the cloud scheduler in the present invention is to allocate sufficient resources to the cloud users, even if the bottleneck resources are not less than that required by the cloud users.
  • Multi-resource allocation is a nonlinear optimization problem, expressed as an inequality in the form of:
  • f(x) represents the fairness function corresponding to the optimization target
  • Indicates minimization means value, a formula that represents the fairness of the primary resource, A formula for calculating the general fairness of a task, st represents a constraint condition, m represents the number of types of resources, C i represents the total amount of the i-th resource, and R i represents a demand of the user task for the i-th resource;
  • R i [R i1 ,R i2 ,...,R ij ]
  • R ij represents the number of needs of the jth cloud user for the i-th resource
  • x j represents the number of tasks assigned to the j-th cloud user
  • the robust equivalence is always convex (the feasible region is still the intersection of half space, unlike the original nonlinear optimization problem, the feasible region contains solutions that are infinite unless U i is finite ).
  • a feasible set of robust equivalence is the intersection of a single constraint, called a robust half-space constraint. Given a subset U of R n , the robust half-space constraints are expressed as follows:
  • R represents a certain resource requirement
  • R n represents all possible resource requirements
  • C represents the total amount of resources.
  • U represents a collection of resource requirements
  • the maximization problem can be solved explicitly.
  • the ellipsoid uncertainty model has the following form:
  • U represents a set of resource requirements
  • R n ⁇ p represents a set of ellipsoidal resource demand changes. Represents the resource requirement corresponding to the center point of the ellipsoid
  • u represents the obstacle factor vector
  • I represents the imaginary set
  • P represents the uncertain parameter;
  • the ellipsoid uncertainty model can handle the uncertainty of each part of the coefficient vector.
  • the robust equivalents associated with the ellipsoidal model are expressed as follows:
  • This formula uses the Cauchy-Schwartz inequality.
  • U i represents a set of requirements corresponding to the i-th resource, Indicates the i-th resource requirement corresponding to the center point of the ellipsoid
  • P i represents the demand change vector corresponding to the i-th resource
  • u represents a vector
  • m represents the number of types of resources
  • SOCP is a linear programming equivalent to a convex quadratic constraint.
  • the quadratic constraint quadratic programming can be written as SOCP.
  • Semi-definite programming including SOCPs as SOCP constraints can be written as linear matrices (LMI), and can also be rewritten as an example of a semi-determined program. Through the interior point method, SOCPs can be efficiently solved.
  • the multi-resource scheduling method for responding to uncertain requirements in the cloud scheduler provided by the present invention includes:
  • Step 1 Set the following parameters:
  • the initial value of the obstacle factor ⁇ is ⁇ 0 ;
  • f(X) represents a fairness function corresponding to the optimization target
  • X represents the number of task assignments
  • ⁇ k represents a value corresponding to the obstacle factor ⁇ at the kth outer loop
  • h i (X) represents a slack variable corresponding to the i-th resource
  • N represents a natural number set
  • P i represents a demand change vector corresponding to the i-th resource
  • x j represents the number of tasks assigned to the jth cloud user
  • the initial search point X 0 * of f(X) is set to X 0 ;
  • the initial search point x 0 * is set to X 0 ;
  • the outer loop iteration number k is assigned a value of 0;
  • Step 3 Assign the flag newtonFlag to fail, where the flag newtonFlag is a flag indicating whether the Newton iteration method can calculate the result, and fail indicates no;
  • Step 4 If it is a positive definite matrix, go to step 5, otherwise go to step 8; Indicates secondary derivation, Representing the k-th outer loop, corresponding to the obstacle function whose task allocation quantity is X k , and X k represents the task allocation quantity calculated by the k-th iteration;
  • Step 5 Assign the inner loop iteration number s to 0;
  • Step 6 Assign the search point x s+1 * of the s+1th inner loop of f(X) Increase the value of s by one; Indicates a derivation;
  • Step 7 If Or s>MaxIter, go to step 8, otherwise go to step 6; Indicates the number of task assignments calculated in the sth inner loop. Indicates the number of task assignments calculated in the s-1th inner loop. Indicates that in the kth outer loop, the number of tasks assigned is Barrier function, Indicates that in the kth outer loop, the number of tasks assigned is The obstacle function, TolFun represents the function tolerance value of the iterative termination, TolX represents the number of task assignments for the iteration termination X tolerance value, and MaxIter represents the maximum number of iterations;
  • Step 8 If s ⁇ MaxIter, assign newtonFlag to success, go to step 9; if s>MaxIter, then go to step 9; where success is YES;
  • Step 9 If newtonFlag is equal to fail, go to step 10, otherwise go to step 15;
  • Step 10 Assign the inner loop iteration number s to 0, and assign the search direction d k corresponding to the kth outer loop
  • Step 11 If s>0, assign d s to the value Go to step 12; otherwise, go to step 12; where d s-1 represents the search direction corresponding to the s-1th inner loop;
  • Step 12 Calculate the search step size ⁇ s using the exact linear search method and assign x s+1 * Increase the value of s by one; Indicates the number of task assignments calculated in the s+1th inner loop; where d s represents the search direction corresponding to the sth inner loop; for example, the search step size ⁇ s can be calculated by an accurate linear search method, where
  • a linear search method refer to "Newton's method with exact line search for solving the algebraic Riccati equation," Fakult ⁇ at f ⁇ ur Mathematik, TU Chemnitz Zwickau, 09107 Chemnitz, FRG, Tech. Rep. SPC 95-24, 1995; available as SPC95_24 .ps by anonymous ftp from ftp.tu-chemnitz.de, directory/pub/Local/mathematik/Benner;
  • Step 13 If Or s>MaxIter, go to step 14, otherwise go to step 11;
  • Step 14 If s>MaxIter, assign k to 0, go to step 17; if s ⁇ MaxIter, continue to step 15;
  • Step 15 Will Assignment ⁇ k+1 is assigned to r* ⁇ k , which increases the value of k by 1; Indicates the number of task assignments calculated by the k+1th outer loop, ⁇ k+1 represents the obstacle factor in the k+1th outer loop, and ⁇ k represents the obstacle factor in the kth outer loop;
  • Step 16 If Or k>MaxIter, go to step 17, otherwise go to step 3; Indicates the number of task assignments calculated for the kth outer loop. Indicates the number of task assignments calculated for the k-1th outer loop. Representation corresponds to Fair value, Representation corresponds to Fair value
  • Step 17 Output the optimal fair value Optimal number of task assignments for cloud users
  • Step 18 follow the optimal fair value Optimal number of task assignments for cloud users Perform resource scheduling.
  • the multi-resource scheduling method is based on an ellipsoid uncertainty model
  • the ellipsoid uncertainty model namely:
  • U i represents a demand set corresponding to the i-th resource
  • P i represents a demand change vector corresponding to the i-th resource
  • the decreasing coefficient r takes a value of 0.1
  • the function tolerance value TolFun of the iterative termination is equal to 1e-6;
  • the iterative termination task assignment quantity tolerance value TolX is equal to 1e-10;
  • MaxIter The maximum number of iterations, is equal to 1000.
  • the cloud platform provides vCPU and vGPU resources, there are two cloud users who need to run game tasks, and their demand for resources will change. Assume that the task is infinitely separable. To detect performance, we use non-indeterminate parameters as a reference: Cloud User 1 requires 1 vCPU and 4 vGPUs per task, and Cloud User 2 requires 3 vCPUs and 4 vGPUs per task.
  • the cloud platform provides a total of 9 vCPUs and 18 vGPUs.
  • the fairness function is specified as a cost function, namely FDS (formula (1)) and GFJ (formula (2)).
  • Figure 8 shows a three-dimensional diagram of the fair function of the present embodiment, as can be seen from the figure, the FDS and the GFJ are always convex, and The constraints are irrelevant, which means that the optimal FDS and GFJ fair values can always be calculated.
  • Figure 8 shows that when P increases, a better fair value can be obtained. As P increases, the fair value calculated by FDS and GFJ quickly converges to 1.8.

Landscapes

  • Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Data Exchanges In Wide-Area Networks (AREA)

Abstract

本发明提供了一种云调度器中应对不确定需求的多资源调度方法,其使用两个针对公平效率的计算公式,作为优化问题中的成本函数。对于一些资源需求不确定的变化集合,原始非线性优化问题的鲁棒性对等式易于计算,所以本发明对这些资源需求不确定的集合的特征进行了建模,即椭球体不确定模型。该模型将每个系数向量置于一个超椭球形的空间中,并作为测量不确定度大小的一个度量。通过借助于椭球体不确定模型,来解决非线性优化问题,可以得出能够应对于动态变化需求的资源分配方案。

Description

云调度器中应对不确定需求的多资源调度方法 技术领域
本发明涉及云资源调度领域,具体地,涉及云调度器中应对不确定需求的多资源调度方法。
背景技术
云计算采用新型的计算模型和网络服务模型来处理数据中心中的计算任务,这样大量的用户可以远程地获取使用到多样的计算资源,如CPU、GPU、内存、存储空间以及网络带宽等等。通过虚拟化、隔离等技术,这些计算资源可以按照按需供给的策略分配给用户运行的云应用。
多资源分配技术是云计算中的关键技术,因为资源分配的效率和公平性直接影响到整个云计算平台的性能和经济效益。资源分配是通过网络,将可用的物理资源(包括计算、存储、网络)高效公平地分配给远端的云用户的过程。由于资源的需求和供应具有不确定性,设计一个最优的资源分配策略面临很大的挑战。从云服务提供商的角度看,为了最大化资源利用率和经济利润,云计算资源不能被充分地提供给云用户。从云用户的角度考虑,为了使云应用中的任务准时完成,对云资源的估算可能会超过实际的需求。
目前传统的主要有从两种不同角度出发设计的资源分配策略,用来将资源分配给各个云用户运行的云应用:一是只考虑分配单一种类资源,其中有两个被广泛采用的框架,即Hadoop和Dryad,这样的策略将云计算资源分成固定大小的“捆”(或“束”),并以“捆”为分配粒度,这样的“捆”可以看做是单一类型资源的抽象;由于云应用需要多种资源,若只考虑按一种资源类型进行分配,可能会造成严重的效率损失,所以第二类策略考虑到资源的异构性,即按资源的多种类型进行分配。Ghodsi等人提出了基于主资源公平性的多资源分配策略(Dominant Resource Fairness,DRF),其中主资源即分配给云用户的资源中占比最大的资源(A.Ghodsi,M.Zaharia,B.Hindman,A.Konwinski,S.Shenker,andI.Stoica,“Dominant resource fairness:Fair allocation of multipleresource types,”in Proc.of the 8th USENIX Conference on NetworkedSystems Design and Implementation.USENIX Association,2011, pp.323-336)。DRF重点考虑了资源分配的高公平性,但这可能会导致效率的降低,所以一种统一的权衡效率和公平的多资源分配框架被提出(C.Joe-Wong,S.Sen,T.Lan,and M.Chiang,“Multi-resource allocation:Fairness-efficiency tradeoffs in a unifying framework,”in Proc.ofIEEE INFOCOM 2012.IEEE,March 2012,pp.1206-1214.)。这个框架包含两种对公平性进行建模的公式集,即针对主资源的公平性计算公式(Fairnesson Dominant Shares,FDS)和针对任务的一般化公平性计算公式(Generalized Fairness onJobs,GFJ)。
a)FDS:主资源就是分配给云用户的资源中占所提供的资源比率最大的资源,这类资源是我们应该最关注的。
令xj为分配给第j个云用户的任务数量,第i种资源的总量为Ci,则资源数量的约束可以表示为:∑Rijxj≤Ci,其中,Rij表示第j个云用户对第i种资源的需要数量,该不等式可以使用向量改写为:Rix≤Ci,其中Ri=[Ri1,Ri2,...,Rij],x=[x1,x2,...,xj]T,其中,xj表示分配给第j个云用户的任务数量,Ri表示用户任务对第i种资源的需求。对第i种资源,第j个云用户需要数量为Rij的第i种资源来处理一个任务。令
Figure PCTCN2016089898-appb-000001
γij表示为处理一个任务,第j个云用户需要第i种资源的分享比。然后,第j个云用户的每个任务所需要的资源最大分享比μj可以表示为:
Figure PCTCN2016089898-appb-000002
则μjxj即为第j个云用户的主资源分享比。
针对主资源的公平性(FDS)计算公式
Figure PCTCN2016089898-appb-000003
如下所示:
Figure PCTCN2016089898-appb-000004
其中,β∈□和λ∈□是被选中的参数。参数β给出了公平的类型,参数λ代表对效率的关注程度,较大的λ绝对值表示对效率的强调多于公平性。
Figure PCTCN2016089898-appb-000005
表示针对主资源的公平性的计算公式,sgn(·)表示符号函数,□表示实数,μj表示第j个云用户的每个任务所需要的资源最大分享比,xj表示分配给第j个云用户的任务数量,μk表示第k个云用户的每个任务所需要的资源最大分享比,xk表示分配给第k个云用户的任务数量, n表示云用户的数量。
b)GFJ:假设有两个云用户从一个云平台中获取服务,云用户1的每个任务需要的资源多于云用户2的每个任务。从FDS角度看,相比于云用户2,云用户1被分配较少的任务数量会更加公平,并且会提高效率。但是,一般云用户通常只关系自己被分配的任务的数量,而不是不同的资源需求,云用户对公平性的理解即为自己被分配的任务的数量。因此,另一种对公平性的计算公式GFJ被提出,GFJ只使用云用户被分配的任务的数量来测量公平性。
针对任务的一般化公平性(GFJ)计算公式
Figure PCTCN2016089898-appb-000006
如下所示:
Figure PCTCN2016089898-appb-000007
其中,
Figure PCTCN2016089898-appb-000008
表示针对任务的一般化公平性的计算公式。
已有的传统资源分配策略没有考虑到云用户对资源的需求是具有不确定性的。在真实的云平台上,由于工作负载的变化,云应用对资源的需求是动态变化的,但是,目前存在的多资源分配策略的设计基于这样一个假设,即云用户的任务对资源的需求量是静态不变的。因此,已有的这些策略缺乏鲁棒性,在真实的应用环境中,这些策略会计算出不实用的分配方案,造成很大的性能损失。这里举一个例子来进一步说明该问题:
假设有两个云用户,他们需要vCPU和vGPU来运行云游戏服务,如图1所示。云平台总共可以提供9个vCPU和18个vGPU给两个云用户共享,云用户1的每个任务需要1个vCPU和4个vGPU,云用户2的每个任务需要3个vCPU和1个vGPU。我们将这样的资源需求为静态的情况称作为场景1,这里使用DRF作为资源分配策略。
在场景1中,云用户1的每个任务需要使用占总量1/9的vCPU和2/9的vGPU,所以云用户1的主资源为vGPU;同样,对应于云用户2的主资源是vCPU,因为他的每个任务需要使用占总量1/3的vCPU和1/18的vGPU。因此,DRF会分配3个任务给云用户1,包含3个vCPU和12个vGPU,分配2个任务给云用户2,包含6个vCPU和2个vGPU,这样总共可以处理5个任务(参见图1(a)),这里两个云用户的主资源分享比均为2/3。但是,考虑到云用户的资源需求具有不确定性的本质,云用户的资源需求会变化,如云用户2的每个任务需要使用的资源变为3个vCPU和2个vGPU,我们称该资源需求变化后的情况为场景2。在这样的情况中,云用户2的每个任务需要使用占总量1/3的vCPU 和1/9的vGPU,所以云用户2的主资源依旧是vCPU。由于DRF不考虑资源需求的变化,所以DRF不会改变分配给两个云用户的任务数量(参见图1(b)),这样会导致云用户2的任务不能够被分配到足够的资源,即得到的资源少于实际的需求。从而,这样的静态分配策略会给系统的鲁棒性带来严重的损害。
因此,需要设计一个考虑到资源需求动态变化的具有鲁棒性的多资源分配策略,以避免资源的过多或过少配置,从而满足真实的资源需求约束。
发明内容
针对现有技术中的缺陷,本发明的目的是提供一种云调度器中应对不确定需求的多资源调度方法。
根据本发明提供的一种云调度器中应对不确定需求的多资源调度方法,包括:
步骤1:设定如下参数:
第i种资源的总量Ci,i=1,2,...,m,m表示资源的种类数目;
任务分配数量X的初始值X0
定障碍因子μ的初始值μ0
递减系数r;
迭代终止的函数公差值TolFun;
迭代终止的任务分配数量公差值TolX;
最大迭代次数MaxIter;
步骤2:令:
Figure PCTCN2016089898-appb-000009
Figure PCTCN2016089898-appb-000010
x=[x1,x2,...,xj]T
其中,
Figure PCTCN2016089898-appb-000011
表示第k次外循环对应的障碍函数;
f(X)表示优化目标对应的公平性函数;
X表示任务分配数量;
μk表示障碍因子μ在第k次外循环对应的值;
hi(X)表示第i种资源对应的松弛变量;
N表示自然数集;
Figure PCTCN2016089898-appb-000012
表示椭球体中心点对应的第i种资源需求;
Pi表示第i种资源对应的需求变化向量;
xj表示分配给第j个云用户的任务数量;
将f(X)的初始查找点X0 *设为X0
Figure PCTCN2016089898-appb-000013
的初始查找点x0 *设为X0
将外循环迭代次数k赋值为0;
步骤3:将标志newtonFlag赋值为fail,其中,标志newtonFlag是表示牛顿迭代法是否能够计算出结果的标志,fail表示否;
步骤4:如果
Figure PCTCN2016089898-appb-000014
是正定矩阵,则进入步骤5,否则进入步骤8;其中,
Figure PCTCN2016089898-appb-000015
表示二次求导,
Figure PCTCN2016089898-appb-000016
表示第k次外循环中,对应于任务分配数量为Xk的障碍函数,Xk表示第k次迭代计算出的任务分配数量;
步骤5:将内循环迭代次数s赋值为0;
步骤6:将f(X)的第s+1次内循环的查找点xs+1 *赋值为
Figure PCTCN2016089898-appb-000017
令s的值增加1;
Figure PCTCN2016089898-appb-000018
表示一次求导;
步骤7:如果
Figure PCTCN2016089898-appb-000019
或者s>MaxIter,则进入步骤8,否则进入步骤6;其中,
Figure PCTCN2016089898-appb-000020
表示第s次内循环中计算出的任务分配数量,
Figure PCTCN2016089898-appb-000021
表示第s-1次内循环中计算出的任务分配数量,
Figure PCTCN2016089898-appb-000022
表示在第k次外循环中,对应于任务分配数量为
Figure PCTCN2016089898-appb-000023
的障碍函数,
Figure PCTCN2016089898-appb-000024
表示在第k次外循环中,对应于任务分配数量为
Figure PCTCN2016089898-appb-000025
的障碍函数,TolFun表示迭代终止的函数公差值,TolX表示迭代终止的任务分配数量X公差值,MaxIter表示最大迭代次数;
步骤8:如果s≤MaxIter,则将newtonFlag赋值为success,进入步骤9;如果s>MaxIter,则接着执行步骤9;其中,success表示是;
步骤9:如果newtonFlag等于fail,则转入步骤10,否则转入步骤15;
步骤10:将内循环迭代次数s赋值为0,将第k次外循环对应的搜索方向dk赋值为
Figure PCTCN2016089898-appb-000026
步骤11:如果s>0,则将ds赋值为
Figure PCTCN2016089898-appb-000027
进入步骤12;否则,则进入步骤12;其中,ds-1表示第s-1次内循环对应的搜索方向;
步骤12:用精确线性搜索方法计算出搜索步长αs,将xs+1 *赋值为
Figure PCTCN2016089898-appb-000028
令s的值增加1;
Figure PCTCN2016089898-appb-000029
表示第s+1次内循环中计算出的任务分配数量;其中,ds表示第s次内循环对应的搜索方向;例如,可以用精确线性搜索方法计算出搜索步长αs,其中,精确线性搜索方法可参考“Newton’s method with exact line search for solving the algebraic Riccati equation,”Fakult¨at f¨ur Mathematik,TU ChemnitzZwickau,09107 Chemnitz,FRG,Tech.Rep.SPC 95-24,1995;available as SPC95_24.ps by anonymous ftp from ftp.tu-chemnitz.de,directory/pub/Local/mathematik/Benner;
步骤13:如果
Figure PCTCN2016089898-appb-000030
或者s>MaxIter,则进入步骤14,否则进入步骤11;
步骤14:如果s>MaxIter,则将k赋值为0,进入步骤17;如果s≤MaxIter,则继续执行步骤15;
步骤15:将
Figure PCTCN2016089898-appb-000031
赋值为
Figure PCTCN2016089898-appb-000032
μk+1赋值为r*μk,令k的值增加1;其中,
Figure PCTCN2016089898-appb-000033
表示第k+1次外循环计算出的任务分配数量,μk+1表示第k+1次外循环中的障碍因子,μk表示第k次外循环中的障碍因子;
步骤16:如果
Figure PCTCN2016089898-appb-000034
或者k>MaxIter,则进入步骤17,否则进入步骤3;其中,
Figure PCTCN2016089898-appb-000035
表示第k次外循环计算出的任务分配数量,
Figure PCTCN2016089898-appb-000036
表示第k-1次外循环计算出的任务分配数量,
Figure PCTCN2016089898-appb-000037
表示对应于
Figure PCTCN2016089898-appb-000038
的公平值,
Figure PCTCN2016089898-appb-000039
表示对应于
Figure PCTCN2016089898-appb-000040
的公平值;
步骤17:输出最优的公平值
Figure PCTCN2016089898-appb-000041
云用户的最优任务分配数量
Figure PCTCN2016089898-appb-000042
优选地,所述多资源调度方法基于椭球体不确定模型;
椭球体不确定模型,即:
Figure PCTCN2016089898-appb-000043
其中,Ui表示第i种资源对应的需求集合,Pi表示第i种资源对应的需求变化向量。
优选地,还包括步骤:
步骤18:按照最优的公平值
Figure PCTCN2016089898-appb-000044
云用户的最优任务分配数量
Figure PCTCN2016089898-appb-000045
进行资源调度。
优选地,递减系数r取值0.1;
迭代终止的函数公差值TolFun等于1e-6;
迭代终止的任务分配数量公差值TolX等于1e-10;
最大迭代次数MaxIter等于1000。
与现有技术相比,本发明具有显著的有益效果,具体地,本发明应对于动态变化需求的多资源分配策略有如下三点突出贡献:
1、本发明提出了椭球体不确定模型,可以用于捕获不确定资源需求的特征。
2、针对于云调度器中考虑到需求动态变化的多资源分配这样的非线性优化问题,本发明列出了表示公式。
3、本发明计算出的资源分配方案可以在需求动态变化的情况下,依然能够得到最优的公平性和效率,具有良好的鲁棒性。
附图说明
通过阅读参照以下附图对非限制性实施例所作的详细描述,本发明的其它特征、目的和优点将会变得更明显:
图1、图2为资源需求变化前后对比场景中使用DRF计算出的任务分配的数量,图1对应资源需求不变的场景1,图2对应资源需求可变的场景2
图3是图示出本发明鲁棒性的多资源分配管理架构。
图4-12为本发明的一个具体实施例,其中:
图4、图5为公平函数的三维图,图4为FDS,图5为GFJ。
图6、图7为约束方程和最小点,图6为不考虑需求的变化,图7为考虑了椭球形的不确定性。
图8为椭球体不确定模型中,公平值随P变化的曲线图。
图9、图10为椭球体不确定模型中,给两个云用户分配任务的数量随P变化的曲线图,图9为FDS,图10为GFJ。
图11、图12为椭球体不确定模型中,vCPU和vGPU剩余的数量随P变化的曲线图,图11为FDS,图12为GFJ。
具体实施方式
下面结合具体实施例对本发明进行详细说明。以下实施例将有助于本领域的技术人员进一步理解本发明,但不以任何形式限制本发明。应当指出的是,对本领域的普通技 术人员来说,在不脱离本发明构思的前提下,还可以做出若干变化和改进。这些都属于本发明的保护范围。
本发明为云调度器提供了一个新的能够处理动态变化需求的多资源分配策略,该策略使用了两个针对公平效率的计算公式,即FDS和GFJ,作为优化问题中的成本函数。对于一些资源需求不确定的变化集合,原始非线性优化问题的鲁棒性对等式易于计算,所以本发明对这些资源需求不确定的集合的特征进行了建模,即椭球体不确定模型。该模型将每个系数向量置于一个超椭球形的空间中,并作为测量不确定度大小的一个度量。通过借助于椭球体不确定模型,来解决非线性优化问题,可以得出能够应对于动态变化需求的资源分配方案。
本发明主要涉及三个主要的发明点:
(1)鲁棒性的多资源分配管理框架;
(2)对需求动态变化的多资源分配问题规划成非线性优化问题;
(3)椭球体不确定模型。
其中,鲁棒性的多资源分配管理框架由传统的资源分配架构改进而来,在考虑了公平和效率后,加入了鲁棒性的多资源分配模块,可以应对云用户对资源需求动态变化的情况。在规划成的非线性优化问题中,成本函数采用的是FDS和GFJ,将资源分配的公平性和效率都纳入考虑的范围,资源的数量作为约束,避免了资源的过多或过少分配。椭球体不确定模型可以捕获需求变化的特征,使得非线性优化问题的鲁棒对等式易于计算,从而计算出能够应对需求动态变化的资源分配方案。
通过这3个发明点的作用,使得本发明得益于椭球体不确定模型,FDS和GFJ计算出的资源分配方案可以应对云用户对资源需求的动态变化,并可以达到很高的效率和良好的公平性。
在图3的框架中,每个任务(工作负载)都对应有一个监视器,监视器用于监视对应任务的运行参数,并和调度器进行沟通。调度器服务于两个目的:(1)接受从监视器传来的运行参数,将这些运行参数作为调度计算的输入;(2)在复合任务中,自动调度任务,并以对任务的调度作为输出。本发明中云调度器的目标是给云用户分配足够的资源,即使是瓶颈资源也不能少于云用户所需求的。
多资源分配是非线性优化问题,表示成不等式的形式如下:
Figure PCTCN2016089898-appb-000046
Figure PCTCN2016089898-appb-000047
其中,f(x)表示优化目标对应的公平性函数,
Figure PCTCN2016089898-appb-000048
表示最小化,:=表示取值,
Figure PCTCN2016089898-appb-000049
表示针对主资源的公平性的计算公式,
Figure PCTCN2016089898-appb-000050
表示针对任务的一般化公平性的计算公式,s.t.表示约束条件,m表示资源的种类数目,Ci表示第i种资源的总量,Ri表示用户任务对第i种资源的需求;
Ri=[Ri1,Ri2,...,Rij]
x=[x1,x2,...,xj]T
其中,Rij表示第j个云用户对第i种资源的需要数量,xj表示分配给第j个云用户的任务数量;
实际上,资源需求线性约束的数据(包含在向量Ri,i=1,...,m中)一般是未知的。在解决这个非线性优化问题时,如果不把问题数据的不确定性考虑进去将导致计算结果不实用、甚至不可用,因此需要建立鲁棒性的多资源分配优化表达式。
假设非线性优化中的资源需求是未知的,但不确定需求的模型是已知的。这里先列出鲁棒性的多资源分配优化表达式的最简单版本。假设Ri属于第i个给定的集合Ui
Figure PCTCN2016089898-appb-000051
集合Ui可以视为式(3)中线性约束的系数子集。多资源分配的原始非线性优化问题的鲁棒对等式可以有如下表示:
Figure PCTCN2016089898-appb-000052
不论Ui的形状如何,鲁棒对等式总是凸的(可行区域依然是半空间的交集,不同于原始的非线性优化问题,可行区域包含的解是无限的,除非Ui是有限的)。
对于一些类型的不确定集合Ui,鲁棒对等式易于计算,可以对这些类型进行建模。
鲁棒对等式的可行集合是单约束的交集,称为鲁棒半空间约束。给定Rn的子集U,鲁棒半空间约束表示如下:
Figure PCTCN2016089898-appb-000053
不论集合U的选择如何,x上的条件总是凸形的。
其中,R表示某一种资源需求,Rn表示所有可能的资源需求,C表示资源总量,
Figure PCTCN2016089898-appb-000054
表示任意;U表示资源需求集合;
为了分析鲁棒半空间约束,可将其建立成如下的优化问题表达式:
Figure PCTCN2016089898-appb-000055
对于一些集合U的选择,可以明确地解决这个最大化问题。
椭球体不确定模型:
椭球体不确定模型有如下形式:
Figure PCTCN2016089898-appb-000056
其中,P属于Rn×p,是一个矩阵,用于描述椭球体围绕椭球体的中心的形状。如果对于一些ρ≥0,有P=ρI,那么式(7)就是一个简单的半径ρ围绕
Figure PCTCN2016089898-appb-000057
的球体。这样的特殊情况可以称为球形不确定性模型。U表示资源需求集合,Rn×p表示椭球形资源需求变化集合,
Figure PCTCN2016089898-appb-000058
表示椭球体中心点对应的资源需求,u表示障碍因子向量,I表示虚数集;P表示不确定参数;
椭球体不确定模型可以处理系数向量各个部分的不确定性。与椭球形模型相关的鲁棒对等式表示如下:
Figure PCTCN2016089898-appb-000059
这个公式使用到了Cauchy-Schwartz不等式。
有了椭球体不确定模型,即:
Figure PCTCN2016089898-appb-000060
其中,Ui表示第i种资源对应的需求集合,
Figure PCTCN2016089898-appb-000061
表示椭球体中心点对应的第i种资源需求,Pi表示第i种资源对应的需求变化向量,u表示向量,m表示资源的种类数量;
原始非线性优化问题的鲁棒对等式变为:
Figure PCTCN2016089898-appb-000062
变为二阶锥规划(second-order cone program,SOCP):
Figure PCTCN2016089898-appb-000063
SOCP是相当于一个凸二次约束的线性规划,通过调整目标函数作为约束条件,二次约束的二次规划可以写成SOCP。半定规划包括SOCPs作为SOCP约束可以写为线性矩阵不等 式(LMI),并且也可以改写成半定程序的实例。通过内点方法,SOCPs可以被高效地解决。
更为具体地,本发明提供的云调度器中应对不确定需求的多资源调度方法,包括:
步骤1:设定如下参数:
第i种资源的总量Ci,i=1,2,...,m,m表示资源的种类数目;
任务分配数量X的初始值X0
定障碍因子μ的初始值μ0
递减系数r;
迭代终止的函数公差值TolFun;
迭代终止的任务分配数量公差值TolX;
最大迭代次数MaxIter;
步骤2:令:
Figure PCTCN2016089898-appb-000064
Figure PCTCN2016089898-appb-000065
x=[x1,x2,...,xj]T
其中,
Figure PCTCN2016089898-appb-000066
表示第k次外循环对应的障碍函数;
f(X)表示优化目标对应的公平性函数;
X表示任务分配数量;
μk表示障碍因子μ在第k次外循环对应的值;
hi(X)表示第i种资源对应的松弛变量;
N表示自然数集;
Figure PCTCN2016089898-appb-000067
表示椭球体中心点对应的第i种资源需求;
Pi表示第i种资源对应的需求变化向量;
xj表示分配给第j个云用户的任务数量;
将f(X)的初始查找点X0 *设为X0
Figure PCTCN2016089898-appb-000068
的初始查找点x0 *设为X0
将外循环迭代次数k赋值为0;
步骤3:将标志newtonFlag赋值为fail,其中,标志newtonFlag是表示牛顿迭代法是否能够计算出结果的标志,fail表示否;
步骤4:如果
Figure PCTCN2016089898-appb-000069
是正定矩阵,则进入步骤5,否则进入步骤8;其中,
Figure PCTCN2016089898-appb-000070
表示二次求导,
Figure PCTCN2016089898-appb-000071
表示第k次外循环中,对应于任务分配数量为Xk的障碍函数,Xk表示第k次迭代计算出的任务分配数量;
步骤5:将内循环迭代次数s赋值为0;
步骤6:将f(X)的第s+1次内循环的查找点xs+1 *赋值为
Figure PCTCN2016089898-appb-000072
令s的值增加1;
Figure PCTCN2016089898-appb-000073
表示一次求导;
步骤7:如果
Figure PCTCN2016089898-appb-000074
或者s>MaxIter,则进入步骤8,否则进入步骤6;其中,
Figure PCTCN2016089898-appb-000075
表示第s次内循环中计算出的任务分配数量,
Figure PCTCN2016089898-appb-000076
表示第s-1次内循环中计算出的任务分配数量,
Figure PCTCN2016089898-appb-000077
表示在第k次外循环中,对应于任务分配数量为
Figure PCTCN2016089898-appb-000078
的障碍函数,
Figure PCTCN2016089898-appb-000079
表示在第k次外循环中,对应于任务分配数量为
Figure PCTCN2016089898-appb-000080
的障碍函数,TolFun表示迭代终止的函数公差值,TolX表示迭代终止的任务分配数量X公差值,MaxIter表示最大迭代次数;
步骤8:如果s≤MaxIter,则将newtonFlag赋值为success,进入步骤9;如果s>MaxIter,则接着执行步骤9;其中,success表示是;
步骤9:如果newtonFlag等于fail,则转入步骤10,否则转入步骤15;
步骤10:将内循环迭代次数s赋值为0,将第k次外循环对应的搜索方向dk赋值为
Figure PCTCN2016089898-appb-000081
步骤11:如果s>0,则将ds赋值为
Figure PCTCN2016089898-appb-000082
进入步骤12;否则,则进入步骤12;其中,ds-1表示第s-1次内循环对应的搜索方向;
步骤12:用精确线性搜索方法计算出搜索步长αs,将xs+1 *赋值为
Figure PCTCN2016089898-appb-000083
令s的值增加1;
Figure PCTCN2016089898-appb-000084
表示第s+1次内循环中计算出的任务分配数量;其中,ds表示第s次内循环对应的搜索方向;例如,可以用精确线性搜索方法计算出搜索步长αs,其中,精确线性搜索方法可参考“Newton’s method with exact line search for solving the algebraic Riccati equation,”Fakult¨at f¨ur Mathematik,TU ChemnitzZwickau, 09107 Chemnitz,FRG,Tech.Rep.SPC 95-24,1995;available as SPC95_24.ps by anonymous ftp from ftp.tu-chemnitz.de,directory/pub/Local/mathematik/Benner;
步骤13:如果
Figure PCTCN2016089898-appb-000085
或者s>MaxIter,则进入步骤14,否则进入步骤11;
步骤14:如果s>MaxIter,则将k赋值为0,进入步骤17;如果s≤MaxIter,则继续执行步骤15;
步骤15:将
Figure PCTCN2016089898-appb-000086
赋值为
Figure PCTCN2016089898-appb-000087
μk+1赋值为r*μk,令k的值增加1;其中,
Figure PCTCN2016089898-appb-000088
表示第k+1次外循环计算出的任务分配数量,μk+1表示第k+1次外循环中的障碍因子,μk表示第k次外循环中的障碍因子;
步骤16:如果
Figure PCTCN2016089898-appb-000089
或者k>MaxIter,则进入步骤17,否则进入步骤3;其中,
Figure PCTCN2016089898-appb-000090
表示第k次外循环计算出的任务分配数量,
Figure PCTCN2016089898-appb-000091
表示第k-1次外循环计算出的任务分配数量,
Figure PCTCN2016089898-appb-000092
表示对应于
Figure PCTCN2016089898-appb-000093
的公平值,
Figure PCTCN2016089898-appb-000094
表示对应于
Figure PCTCN2016089898-appb-000095
的公平值;
步骤17:输出最优的公平值
Figure PCTCN2016089898-appb-000096
云用户的最优任务分配数量
Figure PCTCN2016089898-appb-000097
步骤18:按照最优的公平值
Figure PCTCN2016089898-appb-000098
云用户的最优任务分配数量
Figure PCTCN2016089898-appb-000099
进行资源调度。
在优选例中,所述多资源调度方法基于椭球体不确定模型;
椭球体不确定模型,即:
Figure PCTCN2016089898-appb-000100
其中,Ui表示第i种资源对应的需求集合,Pi表示第i种资源对应的需求变化向量。
递减系数r取值0.1;
迭代终止的函数公差值TolFun等于1e-6;
迭代终止的任务分配数量公差值TolX等于1e-10;
最大迭代次数MaxIter等于1000。
在本实施例中,云平台提供vCPU和vGPU资源,有两个需要运行游戏任务的云用户,并且他们对资源的需求是会变化的。假设任务是无限可分的。为了检测性能,我们使用非不确定的参数作为参照:云用户1的每个任务需要1个vCPU和4个vGPU,云用户2的每个任务需要3个vCPU和4个vGPU。云平台总共提供9个vCPU和18个vGPU。公平函数指定为成本函数,即FDS(式(1))和GFJ(式(2))。
图8示出本实施例的公平函数三维图,从图中可以看出,FDS和GFJ总是凸的,和 约束无关,这意味着最优的FDS和GFJ公平值总能够被计算出。在本实施例中,参数β=2,λ=-0.5。
椭球形不确定选为Pi=Pi,i=1,2,...,21,其中参数
Figure PCTCN2016089898-appb-000101
其中,Pi表示P的i次方;
通过图6、图7的比较可以发现,相比于传统的资源分配方法,在椭球体不确定模型中,更多的任务可以被分配给云用户。
图8显示当P增大时,更好的公平值可以被获取,随着P的增大,FDS和GFJ计算出的公平值迅速收敛到1.8。
在图9-图12中,随着P的增大,分配给两个云用户的任务数量都逐渐增多,同时剩余的vCPU和vGPU在减少,这意味着在椭球体不确定模型中,FDS和GFJ分配资源策略都可以达到高效率。
以上对本发明的具体实施例进行了描述。需要理解的是,本发明并不局限于上述特定实施方式,本领域技术人员可以在权利要求的范围内做出各种变化或修改,这并不影响本发明的实质内容。在不冲突的情况下,本申请的实施例和实施例中的特征可以任意相互组合。

Claims (4)

  1. 一种云调度器中应对不确定需求的多资源调度方法,其特征在于,包括:
    步骤1:设定如下参数:
    第i种资源的总量Ci,i=1,2,…,m,m表示资源的种类数目;
    任务分配数量X的初始值X0
    定障碍因子μ的初始值μ0
    递减系数r;
    迭代终止的函数公差值TolFun;
    迭代终止的任务分配数量公差值TolX;
    最大迭代次数MaxIter;
    步骤2:令:
    Figure PCTCN2016089898-appb-100001
    Figure PCTCN2016089898-appb-100002
    x=[x1,x2,...,xj]T
    其中,
    Figure PCTCN2016089898-appb-100003
    表示第k次外循环对应的障碍函数;
    f(X)表示优化目标对应的公平性函数;
    X表示任务分配数量;
    μk表示障碍因子μ在第k次外循环对应的值;
    hi(X)表示第i种资源对应的松弛变量;
    N表示自然数集;
    Figure PCTCN2016089898-appb-100004
    表示椭球体中心点对应的第i种资源需求;
    Pi表示第i种资源对应的需求变化向量;
    xj表示分配给第j个云用户的任务数量;
    将f(X)的初始查找点X0 *设为X0
    Figure PCTCN2016089898-appb-100005
    的初始查找点x0 *设为X0
    将外循环迭代次数k赋值为0;
    步骤3:将标志newtonFlag赋值为fail,其中,标志newtonFlag是表示牛顿迭代法是否能够计算出结果的标志,fail表示否;
    步骤4:如果
    Figure PCTCN2016089898-appb-100006
    是正定矩阵,则进入步骤5,否则进入步骤8;其中,
    Figure PCTCN2016089898-appb-100007
    表示二次求导,
    Figure PCTCN2016089898-appb-100008
    表示第k次外循环中,对应于任务分配数量为Xk的障碍函数,Xk表示第k次迭代计算出的任务分配数量;
    步骤5:将内循环迭代次数s赋值为0;
    步骤6:将f(X)的第s+1次内循环的查找点xs+1 *赋值为
    Figure PCTCN2016089898-appb-100009
    令s的值增加1;
    Figure PCTCN2016089898-appb-100010
    表示一次求导;
    步骤7:如果
    Figure PCTCN2016089898-appb-100011
    或者s>MaxIter,则进入步骤8,否则进入步骤6;其中,
    Figure PCTCN2016089898-appb-100012
    表示第s次内循环中计算出的任务分配数量,
    Figure PCTCN2016089898-appb-100013
    表示第s-1次内循环中计算出的任务分配数量,
    Figure PCTCN2016089898-appb-100014
    表示在第k次外循环中,对应于任务分配数量为
    Figure PCTCN2016089898-appb-100015
    的障碍函数,
    Figure PCTCN2016089898-appb-100016
    表示在第k次外循环中,对应于任务分配数量为
    Figure PCTCN2016089898-appb-100017
    的障碍函数,TolFun表示迭代终止的函数公差值,TolX表示迭代终止的任务分配数量X公差值,MaxIter表示最大迭代次数;
    步骤8:如果s≤MaxIter,则将newtonFlag赋值为success,进入步骤9;如果s>MaxIter,则接着执行步骤9;其中,success表示是;
    步骤9:如果newtonFlag等于fail,则转入步骤10,否则转入步骤15;
    步骤10:将内循环迭代次数s赋值为0,将第k次外循环对应的搜索方向dk赋值为
    Figure PCTCN2016089898-appb-100018
    步骤11:如果s>0,则将ds赋值为
    Figure PCTCN2016089898-appb-100019
    进入步骤12;否则,则进入步骤12;其中,ds-1表示第s-1次内循环对应的搜索方向;
    步骤12:用精确线性搜索方法计算出搜索步长αs,将xs+1 *赋值为
    Figure PCTCN2016089898-appb-100020
    令s的值增加1;
    Figure PCTCN2016089898-appb-100021
    表示第s+1次内循环中计算出的任务分配数量;其中,ds表示第s次内循环对应的搜索方向;
    步骤13:如果
    Figure PCTCN2016089898-appb-100022
    或者s>MaxIter,则进入步骤14,否则进入步骤11;
    步骤14:如果s>MaxIter,则将k赋值为0,进入步骤17;如果s≤MaxIter,则继 续执行步骤15;
    步骤15:将
    Figure PCTCN2016089898-appb-100023
    赋值为
    Figure PCTCN2016089898-appb-100024
    μk+1赋值为r*μk,令k的值增加1;其中,
    Figure PCTCN2016089898-appb-100025
    表示第k+1次外循环计算出的任务分配数量,μk+1表示第k+1次外循环中的障碍因子,μk表示第k次外循环中的障碍因子;
    步骤16:如果
    Figure PCTCN2016089898-appb-100026
    或者k>MaxIter,则进入步骤17,否则进入步骤3;其中,
    Figure PCTCN2016089898-appb-100027
    表示第k次外循环计算出的任务分配数量,
    Figure PCTCN2016089898-appb-100028
    表示第k-1次外循环计算出的任务分配数量,
    Figure PCTCN2016089898-appb-100029
    表示对应于
    Figure PCTCN2016089898-appb-100030
    的公平值,
    Figure PCTCN2016089898-appb-100031
    表示对应于
    Figure PCTCN2016089898-appb-100032
    的公平值;
    步骤17:输出最优的公平值
    Figure PCTCN2016089898-appb-100033
    云用户的最优任务分配数量
    Figure PCTCN2016089898-appb-100034
  2. 根据权利要求1所述的云调度器中应对不确定需求的多资源调度方法,其特征在于,所述多资源调度方法基于椭球体不确定模型;
    椭球体不确定模型,即:
    Figure PCTCN2016089898-appb-100035
    其中,Ui表示第i种资源对应的需求集合,Pi表示第i种资源对应的需求变化向量。
  3. 根据权利要求1所述的云调度器中应对不确定需求的多资源调度方法,其特征在于,还包括步骤:
    步骤18:按照最优的公平值
    Figure PCTCN2016089898-appb-100036
    云用户的最优任务分配数量
    Figure PCTCN2016089898-appb-100037
    进行资源调度。
  4. 根据权利要求1所述的云调度器中应对不确定需求的多资源调度方法,其特征在于,递减系数r取值0.1;
    迭代终止的函数公差值TolFun等于1e-6;
    迭代终止的任务分配数量公差值TolX等于1e-10;
    最大迭代次数MaxIter等于1000。
PCT/CN2016/089898 2016-04-13 2016-07-13 云调度器中应对不确定需求的多资源调度方法 Ceased WO2017177567A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US16/093,585 US11157327B2 (en) 2016-04-13 2016-07-13 Multi-resource scheduling method responding to uncertain demand in cloud scheduler

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201610228390.7A CN105871618B (zh) 2016-04-13 2016-04-13 云调度器中应对不确定需求的多资源调度方法
CN201610228390.7 2016-04-13

Publications (1)

Publication Number Publication Date
WO2017177567A1 true WO2017177567A1 (zh) 2017-10-19

Family

ID=56637730

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/089898 Ceased WO2017177567A1 (zh) 2016-04-13 2016-07-13 云调度器中应对不确定需求的多资源调度方法

Country Status (3)

Country Link
US (1) US11157327B2 (zh)
CN (1) CN105871618B (zh)
WO (1) WO2017177567A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118897739A (zh) * 2024-10-08 2024-11-05 山东蓝海领航大数据发展有限公司 一种多云异构的算力调度系统及调度方法

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108234151B (zh) * 2016-12-09 2021-02-23 河南工业大学 一种云平台资源分配方法
CN108449411B (zh) * 2018-03-19 2020-09-11 河南工业大学 一种随机需求下面向异质费用的云资源调度方法
CN113946429A (zh) * 2021-11-03 2022-01-18 重庆邮电大学 一种基于成本效益的Kubernetes Pod调度方法
CN117793180B (zh) * 2024-02-01 2024-12-06 广东天耘科技有限公司 云手机分配方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110125539A1 (en) * 2009-11-25 2011-05-26 General Electric Company Systems and methods for multi-resource scheduling
CN103220337A (zh) * 2013-03-22 2013-07-24 合肥工业大学 基于自适应弹性控制的云计算资源优化配置方法
CN103458052A (zh) * 2013-09-16 2013-12-18 北京搜狐新媒体信息技术有限公司 一种基于IaaS云平台的资源调度方法和装置
CN105302650A (zh) * 2015-12-10 2016-02-03 云南大学 一种面向云计算环境下的动态多资源公平分配方法
CN105446817A (zh) * 2015-11-23 2016-03-30 重庆邮电大学 移动云计算中一种基于鲁棒优化的联合资源预留配置算法

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102004671B (zh) * 2010-11-15 2013-03-13 北京航空航天大学 一种云计算环境下数据中心基于统计模型的资源管理方法
CN104932938B (zh) * 2015-06-16 2019-08-23 中电科软件信息服务有限公司 一种基于遗传算法的云资源调度方法
US9672064B2 (en) * 2015-07-13 2017-06-06 Palo Alto Research Center Incorporated Dynamically adaptive, resource aware system and method for scheduling

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110125539A1 (en) * 2009-11-25 2011-05-26 General Electric Company Systems and methods for multi-resource scheduling
CN103220337A (zh) * 2013-03-22 2013-07-24 合肥工业大学 基于自适应弹性控制的云计算资源优化配置方法
CN103458052A (zh) * 2013-09-16 2013-12-18 北京搜狐新媒体信息技术有限公司 一种基于IaaS云平台的资源调度方法和装置
CN105446817A (zh) * 2015-11-23 2016-03-30 重庆邮电大学 移动云计算中一种基于鲁棒优化的联合资源预留配置算法
CN105302650A (zh) * 2015-12-10 2016-02-03 云南大学 一种面向云计算环境下的动态多资源公平分配方法

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
ARMIN, A. ET AL.: "An Ellipsoid Algorithm for Linear Optimization with Uncertain LMI Constraints", 2012 AMERICAN CONTROL CONFERENCE, 29 June 2012 (2012-06-29), pages 857 - 862, XP032245045 *
WANG, JUN: "Model and Hybrid Genetic Algorithm for Robust Multi-objective Linear Programming", APPLICATION RESEARCH OF COMPUTERS, vol. 30, no. 9, 30 September 2013 (2013-09-30) *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118897739A (zh) * 2024-10-08 2024-11-05 山东蓝海领航大数据发展有限公司 一种多云异构的算力调度系统及调度方法

Also Published As

Publication number Publication date
US11157327B2 (en) 2021-10-26
US20210224135A1 (en) 2021-07-22
CN105871618A (zh) 2016-08-17
CN105871618B (zh) 2019-05-24

Similar Documents

Publication Publication Date Title
Mohan et al. Edge-Fog cloud: A distributed cloud for Internet of Things computations
Bruneo A stochastic model to investigate data center performance and QoS in IaaS cloud computing systems
Wang et al. Adaptive scheduling for parallel tasks with QoS satisfaction for hybrid cloud environments
Nan et al. Optimal resource allocation for multimedia cloud based on queuing model
Singh et al. Resource provisioning and scheduling in clouds: QoS perspective
Feldman et al. The proportional-share allocation market for computational resources
Nguyen et al. Two-stage robust edge service placement and sizing under demand uncertainty
WO2017177567A1 (zh) 云调度器中应对不确定需求的多资源调度方法
CN110198339B (zh) 一种基于QoE感知的边缘计算任务调度方法
Fan et al. Modeling and analyzing dynamic fault-tolerant strategy for deadline constrained task scheduling in cloud computing
Genez et al. Estimation of the available bandwidth in inter-cloud links for task scheduling in hybrid clouds
CN108021435B (zh) 一种基于截止时间的具有容错能力的云计算任务流调度方法
CN113037877A (zh) 云边端架构下时空数据及资源调度的优化方法
Nguyen et al. Monad: Self-adaptive micro-service infrastructure for heterogeneous scientific workflows
Xu et al. Optimized contract-based model for resource allocation in federated geo-distributed clouds
US20120324466A1 (en) Scheduling Execution Requests to Allow Partial Results
Yang et al. Trust-based scheduling strategy for workflow applications in cloud environment
Keshavarzi et al. Adaptive Resource Management and Provisioning in the Cloud Computing: A Survey of Definitions, Standards and Research Roadmaps.
Ji et al. Adaptive workflow scheduling for diverse objectives in cloud environments
CN119668886A (zh) 基于容器和Kubernetes的多租户管理方法及系统
Liu et al. DPBalance: Efficient and fair privacy budget scheduling for federated learning as a service
Tiwari et al. An optimized scheduling algorithm for cloud broker using adaptive cost model
Oliveira et al. Optimizing query prices for data-as-a-service
Wen et al. Pargen: A parallel method for partitioning data stream applications in mobile edge computing
Zhan et al. Cost-aware traffic management under demand uncertainty from a colocation data center user’s perspective

Legal Events

Date Code Title Description
NENP Non-entry into the national phase

Ref country code: DE

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16898382

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 16898382

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 11.06.2019)

122 Ep: pct application non-entry in european phase

Ref document number: 16898382

Country of ref document: EP

Kind code of ref document: A1