WO2017139002A2 - Mitigating component performance variation - Google Patents

Mitigating component performance variation Download PDF

Info

Publication number
WO2017139002A2
WO2017139002A2 PCT/US2016/063554 US2016063554W WO2017139002A2 WO 2017139002 A2 WO2017139002 A2 WO 2017139002A2 US 2016063554 W US2016063554 W US 2016063554W WO 2017139002 A2 WO2017139002 A2 WO 2017139002A2
Authority
WO
WIPO (PCT)
Prior art keywords
power
components
safe operation
component
maximum safe
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2016/063554
Other languages
French (fr)
Other versions
WO2017139002A3 (en
Inventor
Alan G. Gara
Steve S. SYLVESTER
Jonathan M. Eastep
Ramkumar NAGAPPAN
Christopher M. Cantalupo
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Intel Corp
Original Assignee
Intel Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Intel Corp filed Critical Intel Corp
Publication of WO2017139002A2 publication Critical patent/WO2017139002A2/en
Publication of WO2017139002A3 publication Critical patent/WO2017139002A3/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/16Constructional details or arrangements
    • G06F1/20Cooling means
    • G06F1/206Cooling means comprising thermal management
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • G06F1/3234Power saving characterised by the action undertaken
    • G06F1/329Power saving characterised by the action undertaken by task scheduling
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • G06F1/3234Power saving characterised by the action undertaken
    • G06F1/3296Power saving characterised by the action undertaken by lowering the supply or operating voltage
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/50Allocation of resources, e.g. of the central processing unit [CPU]
    • G06F9/5094Allocation of resources, e.g. of the central processing unit [CPU] where the allocation takes into account power or heat criteria
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02BCLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO BUILDINGS, e.g. HOUSING, HOUSE APPLIANCES OR RELATED END-USER APPLICATIONS
    • Y02B70/00Technologies for an efficient end-user side electric power management and consumption
    • Y02B70/10Technologies improving the efficiency by using switched-mode power supplies [SMPS], i.e. efficient power electronics conversion e.g. power factor correction or reduction of losses in power supplies or efficient standby modes
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • Embodiments generally relate to mitigating component performance variation. More particularly, embodiments relate to mitigating performance variation of components such as processors in distributed computing systems.
  • the scale and performance of large distributed computing systems may be limited by power and thermal constraints. For example, there may be constraints at the site level as well as at the component level. At the component level, components may tend to throttle their performance to reduce temperature and power and avoid damage when workloads cause them to exceed safe thermal and power density operational limits. At the site level, future systems may run under a power boundary to ensure that the site stays within site power limits, wherein the site power limits are derived from constraints on operational costs or limitations of the cooling and power delivery infrastructure.
  • manufacturing process variation may result in higher variance in the voltage that is supplied to a component for its circuits to function correctly at a given performance level.
  • thermal and power limitations in large distributed computing systems may expose these differences in supply voltage requirements, leading to unexpected performance differences across like components.
  • different processors in the system may throttle to different frequencies because they exhaust thermal and power density headroom at different points. This difference may occur even if the processors are selected from the same bin and/or l product SKU because parts from the same bin may still exhibit non-negligible variation in voltage requirements.
  • a uniform partition of power among like components may result in different performance across components when limiting system power to stay within site limits.
  • FIG. 1 is an illustration of an example of power allocation according to an embodiment
  • FIG. 2 is a flowchart of an example of a method of power allocation according to an embodiment
  • FIG. 3 is an illustration of an example of a computing system according to an embodiment.
  • FIG. 1 depicts an apparatus 10 that allocates various powers Pi, P 2 , and P n -i to components 20, 22, and 24 (e.g., processors) within a distributed computing system. While components 20, 22, and 24 may be processors, various other components may have power allocated to them according to the described embodiments. Other components include memory components, network interface components, chipsets, or any other component that draws power in apparatus 10.
  • distributed computing system may relate to a system with networked computers that communicate and coordinate their actions in order to achieve a common goal. In one example, a distributed computing system may utilize component groups to handle work according to various computational topologies and architectures.
  • an application may be divided into various tasks that may be subdivided into groups of related subtasks (e.g., threads), which may be run in parallel on a computational resource.
  • Related threads may be processed in parallel with one another on different components as "parallel threads," and the completion of a given task may entail the completion of all of the related threads that form the task. Similar performance of each component reduces wait time for completion of related threads.
  • the components 20, 22, and 24 may be characterized by a component characterization module 30, which may include fixed-functionality logic hardware, configurable logic, logic instructions (e.g., software), etc., or any combination thereof.
  • the characterization data may be obtained from a component database 40 employing, for example, software and/or hardware.
  • the database 40 may include a measurement of the component thermal design power.
  • the thermal design power (TDP) may be the maximum power at which it is safe to operate the component before there is risk of damage.
  • the TDP might represent the maximum amount of power the cooling system is required to dissipate.
  • the maximum amount of power may therefore be the power budget under which the system typically operates.
  • the maximum amount of power may not be, however, the same as the maximum power that the component can consume. For example, it is possible for the component to consume more than the TDP power for a short period of time without it being "thermally significant". Using basic physics, heat will take some time to propagate, so a short burst may not necessarily violate TDP.
  • the illustrated power allocator 50 allocates power among the components 20,
  • PMC Power Monitoring and Control
  • PMC may provide energy measurements or estimates through system software, firmware, or hardware, and it may enforce configurable limits on power consumption.
  • PMC may report either a digital estimation of power, a measurement of power from sensors (external to the component 20, 22, 24 or internal), or some hybrid. In one embodiment, PMC reports measurements from sensors at voltage regulators and the voltage regulator and sensors are external to the component (e.g., located on a motherboard).
  • PMC may report measurements from sensors at voltage regulators that may be a mixture of external and internal voltage regulators.
  • PMC may report digital estimations of power from a model based on activity in different component resources with activity being measured through event counters. Note that any of the PMC techniques, whether power is reported as an estimate, a measurement from one or more sensors, or a hybrid are applicable for power allocation in the various embodiments set forth herein.
  • RAPL Heating Average Power Limiting
  • the power allocator 50 may compute the best allocation of power among the components 20, 22, and 24 in various ways.
  • the power allocation of each component (denoted as Pi), may be computed by taking the ratio of that component's TDP measurement (JDP t ) to the sum of TDP for all components and scaling that ratio by a factor equal to the total power (P to tai) f° r all components.
  • a more general model of the relationship between the power and performance of a component may be employed.
  • linear relationships may be assumed, wherein the power allocation decision may be cast as a linear optimization problem and the best allocation may be obtained through linear programming.
  • non-linear models may be assumed, and the best allocation may be obtained through a numerical solver.
  • the power allocator 50 and component TDP database 40 may be integrated into a system resource manager and/or scheduler or they may be separate components that provide software APIs to the system resource manager and/or scheduler enabling them to select a total power budget and query component TDP data.
  • the system resource manager and/or scheduler may likewise be provided with interfaces to the hardware to specify a total power budget or query TDP data. Control registers or model specific registers may be used for the hardware interface.
  • the power may be selected for each component differently: as before, the component's TDP is read from the database 40 but it is scaled by a factor equal to the factor that total power should be reduced to from maximum power (a):
  • FIG. 2 shows an overview of a method 100 of mitigating performance variation among components in a distributed computing system.
  • the method 100 may generally be implemented in an apparatus such as, for example, the apparatus 10 (FIG. 1), already discussed. More particularly, the method 100 may be implemented in one or more modules as a set of logic instructions stored in a machine- or computer-readable storage medium such as random access memory (RAM), read only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc., in configurable logic such as, for example, programmable logic arrays (PLAs), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), in fixed-functionality logic hardware using circuit technology such as, for example, application specific integrated circuit (ASIC), complementary metal oxide semiconductor (CMOS) or transistor- transistor logic (TTL) technology, or any combination thereof.
  • ASIC application specific integrated circuit
  • CMOS complementary metal oxide semiconductor
  • TTL transistor- transistor logic
  • computer program code to carry out operations shown in method 100 may be written in any combination of one or more programming languages, including an object oriented programming language such as JAVA, SMALLTALK, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages.
  • object oriented programming language such as JAVA, SMALLTALK, C++ or the like
  • conventional procedural programming languages such as the "C" programming language or similar programming languages.
  • Illustrated processing block 110 provides for characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation power level associated with each component.
  • the characterization data is stored in a database.
  • the components may be components 20, 22, and 24 that, in one embodiment, may be similar processors having manufacturing variations from processor to processor that affect power levels required by each processor for correct functioning.
  • the database may be, for example, the component TDP database 40.
  • the database 40 may update the maximum safe operation power level while the system is running or may draw on data established at the time of manufacture of the components.
  • Illustrated processing block 120 allocates non-uniform power to each similar component (e.g., components 20, 22, and 24 of FIG. 1) based at least in part on the characterization data in the database to substantially equalize performance of the components. As discussed above, nominally "identical" components may have different power needs to achieve a similar performance due to manufacturing variations. Thus, block 120 might allocate non-uniform power to components 20, 22, and 24 to ensure that the performance of the components 20, 22, and 24 will be substantially similar.
  • each similar component e.g., components 20, 22, and 24 of FIG. 1
  • block 120 might allocate non-uniform power to components 20, 22, and 24 to ensure that the performance of the components 20, 22, and 24 will be substantially similar.
  • FIG. 3 shows a computing system 200.
  • the computing system may be part of an electronic device/platform having computing functionality and may also form part of a distributed computing system, as described above.
  • the system 100 includes a power source 292 to provide power to the system.
  • the power source 292 may provide power to plural computing systems 200 as part of a distributed computing system.
  • the computing system 200 may include a die 282 that has a processor 220 with a PMC module 260.
  • the processor 220 may communicate with system memory 288.
  • the system memory 288 may include, for example, dynamic random access memory (DRAM) configured as one or more memory modules such as, for example, dual inline memory modules (DIMMs), small outline DIMMs (SODIMMs), etc.
  • DRAM dynamic random access memory
  • the illustrated system 80 also includes an input output (I/O) module 290 implemented together with the processor 220 on semiconductor die 282 as a system on a chip (SoC), wherein the I/O module 290 functions as a host device and may communicate with peripheral devices such as a display (not shown), network controller 296 and mass storage (not shown) (e.g., hard disk drive HDD, optical disk, flash memory, etc.).
  • SoC system on a chip
  • the illustrated I/O module 290 may execute logic 300 that forms a part of the power allocation as set forth in FIG. 1.
  • FIG. 3 Alternatively hardware and/or software logic that performs the power allocation of FIG. 1 may be found in a management controller 280 as seen in FIG. 3.
  • the power allocator 250 communicates with the TDP database 240 that, in turn, communicates with component characterization element 230, in a manner substantially similar to that described for elements 50, 40, and 30 of FIG. 1.
  • the hardware and/or software logic that performs the power allocation may reside elsewhere within a distributed computing network.
  • power allocator 250, database 240, and component characterization 230 may be implemented in software with the allocator running as a process on one or more of the processors (a distributed software runtime).
  • the allocator may by implemented in firmware in the management controller on one or more of the motherboards (a distributed firmware solution).
  • the allocator may be implemented in firmware in one or more of the processors (a distributed firmware solution).
  • the allocator may be implemented in hardware logic on one or more of the processors (a distributed hardware solution).
  • Example 1 may include an apparatus to mitigate performance variation among components in a distributed computing system comprising a plurality of computational resources that are connected to one another to form the distributed computing system, logic, implemented at least partly in one or more of configurable logic or fixed- functionality logic hardware, to characterize a plurality of similar components of the distributed computing system based on a maximum safe operation level associated with each component and store characterization data in a database, and allocate non-uniform power to each similar component based at least in part on characterization data in the database to substantially equalize performance of the components.
  • Example 2 may include the apparatus of example 1, wherein the components are processors.
  • Example 3 may include the apparatus of examples 1 or 2, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
  • Example 4 may include the apparatus of examples 1 or 2, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
  • Example 5 may include the apparatus of examples 1 or 2, further including logic to determine, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scale the ratio by a factor equal to total power for all components.
  • Example 6 may include the apparatus of examples 1 or 2, further including logic to scale maximum safe operation power level by a factor that total power is to be reduced from maximum power.
  • Example 7 may include a method of mitigating performance variation among components in a distributed computing system comprising characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database and allocating non-uniform power to each similar component based at least in part on the characterization data in the database to substantially equalize performance of the components.
  • Example 8 may include the method of example 7, wherein processors are characterized.
  • Example 9 may include the method of examples 7 or 8, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
  • Example 10 may include the method of examples 7 or 8 wherein the plurality of similar components are to be characterized at the time of manufacture of the components.
  • Example 11 may include the method of examples 7 or 8, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
  • Example 12 may include the method of examples 7 or 8, further including determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scaling the ratio by a factor equal to total power for all components.
  • Example 13 may include the method of examples 7 or 8, further including scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
  • Example 14 may include at least one computer readable storage medium comprising a set of instructions, wherein the instructions, when executed, cause a computing device in a distributing computing system to characterize a plurality of similar components of the distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database, and allocate non-uniform power to each similar component based at least in part on the characterization data in the database to substantially equalize performance of the components.
  • Example 15 may include the medium of example 14, wherein processors are to be characterized.
  • Example 16 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to characterize the components while the distributed computing system is running.
  • Example 17 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power through running average power limiting firmware and hardware.
  • Example 18 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power by determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scaling the ratio by a factor equal to total power for all components.
  • Example 19 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power by scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
  • Example 20 may include an apparatus to mitigate performance variation among components in a distributed computing system comprising means for characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database and means for allocating non-uniform power to each similar component based at least in part on the characterization data to substantially equalize performance of the components.
  • Example 21 may include the apparatus of example 20, wherein processors are to be characterized.
  • Example 22 may include the apparatus of examples 20 or 21, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
  • Example 23 may include the apparatus of examples 20 or 21, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
  • Example 24 may include the apparatus of examples 20 or 21 further including means for determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scaling the ratio by a factor equal to total power for all components.
  • Example 25 may include the apparatus of examples 20 or 21 further comprising means to cause the computing device to allocate power by scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
  • the embodiments described above allocate non-uniform power to each component from a total budget.
  • the embodiments may correct for manufacturing variation and the performance variation it induces across components by selecting the power allocation that equalizes performance of all components.
  • the allocation decision may be based on detailed characterization of each component either at manufacturing time or while the system is running.
  • the allocation decision may also be based on averages and characterization over a set of representative benchmarks or adapted based on characterization of individual workloads.
  • use of the apparatus and methods described above may mitigate performance variation due to manufacturing variation across components in a distributed computing system.
  • the embodiments may significantly improve performance determinism across workload runs in typical systems that employ resource managers or schedulers which allocate different nodes for the workload on each run.
  • the embodiments may also significantly improve performance and efficiency of high performance computing workloads which tend to employ distributed computation that may be tightly synchronized across nodes through collective operations or other global synchronization events; use of the embodiments reduces wait times at the synchronization events which reduces power wasted while waiting and reduces overall time to solution.
  • the embodiments may provide several advantages over other approaches. They may mitigate rather than aggravate performance variation across components. They may provide a substantially less complex solution for mitigating performance variation (by allocating power in proportion to TDP to equalize performance differences deriving from TDP differences). They may also provide substantially higher performance when applied to mitigating processor performance differences because processor frequency is controlled through a different mechanism than through discrete voltage-frequency steps at which the component may operate. Processor frequency may be managed by power limiting features such as PMC. Power limiting features like PMC enable each processor to run at the maximum safe operating frequency for each individual workload rather than a worst-case power virus workload. PMC provides energy measurements or estimates through system software, firmware, or hardware that enforce configurable limits on power consumption.
  • Embodiments are applicable for use with all types of semiconductor integrated circuit (“IC") chips.
  • IC semiconductor integrated circuit
  • Examples of these IC chips include but are not limited to processors, controllers, chipset components, programmable logic arrays (PLAs), memory chips, network chips, systems on chip (SoCs), SSD/NAND controller ASICs, and the like.
  • PLAs programmable logic arrays
  • SoCs systems on chip
  • SSD/NAND controller ASICs solid state drive/NAND controller ASICs
  • signal conductor lines are represented with lines. Some may be different, to indicate more constituent signal paths, have a number label, to indicate a number of constituent signal paths, and/or have arrows at one or more ends, to indicate primary information flow direction. This, however, should not be construed in a limiting manner.
  • Any represented signal lines may actually comprise one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, optical fiber lines, and/or single-ended lines.
  • Example sizes/models/values/ranges may have been given, although embodiments are not limited to the same. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices of smaller size could be manufactured.
  • well known power/ground connections to IC chips and other components may or may not be shown within the figures, for simplicity of illustration and discussion, and so as not to obscure certain aspects of the embodiments. Further, arrangements may be shown in block diagram form in order to avoid obscuring embodiments, and also in view of the fact that specifics with respect to implementation of such block diagram arrangements are highly dependent upon the platform within which the embodiment is to be implemented, i.e., such specifics should be well within purview of one skilled in the art.
  • Coupled and “communicating” may be used herein to refer to any type of relationship, direct or indirect, between the components in question, and may apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections.
  • first”, “second”, etc. may be used herein only to facilitate discussion, and carry no particular temporal or chronological significance unless otherwise indicated.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Human Computer Interaction (AREA)
  • Power Sources (AREA)
  • Supply And Distribution Of Alternating Current (AREA)

Abstract

Apparatus and methods may provide for characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database and allocating non-uniform power to each similar component based at least in part on the characterization data in the database to substantially equalize performance of the components.

Description

MITIGATING COMPONENT PERFORMANCE VARIATION
CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims the benefit of priority to U. S. Non-Provisional
Patent Application No. 14/998,082 filed on December 24, 2015.
GOVERNMENT INTEREST STATEMENT
This invention was made with Government support under contract number B609815 awarded by the Department of Energy. The Government has certain rights in this invention.
TECHNICAL FIELD
Embodiments generally relate to mitigating component performance variation. More particularly, embodiments relate to mitigating performance variation of components such as processors in distributed computing systems.
BACKGROUND
The scale and performance of large distributed computing systems may be limited by power and thermal constraints. For example, there may be constraints at the site level as well as at the component level. At the component level, components may tend to throttle their performance to reduce temperature and power and avoid damage when workloads cause them to exceed safe thermal and power density operational limits. At the site level, future systems may run under a power boundary to ensure that the site stays within site power limits, wherein the site power limits are derived from constraints on operational costs or limitations of the cooling and power delivery infrastructure.
Concomitantly, manufacturing process variation may result in higher variance in the voltage that is supplied to a component for its circuits to function correctly at a given performance level. Unfortunately, the thermal and power limitations in large distributed computing systems may expose these differences in supply voltage requirements, leading to unexpected performance differences across like components. For example, different processors in the system may throttle to different frequencies because they exhaust thermal and power density headroom at different points. This difference may occur even if the processors are selected from the same bin and/or l product SKU because parts from the same bin may still exhibit non-negligible variation in voltage requirements. As another example, a uniform partition of power among like components may result in different performance across components when limiting system power to stay within site limits.
BRIEF DESCRIPTION OF THE DRAWINGS
The various advantages of the embodiments will become apparent to one skilled in the art by reading the following specification and appended claims, and by referencing the following drawings, in which:
FIG. 1 is an illustration of an example of power allocation according to an embodiment;
FIG. 2 is a flowchart of an example of a method of power allocation according to an embodiment;
FIG. 3 is an illustration of an example of a computing system according to an embodiment.
DESCRIPTION OF EMBODIMENTS
Turning now to the drawings in detail, FIG. 1 depicts an apparatus 10 that allocates various powers Pi, P2, and Pn-i to components 20, 22, and 24 (e.g., processors) within a distributed computing system. While components 20, 22, and 24 may be processors, various other components may have power allocated to them according to the described embodiments. Other components include memory components, network interface components, chipsets, or any other component that draws power in apparatus 10. The expression "distributed computing system" may relate to a system with networked computers that communicate and coordinate their actions in order to achieve a common goal. In one example, a distributed computing system may utilize component groups to handle work according to various computational topologies and architectures. For example, an application may be divided into various tasks that may be subdivided into groups of related subtasks (e.g., threads), which may be run in parallel on a computational resource. Related threads may be processed in parallel with one another on different components as "parallel threads," and the completion of a given task may entail the completion of all of the related threads that form the task. Similar performance of each component reduces wait time for completion of related threads.
At the time a new workload such as an application is launched in the distributed computing system, power may be allocated to the various parallel components 20, 22, and 24. The components 20, 22, and 24 may be characterized by a component characterization module 30, which may include fixed-functionality logic hardware, configurable logic, logic instructions (e.g., software), etc., or any combination thereof. The characterization data may be obtained from a component database 40 employing, for example, software and/or hardware. For each of the components 20, 22, and 24, the database 40 may include a measurement of the component thermal design power. The thermal design power (TDP) may be the maximum power at which it is safe to operate the component before there is risk of damage. For example, the TDP might represent the maximum amount of power the cooling system is required to dissipate. The maximum amount of power may therefore be the power budget under which the system typically operates. The maximum amount of power may not be, however, the same as the maximum power that the component can consume. For example, it is possible for the component to consume more than the TDP power for a short period of time without it being "thermally significant". Using basic physics, heat will take some time to propagate, so a short burst may not necessarily violate TDP.
The illustrated power allocator 50 allocates power among the components 20,
22, and 24, then utilizes mechanisms provided by the system software, firmware, and/or hardware to enforce the power allocation. One example of a mechanism of enforcing a power allocation is a Power Monitoring and Control (PMC) module 60 that may be included with each of the components 20, 22, and 24. PMC may provide energy measurements or estimates through system software, firmware, or hardware, and it may enforce configurable limits on power consumption. PMC may report either a digital estimation of power, a measurement of power from sensors (external to the component 20, 22, 24 or internal), or some hybrid. In one embodiment, PMC reports measurements from sensors at voltage regulators and the voltage regulator and sensors are external to the component (e.g., located on a motherboard). However, PMC may report measurements from sensors at voltage regulators that may be a mixture of external and internal voltage regulators. In other aspects, PMC may report digital estimations of power from a model based on activity in different component resources with activity being measured through event counters. Note that any of the PMC techniques, whether power is reported as an estimate, a measurement from one or more sensors, or a hybrid are applicable for power allocation in the various embodiments set forth herein. RAPL (Running Average Power Limiting) is an example of a PMC technique that may be used in the embodiments.
The power allocator 50 may compute the best allocation of power among the components 20, 22, and 24 in various ways. In one embodiment, the power allocation of each component (denoted as Pi), may be computed by taking the ratio of that component's TDP measurement (JDPt) to the sum of TDP for all components and scaling that ratio by a factor equal to the total power (Ptotai) f°r all components. έ may represent a proportional allocation of power based on TDP and assumes that component performance is proportional to power with constant of proportionality m=l ; i.e., adding one unit of power to a component increases its performance by one unit:
TDPi
Pi - Ptotal *n-i TDpk
In another embodiment, a more general model of the relationship between the power and performance of a component may be employed. For example, linear relationships may be assumed, wherein the power allocation decision may be cast as a linear optimization problem and the best allocation may be obtained through linear programming. In another example, non-linear models may be assumed, and the best allocation may be obtained through a numerical solver.
In various embodiments, the power allocator 50 and component TDP database
40 may be implemented in different combinations of software, firmware, and hardware. For embodiments with a software implementation, the power allocator 50 and component TDP database 40 may be integrated into a system resource manager and/or scheduler or they may be separate components that provide software APIs to the system resource manager and/or scheduler enabling them to select a total power budget and query component TDP data. In embodiments with a firmware or hardware implementation, the system resource manager and/or scheduler may likewise be provided with interfaces to the hardware to specify a total power budget or query TDP data. Control registers or model specific registers may be used for the hardware interface.
In an alternative embodiment where the system or workload power may be reduced to a percentage of its maximum value rather than reduced to a specific absolute power, the power may be selected for each component differently: as before, the component's TDP is read from the database 40 but it is scaled by a factor equal to the factor that total power should be reduced to from maximum power (a):
Pi = a * TDPi
FIG. 2 shows an overview of a method 100 of mitigating performance variation among components in a distributed computing system. The method 100 may generally be implemented in an apparatus such as, for example, the apparatus 10 (FIG. 1), already discussed. More particularly, the method 100 may be implemented in one or more modules as a set of logic instructions stored in a machine- or computer-readable storage medium such as random access memory (RAM), read only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc., in configurable logic such as, for example, programmable logic arrays (PLAs), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), in fixed-functionality logic hardware using circuit technology such as, for example, application specific integrated circuit (ASIC), complementary metal oxide semiconductor (CMOS) or transistor- transistor logic (TTL) technology, or any combination thereof. For example, computer program code to carry out operations shown in method 100 may be written in any combination of one or more programming languages, including an object oriented programming language such as JAVA, SMALLTALK, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages.
Illustrated processing block 110 provides for characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation power level associated with each component. The characterization data is stored in a database. With continuing reference to FIGs. 1 and 2, the components may be components 20, 22, and 24 that, in one embodiment, may be similar processors having manufacturing variations from processor to processor that affect power levels required by each processor for correct functioning. The database may be, for example, the component TDP database 40. The database 40 may update the maximum safe operation power level while the system is running or may draw on data established at the time of manufacture of the components.
Illustrated processing block 120 allocates non-uniform power to each similar component (e.g., components 20, 22, and 24 of FIG. 1) based at least in part on the characterization data in the database to substantially equalize performance of the components. As discussed above, nominally "identical" components may have different power needs to achieve a similar performance due to manufacturing variations. Thus, block 120 might allocate non-uniform power to components 20, 22, and 24 to ensure that the performance of the components 20, 22, and 24 will be substantially similar.
FIG. 3 shows a computing system 200. The computing system may be part of an electronic device/platform having computing functionality and may also form part of a distributed computing system, as described above. In the illustrated example, the system 100 includes a power source 292 to provide power to the system. The power source 292 may provide power to plural computing systems 200 as part of a distributed computing system. The computing system 200 may include a die 282 that has a processor 220 with a PMC module 260. The processor 220 may communicate with system memory 288. The system memory 288 may include, for example, dynamic random access memory (DRAM) configured as one or more memory modules such as, for example, dual inline memory modules (DIMMs), small outline DIMMs (SODIMMs), etc.
The illustrated system 80 also includes an input output (I/O) module 290 implemented together with the processor 220 on semiconductor die 282 as a system on a chip (SoC), wherein the I/O module 290 functions as a host device and may communicate with peripheral devices such as a display (not shown), network controller 296 and mass storage (not shown) (e.g., hard disk drive HDD, optical disk, flash memory, etc.). The illustrated I/O module 290 may execute logic 300 that forms a part of the power allocation as set forth in FIG. 1.
Alternatively hardware and/or software logic that performs the power allocation of FIG. 1 may be found in a management controller 280 as seen in FIG. 3. The power allocator 250 communicates with the TDP database 240 that, in turn, communicates with component characterization element 230, in a manner substantially similar to that described for elements 50, 40, and 30 of FIG. 1. The hardware and/or software logic that performs the power allocation may reside elsewhere within a distributed computing network. For example, power allocator 250, database 240, and component characterization 230 may be implemented in software with the allocator running as a process on one or more of the processors (a distributed software runtime). Alternatively, the allocator may by implemented in firmware in the management controller on one or more of the motherboards (a distributed firmware solution). The allocator may be implemented in firmware in one or more of the processors (a distributed firmware solution). Moreover, the allocator may be implemented in hardware logic on one or more of the processors (a distributed hardware solution).
Additional Notes and Examples:
Example 1 may include an apparatus to mitigate performance variation among components in a distributed computing system comprising a plurality of computational resources that are connected to one another to form the distributed computing system, logic, implemented at least partly in one or more of configurable logic or fixed- functionality logic hardware, to characterize a plurality of similar components of the distributed computing system based on a maximum safe operation level associated with each component and store characterization data in a database, and allocate non-uniform power to each similar component based at least in part on characterization data in the database to substantially equalize performance of the components.
Example 2 may include the apparatus of example 1, wherein the components are processors.
Example 3 may include the apparatus of examples 1 or 2, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
Example 4 may include the apparatus of examples 1 or 2, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
Example 5 may include the apparatus of examples 1 or 2, further including logic to determine, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scale the ratio by a factor equal to total power for all components.
Example 6 may include the apparatus of examples 1 or 2, further including logic to scale maximum safe operation power level by a factor that total power is to be reduced from maximum power.
Example 7 may include a method of mitigating performance variation among components in a distributed computing system comprising characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database and allocating non-uniform power to each similar component based at least in part on the characterization data in the database to substantially equalize performance of the components. Example 8 may include the method of example 7, wherein processors are characterized.
Example 9 may include the method of examples 7 or 8, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
Example 10 may include the method of examples 7 or 8 wherein the plurality of similar components are to be characterized at the time of manufacture of the components.
Example 11 may include the method of examples 7 or 8, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
Example 12 may include the method of examples 7 or 8, further including determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scaling the ratio by a factor equal to total power for all components.
Example 13 may include the method of examples 7 or 8, further including scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
Example 14 may include at least one computer readable storage medium comprising a set of instructions, wherein the instructions, when executed, cause a computing device in a distributing computing system to characterize a plurality of similar components of the distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database, and allocate non-uniform power to each similar component based at least in part on the characterization data in the database to substantially equalize performance of the components.
Example 15 may include the medium of example 14, wherein processors are to be characterized.
Example 16 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to characterize the components while the distributed computing system is running.
Example 17 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power through running average power limiting firmware and hardware. Example 18 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power by determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scaling the ratio by a factor equal to total power for all components.
Example 19 may include the medium of examples 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power by scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
Example 20 may include an apparatus to mitigate performance variation among components in a distributed computing system comprising means for characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database and means for allocating non-uniform power to each similar component based at least in part on the characterization data to substantially equalize performance of the components.
Example 21 may include the apparatus of example 20, wherein processors are to be characterized.
Example 22 may include the apparatus of examples 20 or 21, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
Example 23 may include the apparatus of examples 20 or 21, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
Example 24 may include the apparatus of examples 20 or 21 further including means for determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components, and scaling the ratio by a factor equal to total power for all components.
Example 25 may include the apparatus of examples 20 or 21 further comprising means to cause the computing device to allocate power by scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
Given a distributed computing system with homogeneous components, the embodiments described above allocate non-uniform power to each component from a total budget. The embodiments may correct for manufacturing variation and the performance variation it induces across components by selecting the power allocation that equalizes performance of all components. The allocation decision may be based on detailed characterization of each component either at manufacturing time or while the system is running. The allocation decision may also be based on averages and characterization over a set of representative benchmarks or adapted based on characterization of individual workloads.
Advantageously, use of the apparatus and methods described above may mitigate performance variation due to manufacturing variation across components in a distributed computing system. By mitigating manufacturing variation, the embodiments may significantly improve performance determinism across workload runs in typical systems that employ resource managers or schedulers which allocate different nodes for the workload on each run. The embodiments may also significantly improve performance and efficiency of high performance computing workloads which tend to employ distributed computation that may be tightly synchronized across nodes through collective operations or other global synchronization events; use of the embodiments reduces wait times at the synchronization events which reduces power wasted while waiting and reduces overall time to solution.
The embodiments may provide several advantages over other approaches. They may mitigate rather than aggravate performance variation across components. They may provide a substantially less complex solution for mitigating performance variation (by allocating power in proportion to TDP to equalize performance differences deriving from TDP differences). They may also provide substantially higher performance when applied to mitigating processor performance differences because processor frequency is controlled through a different mechanism than through discrete voltage-frequency steps at which the component may operate. Processor frequency may be managed by power limiting features such as PMC. Power limiting features like PMC enable each processor to run at the maximum safe operating frequency for each individual workload rather than a worst-case power virus workload. PMC provides energy measurements or estimates through system software, firmware, or hardware that enforce configurable limits on power consumption.
Embodiments are applicable for use with all types of semiconductor integrated circuit ("IC") chips. Examples of these IC chips include but are not limited to processors, controllers, chipset components, programmable logic arrays (PLAs), memory chips, network chips, systems on chip (SoCs), SSD/NAND controller ASICs, and the like. In addition, in some of the drawings, signal conductor lines are represented with lines. Some may be different, to indicate more constituent signal paths, have a number label, to indicate a number of constituent signal paths, and/or have arrows at one or more ends, to indicate primary information flow direction. This, however, should not be construed in a limiting manner. Rather, such added detail may be used in connection with one or more exemplary embodiments to facilitate easier understanding of a circuit. Any represented signal lines, whether or not having additional information, may actually comprise one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, optical fiber lines, and/or single-ended lines.
Example sizes/models/values/ranges may have been given, although embodiments are not limited to the same. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices of smaller size could be manufactured. In addition, well known power/ground connections to IC chips and other components may or may not be shown within the figures, for simplicity of illustration and discussion, and so as not to obscure certain aspects of the embodiments. Further, arrangements may be shown in block diagram form in order to avoid obscuring embodiments, and also in view of the fact that specifics with respect to implementation of such block diagram arrangements are highly dependent upon the platform within which the embodiment is to be implemented, i.e., such specifics should be well within purview of one skilled in the art. Where specific details (e.g., circuits) are set forth in order to describe example embodiments, it should be apparent to one skilled in the art that embodiments can be practiced without, or with variation of, these specific details. The description is thus to be regarded as illustrative instead of limiting.
The terms "coupled" and "communicating" may be used herein to refer to any type of relationship, direct or indirect, between the components in question, and may apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections. In addition, the terms "first", "second", etc. may be used herein only to facilitate discussion, and carry no particular temporal or chronological significance unless otherwise indicated.
Those skilled in the art will appreciate from the foregoing description that the broad techniques of the embodiments can be implemented in a variety of forms. Therefore, while the embodiments have been described in connection with particular examples thereof, the true scope of the embodiments should not be so limited since other modifications will become apparent to the skilled practitioner upon a study of the drawings, specification, and following claims.

Claims

CLAIMS We claim:
1. An apparatus comprising:
a plurality of computational resources that are connected to one another to form a distributed computing system;
logic, implemented at least partly in one or more of configurable logic or fixed- functionality logic hardware, to:
characterize a plurality of similar components of the distributed computing system based on a maximum safe operation level associated with each component and store characterization data in a database;
allocate non-uniform power based at least in part on the characterization data in the database to each similar component to substantially equalize performance of the components.
2. The apparatus of claim 1, wherein the components are processors.
3. The apparatus of claims 1 or 2, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
4. The apparatus of claims 1 or 2, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
5. The apparatus of claims 1 or 2, further including logic to:
determine, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components; and
scale the ratio by a factor equal to total power for all components.
6. The apparatus of claims 1 or 2, further including logic to scale maximum safe operation power level by a factor that total power is to be reduced from maximum power.
7. A method comprising: characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database;
allocating non-uniform power based at least in part on the characterization data in the database to each similar component to substantially equalize performance of the components.
8. The method of claim 7, wherein processors are characterized.
9. The method of claims 7 or 8, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
10. The method of claims 7 or 8, wherein the plurality of similar components are to be characterized at the time of manufacture of the components.
11. The method of claims 7 or 8, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
12. The method of claims 7 or 8, further including determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components; and
scaling the ratio by a factor equal to total power for all components.
13. The method of claims 7 or 8, further including scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
14. At least one computer readable storage medium comprising a set of instructions, wherein the instructions, when executed, cause a computing device in a distributing computing system to:
characterize a plurality of similar components of the distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database; allocate non-uniform power to each similar component based at least in part on the characterization data in the database to substantially equalize performance of the components.
15. The medium of claim 14, wherein processors are to be characterized.
16. The medium of claims 14 or 15, wherein the instructions, when executed, cause the computing device to characterize the components while the distributed computing system is running.
17. The medium of claims 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power through running average power limiting firmware and hardware.
18. The medium of claims 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power by determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components; and
scaling the ratio by a factor equal to total power for all components.
19. The medium of claims 14 or 15, wherein the instructions, when executed, cause the computing device to allocate power by scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
20. An apparatus to mitigate performance variation among components in a distributed computing system comprising:
means for characterizing a plurality of similar components of a distributed computing system based on a maximum safe operation level associated with each component and storing characterization data in a database; and
means for allocating non-uniform power based at least in part on the characterization data in the database to each similar component to substantially equalize performance of the components.
21. The apparatus of claim 20, wherein processors are to be characterized.
22. The apparatus of claims 20 or 21, wherein the plurality of similar components are to be characterized while the distributed computing system is running.
23. The apparatus of claims 20 or 21, further including running average power limiting firmware and hardware to enforce non-uniform power allocations.
24. The apparatus of claims 20 or 21 further including means for determining, for each component, a ratio of maximum safe operation power level to a sum of maximum safe operation power level for all components; and
scaling the ratio by a factor equal to total power for all components.
25. The apparatus of claims 20 or 21 further comprising means to cause the computing device to allocate power by scaling maximum safe operation power level by a factor that total power is to be reduced from maximum power.
PCT/US2016/063554 2015-12-24 2016-11-23 Mitigating component performance variation Ceased WO2017139002A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US14/998,082 2015-12-24
US14/998,082 US9864423B2 (en) 2015-12-24 2015-12-24 Mitigating component performance variation

Publications (2)

Publication Number Publication Date
WO2017139002A2 true WO2017139002A2 (en) 2017-08-17
WO2017139002A3 WO2017139002A3 (en) 2017-09-28

Family

ID=59087070

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2016/063554 Ceased WO2017139002A2 (en) 2015-12-24 2016-11-23 Mitigating component performance variation

Country Status (3)

Country Link
US (1) US9864423B2 (en)
SG (1) SG10201609839PA (en)
WO (1) WO2017139002A2 (en)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11449245B2 (en) 2019-06-13 2022-09-20 Western Digital Technologies, Inc. Power target calibration for controlling drive-to-drive performance variations in solid state drives (SSDs)
CN114816025B (en) * 2021-01-19 2024-10-22 联想企业解决方案(新加坡)有限公司 Power management method and system

Family Cites Families (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7661003B2 (en) * 2005-01-21 2010-02-09 Hewlett-Packard Development Company, L.P. Systems and methods for maintaining performance of an integrated circuit within a working power limit
US7464278B2 (en) 2005-09-12 2008-12-09 Intel Corporation Combining power prediction and optimal control approaches for performance optimization in thermally limited designs
US7844838B2 (en) 2006-10-30 2010-11-30 Hewlett-Packard Development Company, L.P. Inter-die power manager and power management method
US8478451B2 (en) 2009-12-14 2013-07-02 Intel Corporation Method and apparatus for dynamically allocating power in a data center
US8627118B2 (en) * 2010-05-24 2014-01-07 International Business Machines Corporation Chassis power allocation using expedited power permissions
US9304570B2 (en) * 2011-12-15 2016-04-05 Intel Corporation Method, apparatus, and system for energy efficiency and energy conservation including power and performance workload-based balancing between multiple processing elements
US20130173946A1 (en) * 2011-12-29 2013-07-04 Efraim Rotem Controlling power consumption through multiple power limits over multiple time intervals
US9043170B2 (en) * 2012-02-23 2015-05-26 Dell Products L.P. Systems and methods for providing component characteristics
WO2013147801A1 (en) * 2012-03-29 2013-10-03 Intel Corporation Dynamic power limit sharing in a platform
US9261945B2 (en) * 2012-08-30 2016-02-16 Dell Products, L.P. Dynanmic peak power limiting to processing nodes in an information handling system
US9229503B2 (en) * 2012-11-27 2016-01-05 Qualcomm Incorporated Thermal power budget allocation for maximum user experience
JP6042217B2 (en) * 2013-01-28 2016-12-14 ルネサスエレクトロニクス株式会社 Semiconductor device, electronic device, and control method of semiconductor device
US9442559B2 (en) 2013-03-14 2016-09-13 Intel Corporation Exploiting process variation in a multicore processor
US9389924B2 (en) 2014-01-30 2016-07-12 Vmware, Inc. System and method for performing resource allocation for a host computer cluster
US10466754B2 (en) 2014-12-26 2019-11-05 Intel Corporation Dynamic hierarchical performance balancing of computational resources

Also Published As

Publication number Publication date
SG10201609839PA (en) 2017-07-28
US9864423B2 (en) 2018-01-09
WO2017139002A3 (en) 2017-09-28
US20170185129A1 (en) 2017-06-29

Similar Documents

Publication Publication Date Title
US10025361B2 (en) Power management across heterogeneous processing units
US10452437B2 (en) Temperature-aware task scheduling and proactive power management
JP5770300B2 (en) Method and apparatus for thermal control of processing nodes
US10877533B2 (en) Energy efficient workload placement management using predetermined server efficiency data
KR101748747B1 (en) Controlling configurable peak performance limits of a processor
TWI735531B (en) Processor power monitoring and control with dynamic load balancing
Colin et al. Energy-efficient allocation of real-time applications onto single-ISA heterogeneous multi-core processors
US20170371761A1 (en) Real-time performance tracking using dynamic compilation
US9846475B2 (en) Controlling power consumption in multi-core environments
WO2017151276A1 (en) Hierarchical autonomous capacitance management
US20150095622A1 (en) Apparatus and method for controlling execution of processes in a parallel computing system
US8457805B2 (en) Power distribution considering cooling nodes
US9864423B2 (en) Mitigating component performance variation
US9367422B2 (en) Determining and using power utilization indexes for servers
US10223171B2 (en) Mitigating load imbalances through hierarchical performance balancing
US20220147127A1 (en) Power level of central processing units at run time
US20190317461A1 (en) Distributed multi-input multi-output control theoretic method to manage heterogeneous systems
US12443877B2 (en) Calculator, deep learning method and computer-readable recording medium storing program for deep learning
US20240004448A1 (en) Platform efficiency tracker
Lösch et al. A highly accurate energy model for task execution on heterogeneous compute nodes
KR20240044533A (en) Global integrated circuit power control
Anupindi Proactive thermal-aware scheduling
CN117980862A (en) Global IC Power Control

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16890101

Country of ref document: EP

Kind code of ref document: A2

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16890101

Country of ref document: EP

Kind code of ref document: A2