WO2017172050A1 - Method and apparatus to improve energy efficiency of parallel tasks - Google Patents

Method and apparatus to improve energy efficiency of parallel tasks Download PDF

Info

Publication number
WO2017172050A1
WO2017172050A1 PCT/US2017/017023 US2017017023W WO2017172050A1 WO 2017172050 A1 WO2017172050 A1 WO 2017172050A1 US 2017017023 W US2017017023 W US 2017017023W WO 2017172050 A1 WO2017172050 A1 WO 2017172050A1
Authority
WO
WIPO (PCT)
Prior art keywords
core
submessage
task
message
power state
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2017/017023
Other languages
French (fr)
Inventor
Devadatta V. Bodas
Muralidhar Rajappa
Justin J. Song
Andy Hoffman
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Intel Corp
Original Assignee
Intel Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Intel Corp filed Critical Intel Corp
Publication of WO2017172050A1 publication Critical patent/WO2017172050A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • G06F1/3234Power saving characterised by the action undertaken
    • G06F1/324Power saving characterised by the action undertaken by lowering clock frequency
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • G06F1/3206Monitoring of events, devices or parameters that trigger a change in power modality
    • G06F1/3228Monitoring task completion, e.g. by use of idle timers, stop commands or wait commands
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • Embodiments of the invention relate to high-performance computing. More specifically, embodiments of the invention relate to improving power consumption characteristics in system executing parallel tasks.
  • HPC high- performance computing
  • Typical HPC environments divide the processing task between a number of different computing cores so these tasks can be performed in parallel.
  • data is required to be exchanged between the different tasks.
  • Such times are generally referred to as "synchronization points" because they require that the tasks be synchronized, that is, have reached the same point in execution so that the exchanged data is valid.
  • synchronization points because they require that the tasks be synchronized, that is, have reached the same point in execution so that the exchanged data is valid.
  • early arrivers must wait for the other tasks to get to that synchronization point.
  • the task calls a wait routine and executes a spin loop until other tasks arrive at the same synchronization point.
  • the core continues to consume significant energy.
  • Figure 1 is a block diagram of a system according to one embodiment of the invention.
  • Figures 2A-2D show timing diagrams of operation according to embodiments of the invention.
  • Figure 3 is a flow diagram of operation of a system according to one embodiment of the invention.
  • FIG. 1 is a block diagram of a system according to one embodiment of the invention.
  • a plurality of processing cores 102-1, 102-2, 102-n are provided to process tasks in parallel.
  • core 102-1 processes task 112-1
  • core 102-2 processes task 112-2
  • core 102-N processes task 112-N.
  • "Task,” “thread” and “process” are use herein interchangeably to refer to and instance of executable software or hardware.
  • the number of cores 102 can be arbitrarily large.
  • Each core 120 includes a corresponding power management agent 114-1, 114-2-, 114-N (generically, power management agent 114).
  • Power management agent 114 may be implemented as software, hardware, microcode etc.
  • the power management agent 114 is used to place its core 120 in a lower power state when it reaches a synchronization point before other cores 120 processing other tasks 112.
  • synchronization point refers to any point in the processing where the further processing is dependent on receipt of data from another core in the system.
  • An inter-core messaging unit 104 provides messaging services between the different processing cores 120.
  • inter-core messaging unit 104 adheres to the message passing interface (MPI) protocol.
  • MPI message passing interface
  • core 102-1 When core 102-1 reaches a synchronization point, it calls a wait routine. For example, it may call MPI- wait from the inter-core messaging unit 104.
  • the power management agent 114-1 responsive to the call of the wait messaging routine, transitions core 102-1 into a lower power state. This may take the form of reducing core and/or its power domain power by employing whatever applicable power saving technology such as DVFS (dynamic voltage frequency scaling), gating, parking, offlining, throttling, non-active states, or standby states.
  • DVFS dynamic voltage frequency scaling
  • the power management agent 114 includes or has access to a timer 116-1, 116-2, 116-N, respectively (generically, timer 116) that delays entry into the low power state for a threshold period.
  • timer 116 Generally, there is a certain amount of overhead in entering and leaving the low power state. It has been found empirically that if all cores reach a synchronization point within a relatively short period of time, power consumption characteristics are not meaningfully improved, and in some cases, are diminished by immediate transition upon the call of the wait routine.
  • spin loop refers to either a legacy spin-loop (checking one flag and immediately going to itself) or any other low latency state which can be immediately exited once a condition that caused a thread/core to wait has been met.
  • core 102- 1 may be waiting for a message Mi from core 102-2.
  • message Mi is sent to inter-core messaging unit 104. If message Mi exceeds a threshold length, inter-core messaging unit 104 subdivides the message into two submessages using a message subdivision unit 124. Submessages Mi 'and Mi "are sent sequentially to core 102- 1. Each submessage includes its own message validation value, such as cyclic redundancy check (CRC) values, to allow the submessage to be validated individually. By subdividing the message, power savings can be achieved while improving processing performance.
  • CRC cyclic redundancy check
  • power management agent 114-1 can transition core 102- 1 into a higher power state once message Mi ' is received and validated without waiting for the entire message (the remainder Mi ") to be received.
  • core 102- 1 exits the spin loop or enters the higher power state sooner so there is less power churn, and begins processing message Mi ' while receiving message Mi " .
  • Mi " fails to validate, core 102- 1 will need to invalidate message Mi ' and request retransmission of the entire Mi message, but as message failure transmissions are relatively infrequent, improved power savings and execution by virtue of the message subdivision generally results.
  • FIGS. 2A-2D show timing diagrams of operation according to
  • FIG. 2A four tasks, task 1, task 2, task 3 and task 4 are shown as part of the execution environment. As shown, each of tasks 1 , 2 and 3 finish (reach a synchronization point) before task 4. In this example, each thread executing the
  • FIG. 2B is the same as Figure 2A, except that tasks 1 , 2 and 3 each wait a hold off delay before entering the lower power state. During the delay, each task enters a spin loop during the delay.
  • Figure 2C and 2D show behavior of the system with short and long messages respectively. Empirically message traffic tends to be bimodal characterized by either short or long messages. Where the messages are long message division can provide additional power savings.
  • Figure 2C shows a receiving task R waiting for a sending task S to send a message. When the message is receiving and validated, it exits the low power state, and the receiving task, task R, resumes. However, there is a finite delay to exit the low power state after the message has been received and validated. This is reflective of appropriate behavior when the message is relatively short.
  • Figure 2D shows an embodiment for messages subdivided into two submessages, MSG 1 and MSG 2. This allows task R to initiate the transition from the low power state upon validation of MSG 1.
  • FIG. 3 is a flow diagram of operation of a system according to one embodiment of the invention.
  • a core completes its task (arrives at a
  • a delay timer is triggered to hold off entry into a lower power state. Some embodiments may omit the delay timer or have the delay set to zero.
  • a determination is made if the delay threshold has been achieved. As noted previously, system designers may select the delay threshold based on the overhead of entry and leaving the low power state. If the delay threshold has been achieved, the core is transitioned into the lower power state (LPS) at block 310.
  • LPS lower power state
  • Some embodiments pertain to a multi-processing-core system in which a plurality of processing cores are each to execute a respective task.
  • An inter-core messaging unit is to convey messages between the cores.
  • a power management agent in use, transitions a first core into a lower power state responsive to the first core waiting for a second core to complete a second task.
  • the system uses a delay timer to hold off a transitioning into the lower power state for a defined time after the first core begins waiting.
  • the messaging unit segments a message into a first submessage and a second submessage for transmission to the first core and the power management agent transitions the first core to a higher power state responsive to receipt and validation of the first submessage.
  • the first core begins processing the first submessage while the second submessage is being received.
  • Some embodiments pertain to a method to reduce power consumption in a multi-core execution environment.
  • a first core is transitioned into a reduced power state responsive to a task executing on the core reaching a synchronization point, a threshold before a second task executing on a second core reaches the synchronization point.
  • the first core is returned to a higher power state responsive to the second task reaching the synchronization point.
  • a delay timer is initiated to hold off the transition until after the threshold.
  • the delay timer is triggered based on a call of a messaging wait routine from the first task.
  • the message when a message is to be sent to the first task in the reduced power state, the message is subdivided into a first submessage and a second submessage, each with a correction code value.
  • the first submessage and the second submessage are sentto the first core.
  • the return to the higher power state is initiated once the first submessage is validated.
  • the first submessage begins processing while receiving the second submessage and the first submessage is invalidated responsive to a validation failure of the second submessage.
  • Some embodiments pertain to a method of reducing power consumption into a multi-core messaging.
  • a first task is placed in a spin loop to wait for a message from a second task.
  • the message is subdivided into a first submessage and a second submessage.
  • the first task exits the spin loop responsive to a validation of the first submessage.
  • the first submessage is processed in the first task while receiving the second submessage and invalidate responsive to a validation failure in the second submessage.
  • Some embodiments pertain to a non-transitory computer-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform a set of operations to reduce power consumption in a multi-core execution environment.
  • a first core is transitioned into a reduced power state responsive to a task executing on the core reaching a synchronization point, a threshold before a second task executing on a second core reaches the synchronization point.
  • the first core is returned to a higher power state responsive to the second task reaching the synchronization point.
  • a delay timer is initiated to hold off the transition until after the threshold. In further embodiments, the delay timer is triggered based on a call of a messaging wait routine from the first task.
  • the message when a message is to be sent to the first task in the reduced power state, the message is subdivided into a first submessage and a second submessage, each with a correction code value.
  • the first submessage and the second submessage are sentto the first core.
  • the return to the higher power state is initiated once the first submessage is validated.
  • the first submessage begins processing while receiving the second submessage and the first submessage is invalidated responsive to a validation failure of the second submessage.
  • Some embodiments pertain to high-performance computing system having a plurality of processing cores.
  • the system includes means for inter-core messaging and means for reducing power consumption on a first core when the first core is waiting for a second core.
  • the means for inter-core messaging has means for subdividing a message into a first submessage and a second submessage, and wherein a receiving core begins processing of the first submessage before the second submessage is fully received.
  • the means for reducing power consumption includes means for reducing a clock frequency in the processing core.
  • the means for reducing power consumption includes means for transitioning the waiting core into a lower power state.
  • Elements of embodiments of the present invention may also be provided as a machine-readable medium for storing the machine-executable instructions.
  • the machine- readable medium may include, but is not limited to, flash memory, optical disks, compact disks read only memory (CD-ROM), digital versatile/video disks (DVD) ROM, random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Power Sources (AREA)
  • Computer Hardware Design (AREA)
  • Computing Systems (AREA)

Abstract

A system with improved power performance for task executed in parallel. A plurality of processing cores each to execute tasks. An inter-core messaging unit to conveys messages between the cores. A power management agent transitions a first core into a lower power state responsive to the first core waiting for a second core to complete a second task. In some embodiments long messages are subdivided to allow a receiving core to resume useful work sooner.

Description

METHOD AND APPARATUS TO IMPROVE ENERGY EFFICIENCY OF
PARALLEL TASKS
FIELD
Embodiments of the invention relate to high-performance computing. More specifically, embodiments of the invention relate to improving power consumption characteristics in system executing parallel tasks.
BACKGROUND
Generally, power consumption has become an important issue in high- performance computing (HPC). Typical HPC environments divide the processing task between a number of different computing cores so these tasks can be performed in parallel. At different points, data is required to be exchanged between the different tasks. Such times are generally referred to as "synchronization points" because they require that the tasks be synchronized, that is, have reached the same point in execution so that the exchanged data is valid. Because all tasks do not require the same amount of time to reach the synchronization point, early arrivers must wait for the other tasks to get to that synchronization point.
Generally, the task calls a wait routine and executes a spin loop until other tasks arrive at the same synchronization point. Unfortunately, in the spin loop, the core continues to consume significant energy.
BRIEF DESCRIPTION OF THE DRAWINGS
Embodiments of the invention are illustrated by way of example and not by way of limitation in the figures of the accompanying drawings in which like references indicate similar elements. It should be noted that different references to "an" or "one" embodiment in this disclosure are not necessarily to the same embodiment, and such references mean at least one.
Figure 1 is a block diagram of a system according to one embodiment of the invention.
Figures 2A-2D show timing diagrams of operation according to embodiments of the invention. Figure 3 is a flow diagram of operation of a system according to one embodiment of the invention.
DETAILED DESCRIPTION
Figure 1 is a block diagram of a system according to one embodiment of the invention. A plurality of processing cores 102-1, 102-2, 102-n (generically, core 102) are provided to process tasks in parallel. For example, core 102-1 processes task 112-1, core 102-2 processes task 112-2, and core 102-N processes task 112-N. "Task," "thread" and "process" are use herein interchangeably to refer to and instance of executable software or hardware. The number of cores 102 can be arbitrarily large. Each core 120 includes a corresponding power management agent 114-1, 114-2-, 114-N (generically, power management agent 114). Power management agent 114 may be implemented as software, hardware, microcode etc. The power management agent 114 is used to place its core 120 in a lower power state when it reaches a synchronization point before other cores 120 processing other tasks 112. As used herein, "synchronization point" refers to any point in the processing where the further processing is dependent on receipt of data from another core in the system.
An inter-core messaging unit 104 provides messaging services between the different processing cores 120. In one embodiment, inter-core messaging unit 104 adheres to the message passing interface (MPI) protocol. When core 102-1 reaches a synchronization point, it calls a wait routine. For example, it may call MPI- wait from the inter-core messaging unit 104. In one embodiment, responsive to the call of the wait messaging routine, the power management agent 114-1 transitions core 102-1 into a lower power state. This may take the form of reducing core and/or its power domain power by employing whatever applicable power saving technology such as DVFS (dynamic voltage frequency scaling), gating, parking, offlining, throttling, non-active states, or standby states. In other embodiments, the power management agent 114 includes or has access to a timer 116-1, 116-2, 116-N, respectively (generically, timer 116) that delays entry into the low power state for a threshold period. Generally, there is a certain amount of overhead in entering and leaving the low power state. It has been found empirically that if all cores reach a synchronization point within a relatively short period of time, power consumption characteristics are not meaningfully improved, and in some cases, are diminished by immediate transition upon the call of the wait routine.
However, since tasks 120 may execute for minutes, hours, or even longer beyond some relatively short threshold, the power savings of transitioning to a lower power state are quite significant. In embodiments where the timer 116-1 is present, the core 102-1 will still enter the spin loop until transitioned into a lower power state. As used herein, "spin loop" refers to either a legacy spin-loop (checking one flag and immediately going to itself) or any other low latency state which can be immediately exited once a condition that caused a thread/core to wait has been met.
In one example, core 102- 1 may be waiting for a message Mi from core 102-2. In one embodiment, message Mi is sent to inter-core messaging unit 104. If message Mi exceeds a threshold length, inter-core messaging unit 104 subdivides the message into two submessages using a message subdivision unit 124. Submessages Mi 'and Mi "are sent sequentially to core 102- 1. Each submessage includes its own message validation value, such as cyclic redundancy check (CRC) values, to allow the submessage to be validated individually. By subdividing the message, power savings can be achieved while improving processing performance. This is because power management agent 114-1 can transition core 102- 1 into a higher power state once message Mi ' is received and validated without waiting for the entire message (the remainder Mi ") to be received. Thus, core 102- 1 exits the spin loop or enters the higher power state sooner so there is less power churn, and begins processing message Mi ' while receiving message Mi " . Of course, if Mi " fails to validate, core 102- 1 will need to invalidate message Mi ' and request retransmission of the entire Mi message, but as message failure transmissions are relatively infrequent, improved power savings and execution by virtue of the message subdivision generally results.
Figures 2A-2D show timing diagrams of operation according to
embodiments of the invention. In Figure 2A, four tasks, task 1, task 2, task 3 and task 4 are shown as part of the execution environment. As shown, each of tasks 1 , 2 and 3 finish (reach a synchronization point) before task 4. In this example, each thread executing the
corresponding task transitions to a lower power state immediately when it reaches a respective synchronization point. Figure 2B is the same as Figure 2A, except that tasks 1 , 2 and 3 each wait a hold off delay before entering the lower power state. During the delay, each task enters a spin loop during the delay.
Figure 2C and 2D show behavior of the system with short and long messages respectively. Empirically message traffic tends to be bimodal characterized by either short or long messages. Where the messages are long message division can provide additional power savings. Figure 2C shows a receiving task R waiting for a sending task S to send a message. When the message is receiving and validated, it exits the low power state, and the receiving task, task R, resumes. However, there is a finite delay to exit the low power state after the message has been received and validated. This is reflective of appropriate behavior when the message is relatively short. Figure 2D shows an embodiment for messages subdivided into two submessages, MSG 1 and MSG 2. This allows task R to initiate the transition from the low power state upon validation of MSG 1. This allows task R to resume processing and begin processing of MSG 1 while receiving MSG 2, thereby improving performance. Even in a system where task R is merely residing in a spin loop, this message subdivision can improve power because the time spent in the spin loop (not performing any useful work) is reduced over systems in which task R waits in a spin loop for the receipt of the entire lengthy message (here, the composition of MSG 1 and MSG2). This behavior is suitable where the message is long.
Figure 3 is a flow diagram of operation of a system according to one embodiment of the invention. At block 302, a core completes its task (arrives at a
synchronization point) and enters a wait condition. At block 304, the core notifies a messaging unit that it is waiting. At block 306, a delay timer is triggered to hold off entry into a lower power state. Some embodiments may omit the delay timer or have the delay set to zero. At decision block 308, a determination is made if the delay threshold has been achieved. As noted previously, system designers may select the delay threshold based on the overhead of entry and leaving the low power state. If the delay threshold has been achieved, the core is transitioned into the lower power state (LPS) at block 310.
At decision block 312, a determination is made whether there is a long message directed at the waiting core. Empirically, as noted above, it has been found that most messages fall into a bimodal length distribution, that is, most messages are either very short, or quite long. If the message is not long, it is processed normally at block 314, that is, the message is not subdivided and is merely sent as a unit. Then, at block 316, the core transitions to the higher power state or active state at block 316 once the entire message has been validated. Conversely, at block 318, if the message is a "long message," the message is subdividing into a first and second submessage, each with its validation values. At block 320, the core receives and validates the first submessage. A determination is made at block 322 if the first submessage is valid. If the first submessage is not valid, the core remains in a lower power state and waits for subsequent valid message receipt. If the first message is valid, upon validating that first message, the core transitions into a higher power/active state, that is, it goes to a higher power, possibly CO state, or exits a spin loop, for example. Then, at block 326, the core processes the submessage while receiving the second submessage. At block 328, a determination is made if the second submessage is valid. If the second submessage is not valid, the core invalidates both submessages and requests they be resent at block 330. If, however, the second submessage is valid (the usual case), it continues processing in the normal manner at block 332.
The following examples pertain to further embodiments. The various features of the different embodiments may be variously combined with some features included and others excluded to suit a variety of different applications. Some embodiments pertain to a multi-processing-core system in which a plurality of processing cores are each to execute a respective task. An inter-core messaging unit is to convey messages between the cores. A power management agent, in use, transitions a first core into a lower power state responsive to the first core waiting for a second core to complete a second task.
In further embodiments, the system uses a delay timer to hold off a transitioning into the lower power state for a defined time after the first core begins waiting.
In further embodiments, the the messaging unit segments a message into a first submessage and a second submessage for transmission to the first core and the power management agent transitions the first core to a higher power state responsive to receipt and validation of the first submessage.
In further embodiments, the first core begins processing the first submessage while the second submessage is being received.
Some embodiments pertain to a method to reduce power consumption in a multi-core execution environment. A first core is transitioned into a reduced power state responsive to a task executing on the core reaching a synchronization point, a threshold before a second task executing on a second core reaches the synchronization point. The first core is returned to a higher power state responsive to the second task reaching the synchronization point. In further embodiments, a delay timer is initiated to hold off the transition until after the threshold.
In further embodiments, the delay timer is triggered based on a call of a messaging wait routine from the first task.
In further embodiments, when a message is to be sent to the first task in the reduced power state, the message is subdivided into a first submessage and a second submessage, each with a correction code value. The first submessage and the second submessage are sentto the first core. The return to the higher power state is initiated once the first submessage is validated.
In further embodiments, the first submessage begins processing while receiving the second submessage and the first submessage is invalidated responsive to a validation failure of the second submessage.
Some embodiments pertain to a method of reducing power consumption into a multi-core messaging. A first task is placed in a spin loop to wait for a message from a second task. The message is subdivided into a first submessage and a second submessage. The first task exits the spin loop responsive to a validation of the first submessage.
In further embodiments, the first submessage is processed in the first task while receiving the second submessage and invalidate responsive to a validation failure in the second submessage.
Some embodiments pertain to a non-transitory computer-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform a set of operations to reduce power consumption in a multi-core execution environment. A first core is transitioned into a reduced power state responsive to a task executing on the core reaching a synchronization point, a threshold before a second task executing on a second core reaches the synchronization point. The first core is returned to a higher power state responsive to the second task reaching the synchronization point.
In further embodiments, a delay timer is initiated to hold off the transition until after the threshold. In further embodiments, the delay timer is triggered based on a call of a messaging wait routine from the first task.
In further embodiments, when a message is to be sent to the first task in the reduced power state, the message is subdivided into a first submessage and a second submessage, each with a correction code value. The first submessage and the second submessage are sentto the first core. The return to the higher power state is initiated once the first submessage is validated.
In further embodiments, the first submessage begins processing while receiving the second submessage and the first submessage is invalidated responsive to a validation failure of the second submessage.
Some embodiments pertain to high-performance computing system having a plurality of processing cores. The system includes means for inter-core messaging and means for reducing power consumption on a first core when the first core is waiting for a second core.
In further embodiments, the means for inter-core messaging has means for subdividing a message into a first submessage and a second submessage, and wherein a receiving core begins processing of the first submessage before the second submessage is fully received.
In further embodiments, the means for reducing power consumption includes means for reducing a clock frequency in the processing core.
In further embodiments, the means for reducing power consumption includes means for transitioning the waiting core into a lower power state.
While embodiments of the invention are discussed above in the context of flow diagrams reflecting a particular linear order, this is for convenience only. In some cases, various operations may be performed in a different order than shown or various operations may occur in parallel. It should also be recognized that some operations described with respect to one embodiment may be advantageously incorporated into another embodiment. Such incorporation is expressly contemplated.
Elements of embodiments of the present invention may also be provided as a machine-readable medium for storing the machine-executable instructions. The machine- readable medium may include, but is not limited to, flash memory, optical disks, compact disks read only memory (CD-ROM), digital versatile/video disks (DVD) ROM, random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards.
In the foregoing specification, the invention has been described with reference to the specific embodiments thereof. It will, however, be evident that various modifications and changes can be made thereto without departing from the broader spirit and scope of the invention as set forth in the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.

Claims

CLAIMS What is claimed is:
1. A multi-processing-core system:
a plurality of processing cores, each core to execute a task;
an inter-core messaging unit to convey messages between the cores; and
a power management agent to transition a first core into a lower power state responsive to the first core waiting for a second core to complete a second task.
2. The system of claim 1, further comprising:
a delay timer to hold off a transitioning into the lower power state for a defined time after the first core begins waiting.
3. The system of claim 2, wherein the messaging unit segments a message into a first submessage and a second submessage for transmission to the first core, and wherein the power management agent transitions the first core to a higher power state responsive to receipt and validation of the first submessage.
4. The system of claim 3, wherein the first core begins processing the first submessage while the second submessage is being received.
5. A method to reduce power consumption in a multi-core execution environment, the method comprising:
transitioning a first core into a reduced power state responsive to a task executing on the core reaching a synchronization point, a threshold before a second task executing on a second core reaches the synchronization point; and
returning the first core to a higher power state responsive to the second task reaching the synchronization point.
6. The method of claim 5, further comprising:
initiating a delay timer to hold off the transition until after the threshold.
7. The method of claim 6, further comprising:
triggering the delay timer based on a call of a messaging wait routine from the first task.
8. The method of any of claims 5-7, wherein a message is to be sent to the first task in the reduced power state, the method further comprising:
subdividing the message into a first submessage and a second submessage, each with a correction code value;
sending the first submessage and the second submessage to the first core; and initiating the return to the higher power state once the first submessage is validated.
9. The method of claim 8, further comprising:
processing the first submessage while receiving the second submessage; and invalidating the first submessage responsive to a validation failure of the second submessage.
10. A method of reducing power consumption into a multi-core messaging system:
placing a first task in a spin loop to wait for a message from a second task;
subdividing the message into a first submessage and a second submessage; and exiting the spin loop in the first task responsive to a validation of the first submessage.
11. The method of claim 10, further comprising:
processing the first submessage in the first task while receiving the second submessage; and
invalidating the first submessage responsive to a validation failure in the second submessage.
12. A non-transitory computer-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform a set of operations to reduce power consumption in a multi-core execution environment, comprising:
transitioning a first core into a reduced power state responsive to a task executing on the core reaching a synchronization point, a threshold before a second task executing on a second core reaches the synchronization point; and
returning the first core to a higher power state responsive to the second task reaching the synchronization point.
13. The non-transitory computer-readable medium of claim 12, wherein the instructions cause the processor to perform a set of operations further comprising:
initiating a delay timer to hold off the transition until after the threshold.
14. The non-transitory computer-readable medium of claim 13, wherein the instructions cause the processor to perform a set of operations further comprising:
triggering the delay timer based on a call of a messaging wait routine from the first task.
15. The non-transitory computer-readable medium of any of claim 12-14, wherein the instructions cause the processor to perform a set of operations further comprising:
sending a message to the first task in the reduced power state;
subdividing the message to a first submessage and a second submessage, each with a correction code value;
sending the first submessage and the second submessage to the first core; and initiating the return to the higher power state once the first submessage is validated.
16. The non-transitory computer-readable medium of claim 15, wherein the instructions cause the processor to perform a set of operations further comprising:
processing the first submessage while receiving the second submessage; and
invalidating the first submessage responsive to a validation failure of the second submessage.
17. A high-performance computing system comprising:
a plurality of processing cores;
means for inter-core messaging; and
means for reducing power consumption on a first core when the first core is waiting for a second core.
18. The system of claim 17, wherein the means for inter-core messaging comprises:
means for subdividing a message into a first submessage and a second submessage, and wherein a receiving core begins processing of the first submessage before the second submessage is fully received.
19. The system of either of claim 17 or 18, wherein the means for reducing power consumption comprises:
means for reducing a clock frequency in the processing core.
20. The system of either of claim 17 or 18, wherein the means for reducing power consumption comprises:
means for transitioning the waiting core into a lower power state.
PCT/US2017/017023 2016-03-31 2017-02-08 Method and apparatus to improve energy efficiency of parallel tasks Ceased WO2017172050A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US15/087,095 2016-03-31
US15/087,095 US10996737B2 (en) 2016-03-31 2016-03-31 Method and apparatus to improve energy efficiency of parallel tasks

Publications (1)

Publication Number Publication Date
WO2017172050A1 true WO2017172050A1 (en) 2017-10-05

Family

ID=59961478

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2017/017023 Ceased WO2017172050A1 (en) 2016-03-31 2017-02-08 Method and apparatus to improve energy efficiency of parallel tasks

Country Status (2)

Country Link
US (2) US10996737B2 (en)
WO (1) WO2017172050A1 (en)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20250155955A1 (en) * 2023-11-10 2025-05-15 Hewlett Packard Enterprise Development Lp Region-aware power & energy regulation
US20250370757A1 (en) * 2024-05-31 2025-12-04 Dell Products L.P. Power savings during parallel synchronization for distributed memory systems by using different processor states

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20100169528A1 (en) * 2008-12-30 2010-07-01 Amit Kumar Interrupt technicques
US20130047011A1 (en) * 2011-08-19 2013-02-21 David Dice System and Method for Enabling Turbo Mode in a Processor
US8578079B2 (en) * 2009-05-13 2013-11-05 Apple Inc. Power managed lock optimization
US20150067356A1 (en) * 2013-08-30 2015-03-05 Advanced Micro Devices, Inc. Power manager for multi-threaded data processor
US20150113304A1 (en) * 2013-10-22 2015-04-23 Wisconsin Alumni Research Foundation Energy-efficient multicore processor architecture for parallel processing

Family Cites Families (35)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPWO2003083693A1 (en) 2002-04-03 2005-08-04 富士通株式会社 Task scheduling device in distributed processing system
US7210048B2 (en) 2003-02-14 2007-04-24 Intel Corporation Enterprise power and thermal management
US20060241880A1 (en) 2003-07-18 2006-10-26 Forth J B Methods and apparatus for monitoring power flow in a conductor
US7421623B2 (en) 2004-07-08 2008-09-02 International Business Machines Corporation Systems, methods, and media for controlling temperature in a computer system
US20060107262A1 (en) 2004-11-03 2006-05-18 Intel Corporation Power consumption-based thread scheduling
US20080172398A1 (en) 2007-01-12 2008-07-17 Borkenhagen John M Selection of Processors for Job Scheduling Using Measured Power Consumption Ratings
US7739388B2 (en) 2007-05-30 2010-06-15 International Business Machines Corporation Method and system for managing data center power usage based on service commitments
US7724149B2 (en) 2007-06-11 2010-05-25 Hewlett-Packard Development Company, L.P. Apparatus, and associated method, for selecting distribution of processing tasks at a multi-processor data center
US20090037926A1 (en) 2007-08-01 2009-02-05 Peter Dinda Methods and systems for time-sharing parallel applications with performance isolation and control through performance-targeted feedback-controlled real-time scheduling
US7941681B2 (en) 2007-08-17 2011-05-10 International Business Machines Corporation Proactive power management in a parallel computer
US20090070611A1 (en) 2007-09-12 2009-03-12 International Business Machines Corporation Managing Computer Power Consumption In A Data Center
US8555283B2 (en) 2007-10-12 2013-10-08 Oracle America, Inc. Temperature-aware and energy-aware scheduling in a computer system
WO2009052121A2 (en) 2007-10-14 2009-04-23 Gilbert Masters Electrical energy usage monitoring system
US7979729B2 (en) 2007-11-29 2011-07-12 International Business Machines Corporation Method for equalizing performance of computing components
US8001403B2 (en) 2008-03-14 2011-08-16 Microsoft Corporation Data center power management utilizing a power policy and a load factor
CN102395937B (en) 2009-04-17 2014-06-11 惠普开发有限公司 Power capping system and method
JP5549131B2 (en) 2009-07-07 2014-07-16 富士通株式会社 Job allocation apparatus, job allocation method, and job allocation program
US8397088B1 (en) 2009-07-21 2013-03-12 The Research Foundation Of State University Of New York Apparatus and method for efficient estimation of the energy dissipation of processor based systems
US8793348B2 (en) 2009-09-18 2014-07-29 Group Business Software Ag Process for installing software application and platform operating system
US20110138395A1 (en) 2009-12-08 2011-06-09 Empire Technology Development Llc Thermal management in multi-core processor
US9292662B2 (en) 2009-12-17 2016-03-22 International Business Machines Corporation Method of exploiting spare processors to reduce energy consumption
US8443373B2 (en) 2010-01-26 2013-05-14 Microsoft Corporation Efficient utilization of idle resources in a resource manager
JP5621287B2 (en) 2010-03-17 2014-11-12 富士通株式会社 Load balancing system and computer program
US8738195B2 (en) 2010-09-21 2014-05-27 Intel Corporation Inferencing energy usage from voltage droop
US8849469B2 (en) 2010-10-28 2014-09-30 Microsoft Corporation Data center system that accommodates episodic computation
US8677158B2 (en) 2011-08-10 2014-03-18 Microsoft Corporation System and method for assigning a power management classification including exempt, suspend, and throttling to an process based upon various factors of the process
US20130054987A1 (en) 2011-08-29 2013-02-28 Clemens Pfeiffer System and method for forcing data center power consumption to specific levels by dynamically adjusting equipment utilization
US9003216B2 (en) 2011-10-03 2015-04-07 Microsoft Technology Licensing, Llc Power regulation of power grid via datacenter
WO2013119195A1 (en) 2012-02-06 2013-08-15 Empire Technology Development Llc Multicore computer system with cache use based adaptive scheduling
KR20140140636A (en) 2012-05-14 2014-12-09 인텔 코오퍼레이션 Managing the operation of a computing system
US10162687B2 (en) 2012-12-28 2018-12-25 Intel Corporation Selective migration of workloads between heterogeneous compute elements based on evaluation of migration performance benefit and available energy and thermal budgets
US9152469B2 (en) 2013-01-28 2015-10-06 Hewlett-Packard Development Company, L.P. Optimizing execution and resource usage in large scale computing
US9529642B2 (en) 2013-03-28 2016-12-27 Vmware, Inc. Power budget allocation in a cluster infrastructure
CN104252391B (en) * 2013-06-28 2017-09-12 国际商业机器公司 Method and apparatus for managing multiple operations in distributed computing system
EP2937783B1 (en) * 2014-04-24 2018-08-15 Fujitsu Limited A synchronisation method

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20100169528A1 (en) * 2008-12-30 2010-07-01 Amit Kumar Interrupt technicques
US8578079B2 (en) * 2009-05-13 2013-11-05 Apple Inc. Power managed lock optimization
US20130047011A1 (en) * 2011-08-19 2013-02-21 David Dice System and Method for Enabling Turbo Mode in a Processor
US20150067356A1 (en) * 2013-08-30 2015-03-05 Advanced Micro Devices, Inc. Power manager for multi-threaded data processor
US20150113304A1 (en) * 2013-10-22 2015-04-23 Wisconsin Alumni Research Foundation Energy-efficient multicore processor architecture for parallel processing

Also Published As

Publication number Publication date
US20210365096A1 (en) 2021-11-25
US11435809B2 (en) 2022-09-06
US20170285717A1 (en) 2017-10-05
US10996737B2 (en) 2021-05-04

Similar Documents

Publication Publication Date Title
US20190042331A1 (en) Power aware load balancing using a hardware queue manager
KR101569610B1 (en) Platform agnostic power management
CN108885486B (en) Enhanced Dynamic Clock and Voltage Scaling (DCVS) Scheme
RU2651238C2 (en) Synchronization of interrupt processing for energy consumption reduction
US20150205671A1 (en) Dynamic Checkpointing Systems and Methods
WO2017014913A1 (en) Systems and methods for scheduling tasks in a heterogeneous processor cluster architecture using cache demand monitoring
JP2022500777A (en) Accelerate or suppress processor loop mode using loop end prediction
US10467054B2 (en) Resource management method and system, and computer storage medium
US11435809B2 (en) Method and apparatus to improve energy efficiency of parallel tasks
US12164450B2 (en) Managing network interface controller-generated interrupts
US20160004654A1 (en) System for migrating stash transactions
US10127076B1 (en) Low latency thread context caching
EP2846217B1 (en) Controlling reduced power states using platform latency tolerance
CN109491780B (en) Multi-task scheduling method and device
JP2013149221A (en) Control device for processor and method for controlling the same
WO2013159464A1 (en) Multiple core processor clock control device and control method
US9904582B2 (en) Method and apparatus for executing software in electronic device
US20130346701A1 (en) Replacement method and apparatus for cache
CN105068872B (en) The control method and system of arithmetic element
JP6236996B2 (en) Information processing apparatus and information processing apparatus control method
US20240213987A1 (en) Ip frequency adaptive same-cycle clock gating
EP3702911A3 (en) Hardware for supporting os driven load anticipation based on variable sized load units
US9766885B2 (en) System, method, and storage medium
WO2013147878A1 (en) Prediction-based thread selection in a multithreading processor
CN116627622A (en) Data processing method, device, equipment and storage medium

Legal Events

Date Code Title Description
NENP Non-entry into the national phase

Ref country code: DE

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17776088

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 17776088

Country of ref document: EP

Kind code of ref document: A1