WO2006016451A1 - 演算パイプライン、演算パイプラインによる処理方法、半導体装置、コンピュータプログラム - Google Patents
演算パイプライン、演算パイプラインによる処理方法、半導体装置、コンピュータプログラム Download PDFInfo
- Publication number
- WO2006016451A1 WO2006016451A1 PCT/JP2005/011482 JP2005011482W WO2006016451A1 WO 2006016451 A1 WO2006016451 A1 WO 2006016451A1 JP 2005011482 W JP2005011482 W JP 2005011482W WO 2006016451 A1 WO2006016451 A1 WO 2006016451A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- retiming
- data
- processing
- control data
- arithmetic
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3867—Concurrent instruction execution, e.g. pipeline or look ahead using instruction pipelines
- G06F9/3869—Implementation aspects, e.g. pipeline latches; pipeline synchronisation and clocking
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F1/00—Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
- G06F1/26—Power supply means, e.g. regulation thereof
- G06F1/32—Means for saving power
Definitions
- Arithmetic pipeline processing method by arithmetic pipeline, semiconductor device, computer program
- the present invention relates to a semiconductor device, and more particularly to a technique for suppressing power consumption of an arithmetic pipeline by the semiconductor device.
- Such heat generation due to increased power consumption of semiconductor devices may be dealt with by reducing the operating voltage.
- reducing the operating voltage it is necessary to lower the threshold voltage to guarantee the operation of the transistor itself.
- lowering the threshold voltage has another problem of increasing leakage current.
- leakage current it is predicted that further miniaturization will progress in the future, and it is predicted that the leakage current will be larger than the operating current in the countermeasures that reduce the operating voltage.
- serial operation pipelines are suitable for applications that require discrete operations that suddenly and explode, such as computer graphics that represent frequently moving objects, such as instruction fetch (fe tch), It is intended to speed up processing by sequentially performing separate tasks such as instruction decoding and execution, and cascades arithmetic units that handle a small number of instructions. Composed.
- Various operations such as addition / subtraction, floating-point operation, comparison, Boolean algebra, selection (IF statement), etc. can be realized by appropriately changing the combination of multiple arithmetic units connected in cascade.
- serial computation pipelines tend to consume more power because they use a large number of computing units.
- An object of the present invention is to provide a technique for suppressing the power consumption of such a conventional arithmetic pipeline while minimizing the deterioration of the performance. Disclosure of the invention
- An arithmetic pipeline that solves the above-described problems includes a processing unit capable of dynamically changing processing capacity and a retiming unit for retiming or squeezing input data alternately. And a pipeline controller that supplies control data for retiming or passing through the processing results of the processing means connected to the previous stage to each retiming means.
- the pipeline controller supplies control data to each retiming means so that a predetermined number of retiming means perform retiming of processing results and the remaining retiming means pass through the processing results.
- the latency of the deductive mechanism is determined according to the number of retiming means that perform retiming. It is configured to Mel so.
- the number of retiming means for performing retiming is determined by control data supplied from the pipeline controller to the arithmetic mechanism. Retiming increases latency. By controlling the number of retiming means that perform retiming, the latency of the calculation mechanism can be controlled, so that power consumption can be reduced while minimizing performance changes.
- the pipeline controller is configured to generate a plurality of the control data according to the processing capability of each processing means.
- the latency of the calculation mechanism changes according to the change in the processing capability of the processing means.
- the pipeline controller The plurality of control data is generated according to Z or the substrate bias value.
- the pipeline controller dynamically changes the latency and threshold voltage according to the operating voltage value and Z or substrate bias value actually supplied to the processing means, thereby changing the performance of the processing device. Power consumption can be minimized while suppressing
- the pipeline controller For example, when each processing means is configured by an element capable of changing performance, and the processing capacity is dynamically changed by the change in the performance of the element, the pipeline controller For example, the control data is configured to be generated according to the performance of the element.
- the pipeline controller In the case where the element constituting the processing unit is configured to change its processing capability in accordance with a change in supplied operating voltage value and Z or substrate bias value, the pipeline controller The plurality of control data is generated according to the operating voltage value and Z or the substrate bias value. Even in such a configuration, the pipeline controller dynamically changes the latency and the threshold voltage in accordance with the operating voltage value and / or the substrate bias value that is actually supplied to the elements of the processing means. Various costs Electric power can be minimized while suppressing changes in performance.
- the pipeline controller is configured to generate the plurality of control data as alternating data or data having a constant value.
- the retiming means is configured to retime the processing result when the control data is alternating, and to slew the processing result when the control data takes a constant value.
- the retiming means retimes the processing result according to the change of “1” and “0” of the control data.
- a processing method using an arithmetic pipeline according to the present invention includes an element that can change performance, and a processing capacity that dynamically changes in accordance with the change of the element, and a processing result by the processing apparatus.
- Retiming or slewing retiming circuits are connected to arithmetic mechanisms that are cascaded alternately, and each retiming circuit is to retime or slew processing results from the processor connected in the previous stage.
- This method is executed by an apparatus that supplies the control data.
- the control for each retiming means is performed so that a predetermined number of retiming circuits perform retiming of processing results and the remaining retiming circuits slew processing results according to the performance of the element. Generate data.
- the semiconductor device of the present invention is configured by an element capable of changing performance, and a processing unit whose processing capability dynamically changes according to the change of the element, and a processing result by the processing unit is retimed.
- each retiming circuit is connected to an arithmetic mechanism that is cascade-connected with each other, and each retiming circuit re-times or slews the processing result of the processor connected in the previous stage.
- a predetermined number of retiming circuits perform retiming of processing results, and the remaining retiming circuits pass through the processing results.
- means for generating the control data for each retiming means are examples of control data for each retiming means.
- the computer program of the present invention is composed of an element capable of changing performance, and a processor whose processing capability dynamically changes in accordance with the change of the element, and a processing result by the processor is retimed or And a retiming circuit that passes through is connected to an arithmetic mechanism that is cascade-connected alternately, Depending on the performance of the element, a predetermined number of retiming circuits reprocess the processing results to a computer that supplies control data for retiming or slewing the processing results of the processing unit connected to the previous stage to the processing circuit
- This is a computer program for executing the processing for generating the control data for each retiming means so that the timing is executed and the remaining retiming circuit passes through the processing result.
- FIG. 1 is a diagram showing a configuration example of an arithmetic pipeline to which the present invention is applied.
- FIG. 2 is a diagram showing SALP when SALC is connected in 8-stage cascade.
- FIG. 3 is a diagram showing S A L P when S A L C is connected in a four-stage cascade connection.
- Fig. 4 is an example of SAL C.
- FIG. 5 is an illustration of a pipeline controller.
- FIG. 6 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 7 is a diagram for explaining processing by SALC.
- FIG. 8 is a diagram for explaining processing by SALC.
- FIG. 9 is a diagram for explaining processing by SALC.
- FIG. 10 is a diagram for explaining processing by SALC.
- FIG. 11 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 12 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 13 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 14 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 15 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 16 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 17 is a diagram for explaining processing using the arithmetic pipeline.
- FIG. 18 is a diagram for explaining processing using the arithmetic pipeline. BEST MODE FOR CARRYING OUT THE INVENTION
- a serial operation pipeline which is an example of an operation mechanism, is configured with an architecture that is resistant to latency fluctuations, and the latency is dynamically controlled.
- the threshold voltage is dynamically controlled by substrate bias technology.
- the latency of the serial operation pipeline is determined by the number of times retiming is performed on the output of the cascade connected arithmetic units in the serial operation pipeline. For example, if retiming is performed once with cascaded computing units, the latency is 2. The latency is 3 if you retime twice. In the present invention, the latency is varied by dynamically changing the number of times retiming is performed in the serial operation pipeline.
- FIG. 1 is a diagram showing a configuration example of an arithmetic pipeline 1 to which the present invention is applied.
- This arithmetic pipeline 1 includes a serial arithmetic pipeline (Serial-ALPipeline, hereinafter referred to as “SALP”) in which a plurality of serial arithmetic logic operation circuits (Serial ALCelK, hereinafter referred to as “SALC”) 3 are cascade-connected. 2 is installed, and a pipeline controller 4 that supplies control data for adjusting the number of stages of SAL C 3 connected in cascade is connected to SALP2.
- SALP 2 is supplied with the operating voltage of each element in the SAL C 3 from the operating voltage supply device 6 and is supplied with the substrate bias of each element in the SAL C 3 from the substrate bias supply device 7.
- a single SAL P 2 is provided for ease of explanation, but a plurality of SALP 2 may be provided.
- control data is supplied from the pipeline controller 4 to each SAL P 2
- operating voltage is supplied from the operating voltage supply device 6 to each SAL P 2.
- the substrate bias is supplied to the SALP 2 from the substrate bias supply device 7.
- SALP 2 implements various multi-bit arithmetic instructions by combining simple serial instructions.
- SALP2 has multiple SAL Cs connected in cascade 3 and a retiming circuit 5 provided after the SALC 3 for latching and retiming output data from the SALC 3.
- the SA LC 3 and the retiming circuit 5 are cascade-connected alternately.
- SALP 2 shown in FIG. 1 has nine SALCs 3 that can be connected to a nine-stage cascade, and a retiming circuit 5 is provided after each S ALC 3.
- the retiming circuit 5 is configured to determine whether the output data from the SALC 3 is latched and retimed or through according to the control data supplied from the pipeline controller 4. When retiming the output data of SALC 3, the retiming circuit 5 latches the output data of SALC 3 and outputs the latched output data at the next operation timing.
- SALC 3 output data When SALC 3 output data is passed through, SALC 3 output data is output as it is without being latched. To the same operation timing.
- a retiming circuit 5 can realize the same function as the latch circuit, for example, can be configured using FF (Flip-flop). It is configured by using the same number of FFs as the number of SALC3 output terminals.
- the SALP 2 is configured by connecting the 9 stages of SALC 3 in cascade.
- the SALP 2 when data is input to SALC 3 at the 1st stage, the data is output from SALC 3 at the 9th stage, so the latency is “1”.
- the number of retiming circuits 5 that perform retiming in this way is “0”, the latency is “1”.
- this SAL P 2 is connected to the SALC with 4 stage power scale connection. It consists of SALC 3 connected in cascade with 3 and 5 stages.
- the data When data is input to the first-stage SALC 3, the data is output from the fourth-stage SALC 3, and the output data is latched by the retiming circuit 5.
- the data When data is input from the retiming circuit 5 to the fifth stage SALC 3 at the next operation timing, the data is output from the ninth stage SAL C 3. In this way, from the input of data to the first stage SALC 3, the ninth stage S The latency is “2” because two operation timings are required until data is output from ALC3.
- the number of retiming circuits 5 that perform retiming is “1”, the latency is “2”.
- the critical path will be extended, and the elements used for SALC 3 will need high performance.
- the interval of the retiming circuit 5 that performs retiming by reducing the number of stages of SALC3 is shortened (the latency is increased)
- the elements used for SAL C 3 need only have low performance. Power consumption can be reduced by using low-performance elements.
- the performance of S A L C 3 is varied by varying the performance of the device by varying the operating voltage and / or substrate bias.
- the operating voltage supply device 6 is a device that supplies an operating voltage to the elements constituting the S AL C 3.
- the supplied operating voltage value is variable, and can be operated from outside the operating voltage supply device 6.
- the operating voltage value may be manipulated artificially, but for example, it may be manipulated by a control device not shown in the figure so that appropriate element performance can be obtained according to the content processed by SAL P 2. Also good.
- the substrate bias supply device 7 is a device that supplies substrate bias to the elements constituting the S A L C 3.
- the substrate bias value to be supplied is variable and can be operated from the outside of the substrate bias supply device 6.
- the operation of the substrate bias value may be performed artificially as with the operating voltage supply device 6. For example, control outside the figure is performed so that appropriate element performance can be obtained according to the content processed by SALP 2. You may make it operate with an apparatus.
- the performance of the element used for the SALC 3 can be changed by changing the operating voltage value with the operating voltage supply device 6 or changing the substrate bias value with the substrate bias supply device 7. Board via in forward or reverse direction By changing the bias voltage or changing the substrate bias value, the threshold voltage of the element used for SALC3, the amount of leakage current, etc. are changed, thereby changing the device performance.
- FIGS. 2 and 3 are diagrams respectively showing the case where S ALC 3 is connected in 8-stage cascade and the case where 4-stage cascade is used.
- SALC3 is in retiming circuit 5 for every 8 stages.
- SALC3 is in the retiming state of retiming circuit 5 every four stages.
- the other retiming circuit 5 is in the through state.
- the number of stages where SAL C 3 is cascaded is larger than in Fig. 3, and the number of retiming times is reduced, resulting in shorter latency. 2 has a longer retiming interval than FIG. 3, so the element used in SAL C 3 in FIG. 2 has higher performance than the element used in SAL C 3 in FIG.
- FIG. 4 shows an example of SAL C 3 configuration.
- SAL C 3 according to this embodiment has three data input terminals Dli, D2i, D3i and three data output terminals Dlo, D2o, D3o.
- a forward line is formed to output three lines of data from to the next stage (right side of the figure). Since all lines are connected in the forward direction, there is no need to retiming between all SAL P 3, and as many SAL P 3 as can be processed in one cycle can be inserted.
- the data on the line output from the data output terminal Dlo is referred to as “output data”, and the data on the line output from the data output terminal D2o
- the evening is referred to as “reference data”, and the data on the line output from the data output terminal D3o is referred to as “enable evening”.
- the SALC 3 also decodes the contents of the instruction input from the instruction input terminal CON, executes a process according to the decoded result, and selects a decoder 31 for selecting a line for outputting the execution result.
- Examples of processing include control processing such as path control, latch control, and conditional instruction in addition to arithmetic processing such as four arithmetic operations and logical operations.
- macro instructions can be executed.
- the command is input from a command input device not shown. This command input device inputs a predetermined command to each S A L C 3 according to the processing executed in S A L P 2.
- the decoder 31 is connected to various latch circuits for facilitating the above operations, that is, a shift latch circuit 33, an enable latch circuit 34, and a carrier latch circuit 32. Yes.
- the shift latch circuit 33 latches the reference data so that the reference data line is delayed by a predetermined time from the output line, and outputs this in the next digit, for example, in the operation. To work.
- the carry latch circuit 32 latches the carry of the calculation result until the next digit is calculated.
- the enable latch circuit 34 latches the enable data input from the previous stage S A L C 3.
- Enable Day is a data for instructing Enable / Disable for the processing executed by the decoder 31.
- the pipeline controller 4 in FIG. 1 receives latency data and a clock for determining the retiming interval of S A L P 2. Latency overnight depends on the contents of the processing by SALP 2 and the performance of the elements used in SALC 3.For example, it depends on the substrate bias value supplied from the substrate bias supply device 7 and the operation voltage supply device 6, the operation voltage value, etc. Determined.
- the pipeline controller 4 generates the same number of control data as the retiming circuit 5 based on these data. Each control data is supplied to the corresponding retiming circuit 5.
- FIG. 5 is an exemplary diagram of such a pipeline controller 4.
- the pipeline controller 4 includes a latency register 41 and the same number of control data generators 42 as the retiming circuit 5.
- the pipeline Although an example in which the in-controller 4 is configured by hardware such as a semiconductor device is shown, it may be configured in software by causing a predetermined CPU to execute the computer program of the present invention.
- the latency register 41 derives a stage value representing the number of stages of S A L C 3 connected in cascade from the latency data.
- the latency register 4 1 sends the derived stage value to the control data generation unit 4 2.
- the latency data may be used as a step value as it is.
- the control data generation unit 4 2 includes a subtractor 4 3, a selector 4 4, a discriminator 4 5, and an OR circuit 4 6, and generates control data from the stage values sent from the clock and latency register 4 1. To do.
- the subtractor 4 3 is configured to subtract 1 from the input value.
- the step value is input from the latency register 4 1 to the subtracter 4 3 of the control data generation unit 42 in the first stage.
- the output of the selector 4 4 of the control data generation unit 42 in the previous stage is input to the subtracter 4 3 of the control data generation unit 42 in the second and subsequent stages.
- the discriminator 44 is configured to discriminate whether or not the subtraction result by the subtractor 43 is “0” and input the discrimination result to the selector 45 and the OR circuit 46.
- the selector 45 is configured to output either the subtraction result from the subtractor 43 or the stage value sent from the latency register 41 according to the determination result from the discriminator 44.
- the selector 4 5 outputs the stage value sent from the latency register 4 1 when the discrimination result from the discriminator 4 4 indicates that the subtraction result by the subtractor 4 3 is “0”. When it indicates that it is not “0”, the result of subtraction by the subtractor 4 3 is output.
- the 0R circuit 46 is configured to perform 0R operation on the clock and the discrimination result from the discriminator 44 and to input the output to the retiming circuit 5 as control data.
- the OR circuit 4 6 outputs control data as alternating data when the discrimination result from the discriminator 4 4 indicates that the subtraction result by the subtracter 4 3 is “0”. When it indicates that it is not, a constant value of control data is output.
- the selector 45 and the OR circuit 46 operate as follows.
- the selector 45 outputs the stage value output from the latency register 41 when the determination result from the determiner 44 is “1”, and outputs the subtraction result from the subtractor 43 when it is “0”.
- the OR circuit 46 inverts and receives the discrimination result from the discriminator 44. In other words, “1” is input to the OR circuit 46 when the determination result is “0”, and “0” is input when the determination result is “1”. When the discrimination result from the discriminator 44 is “1”, “0” is inverted and “0” and “1” of the clock are output as control data as they are. When the discrimination result from the discriminator 44 is “0”, it is inverted and “1” is inputted, and a constant value “1” is outputted as control data.
- control data such that SAL C 3 having the number of stages corresponding to the stage value is cascade-connected is supplied from the pipeline controller 4 to the SALP 2.
- the control data output from the control data generator 42 up to the third level is a constant value, and is output from the control data generator 42 of the fourth level.
- the control day will be a police box evening.
- the SALP2 retiming circuit 5 is in the through state when the control data is a constant value (for example, “1”), and is in the retiming state when the control data is alternating data.
- the stage value is “4”
- the control data input to every four retiming circuits 5 is alternating data, and the others are constant control data.
- FIG. 6 shows SALP 2 for the compute pipeline 1 used.
- S ALC 3 has 6 stages, the retiming circuit 5 between the 4th stage and the 5th stage SALC 3 is in the retiming state, and the other retiming circuit 5 is in the through state.
- the first SALC is S ALC3A
- the second and subsequent stages are S ALC 3 B, SALC 3C, SAL C3D, SALC 3 E, and SALC 3F.
- SALC 3 uses the command input from the command input terminal CON as follows. Do it.
- the decoder 31 adds the data A inputted to the data input terminal Dl i and the reference data B inputted to the data input terminal D2i, and sends the result to the data output terminal Dlo.
- the enable signal C input to the data input terminal D3i is sent to the data output terminal D3o and is latched by the enable latch circuit 34.
- the decoder 31 performs the above addition process according to the enable data C sent to the data output terminal D3o.
- the reference data B 1 latched in the shift latch circuit 3 3 is sent to the data output terminal D2o, and the new reference data B input to the data input terminal D2i is transferred to the shift latch circuit 3 3 via the decoder 31. Sent.
- the shift latch circuit 33 latches the new reference data B that has been sent (FIG. 7).
- the decoder 31 adds the data A input to the data input terminal Dli and the reference data B input to the data input terminal D2i, and sends the result to the data output terminal Dlo.
- the enable signal C input to the data input terminal D3i is sent to the data output terminal D3o.
- the decoder 31 performs the above addition process according to the enable data C 1 latched in the enable latch circuit 34.
- the reference data B 1 latched in the shift latch circuit 33 is sent to the data output terminal D2o, and the new reference data B input to the data input terminal D2i is transferred to the shift latch circuit 33 via the coder 31. Sent.
- the shift latch circuit 33 latches the new reference data B that has been sent (FIG. 8).
- the decoder 31 sends the data A input to the data input terminal Dl i as it is to the data output terminal Dlo.
- Enable device C input to data input terminal D3i is sent to data output terminal D3o.
- the decoder 31 sends the data A input to the data input terminal Dli as described above to the data output terminal Dlo.
- the reference data B 1 latched in the shift latch circuit 3 3 is sent to the data output terminal D2o, and the new reference data B input to the data input terminal D2 i is transferred to the shift latch circuit 3 3 via the decoder 31. Sent to.
- the shift latch circuit 33 latches the new reference data B that has been sent (FIG.
- the decoder 31 sends the data A input to the data input terminal Dli as it is to the output terminal Dlo.
- the enable data C input to the data input terminal D3i is sent to the data output terminal D3o and is latched by the enable latch circuit 34.
- the decoder 31 sends the data A input to the data input terminal Dli to the data output terminal Dlo as described above according to the enable signal C sent to the data output terminal D3o.
- the reference data B 1 latched in the shift latch circuit 33 is sent to the data output terminal D2o, and the new reference data B input to the data input terminal D2i is sent to the shift latch circuit 33 via the decoder 31. It is done.
- the shift latch circuit 33 latches the new reference data B that has been sent (FIG. 10).
- the decoder 31 adds the data input to the data output terminal Dli or the data input terminal Dlo according to the data output terminal D3o or the enable data of the enable latch circuit 34. Will be sent to.
- the enable data is “1”
- the decoder 31 performs an addition process, and when it is “0”, the decoder sends the data input to the data input terminal Dli as it is to the data output terminal Dlo. To do.
- Figure 11 shows the data input to the data input terminals Dli, D2i, and D3i.
- Data “100000” is input to the data input terminal Dli
- reference data “010000” is input to the data input terminal D2i
- enable data “110000” is input to the data input terminal D3i in this order.
- the operation (2X 3 + 1) can be executed.
- Each S ALC 3 A to F executes any of the above processes i) to iv) according to the instruction input from the instruction input terminal C0N, and the solution of the operation is the data output terminal of SAL C 3 at the final stage. Output from Dlo.
- Figure 12 shows the state in which the first data (1,0,1) is input to the first SAL C 3 A in the first cycle (0 cycle).
- SALC3A executes the process i) according to the instruction input from the instruction input terminal C0N.
- (1,0,1) is output to the data output terminals Dlo, D2o, and D3o of S ALC 3 A.
- SALC3B executes the process of iii) according to the instruction input from instruction input terminal C0N.
- SALC3B Day (1,0,1) is output to the evening output terminals Dlo, D2o, and D3o.
- the third and fourth stage SA LC3 C and D perform the same processing as SALC 3 B. In this cycle, no data is sent after the retiming circuit 5, and the initial state remains unchanged.
- FIG 13 shows the state in which the second delay (0,1,1) is input to the first stage SALC 3 A in the first cycle.
- SALC 3 A executes the process ii) according to the instruction input from the instruction input terminal CON. (1,0,1) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 A.
- SALC 3 B executes the process i) according to the instruction input from the instruction input terminal CON. (1,0,1) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 B.
- SALC 3 C executes the process of iii) according to the command input from the command input terminal CON when the data (1, 0, 1) is input from SALC3B. (1, 0, 1) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 C.
- SALC 3D performs the same processing as SALC 3 C.
- the retiming circuit 5 is supplied with data (1,0,1) from SALC 3D.
- the retiming circuit 5 sends the supplied data to the SALC 3 E at the next stage according to the control data supplied from the pipeline controller 4.
- data (1,0, 1) is input from the retiming circuit 5
- the SALC 3 E executes the process of iii) according to the instruction input from the instruction input terminal C0N.
- (1,0,1) is output to the output terminals Dlo, D2o, and D3o of SALC 3 E.
- SALC 3 F performs the same processing as SALC 3 E. “1” is output from the data output terminal Dlo of S ALC 3 E.
- Figure 14 shows the state in which the third data (0, 0, 0) is input to the first stage SALC 3 A in the second cycle.
- SALC 3 A executes the process ii) according to the instruction input from the instruction input terminal C0N. (0,1,0) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 A.
- SALC 3 B executes the process ii) according to the instruction input from the instruction input terminal C0N. (1,0,0) is output to the data output terminals Dlo, D2o and D3o of SALC 3 B.
- SALC 3 C data (1,0,0) is input from SALC 3 B force, it is determined by the command input from command input terminal C0N. Perform step iv). (1, 0, 0) is output to the data output terminals Dlo, D2o, and D3o of S ALC 3 C.
- SALC3D executes the processing of iii) according to the command input from the command input terminal CON when the data (1,0,0) is input from SALC 3C. (1,0,0) is output to the data output terminals Dlo, D2o, and D3o of SA LC 3D.
- the output of the SALC3D of the previous cycle is latched as it is.
- (1,0,1) is latched.
- (1,0,1) latched in the retiming circuit 5 is input to the data input terminals Dli, D2i, and D3i of the SALC 3 E subsequent to the retiming circuit 5.
- SALC3E executes the processing of iii) according to the command input from the command input terminal CON when the data (1,0,1) is input from the retiming circuit 5.
- (1, 0, 1) is output to the output terminals Dlo, D2o, and D3o of S ALC 3 E.
- SALC3F executes the process iii) according to the instruction input from the instruction input terminal CON. “1” is output from the data output terminal Dlo of SALC 3 E.
- Figure 15 shows the state in which the fourth data (0, 0, 0) is input to the first stage SALC3A in the third cycle.
- SALC 3 A executes the process ii) according to the instruction input from the instruction input terminal C0N. (0,0,0) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 A.
- SALC 3 B executes the processing of ii) according to the instruction input from the instruction input terminal C0N when the data (0,0,0) is input from SALC 3 A. (0,1,0) is output to the data output terminals Dio, D2o, D3o of SALC 3 B.
- SALC3C executes the process of iii) according to the command input from the command input terminal C0N when the data (0, 1, 0) is input from SALC3B. (0,0,0) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 C.
- SALC 3D executes the process of iv) according to the instruction input from instruction input terminal C0N. (0,0,0) is output to the data output terminals Dlo, D2o, D3o of SALC 3D.
- the SALC 3D output of the previous cycle is latched as it is. In this case, (1,0,0) is latched.
- SALC3E is a retiming circuit
- data (1,0,0) is input from 5
- the process of iii) is executed by the instruction input from the instruction input terminal CON.
- (1,0,0) is output to the output terminals Dlo, D2o, and D3o of S ALC 3 E.
- SALC3F executes the process of iii) according to the instruction input from the instruction input terminal CON. 3 Eight. “1” is output from the data output terminal 010 of 3 £.
- Figure 16 shows the state in which the fifth data (0, 0, 0) is input to the first stage SALC3A in the fourth cycle.
- SAL C 3 A executes the process ii) according to the instruction input from the instruction input terminal CON. (0,0,0) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 A.
- SALC 3 B executes the process ii) according to the instruction input from the instruction input terminal CON. (0, 0, 0) is output to the output terminals Dlo, D2o, D3o of SALC3B.
- SALC3C executes the process of iii) according to the instruction input from instruction input terminal C0N. (0,1,0) is output to the data output terminals Dlo, D2o, and D3o of S AL C 3 C.
- SALC 3D executes the process iii) according to the command input from command input terminal C0N. (0, 0, 0) is output to the output terminals Dlo, D2o, and D3o of S ALC 3D.
- the output of the SAL C 3D of the previous cycle is latched as it is.
- (0,0,0) is latched.
- (0,0,0) latched by the retiming circuit 5 is input to the data input terminals Dli, D2i, and D3i of the SALC3E at the subsequent stage of the retiming circuit 5.
- the SALC3E executes the process iv) according to the instruction input from the instruction input terminal C0N.
- (0,0,0) is output to the data output terminals Dlo, D2o, and D3o of S ALC 3 E.
- SALC3F executes the process iii) according to the instruction input from the instruction input terminal C0N. “0” is output from the data output terminal Dlo of SAL C 3 E.
- FIG 17 shows the state in which the 6th data (0,0,0) is input to the first stage SALC3A in the 5th cycle.
- SAL C 3 A receives data (0, 0, 0)
- process ii) is executed by the command input from command input terminal CON. (0, 0, 0) is output to the data output terminals Dlo, D2o, and D3o of SALC 3 A.
- SALC3 B executes the process ii) according to the instruction input from the instruction input terminal CON. (0,0,0) is output to the data output terminals Dlo, D2o, D3o of SALC 3B.
- SALC3C executes the process iii) according to the instruction input from the instruction input terminal CON.
- S ALC 3 C data output terminal Dlo, D2o, D3o
- S ALC 3D executes the process of iii) according to the instruction input from the instruction input terminal CON. (0, 1, 0) is output to the data output terminals Dlo, D2o, and D3o of SALC 3D.
- the output of the SAL C 3D of the previous cycle is latched as it is.
- (0,0,0) is latched.
- (0,0,0) latched in the retiming circuit 5 is input to the data input terminals Dli, D2i, and D3i of the SALC3E at the subsequent stage of the retiming circuit 5.
- the SALC3E executes the process iii) according to the instruction input from the instruction input terminal C0N.
- (0,0,0) is output to the output terminals Dlo, D2o, and D3o of S ALC 3 E.
- SALC3F executes the process iv) according to the instruction input from the instruction input terminal C0N. “0” is output from the data output terminal Dlo of SALC3E.
- FIG 18 shows the state after all data has been input in the sixth cycle.
- S ALC 3 A to S ALC 3 D have been processed because there is no input.
- the output of the SAL C 3D of the previous cycle is latched as it is.
- (0,1,0) is latched.
- (0,1,0) latched by the retiming circuit 5 is input to the data input terminals Dli, D2i, and D3i of the SALC 3 E at the subsequent stage of the retiming circuit 5.
- the SALC 3 E executes the process iii) according to the instruction input from the instruction input terminal C0N.
- S ALC 3 E output terminal Dlo, D2o, D3o Outputs (0, 0, 0).
- SALC3F executes the process of iii) according to the instruction input from the instruction input terminal CON. “0” is output from the data output terminal 010 of 3803 £.
- the arithmetic pipeline 1 of this embodiment can adjust the number of stages of S A LC 3 cascade-connected in the SAL P 2 by the control data from the pipeline controller 4. Therefore, the latency of SALP 2 can be changed dynamically.
- each element constituting the SAL C 3 can dynamically change the threshold voltage by the operation voltage supplied from the operation voltage supply device 6 and the substrate bias supplied from the substrate bias supply device 7. The performance of each element is variable.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Software Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Advance Control (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2004-233366 | 2004-08-10 | ||
| JP2004233366A JP2006053652A (ja) | 2004-08-10 | 2004-08-10 | 演算パイプライン、演算パイプラインによる処理方法、半導体装置、コンピュータプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2006016451A1 true WO2006016451A1 (ja) | 2006-02-16 |
Family
ID=35839227
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2005/011482 Ceased WO2006016451A1 (ja) | 2004-08-10 | 2005-06-16 | 演算パイプライン、演算パイプラインによる処理方法、半導体装置、コンピュータプログラム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP2006053652A (ja) |
| WO (1) | WO2006016451A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2011108020A (ja) * | 2009-11-18 | 2011-06-02 | Mitsubishi Electric Corp | 信号処理装置 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH08274620A (ja) * | 1995-03-29 | 1996-10-18 | Hitachi Ltd | 半導体集積回路装置及びマイクロコンピュータ |
| JP2003316566A (ja) * | 2002-04-24 | 2003-11-07 | Matsushita Electric Ind Co Ltd | パイプラインプロセッサ |
| JP2004062281A (ja) * | 2002-07-25 | 2004-02-26 | Nec Micro Systems Ltd | パイプライン演算処理装置及びパイプライン演算制御方法 |
-
2004
- 2004-08-10 JP JP2004233366A patent/JP2006053652A/ja active Pending
-
2005
- 2005-06-16 WO PCT/JP2005/011482 patent/WO2006016451A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH08274620A (ja) * | 1995-03-29 | 1996-10-18 | Hitachi Ltd | 半導体集積回路装置及びマイクロコンピュータ |
| JP2003316566A (ja) * | 2002-04-24 | 2003-11-07 | Matsushita Electric Ind Co Ltd | パイプラインプロセッサ |
| JP2004062281A (ja) * | 2002-07-25 | 2004-02-26 | Nec Micro Systems Ltd | パイプライン演算処理装置及びパイプライン演算制御方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2006053652A (ja) | 2006-02-23 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Abnous et al. | Ultra-low-power domain-specific multimedia processors | |
| Bahadori et al. | High-speed and energy-efficient carry skip adder operating under a wide range of supply voltage levels | |
| US7406588B2 (en) | Dynamically reconfigurable stages pipelined datapath with data valid signal controlled multiplexer | |
| JP4129345B2 (ja) | 電力削減のための複数の等価機能ユニットの制御 | |
| US6065032A (en) | Low power multiplier for CPU and DSP | |
| JP3705022B2 (ja) | 低消費電力マイクロプロセッサおよびマイクロプロセッサシステム | |
| US8150903B2 (en) | Reconfigurable arithmetic unit and high-efficiency processor having the same | |
| JP2004062281A (ja) | パイプライン演算処理装置及びパイプライン演算制御方法 | |
| US5764550A (en) | Arithmetic logic unit with improved critical path performance | |
| JP4566180B2 (ja) | 演算処理装置およびクロック制御方法 | |
| WO2006016451A1 (ja) | 演算パイプライン、演算パイプラインによる処理方法、半導体装置、コンピュータプログラム | |
| Baratalipour et al. | SAMA: Self-adjusting multi-cycle approximate adder | |
| Kulkarni et al. | Implementation of efficient low power sha-256 algorithm | |
| JP3459821B2 (ja) | マイクロプロセッサ | |
| Lee et al. | A low-power implementation of asynchronous 8051 employing adaptive pipeline structure | |
| US7349938B2 (en) | Arithmetic circuit with balanced logic levels for low-power operation | |
| JP2010079841A (ja) | マイクロコンピュータ及びその命令実行方法 | |
| JP4304124B2 (ja) | 半導体装置 | |
| Belkasim | Pipelined Half Adders | |
| JP3906865B2 (ja) | 低消費電力マイクロプロセッサおよびマイクロプロセッサシステム | |
| US7804331B2 (en) | Semiconductor device | |
| JP5187303B2 (ja) | デュアルレイル・ドミノ回路、ドミノ回路及び論理回路 | |
| Kumar et al. | Design and Implementation Of 20T Hybrid Full Adder For Low Power High Performance Computing | |
| JP2009187075A (ja) | デジタル回路 | |
| TW202603568A (zh) | 管線化電路及管線化操作方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| AK | Designated states |
Kind code of ref document: A1 Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS KE KG KM KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NG NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SM SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW |
|
| AL | Designated countries for regional patents |
Kind code of ref document: A1 Designated state(s): BW GH GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| DPE1 | Request for preliminary examination filed after expiration of 19th month from priority date (pct application filed from 20040101) | ||
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |