WO2007089014A1 - デジタルvlsi回路およびそれを組み込んだ画像処理システム - Google Patents

デジタルvlsi回路およびそれを組み込んだ画像処理システム Download PDF

Info

Publication number
WO2007089014A1
WO2007089014A1 PCT/JP2007/051927 JP2007051927W WO2007089014A1 WO 2007089014 A1 WO2007089014 A1 WO 2007089014A1 JP 2007051927 W JP2007051927 W JP 2007051927W WO 2007089014 A1 WO2007089014 A1 WO 2007089014A1
Authority
WO
WIPO (PCT)
Prior art keywords
processing
arithmetic
data
pipeline
clock
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2007/051927
Other languages
English (en)
French (fr)
Inventor
Masahiko Yoshimoto
Kentaro Kawakami
Jun Takemura
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Kobe University NUC
Original Assignee
Kobe University NUC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Kobe University NUC filed Critical Kobe University NUC
Priority to JP2007556947A priority Critical patent/JP4521508B2/ja
Priority to US12/278,015 priority patent/US8291256B2/en
Publication of WO2007089014A1 publication Critical patent/WO2007089014A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/04Generating or distributing clock signals or signals derived directly therefrom
    • G06F1/10Distribution of clock signals, e.g. skew
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • G06F1/3234Power saving characterised by the action undertaken
    • G06F1/3237Power saving characterised by the action undertaken by disabling clock generation or distribution
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the present invention relates to a low power consumption digital VLSI circuit, and more particularly, a digital power consumption reduced by controlling clock supply and power supply for each of a plurality of internal arithmetic units performing pipeline arithmetic processing.
  • the present invention also relates to a VLSI circuit and a digital VLSI circuit that achieves low power consumption by performing feedback control or feedforward control of the operating power supply voltage, substrate bias voltage, and operating frequency.
  • the present invention relates to an image processing system and a portable terminal incorporating a low power consumption digital VLSI circuit.
  • the powerful coding and decoding technology is applied to information terminal equipment such as mobile phones with built-in personal computers and microcomputers.
  • FIG. 20 is a block diagram of the H.264 decoding processing module.
  • H. 264 decoding processing module can be built by dedicated hardware or by running a general-purpose processor based on a program that describes the means of encoding and decoding. is there.
  • FIG. 21 shows an example in which an H.264 decoding processing module is constructed by dedicated hardware using a plurality of arithmetic units based on the block diagram shown in FIG.
  • bit stream as information about the difference image is received in the bit stream buffer 1001
  • entropy decoding 1002 variable length decoding processing
  • inverse quantization processing 1 003 inverse Q processing
  • inverse orthogonal transformation processing 1004 Processing is performed in the order of (Reverse T processing), and a difference image is generated.
  • the predicted image generation process 1005 is executed based on the image developed in the current frame memory.
  • a frame image is generated by the addition processing 1006 of the difference image and the predicted image.
  • the worst cycle number Sn (n is a natural number) so that pipeline processing at each stage does not fail.
  • the worst cycle number Sn is secured as the processing time of each pipeline.
  • the worst cycle number Sn of pipeline processing at each stage may be designed to be as uniform as possible so that the processing performance of the entire pipeline is improved. preferable.
  • Si Sj for anyi, j
  • the ideal video decoding process is that the above-described (Equation 1) is satisfied, the processing performance of all the arithmetic units is exhibited to the upper limit, and the process is always continued at a fixed number of cycles. It will be If the fixed number of cycles is designed as the worst number of cycles, pipeline processing can be performed up to the upper limit of the processing performance of the arithmetic unit, where the arithmetic unit does not play in vain. Force S will be executed.
  • the intensity of change between moving images is small, and the number of processing cycles is smaller than the worst cycle number for a frame. This is because the amount of calculation processing increases in accordance with the intensity of changes between moving pictures (the intensity of movement) in the moving picture encoding and decoding process. This is because the number of processing cycles is small in a frame that is not severe).
  • the number of blocks that need to be written to the frame memory (FM) is less than 24.
  • the number of processing cycles is smaller than the worst cycle number even in a frame where the intensity of change between moving images (the intensity of movement) is large.
  • FIG. 23 is a time chart schematically showing the operation state of the arithmetic unit in the actual knock line processing.
  • the part that is marked with no The period during which the calculator is in operation and the part where no tapping is applied
  • Figure 23 shows a pipeline composed of three arithmetic processing stages, but if the pipeline is composed of three or more arithmetic processing stages, the number of arithmetic units depends on the number of arithmetic stages. Will increase.
  • data n in FIG. 23 is read as macroblock n
  • block n in the case of a digital VLSI circuit that processes in the block pipeline, it is read as block n.
  • the pipeline arithmetic processing is designed to have redundancy (margin) more than the actual arithmetic processing, so that the arithmetic unit operates and the redundant cycle does not occur. Will occur.
  • Clock gating is a technique for reducing the power consumption of the entire digital VLSI circuit by utilizing the redundancy (margin) in the arithmetic processing of the pipeline (Japanese Patent Laid-Open No. 10-020959).
  • Clock gating is a technique for reducing power consumption without supplying a clock to a calculator that does not need to operate during execution of a no-line arithmetic process.
  • An arithmetic unit that operates in synchronism with the clock consumes a large amount of current in the clock system. Therefore, a digital VLSI circuit can be configured by executing arithmetic processing and supplying the clock only to the arithmetic unit. Overall power consumption can be reduced.
  • the period during which the computation in the pipeline is subject to hatching is a period in which the computing unit operates and is in an idle state. The clock supply is stopped and the power consumption of the entire digital VLSI circuit is reduced.
  • FIG. 24 is based on the block diagram shown in FIG. 20, and is operated by a general-purpose processor (hereinafter referred to as software processing) based on a program that describes the means for decoding. It is an example of the flowchart at the time of constructing 64 decoding processing modules. In FIG. 24, one frame processing is described. Entropy decoding, inverse Q, inverse T, intra prediction, inter prediction, calo-calculation of predicted image and difference image, deblocking filter processing, and writing to frame memory are executed sequentially for each macro block constituting the frame. These processes are repeated for the number of macro blocks constituting the frame.
  • FIG. 26 is a diagram schematically showing the number of cycles necessary for processing in software processing.
  • the calculation processing time of one frame is limited to the frame processing time Tf by the definition of the encoding method (MPEG, H.26x, etc.). Therefore, when performing encoding Z decoding processing by software processing, a program is constructed so that the number of cycles required for one frame of arithmetic processing is completed within time Tf for any moving image. Alternatively, select the operating frequency Fmax of the general-purpose processor that operates the program so that the number of cycles required for the computation processing of one frame is within the time Tf.
  • the processing time of one frame is limited to the frame processing time Tf by the definition of the encoding method (MPEG, etc.), and one frame of coding processing is performed within the frame processing time Tf. Is required to be completed. In other words, the encoding operation process only needs to be completed during the frame processing period Tf.
  • Dynamic control of operating power supply voltage, substrate bias voltage, and operating frequency is limited to a predetermined time limit.
  • the processing frequency of the processor is lowered as much as possible and the power supply voltage and substrate bias voltage are dynamically controlled according to the operating frequency. It aims to make it easier.
  • the predetermined constraint time is 1 frame time, for example, 15 (frame Z seconds) for a moving image, 1 frame time is 1/15 second. It can be replaced with a group of macroblocks included in the frame.
  • Patent Document 1 Japanese Patent Application Laid-Open No. 2003-324735
  • Patent Document 2 JP-A-2005-210525
  • Non-Patent Document 1 IEICE Trans. Fundamentals, V0I.E88- A, No.12 December 2005. "Power- Minimum Frequency / Voltage Cooperative Management Method for VLSI Processor in Leakage-Dominant Technology Era.” (K. Kawakamik M Kanamon, Morita, J. Takemura, M. Miyama, and M. Yoshimoto)
  • Non-Patent Document 2 Proceedings of IEEE International Symposium on Circuits and System 2001 (May, 2001) pp918-921 "An LSI for Vdd- Hopping and MPEG4 System Based on the Chip (H. Kawaguchi, G. Zhang, S. Lee , and T. Sakurai)
  • the conventional clock gating method and the method of dynamically controlling the operating frequency and operating power supply voltage of the arithmetic unit are effective technologies for reducing the power consumption of the entire digital VLSI circuit. There is room for further improvement.
  • One of the improvements when using a clock gating method for a hardwired logic circuit such as an ASIC (Application Specific IC) designed specifically for a specific process is to perform operations during the frame processing period.
  • the problem is that the number of changes in the power supply voltage and operating frequency of the calculator increases and the power consumption increases with the number of changes.
  • the method of dynamically controlling the operating power supply voltage, the substrate bias voltage, and the operating frequency that can be applied to software processing cannot be applied to a normal pipelined node wired logic circuit. Can be mentioned. [0020]
  • the conventional clock gating method is used for the pipeline operation in which the operation state and the idle state of the arithmetic unit are repeated, the hatching of FIG. ! /, The power to perform clock gating
  • the clock supply to the computing unit is frequently turned on and off within the specified time limit. There is room for improvement here.
  • the pipeline processing start cycle Is fixed at the cycle determined at the time of design, so it is not possible to advance the nopline processing by omitting the cycle in which all the arithmetic units are in the idle state, as shown in FIG. 25 in the software processing. It was a force that could not significantly reduce the number of cycles.
  • the cycle in which the 10th data is processed by the calculator 1 is always 900 cycles no matter how small the number of cycles required to process the 1st through 9th data is.
  • an HDTV resolution moving image composed of 1920 x 1088 pixels contains 8160 macroblocks in one frame, so the 8160th macroblock requires 10 cycles for processing by the calculator 1.
  • the number of cycles required to process one frame is reduced by (100 X 81 59 + 10) 7 (100 X 8160) X 100 ⁇ 0.01%. Therefore, the operating frequency of the hard-wired logic circuit can only be reduced by 0.01% at the maximum, and it can be said that power consumption reduction by dynamic control of the operating frequency-power supply voltage and substrate bias voltage could not be realized at all.
  • the operating power supply voltage and substrate bias voltage are also controlled in accordance with the control of the operating frequency of the digital VLSI circuit.
  • the operating power supply voltage and the substrate bias voltage are set to voltages that can achieve the operating frequency fmax with appropriate power from time 0 to TfZ2, and the operating frequency 0 is set to appropriate power from time TfZ2 to Tf. Is set to a voltage that can be realized.
  • the present invention reduces the on / off switching of the clock supply to the arithmetic unit in a predetermined constraint time while using the clock gating technique in the actual pipeline arithmetic processing.
  • the purpose is to provide a digital V LSI circuit that can achieve low power consumption.
  • the present invention switches the power supply on / off to the arithmetic unit within a predetermined constraint time while controlling the power supply for each arithmetic unit.
  • An object of the present invention is to provide a digital V LSI circuit that can achieve low power consumption by reducing the power consumption.
  • the present invention enables the computing units constituting the pipeline to be idle by changing the pipeline processing start cycle in accordance with the previous pipeline processing end cycle in the digital VLSI circuit.
  • Dedicated digital VL SI that can achieve low power consumption by controlling the operating frequency, operating power supply voltage, and substrate bias voltage using the cycle margin generated by reducing redundant cycles
  • the first digital VLSI circuit of the present invention includes a pipeline function.
  • a plurality of arithmetic units that are responsible for each stage of arithmetic processing and execute arithmetic processing in synchronization with the clock, detection means for detecting the end of arithmetic processing of the stage in charge in the arithmetic unit, and a clock for each arithmetic unit Clock supply control means for controlling the supply Z stop, the clock supply control means stops the clock supply to the arithmetic unit for which the end of the arithmetic processing is detected by the detection means, and the detection means When the end of the arithmetic processing in all the arithmetic units is detected, the clock supply to all the arithmetic units is resumed for the next pipeline arithmetic processing.
  • the period during which all the arithmetic units are in the idle state can be omitted, and the number of clock gating starts and ends can be reduced by reducing the pipeline arithmetic processing in the arithmetic units. It is possible to achieve lower power consumption than conventional clock gating.
  • the data for the arithmetic processing includes a plurality of macroblock data, and a predetermined processing period (frame processing period) is determined.
  • the arithmetic processing is encoding / decoding of the moving image, and the arithmetic unit executes the pipeline arithmetic processing in units of the macro block data and supplies the clock.
  • the control means stops the clock supply to all the arithmetic units even if the detection means detects the end of the arithmetic processing in all the arithmetic units. After the elapse of the frame processing period, all the arithmetic units are processed for the pipeline operation processing of the next frame data. The clock supply is resumed.
  • the period during which all the arithmetic units are in the idle state can be omitted, and the number of start and end times of clock gating can be reduced by narrowing down the Nopline arithmetic processing in the arithmetic units.
  • the power consumption can be further reduced compared to the conventional clock gating. Note that the knock line calculation processing is performed in units of macroblock data.
  • the data for the arithmetic processing is: This is moving image data including frame data including a plurality of macro block data (the macro block data is composed of a plurality of block data) and having a predetermined processing period (frame processing period).
  • the arithmetic processing is encoding / decoding processing of the moving image
  • the arithmetic unit executes the pipeline arithmetic processing in units of the block data
  • the clock supply control means includes the frame data in the frame data.
  • the end of the arithmetic processing in all the arithmetic units is detected by the detection means for all the arithmetic units. Continue to stop the clock supply, and after the frame processing period has passed, proceed to pipeline operation processing of the next frame data. Characterized by resuming the clock supply to the calculator of Te.
  • the number of blocks included in the macroblock is 24, for example.
  • the period during which all the arithmetic units are in the idle state can be omitted, and the number of clock gating starts and ends can be reduced by reducing the pipeline arithmetic processing in the arithmetic units. It is possible to achieve lower power consumption than conventional clock gating. Note that the knock line calculation processing is performed in units of block data.
  • the second digital VLSI circuit of the present invention is responsible for each stage of pipeline arithmetic processing, and a plurality of arithmetic units that execute arithmetic processing in synchronization with a clock, and in charge of the arithmetic unit Detection means for detecting the end of the arithmetic processing of the stage, and clock supply control means for controlling the supply of the clock Z for each of the arithmetic units, the clock supply control means force, the pipeline arithmetic processing by the detection means When the end of the arithmetic processing of the next-stage computing unit is detected first among the preceding-stage computing unit and the next-stage computing unit, the clock supply to the next-stage computing unit is stopped.
  • the clock supply control means detects the end of the arithmetic processing of the preceding arithmetic unit by the detection means first, Until the preceding arithmetic unit can output processed data to the next arithmetic unit, the clock supply to the preceding arithmetic unit is stopped, and the clock supply to the preceding arithmetic unit is stopped. It is preferable that the clock supply to the preceding stage computing unit is resumed when the preceding stage computing unit is ready to output the computation processing data processed for the next stage computing unit. .
  • the pipeline processing is seamlessly executed, and the pipeline arithmetic processing in the arithmetic units is packed and performed.
  • the number of start and end times of gating can be reduced, and the power consumption can be reduced.
  • the data necessary for the arithmetic processing is composed of a plurality of macroblock data forces, and a restriction time (frame processing time) for completing the processing is determined.
  • the frame data force is also composed of moving image data
  • the arithmetic processing is encoding / decoding of the moving image
  • the arithmetic unit executes the pipeline arithmetic processing in units of the macroblock data.
  • the end of the arithmetic processing of the preceding arithmetic unit is detected after the clock supply to the next arithmetic unit is stopped.
  • the clock supply control means performs an operation process on the last macroblock data in the frame data to the previous stage arithmetic unit. Even after the clock supply is stopped, even if the previous stage computing unit is ready to output processed processing data to the next stage computing unit, the clock supply to the previous stage computing unit is stopped. After the elapse of the frame processing period, the clock supply to all the arithmetic units is resumed for the pipeline arithmetic processing of the next frame data.
  • the data for the arithmetic processing includes a plurality of macroblock data (the macroblock data is composed of a plurality of block data),
  • the moving image data includes frame data that has a predetermined processing period (frame processing period).
  • the arithmetic processing is encoding / decoding processing of the moving image.
  • Pipeline arithmetic processing is executed in units of the block data, and the clock supply control means performs the next stage in the arithmetic processing related to the last block data in the last macro block data in the frame data. After the clock supply to the computing unit is stopped, the clock supply to the next stage computing unit is stopped even if the end of the arithmetic processing of the preceding stage computing unit is detected.
  • the frame processing period it is preferably characterized in that to resume the supply of the clock to all of the operational units towards the pipeline processing of the next frame data.
  • the number of blocks included in the macroblock is 24, for example.
  • the clock supply control means in the arithmetic processing relating to the last block data of the last macroblock data in the frame data, After stopping the clock supply to the preceding arithmetic unit, even if the previous arithmetic unit is ready to output processed arithmetic processing data to the next arithmetic unit, the clock is supplied to the preceding arithmetic unit. It is preferable that after the frame processing period elapses, the clock supply to all the arithmetic units is resumed for the pipeline operation processing of the next frame data.
  • the pipeline arithmetic processing is performed.
  • a feedback control unit that counts the data processing amount of the logic for each predetermined constraint time, and determines an operation power supply voltage, a substrate bias voltage, and an operation frequency of the arithmetic unit at the time of processing of the next predetermined constraint time;
  • a calculation unit adjustment unit for adjusting the operation power supply voltage, the substrate bias voltage, and the operation frequency of the operation unit, and performing dynamic control by feedback control with respect to the operation power supply voltage, the substrate bias voltage, and the operation frequency of the operation unit; It is characterized by.
  • the amount of data included in the predetermined constraint time provided for the pipeline operation processing before being provided for the pipeline operation processing is determined based on the prediction by the processing load prediction unit.
  • a feedforward control unit for controlling the operating power supply voltage, the substrate bias voltage, and the operating frequency of the computing unit. And dynamic control by feedforward control.
  • the first or second digital VLSI circuit described above uses clock gating technology
  • low power consumption can be achieved by controlling on / off of power supply to the arithmetic unit. It is. That is, the power supply to the computing unit may be stopped at the timing when the clock supply to the computing unit is stopped in clock gating.
  • the third digital VLSI circuit of the present invention is responsible for each stage of pipeline arithmetic processing, and detects a plurality of arithmetic units that execute arithmetic processing and the end of arithmetic processing of the assigned stage in the arithmetic unit.
  • power supply control means for controlling power supply Z stop for each of the computing units, wherein the power supply control means supplies power to the computing unit for which the end of computation processing is detected by the detection means.
  • the third digital VLSI circuit of the present invention includes frame data having a predetermined processing period (frame processing period) including a plurality of macro block data and data required for the arithmetic processing.
  • Moving image data wherein the arithmetic processing is encoding / decoding of the moving image, the arithmetic unit executes the pipeline arithmetic processing in units of the macroblock data, and the power supply control unit includes: In the calculation process for the last macroblock data in the frame data, even if the detection means detects the end of the calculation process in all the calculation units, the power supply to all the calculation units is stopped. Continue to supply power to all the arithmetic units for the pipeline arithmetic processing of the next frame data after the elapse of the frame processing period. It is characterized by restarting.
  • the fourth digital VLSI circuit of the present invention is responsible for each stage of pipeline arithmetic processing, and includes a plurality of arithmetic units for executing arithmetic processing, and completion of arithmetic processing for the stage in charge in the arithmetic unit.
  • the power supply control means is a preceding stage in the pipeline arithmetic processing by the detection means
  • the power supply to the next-stage arithmetic unit is stopped and the power supply to the next-stage arithmetic unit is stopped.
  • the power supply to the subsequent arithmetic unit is resumed for the next pipeline arithmetic processing.
  • the pre-stage calculation can output processed data to the next-stage arithmetic unit. Until the power supply to the preceding computing unit is stopped, and after the power supply to the preceding computing unit is stopped, the preceding computing unit can output processed processing data to the next computing unit If so, the power supply is resumed for the preceding arithmetic unit.
  • pipeline processing can be executed seamlessly and the occurrence of idle states of arithmetic units can be suppressed as much as possible. Low power consumption can be achieved.
  • the number of start and end times of power supply can be reduced, and the power consumption can be further reduced.
  • the clock supply is read as power supply
  • the clock supply control means is the power supply control means.
  • the digital VLSI circuit of the present invention by using clock gating technology in actual pipeline arithmetic processing, switching on / off of the clock supply to the arithmetic unit within the restricted time is reduced. Low power consumption can be achieved. Further, according to the digital VLSI circuit of the present invention, in actual pipeline arithmetic processing, the power supply for each arithmetic unit is controlled, and the switching of power supply on / off to the arithmetic unit within the limited time is reduced. To achieve low power consumption
  • the digital VLSI circuit of the present invention in the actual pipeline arithmetic processing, the worst number of cycles required for the arithmetic processing of data that must be completed within the constraint time. Therefore, even if the operating frequency of the digital VLSI circuit within the restricted time is reduced, the predetermined arithmetic processing can be completed within the restricted time. Therefore, the operating frequency and operating power supply of the digital VLSI circuit can be reduced. Low power consumption can be achieved by appropriately controlling the voltage and substrate bias voltage
  • the mobile terminal of the present invention power consumption is reduced in moving image data processing, and moving image encoding / decoding processing is performed even in a small terminal such as a mobile phone. It is possible to use the portable terminal in various ways.
  • the present invention can be widely applied to digital VLSI circuits that execute pipeline processing.
  • the present invention is used for the purpose of encoding / decoding moving images.
  • the active low is described as being active when the logic level is high, and it is active when the logic level is S low. May be hot.
  • FIG. 1 is a diagram schematically showing a configuration of a digital VLSI circuit according to Embodiment 1 of the present invention.
  • the arithmetic units 10a to 10c are connected to the processing end detector 20, and each arithmetic unit 10 and the processing end detector 20 are connected by an end flag line 30 and a processing start flag line 40.
  • Each arithmetic unit 10a: LOc is responsible for each stage of pipeline arithmetic processing, and executes arithmetic processing in synchronization with the clock. For example, when executing pipeline processing for encoding and decoding moving images, computing unit 1 (10a) is responsible for the entropy decoding stage and computing unit 2 (10b) is responsible for the inverse Q processing stage. Assume that arithmetic unit 3 (10c) is in charge of the inverse T processing stage. The subsequent noopline processing stages are not shown. Each computing unit is connected sequentially, and between the preceding and following computing units. The configuration is such that processed data is successively transferred from the previous stage to the next stage.
  • Data that has been subjected to arithmetic processing by each arithmetic unit is stored in a buffer provided in each arithmetic unit, and the preceding and succeeding arithmetic units pass the data through this buffer.
  • the data processed by the arithmetic unit 1 is stored in a buffer provided in the arithmetic unit 1, and this data is also transferred to the arithmetic unit 2 at the next stage.
  • the buffer is composed of a flip-flop and RAM (Random Access Memory).
  • each of the computing units 10a to 10c is connected to the processing end detector 20, and when the computing process for which each computing unit is in charge ends, the computing unit sets an end flag. That is, an active signal is output to the end flag line 30 connected to the processing end detector 20.
  • an active signal is output to the end flag line 30 connected to the processing end detector 20.
  • a high signal is output as high active logic (in the case of low active logic, a mouth signal may be output).
  • the processing end detector 20 detects the end of the arithmetic processing of the assigned stage in the arithmetic unit 10 by detecting that the end flag of each arithmetic unit 10 has been set. Can do.
  • the processing end detector 20 detects the end of the arithmetic processing in the arithmetic unit 10 via an end flag.
  • the processing end detector 20 is a multi-input AND circuit.
  • Clock gating is performed on the arithmetic unit 10 for which the arithmetic processing has been completed.
  • the clock supply control means is configured such that each arithmetic unit 10 outputs an end flag when the arithmetic processing is completed and the supply of the clock is temporarily stopped.
  • FIG. 2 is a diagram showing a configuration example of the clock supply control means.
  • the clock supply ON / OFF is automatically controlled by the state machine, flip-flop, and AND circuit.
  • the state machine 11 is connected to the AND circuit 13 via the flip-flop 12.
  • a part of the output line of the arithmetic unit 10 is connected to the state machine 11 and a part is connected to the end flag line 30.
  • Fig. 2 (a) shows an example of the flow of operations when starting clock supply to the arithmetic unit 10
  • Fig. 2 (b) shows an example of the flow of operations for temporarily stopping the clock supply of the arithmetic unit 10.
  • FIG. 1 shows an example of the flow of operations when starting clock supply to the arithmetic unit 10
  • Fig. 2 (b) shows an example of the flow of operations for temporarily stopping the clock supply of the arithmetic unit 10.
  • the flip-flop 12 is turned on and a clock is supplied from the clock input line 14 through the AND circuit 13 as shown in FIG.
  • the arithmetic unit 10 outputs an end signal when the arithmetic processing of the pipeline is completed.
  • the end signal is input to the flip-flop 12 through the state machine 11, and the flip-flop 12 is inverted and turned off.
  • the AND circuit 13 is turned off by the off signal.
  • the clock supply stop process is performed for each arithmetic unit. For this reason, the clock supply is sequentially stopped from the arithmetic unit 10 that has output the end flag.
  • the processing end detector 20 is a multi-input AND circuit, and when the end flag of all the arithmetic units 10 is detected (when processing is completed in all the arithmetic units), the next pipeline arithmetic processing is performed.
  • the processing start flag signal is output to the processing start flag line 40 so that the clock supply to all the arithmetic units is resumed.
  • the processing start flag line 40 is connected in parallel to all the arithmetic units 10 and the processing start flag signal is notified to all the arithmetic units 10 all at once.
  • each arithmetic unit 10 receives the processing start flag signal, it receives the start of clock supply all at once, and proceeds to the arithmetic processing of the next pipeline processing.
  • the processing start flag signal is input from the processing start flag line 40 to the flip-flop 12 through the state machine 11, and the flip-flop 12 is inverted (the OFF force is also turned on).
  • the flip-flop 12 is off and the clock supply is stopped.
  • the flip-flop 12 is on and the AND circuit 13 is connected to the arithmetic unit 10. The clock supply resumes. This clock supply is restarted all at once in all the arithmetic units 10.
  • the clock supply Z stop control mechanism based on the control of the above configuration is the clock supply control means.
  • the clock supply is stopped in order from the arithmetic unit 10 that has completed the arithmetic processing of the responsible stage in the pipeline processing, and the arithmetic processing of all the arithmetic units 10 is performed.
  • the clock supply to all the arithmetic units 10 is resumed all at once for the next pipeline processing.
  • FIG. 3 is a timing chart showing pipeline processing by the digital VLSI circuit of the first embodiment. It is In this timing chart, only three calculators are shown: calculator 1 (10a), calculator 2 (10b), and calculator 3 (10c).
  • the processing of the data (n + 2) in the arithmetic unit 1 (10a) is completed at the first clock in the figure.
  • the arithmetic unit 1 (10a) sets an end flag at the first clock, notifies the processing end detector 20 from the end flag line 30 of the end of the arithmetic processing, and performs clock gating. In other words, the clock supply is stopped as shown in FIG.
  • the processing of data (n + 1) in the arithmetic unit 2 (10b) is completed in the third clock in the figure.
  • the arithmetic unit 2 sets an end flag at the third clock and notifies the end of processing to the processing end detector 20 from the end flag line 30. In this case, as will be described later, the processing of the next macroblock is transferred without transition to clock gating by the processing start signal of the processing end detector.
  • the processing of the macroblock (n) in the arithmetic unit 3 (10c) is also completed at the third clock in FIG.
  • the arithmetic unit 3 sets an end flag at the third clock and notifies the processing end detector 20 of the end of the arithmetic processing from the end flag line 30. Also in this case, as will be described later, the processing of the next macroblock is started without shifting to clock gating by the processing start signal of the processing end detector.
  • the processing end detector 20 performs AND processing on the end flag signals of the arithmetic unit 1 (10a), the arithmetic unit 2 (10b), and the arithmetic unit 3 (10c). In this example, all the end flags from the arithmetic unit 1 (10a), the arithmetic unit 2 (10b), and the arithmetic unit 3 (10c) are gathered at the third clock, and the AND condition is satisfied.
  • the processing end detector 20 outputs a processing start flag signal to the processing start flag line 40 so that the clock supply to all the arithmetic units 10 is resumed for the next pipeline arithmetic processing.
  • the processing start flag signal Is sent to all of the computing units 1 (10a), 2 (10b), and 3 (10c).
  • computing unit 1 (10a), computing unit 2 (10b), and computing unit 3 (10c) receive the processing start flag signal, the clock supply is resumed all at once, and the processing of the next pipeline processing is started.
  • the arithmetic unit 10 when the arithmetic unit 10 receives the supply of the clock, it first deactivates the end flag. In this example, since it is high active, it is switched to low. Next, the next pipeline process is started. Operation unit 1 (10a) starts processing on data (n + 3), and operation unit 2 (10b) starts processing on data (n + 2), and operation unit 3 (10c ) Starts executing processing on data (n + 1).
  • FIG. 4 is a diagram schematically illustrating the progress of the pipeline operation in the digital VLSI circuit according to the first embodiment.
  • Arithmetic units such as arithmetic units 1, 2, and 3 are typically arranged on the vertical axis.
  • the arithmetically processed data is transferred to the arithmetic unit 2 (10b) and processed by the arithmetic unit 2, and when the processing is completed, the arithmetically processed data is stored.
  • the data is transferred to the arithmetic unit 3 (10c) and processed by the arithmetic unit 3. In this way, the flow of pipeline processing is developed in the vertical axis direction!
  • the horizontal axis is timing.
  • the first row displays the timing (timing 0 to: L000) when pipeline processing is executed while ensuring the worst cycle number according to the prior art.
  • the second level displays the timing (timing 0 to: L000) when pipeline processing of the digital VLSI circuit according to the first embodiment of the present invention is executed.
  • the parts where no and hatching have been performed indicate the period during which the arithmetic processing is being executed, and the parts where no hatching has been performed are terminated until the next data processing is executed. Show the kugging period.
  • the clock supply is stopped by starting clock gating by performing the arithmetic processing in each arithmetic unit 10 as much as possible and seamlessly executing it as continuous processing.
  • the number of clock supply starts due to clock gating stoppage are decreasing.
  • the number of clock gating starts in computing unit 1 is 3 (275 cycles, 425 cycles, 575 cycles).
  • clock gating stops 300 cycles, 450 cycles, 600 cycles.
  • clock gating occurs at each pipeline process, so when looking at the pipeline process until computing unit 1 completes the processing of data 9, clock gating is performed. There are 9 starts and 9 clock gating stops. Similarly, the number of clock gating starts 8 times and the number of clock gating stops 8 times for arithmetic unit 2, the number of clock gating starts 7 times and the number of clock gating stops 7 times for arithmetic unit 3 . Clearly, the number of clock gating starts and the number of clock gating stops are reduced.
  • the difference compared with the number of executions of the pipeline processing of 9 times the difference increases as the number of executions of pipeline processing increases.
  • Both the number of clock gating starts and the number of stops of the digital VLSI circuit of the present invention are the same. It will be understood that the number of clock gating starts and stops of the circuit is smaller. For example, since one frame of an HDTV image (1920 x 1088 pixels) is composed of 8160 macroblocks, when HDTV image processing is performed in the macroblock pipeline, the number of times of executing the noiseline processing is 8160 times. .
  • the arithmetic unit 10 is an arithmetic unit that executes the encoding / decoding processing of moving images in units of macroblock data by pipeline arithmetic processing, the pipeline processing in the arithmetic unit 10 is reduced as described above! /, Therefore, there is time to process the macroblock included in the next frame during the processing time of one frame. For example, assuming that there are 8 macroblocks in one frame in Fig. 4, the computing unit 1 completes the processing of the macroblock for one frame in 575 cycles, and the computing unit 3 in the final stage is the 8th. Completing macroblock processing There are cycles that can process macroblocks included in the next frame up to 750 cycles.
  • the calculator 1 processes the last macroblock included in the frame. After completing the above, the macro block contained in the next frame is not processed. Similarly, the computing unit 2 does not process the macro block included in the next frame (Fig. 28).
  • the computation unit 10 outputs an end signal even when the computation processing of the final macroblock is completed.
  • the processing end detector 20 may separately receive a notification of the end of the frame processing period from the control unit (not shown) and output a processing start flag signal to the processing start flag line.
  • the arithmetic unit 10 outputs an end signal in the same manner when the calculation processing of the final macroblock is completed, but at this time, there is a mechanism for outputting with an attribute signal indicating the end of processing of the final macroblock. Conceivable.
  • the processing end detector 20 waits for the processing start flag signal to be output to the processing start flag line until receiving a notification of the end of the frame processing period from the control unit (not shown). After receiving the end notification, the process start flag signal may be output to the process start flag line.
  • the power consumption can be further reduced by reducing the clock gating start count and the clock gating stop count.
  • the no-line arithmetic processing is performed in units of macro blocks (the macro block data is composed of a plurality of block data, and is often composed of 24 pieces). It is also possible to adopt a configuration that executes in units of force blocks, which was the configuration example executed in step 1.
  • the arithmetic unit executes pipeline operation processing in block data units, and the clock supply control means power is added to the last block data in the last macro block data in the frame data.
  • the clock supply to all the arithmetic units is stopped, and after the frame processing period, the next frame data The clock supply to all the computing units will be resumed for the pipeline operation processing.
  • the arithmetic units before and after the pipeline processing have a handshake-type linkage, and the pipeline processing operation processing is performed to reduce the number of clock gating start times and clock gating stop times.
  • this is an example of a digital VLSI circuit designed to reduce power consumption.
  • a configuration in which the unit of the pipeline calculation processing is described as a configuration in which the unit of the macro block is a macro block is also possible.
  • FIG. 6 is a diagram schematically showing a configuration of a digital VLSI circuit according to the second embodiment of the present invention.
  • the front and rear arithmetic units 10a to 10c are sequentially connected, and the arithmetic units are linked together by handshaking.
  • arithmetic unit 1 10
  • arithmetic unit 2 10
  • arithmetic unit 3 10c
  • the number of arithmetic units increases or decreases depending on the number of pipeline stages. Needless to say, it can be designed!
  • Each arithmetic unit 10a: LOc is responsible for each stage of pipeline arithmetic processing, and executes arithmetic processing in synchronization with the clock.
  • computing unit 1 (10a) is in charge of the entropy decoding stage
  • computing unit 2 (10b) is the inverse Q processing stage
  • computing unit 3 (10 c) is in charge of the inverse T processing stage.
  • the subsequent pipeline processing stages are not shown.
  • Each computing unit is connected sequentially, and the processed data is transferred from the preceding stage to the next stage one after another between the preceding and succeeding computing elements.
  • Data that has been subjected to arithmetic processing by each arithmetic unit is stored in a buffer provided in each arithmetic unit, and the preceding and following arithmetic units pass the data through this buffer.
  • the data processed by the computing unit 1 is stored in a buffer provided in the computing unit 1, and this data is also transferred to the next computing unit 2.
  • the buffer is composed of a flip-flop and RAM (Random Access Memory).
  • Each computing unit is connected to the preceding and following computing units in the pipeline processing sequence through three lines of the receiving signal line 50, the request signal line 60, and the data line 70.
  • the arithmetic unit and clock supply control means of the digital VLSI circuit of the second embodiment are configured to operate according to the following seven rules, for example.
  • the computing units before and after pipeline processing are hand-held.
  • the number of clock gating starts and the number of clock gating stops is reduced to reduce power consumption.
  • the clock supply control means is not particularly limited as long as the circuit configuration realizes the above rule.
  • FIG. 7 shows in detail an example of the configuration of the arithmetic unit 10 and the clock supply control means according to the second embodiment.
  • the clock supply on / off is automatically controlled by four state machines (111 to 114), four flip-flops (121 to 124), and three AND circuits (131 to 133). It has become.
  • connection relation between the present stage computing unit 10 and the next stage computing unit 10 is as follows.
  • the acceptance signal line 50b is connected from the state machine 113 to the AND circuit 132 via the flip-flop 123.
  • the request signal line 60b is connected to the AND circuit 132 from the state machine 114 via the flip-flop 124.
  • the output of the AND circuit 132 is directly input to the arithmetic unit 10, and this stage arithmetic unit 10 establishes a non-shake with the next stage arithmetic unit 10, and data is exchanged between the two.
  • the operation unit 10 is configured to be able to detect that it can be performed, and the arithmetic unit 10 has a configuration example in which processed data can be transferred to the next-stage arithmetic unit 10.
  • connection relation between the present stage arithmetic unit 10 and the previous stage arithmetic unit 10 is as follows.
  • the acceptance signal line 50a is connected from the state machine 111 to the AND circuit 131 via the flip-flop 121. ing.
  • the request signal line 60a is connected from the state machine 112 to the AND circuit 131 via the flip-flop 122.
  • the AND circuit 131 is turned on. That is, when the AND circuit 131 is turned on, a handshake is established between the computing unit 10 and the preceding computing unit 10, and data can be transferred between the two. In this state, the above-mentioned rule 4, rule 5, and rule 6 are established.
  • rule 7 is not satisfied in the computing unit 10 at this stage, actual data is not transferred. For example, if the processed data of the current-stage arithmetic unit 10 has already been output to the next-stage arithmetic unit 10, rule 7 is also satisfied, so that data is transferred between the previous-stage arithmetic unit and the present-stage arithmetic unit.
  • the pipeline processing of the computing unit 10 at this stage is finished first, then the pipeline processing of the computing unit 10 at the next stage is finished, and finally the pipeline processing of the computing unit 10 at the previous stage is finished.
  • FIG. 8 is a diagram showing the flow of processing when the pipeline processing of the current stage computing unit 10 is completed.
  • the processing unit 10 When the processing unit 10 finishes processing, the processing unit 10 issues a termination signal to the state machine 113.
  • the state machine 113 is connected to the reception signal line to the next stage arithmetic unit 10 and sets the reception signal level to active (high).
  • the state machine 113 also outputs a signal to the flip-flop 123 to invert the flip-flop 123 (off ⁇ on). This state machine 113 maintains this state transition and maintains the output state.
  • FIG. 9 is a diagram showing the flow of processing when the pipeline processing of the next stage computing unit 10 is completed and the request signal of the next stage computing unit 10 is detected in the state force of FIG.
  • the next stage computing unit 10 activates (high) the request signal line to the stage computing unit 10 and issues a request signal to the state machine 114 of the stage computing unit 10.
  • the fact that the request signal has been issued means that the next-stage arithmetic unit 10 has already output processed data to the subsequent arithmetic units.
  • the state machine 11 4 of the present stage arithmetic unit 10 outputs a signal to the flip-flop 124 and inverts the flip-flop 124 (OFF ⁇ ON). This state machine 114 maintains this state transition and maintains the output state.
  • the AND circuit 132 is turned on because both inputs are active (high).
  • the output of the AND circuit 132 is input to the present-stage arithmetic unit 10.
  • the present-stage arithmetic unit 10 establishes a no-shake between the present-stage arithmetic unit 10 and the next-stage arithmetic unit 10, and the present-stage arithmetic unit 10 Since it is detected that the processed data of 10 has been output to the next stage, the processed data of the present stage computing unit 10 is output to the next stage computing unit 10.
  • the processed data can be accepted from the previous stage, so that the previous stage computing unit 10 is passed through the request signal line via the state machine 112. Outputs a request signal.
  • the arithmetic unit 10 at this stage stops the clock supply and enters the clock gating state.
  • FIG. 10 is a diagram showing the flow of processing when the data processing of the state power pre-stage computing unit 10 in FIG. 9 is completed.
  • the pre-stage arithmetic unit 10 sets the acceptance signal line to the pre-stage arithmetic unit 10 to active (high), and issues a reception signal to the state machine 111 of the pre-stage arithmetic unit 10.
  • the state machine 111 outputs a signal to the flip-flop 121 and inverts the flip-flop 121 (OFF ⁇ ON). This state machine 111 maintains this state transition and maintains the output state.
  • the AND circuit 131 is active because both inputs are active (high). This means that a non-shake is established between the preceding arithmetic unit 10 and the present arithmetic unit 10. Note that if a non-shake is established between the previous stage computing unit 10 and the current stage computing unit 10, since the current stage computing unit 10 has already output the processed data to the next stage, The processed data is transferred from the unit 10 to the computing unit 10 at this stage.
  • the AND circuit 133 is turned on because both the input from the AND circuit 131 and the input from the AND circuit 132 are active (noisy).
  • the AND circuit 133 operates as a gate with respect to the clock input, and the clock supply is started because the clock gate becomes active.
  • the data processing of the current-stage arithmetic unit 10 ends first, the data processing of the previous-stage arithmetic unit 10 ends, and finally the processing of the next-stage arithmetic unit 10 ends. It is an example of the operation in the case of the flow when finished. The flow will be described with reference to FIGS.
  • FIG. 12 is a diagram showing a processing flow when the data processing of the pre-stage computing unit 10 is completed in the state force of FIG.
  • the reception signal line to the computing unit 10 of the previous stage is made active (high), and a reception signal is issued to the state machine 111 of the computing unit 10 of the previous stage.
  • State machine 111 outputs a signal to flip-flop 121 to invert flip-flop 121 (off ⁇ on). This state machine 111 maintains this state transition and maintains the output state.
  • only one of the AND circuit 131 and the AND circuit 132 is active (noisy) and the other is inactive (low), so it remains off. No handshake is established between 10 and the present stage computing unit 10, and no handshake is established between the next stage computing unit 10 and the present stage computing unit 10.
  • the operation unit 10 at this stage stops the clock supply and enters the clock gating state.
  • FIG. 13 is a diagram showing an operation when the pipeline processing of the next stage computing unit 10 is completed and a request signal is detected from the next stage computing unit 10 in the state force of FIG.
  • the next stage computing unit 10 activates (high) the request signal line to the stage computing unit 10 and issues a request signal to the state machine 114 of the stage computing unit 10.
  • the fact that the request signal has been issued means that the next-stage arithmetic unit 10 has already output processed data to the subsequent arithmetic units.
  • the state machine 11 4 of the computing unit 10 outputs a signal to the flip-flop 124 and inverts the flip-flop 124 (OFF ⁇ on). This state machine 114 maintains this state transition and maintains the output state.
  • the AND circuit 132 is active because both inputs are active (high).
  • the output of the AND circuit 132 is input to the present-stage arithmetic unit 10.
  • the present-stage arithmetic unit 10 establishes a no-shake between the present-stage arithmetic unit 10 and the next-stage arithmetic unit 10, and the present-stage arithmetic unit 10 Since it is detected that the processed data of 10 has been output to the next stage, the processed data of the present stage computing unit 10 is output to the next stage computing unit 10.
  • the present stage computing unit 10 Since the present stage computing unit 10 has output the processed data to the next stage, the processed data can be accepted from the previous stage, so the request signal line is sent to the previous stage computing unit 10 via the state machine 112. The request signal is output via
  • the previous stage computing unit 10 delivers the processed data to the current stage computing unit 10.
  • the AND circuit 131 is active because both inputs are active (high).
  • the AND circuit 133 is turned on because both the input from the AND circuit 131 and the input from the AND circuit 132 are active (noisy).
  • the AND circuit 133 operates as a gate with respect to the clock input, and the clock supply is started because the clock gate becomes active.
  • FIG. 14 is a timing chart showing pipeline processing by the digital VLSI circuit of the second embodiment. In this timing chart, only three calculators are shown: calculator 1 (10a), calculator 2 (10b), and calculator 3 (10c).
  • the calculator 1 (10a) The data (n + 2) is processed, the data (n + 1) is processed in the calculator 2 (10b), and the data (n) is processed in the calculator 3 (10c).
  • the computing unit 1 (10a) completes the pipeline processing of the data (n + 2) in the first clock in the figure and issues a request signal.
  • the acceptance signal is received from the arithmetic unit 2 (10b) at the third clock.
  • Arithmetic unit 2 (10b) has completed the knock-in processing of data (n + 1) in the fourth clock in the figure and issues a request signal.
  • the acceptance signal is received from the arithmetic unit 3 (10c) at the first clock.
  • Arithmetic unit 3 (10c) completes the pipeline processing of data (n) in the second clock in the figure and issues a request signal.
  • the 4th clock is received from the next stage computing unit.
  • the above second operation example shows that the clock gating is the third clock force.
  • arithmetic unit 2 (10b) shows that arithmetic unit 1 (10a) and arithmetic unit 3 (10c) have completed processing before processing in arithmetic unit 2 (10b) is complete, so there is no period for clock gating.
  • the processing of the (n + 2) th data is started at the next clock after the processing of the (n + 1) th data is completed.
  • clock gating is performed from the third clock power to the fourth clock by the above first operation example (the operation examples shown in FIGS. 11 to 13).
  • FIG. 15 is a diagram schematically illustrating the progress of the pipeline operation in the digital VLSI circuit according to the second embodiment.
  • the description of the elements in each figure shown in FIG. 15 is the same as the explanation of the elements in each figure shown in FIG.
  • the pipeline processing in charge of the arithmetic unit 10 is completed early, and the clock gating is not performed until the pipeline processing in the preceding and subsequent arithmetic units 10 is completed. It is done.
  • clock gating is performed between 625 cycles and 650 cycles of the arithmetic unit 3 (10c), for 525 cycles from the 500 cycle force of the arithmetic unit 2 (10b).
  • the clock gating at this stage is released after the processing of the previous arithmetic unit is completed. This is clock gating in the case of the above operation example 1 (operation examples shown in FIGS. 8 to 10).
  • the clock gating period is provided until the processing of all the arithmetic units is completed in the same pipeline stage.
  • the same stage is used in the same pipeline stage. Since the clock gating period is provided until the processing of the computing units before and after the computing unit is completed, the pipeline processing of the final macro block is completed earlier in Example 2, and the clock gating is also performed. It can be seen that the number of starts and the number of stops may be reduced.
  • the nopline calculation processing is performed in units of macro blocks (macro block data is composed of a plurality of block data, and is often composed of 24 pieces). It is also possible to adopt a configuration that executes in units of force blocks, which was the configuration example executed in step 1.
  • the arithmetic unit executes pipeline operation processing in block data units, and the clock supply control means power is added to the last block data in the last macro block data in the frame data.
  • the stop of the clock supply to the next-stage computing unit is continued even if the end of the computation processing of the previous-stage computing unit is detected.
  • the clock supply to all the arithmetic units is resumed for the next frame data to be processed for the knock-in operation.
  • the digital VLSI circuit of the third embodiment performs dynamic control by feedback control or feedforward control with respect to the operation power supply voltage, substrate bias voltage, and operation frequency of the arithmetic unit.
  • the digital VLSI circuit configured by the pipeline of the present invention shown in the first and second embodiments is more effective than the conventional digital VLSI circuit configured by the pipeline in the decoding target bit stream.
  • the number of cycles required for decoding processing varies greatly from frame to frame.
  • the number of cycles required for the encoding process depends on the number of block matching executed in the motion compensation process, the number of effective blocks generated and the effective coefficient. It varies greatly from unit to unit. Therefore, in the case of the digital VLSI circuit of the dedicated hardware configuration of the present invention shown in the embodiment 2, the operation power supply voltage of the arithmetic unit is applied by applying the feedback type dynamic control or the feed forward type dynamic control.
  • power consumption can be reduced by limiting the substrate bias voltage to an appropriate value. It is also effective to reduce the power consumption to keep the operating frequency of the computing unit to an appropriate value.
  • FIG. 16 is a block diagram in which feedback-type dynamic control is applied to a digital VLSI circuit having a dedicated hardware configuration.
  • the processed macroblock counter 80 is a part that counts the data processing amount of the pipeline operation processing for each frame data.
  • the feedback control unit 81 determines the number of unprocessed macroblocks included in the currently processed frame and the processing of the currently processed frame according to the number of processed macroblocks counted by the processed macroblock counter 80. This is the part that calculates the operating frequency of the computing unit from the time when it must be completed.
  • the computing unit adjustment unit 82 is a unit that adjusts the operating power supply voltage, the substrate bias voltage, and the operating frequency of the computing unit based on the operating frequency determined by the feedback control unit 80.
  • the digital VLSI circuit 100 shown in Example 1 or Example 2 is Then, a feedback loop is formed by the processed macroblock counter 80, the feedback control unit 81, and the arithmetic unit adjustment unit 82, thereby providing feedback on the operating power supply voltage, substrate bias voltage, and operating frequency of the arithmetic unit in the digital VLSI circuit 100. Dynamic control by control can be performed.
  • the feedback loop there are various methods for performing feedback control by the cooperation of the processed macroblock counter 80, the feedback control unit 81, and the arithmetic unit adjustment unit 82.
  • the count of the processed macroblock counter 80 indicates the number of processed macroblocks for each elapsed time.
  • the processing time is short as described in the first and second embodiments.
  • a cycle margin for processing is created.
  • the feedback control unit 81 can cause the computing unit adjustment unit 82 to adjust the operating power supply voltage, the substrate voltage, and the operating frequency of the computing unit to nZ (2n ⁇ m).
  • FIG. 17 is a block diagram in which feedforward dynamic control is applied to a digital VLSI circuit with a dedicated hardware configuration.
  • the processing load prediction unit 90 detects the macroblock data amount included in the frame data used for the pipeline operation processing before being subjected to the pipeline operation processing, and reduces the processing load involved in the pipeline operation processing. This is the part to be predicted.
  • the feedforward control unit 91 is a part that determines the operation power supply voltage, the substrate bias voltage, and the operation frequency of the computing unit based on the prediction by the processing load prediction unit 90.
  • the computing unit adjustment unit 92 is a unit that adjusts the operating power supply voltage, the substrate bias voltage, and the operating frequency of the computing unit based on the determination of the feedforward control unit 91.
  • a processing load prediction unit 90, a feedforward control unit 91, and an arithmetic unit adjustment unit 92 perform a feedforward loop.
  • dynamic control by feedforward control can be performed on the operation power supply voltage, the substrate bias voltage, and the operation frequency of the arithmetic unit in the digital VLSI circuit 100.
  • the processing load predicting unit 90 stores the number of processing cycles necessary for processing data to be processed within the past restricted time.
  • the time limit is one frame
  • the data to be processed within the time limit is all macroblocks included in one frame.
  • the processing load prediction unit 90 determines the number of processing cycles for each frame type.
  • the processing load prediction unit 90 examines the frame type of the frame to be processed from now on, predicts the past processing cycle number corresponding to the frame type as the processing load cycle, and sends a signal indicating the predicted cycle number to the feedforward unit 91. Is output. Based on the prediction of the processing load prediction unit 90, the feedforward unit 91 controls the arithmetic unit adjustment unit 92 so as to decrease the operating power supply voltage, the substrate bias voltage, and the operating frequency. If the worst cycle number is n and the predicted cycle number is m, control can be performed so as to decrease to mZn.
  • the feedforward unit 91 adjusts to (1. l) m or (1.2) m and then adjusts to the calculator adjustment unit 92 ( 1. Adjust to l) m Zn and (1.2) mZn.
  • the fourth embodiment has a configuration in which power supply control means is used instead of the clock supply control means of the digital VLSI circuit configuration shown in the second and third embodiments.
  • Low power consumption can be achieved by reducing the number of clock gating starts and stops.
  • the power supply itself is controlled to turn on and off instead of clock gating. The same effect can be obtained by reducing the number of stops.
  • the portion related to the clock supply is replaced with the portion related to the power supply, and in the description of the timing chart, the clock gating period is replaced with the power stop period, and the corresponding drawings can be rewritten and read.
  • the clock input 14 may be read as the power supply line 14).
  • Embodiment 5 is an application example in which the digital VLSI circuit of the present invention shown in Embodiments 1, 2, 3, and 4 is incorporated.
  • FIG. 18 is a diagram showing a configuration example of an image processing system 200 incorporating the digital VLSI circuit of the present invention.
  • the digital VLSI circuit of the present invention may be configured as a personal computer adopting a microprocessor, or the digital VLSI circuit of the present invention may be incorporated into an image processing board as an image processing chip.
  • FIG. 19 is a diagram showing a configuration example of a portable terminal 300 incorporating the digital VLSI circuit of the present invention.
  • This configuration example is an example incorporated in a mobile phone.
  • mobile phones with the ability to handle moving images have been introduced, but the demand for low power consumption is extremely strong.
  • low power consumption It is possible to increase both power and improve the processing speed of moving images.
  • the processed data is exchanged between the arithmetic units, whereby the number of data processing cycles (in the case of moving image processing)
  • the number of cycles for block processing and block processing) can be reduced, and the number of start and stop times for clock gating or power supply can be reduced by reducing the number of pipeline operations in the computing unit. It is possible to further reduce power consumption.
  • the number of cycles required within the constraint time can be greatly reduced according to the number of cycles required for processing actual data, and the operating frequency 'voltage can be reduced using the time margin generated by the reduced cycles. By performing dynamic control, it is possible to reduce the power consumption as much as possible.
  • the mobile terminal of the present invention low power consumption is achieved in moving image data processing, and a moving image encoding / decoding process can be performed even in a small terminal such as a mobile phone.
  • the use of mobile terminals will be expanded.
  • FIG. 1 is a diagram schematically showing a configuration of a digital VLSI circuit according to Embodiment 1 of the present invention.
  • FIG. 2 is a diagram showing a configuration example of clock supply control means
  • FIG. 3 is a timing chart showing the pipeline processing by the digital VLSI circuit of the first embodiment.
  • FIG. 4 is a diagram schematically showing the progress of the pipeline operation in the digital VLSI circuit according to the first embodiment.
  • FIG. 6 is a diagram schematically showing a configuration of a digital VLSI circuit according to the second embodiment of the present invention.
  • FIG. 7 is a diagram showing in detail an example of the configuration of the arithmetic unit 10 and clock supply control means according to the second embodiment.
  • FIG. 8 A diagram showing the flow of processing when the pipeline processing of the current-stage computing unit 10 is completed.
  • FIG. 9 State power of FIG. 8 The pipeline processing of the next-stage computing unit 10 is completed and the next-stage computing unit. 10's Diagram showing the flow of processing when a request signal is detected
  • the state force of Fig. 9 is also a diagram showing the flow of processing when the data processing of the pre-stage computing unit 10 is completed
  • the state force in Fig. 12 is also a diagram showing the operation when the pipeline processing of the next-stage computing unit 10 is completed and a request signal is detected from the next-stage computing unit 10
  • FIG. 14 is a timing chart showing the knock line processing by the digital VLSI circuit of the second embodiment.
  • FIG. 15 is a diagram schematically showing the progress of pipeline operation in the digital VLSI circuit according to the second embodiment.
  • FIG.17 Block diagram of feedforward dynamic control applied to a digital VLSI circuit with dedicated hardware configuration
  • FIG. 18 is a diagram showing a configuration example of an image processing system incorporating a digital VLSI circuit of the present invention.
  • FIG. 19 A diagram showing a configuration example of a portable terminal incorporating the digital VLSI circuit of the present invention.
  • FIG. 21 A diagram showing an example in which an H.264 decoding processing module is constructed by dedicated hardware using a plurality of arithmetic units based on the block diagram shown in FIG.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)
  • Image Processing (AREA)

Abstract

要約  実際のパイプライン演算処理において、演算器ごとの電力供給を制御しつつ、制約時間内での演算器への電力供給オンオフの切り替えを少なくすることにより低消費電力化を達成することのできるデジタルVLSI回路を提供する。本発明のデジタルVLSI回路は、パイプライン演算処理の各ステージを担い、クロックに同期して演算処理を実行する複数の演算器と、演算器における担当ステージの演算処理の終了を検知する検知手段と、演算器ごとにクロックの供給/停止を制御するクロック供給制御手段とを備え、クロック供給制御手段が、検知手段により演算処理の終了が検知された演算器に対するクロック供給を停止し、検知手段によりすべての演算器における演算処理の終了が検知されれば次のパイプライン演算処理に向けてすべての演算器へのクロック供給を再開するように構成する。

Description

明 細 書
デジタル VLSI回路およびそれを組み込んだ画像処理システム 技術分野
[0001] 本発明は、低消費電力デジタル VLSI回路、特に、パイプライン演算処理を行なう 内部の複数の演算器のごとにクロック供給、電力供給を制御することにより低消費電 力化を図ったデジタル VLSI回路、また、動作電源電圧と基板バイアス電圧と動作周 波数をフィードバック制御またはフィードフォワード制御を行なうことにより低消費電力 化を図ったデジタル VLSI回路に関する。
さらに本発明は、低消費電力デジタル VLSI回路を組み込んだ画像処理システム 並びに携帯端末に関する。
背景技術
[0002] 近年、通信ネットワークを通じて動画像の送受信を行うことや、動画像を蓄積メディ ァに蓄積することが広く行なわれている。一般に、動画像は情報量が大きいため、伝 送ビットレートの限られた通信路を用いて動画像を伝送する場合、あるいは蓄積容量 の限られた蓄積メディアに動画像を蓄積する場合には、動画像を符号化'復号化す る技術が必要不可欠である。動画像の符号化'復号ィ匕方式として、 ISO/IECが標準 化を進めている MPEG(Moving Picture Experts Group)や H.26Xがある。これらは動 画像を構成する経時的に連続した複数のフレームの符号化又は復号化を行うもので あり、動画像の時間的相関、空間的相関を利用した冗長性の削減を行うことにより動 画像の情報量を減らして符号化し、符号化された動画像を再度元の動画像に復号 化する技術である。
力かる符号化'復号ィ匕技術はパーソナルコンピュータやマイクロコンピュータを内蔵 する携帯電話等の情報端末機器等に適用されている。
[0003] 図 20は、 H. 264デコード処理モジュールのブロック図である。
H. 264デコード処理モジュールを構築する方法としては専用ハードウェアにより構 築する方法と、符号化'復号化の手段を記述したプログラムに基づいて汎用のプロセ ッサを動作させることにより構築する方法がある。 [0004] 図 21は、図 20に示したブロック図に基づき、複数の演算器を用いた専用ハードウ エアにより H. 264デコード処理モジュールを構築した例である。
差分画像に関する情報としてのビットストリームがビットストリームバッファ 1001に受 け入れられた後、エントロピーデコード 1002 (可変長復号ィ匕処理)、逆量子化処理 1 003 (逆 Q処理)、逆直交変換処理 1004 (逆 T処理)の順に処理が行なわれ、差分画 像が生成される。一方、差分画像生成処理と並行して、現在のフレームメモリに展開 されて ヽる画像を基に予測画像生成処理 1005が実行される。この差分画像と予測 画像との加算処理 1006によりフレーム画像が生成される。
[0005] これら演算処理はシーケンシャルにつながっており、これら複数の演算処理を適切 に分割してパイプライン演算処理とすることができる。図 21に示した演算処理をマク ロブロックレベルのパイプライン処理とする場合、図 22に示すようなパイプライン分割 が考えられる。図 22の例では 7段のパイプラインに分割されている。
ノ ィプライン演算処理のハードウェア設計にぉ ヽてノ ィプラインが破綻しな ヽように 考慮しておくことが重要である。
そこで、パイプライン分割において、各段のパイプライン処理でもっとも時間がかか つてしまった場合のサイクル数を最悪サイクル数 Sn(nは自然数)として設計し、各段 のパイプライン処理が破綻しないようにそれぞれのパイプラインの処理時間として最 悪サイクル数 Snを確保せしめておくことが一般的である。
[0006] なお、パイプライン分割の設計にぉ 、て、パイプライン全体の処理性能が高くなるよ うに各段のパイプライン処理の最悪サイクル数 Snができるだけ均等になるように設計 しておくことが好ましい。
つまり以下の (数式 1)の式が成立するように設計しておくことが好ま 、。
[0007] 〔数 1〕
Si=Sj for anyi, j
[0008] ここで、理想的な動画像のデコード処理とは、上記 (数式 1)を満たし、かつ、すべて の演算器の処理性能が上限まで発揮され、常に一定のサイクル数で処理が継続さ れていくものである。当該一定のサイクル数を最悪サイクル数と設計しておけば、演 算器がまったく無駄に遊ぶことなぐ演算器の処理性能の上限までパイプライン処理 力 S実行されることとなる。
しかし、実際のパイプライン処理ではそのような理想的な状態はなぐ実際のパイプ ライン処理のサイクル数は、常に最悪サイクル数よりも小さ 、ものとなる。
[0009] まず、動画像間の変化の激しさ(動きの激しさ)が小さ 、フレームにつ 、ては、処理 サイクル数は最悪サイクル数よりも小さいものとなる。なぜならば、動画像符号化'復 号化処理は動画像間の変化の激しさ (動きの激しさ)などに従って演算処理量が大き くなるため、あまり動画像間の変化の激しさ(動きの激しさ)が大きくないフレームでは 処理サイクル数は小さ 、ものとなるからである。
また、動画像間の変化の激しさ(動きの激しさ)が大きいフレームであっても以下の 理由力も処理サイクル数は最悪サイクル数よりも小さいものとなる。
[0010] (理由 1) 1マクロブロックに含まれる係数(384個)のうち、逆 Q処理が必要な 0以外 の値を持つ有効係数の数は 384個よりも小さ 、。
(理由 2) 1マクロブロックに含まれるブロック(24個)のうち、逆 T処理が必要な 1っ以 上の有効係数を含む有効ブロックの数は 24個よりも小さい。
(理由 3) 1マクロブロックに含まれるブロックのうち、 Intra予測処理が必要なブロック は 24個よりも小さい。また、 Intra予測が必要なブロックでも、予測モード(画素値のコ ピーのみでよい場合と複数の画素力 計算しなければならない場合)によって、必要 な処理サイクルが変動する。
(理由 4) 1マクロブロックに含まれるブロックのうち、予測画像と差分画像の加算が 必要な有効ブロックの数は 24個よりも小さ!/、。
(理由 5) 1マクロブロックに含まれるブロックのうち、デブロッキングフィルタ処理が必 要なブロックの数は 24個よりも小さ!/、。
(理由 6) 1マクロブロックに含まれるブロックのうち、フレームメモリ(FM)への書き込 みが必要なブロックの数は 24個よりも小さ 、。
上記理由力も動画像間の変化の激しさ(動きの激しさ)が大きいフレームであっても 処理サイクル数は最悪サイクル数よりも小さいものとなる。
[0011] 図 23は実際のノ ィプライン処理における演算器の動作状況を模式的に示すタイム チャートである。図 23において、ノ、ツチングを掛けた部分が、実際に演算器が動作状 態にある期間を示しており、ノ、ツチングが掛力つていない部分が、演算器が動作して
Vヽな 、アイドル状態にある期間を示して 、る。図 23では 3つの演算処理ステージで 構成されるパイプラインを示して ヽるが、パイプラインが 3つ以上の演算処理ステージ で構成される場合は、演算ステージの数に応じて演算器の個数が増えることになる。 動画像データをマクロブロックパイプラインで処理するデジタル VLSI回路の場合、図 23のデータ nをマクロブロック nと読み替え、ブロックパイプラインで処理するデジタル VLSI回路の場合、ブロック nと読み替える。
図 23を見ると、すべての演算器においてハツチングが掛力つていない期間 (演算器 が動作して ヽな 、アイドル状態の期間)が生じて!/、ることが分かる。
[0012] 上記のように、パイプラインの演算処理では、実際の演算処理よりも冗長性 (余裕) を持たせた設計となって 、るため、演算器が動作して 、な 、冗長なサイクルが発生 する。
このパイプラインの演算処理における冗長性 (余裕)を利用してデジタル VLSI回路 全体の消費電力を下げる工夫としてクロックゲーティングがある(特開平 10— 02095 9号公報)。
[0013] クロックゲーティングとは、ノ ィプライン演算処理実行中に、動作する必要のない演 算器にはクロックを供給せず消費電力を低減する手法である。クロック〖こ同期して動 作する演算器においては、クロック系統での電流消費が大きいため、演算処理を実 行して 、る演算器のみにクロックを供給するように構成することでデジタル VLSI回路 全体の消費電力を低減することができる。図 23に示したように、パイプラインの演算 処理におけるハツチングが掛力つて ヽな 、期間は、演算器が動作して ヽな 、アイドル 状態の期間であるので、当該期間においては、演算器へのクロック供給を停止して デジタル VLSI回路全体の消費電力を小さくする。
なお、演算処理を行わな!、冗長なサイクルの電力供給を制御すると!、う観点から、 クロック供給の制御に代え、演算器単位で電力供給 (電流供給)のオンオフを制御す るものも同様である(特開 2005— 235203号公報)
[0014] 図 24は図 20に示したブロック図に基づき、復号化の手段を記述したプログラムに 基づいて汎用プロセッサで動作させること(以下、ソフトウェア処理と記す)により H. 2 64復号化処理モジュールを構築した場合のフローチャートの例である。図 24では 1 フレームの処理にっ 、て記載してある。フレームを構成する各マクロブロックに対し、 エントロピーデコード、逆 Q、逆 T、イントラ予測 Ζインター予測、予測画像と差分画像 のカロ算、デブロッキングフィルタ処理、フレームメモリへの書き込みがシーケンシャル に実行され、これらの処理がフレームを構成するマクロブロックの個数分だけ繰り返さ れる。
[0015] 図 26はソフトウェア処理における処理に必要なサイクル数の状況を模式的に示した 図である。動画像の符号化 Ζ復号化処理は、 1フレームの演算処理時間が符号ィ匕方 式(MPEG、 H. 26xなど)の規定などによりフレーム処理時間 Tfに制約されている。 したがって、ソフトウェア処理によって符号化 Z復号化処理を行う場合、どのような動 画像に対しても 1フレームの演算処理に必要なサイクル数が時間 Tf以内に完了する ようにプログラムを構築する。もしくは、プログラムが 1フレームの演算処理に必要とす るサイクル数が時間 Tf以内に収まるようにプログラムを動作させる汎用プロセッサの 動作周波数 Fmaxを選択する。
[0016] しかし、実際に動画像を符号化 Z復号化処理した場合、前述のごとく説明した理由 力 フレームに含まれる各マクロブロックの演算処理に必要なサイクル数が変動する ため、フレームの演算処理に必要なサイクル数は大きく変動する。このとき、汎用プロ セッサの動作周波数を Fmaxとしてフレームの演算処理を行った場合、図 25に示す ようにプロセッサが演算処理を行わな 、冗長なサイクルが発生する。
[0017] ソフトウェア処理での演算処理においては、冗長なサイクルが発生する特徴を利用 した消費電力削減手法としてプロセッサの動作電源電圧、基板バイアス電圧、動作 周波数を動的に制御する手法がある(例えば IEEE International Symposium on Circu its and System 2001(May,2001)の予稿集 pp918- 921など)。
動画像における符号化処理は、 1フレームの処理時間が符号ィ匕方式 (MPEGなど) の規定などによりフレーム処理時間 Tfに制約されており、そのフレーム処理時間 Tf 内に 1フレームの符号ィヒ処理が完了することが必要とされる。逆に言えば、符号化演 算処理が当該フレーム処理期間 Tf中に完了すれば良いこととなる。
動作電源電圧、基板バイアス電圧、動作周波数の動的制御は、所定の制約時間 内に所定数のデータ群の処理完了を保証しながら、なるべくプロセッサの動作周波 数を下げ、動作周波数に合わせて電源電圧、基板バイアス電圧を動的に制御するこ とで総合的に低消費電力化を図るものである。動画像処理においては所定の制約時 間とは 1フレームの時間、例えば 15 (フレーム Z秒)の動画像であれば 1フレームの 時間は 15分の 1秒となり、所定数のデータ群とは 1フレームに含まれるマクロブロック 群と置き換えるとができる。
[0018] 特許文献 1:特開 2003 - 324735号公報
特許文献 2:特開 2005— 210525号公報
非特許文献 1 : IEICE Trans. Fundamentals, V0I.E88- A, No.12 December 2005. "Po wer- Minimum Frequency/Voltage Cooperative Management Method for VLSI Proce ssor in Leakage-Dominant Technology Era."(K. Kawakamik M. Kanamon, Morita , J. Takemura, M. Miyama, and M. Yoshimoto)
非特許文献 2 : IEEE International Symposium on Circuits and System 2001(May,2001 )の予稿集 pp918- 921 " An LSI for Vdd- Hopping and MPEG4 System Based on the Chip (H. Kawaguchi, G. Zhang, S. Lee, and T. Sakurai)
発明の開示
発明が解決しょうとする課題
[0019] 上記従来のクロックゲーティング手法や、演算器の動作周波数及び動作電源電圧 を動的に制御する手法は、デジタル VLSI回路全体の消費電力を低下させる技術と して有効な技術ではあるが、さらなる改善の余地がある。
特定の処理に特化して設計された ASIC (Application Specific IC)のようなハ ードワイヤドロジック回路に対してクロックゲーティング手法用いた場合の改善点の一 つとして、フレーム処理期間中において演算器へのクロック供給のオンオフ、演算器 の動作電源電圧及び動作周波数の変更回数が多くなり当該変更回数に伴って消費 電力が大きくなるという問題点を挙げることができる。また、ソフトウェア処理に適用で きる動作電源電圧、基板バイアス電圧、動作周波数を動的に制御する手法は、通常 のパイプライン処理ィ匕されたノヽードワイヤドロジック回路には適用できないという問題 点を挙げることができる。 [0020] 演算器の動作状態とアイドル状態が繰り返されるパイプライン動作に対して、従来 のクロックゲーティング手法を用いた場合は、演算器単位に図 23のハツチングが掛 かって ヽな 、サイクルにお!/、てクロックゲーティングを行うこととなる力 所定の制約時 間内において演算器へのクロック供給のオンオフが頻繁に繰り返されることになる。こ こに改善の余地がある。
[0021] また、従来のハードワイヤドロジック回路の設計では、ノ ィプラインを構成する各ス テージの実際に処理に必要なサイクル数が最悪サイクル数より小さくなつたとしても、 パイプライン処理の開始サイクルは設計時に決められたサイクルで固定されているた め、すべての演算器がアイドル状態であるサイクルを省略してノ ィプライン処理を前 倒しすることができず、ソフトウェア処理における図 25に示すようなサイクル数の大幅 な削減ができな力つた。図 23の例では、例えば 10番目のデータが演算器 1での処理 が開始されるサイクルは 1から 9番目のデータの処理に必要なサイクル数がどれだけ 小さくなろうとも必ず 900サイクルとなるため、例えば 10番目のデータの演算器 1での 処理に必要なサイクルが 10サイクルとすれば、演算器 1において 1から 10番目のデ ータの処理に必要な合計サイクル数は 100 X 9 + 10 = 910サイクルとなり、実際には 1から 10番目のデータの処理が最悪サイクル数である場合の 100 X 10= 1000サイ クルと比較して 90サイクルしか削減されないことになる。例えば 1920 X 1088画素で 構成される HDTV解像度の動画像では 1フレームに 8160個のマクロブロックが含ま れるため、 8160番目のマクロブロックが演算器 1での処理に必要なサイクルが 10サ イタルであったとすると、この場合 1フレームの処理に必要なサイクル数は(100 X 81 59 + 10) 7 (100 X 8160) X 100^0. 01%しか削減されない。したがって、ハード ワイヤドロジック回路の動作周波数は最大で 0. 01%しか下げられず、動作周波数- 電源電圧、基板バイアス電圧の動的制御による消費電力削減は全く実現できなかつ たと言える。
[0022] 本出願人らの研究により、デジタル VLSI回路(汎用プロセッサ、ハードワイヤドロジ ック回路を含む専用デジタル VLSI回路)において動作周波数 ·電源電圧、基板バイ ァス電圧の動的制御をもちいて制約時間内に、あるサイクル数の演算処理を実現す る処理の消費電力を削減する場合、高い動作周波数での動作時間を短くすれば短 くするほど消費電力が削減されることが分力 ている(非特許文献 1)。実行されるサ イタル数は数式 2で表されるため、図 5、 28、 30はいずれも制約時間 Tfの間に同一 のサイクル数 Fmax X TfZ2を実現して!/、る。
[0023] 〔数 2〕
(サイクル数) = (動作周波数) X (動作時間)
[0024] 動作周波数'電源電圧、基板バイアス電圧の動的制御を行う場合、デジタル VLSI 回路の動作周波数の制御に合わせて動作電源電圧、基板バイアス電圧も制御され る。例えば図 5においては、時刻 0から TfZ2までは動作電源電圧及び基板バイアス 電圧は動作周波数 fmaxを適切な電力で実現しうる電圧に設定され、時刻 TfZ2か ら Tfまでは動作周波数 0を適切な電力で実現しうる電圧に設定される。
[0025] 上記問題点に鑑み、本発明は、実際のパイプライン演算処理において、クロックゲ 一ティング技術を利用しつつ、所定の制約時間での演算器へのクロック供給のオン オフ切り替えを少なくすることにより低消費電力化を達成することのできるデジタル V LSI回路を提供することを目的とする。
[0026] また、上記問題点に鑑み、本発明は、実際のパイプライン演算処理において、演算 器ごとの電力供給を制御しつつ、所定の制約時間での演算器への電力供給オンォ フの切り替えを少なくすることにより低消費電力化を達成することのできるデジタル V LSI回路を提供することを目的とする。
また、上記問題点に鑑み、本発明は、デジタル VLSI回路においてパイプライン処 理の開始サイクルを一つ前のパイプライン処理の終了サイクルに合わせて変更する ことでパイプラインを構成する演算器がアイドル状態である冗長なサイクルを削減す ることにより生じるサイクル余裕を利用して動作周波数、動作電源電圧、基板バイァ ス電圧の制御を行うことにより低消費電力化を達成することのできる専用デジタル VL SI回路を提供することを目的とする。
また、本発明のデジタル VLSI回路を組み込んだ画像処理システム、携帯端末を提 供することを目的とする。
課題を解決するための手段
[0027] 上記目的を達成するため、本発明の第 1のデジタル VLSI回路は、パイプライン演 算処理の各ステージを担い、クロックに同期して演算処理を実行する複数の演算器 と、前記演算器における担当ステージの演算処理の終了を検知する検知手段と、前 記演算器ごとにクロックの供給 Z停止を制御するクロック供給制御手段とを備え、前 記クロック供給制御手段が、前記検知手段により演算処理の終了が検知された前記 演算器に対するクロック供給を停止し、前記検知手段によりすベての前記演算器に おける演算処理の終了が検知されれば次のパイプライン演算処理に向けてすべて の前記演算器へのクロック供給を再開するように構成されたことを特徴としたものであ る。
上記構成により、すべての演算器がアイドル状態に入っている期間を省略すること ができ、演算器でのパイプライン演算処理を詰めて行なうことによりクロックゲーティン グの開始回数、終了回数を減少させることができ、従来のクロックゲーティングより一 層の低消費電力化を図ることができる。
[0028] 次に、上記第 1のデジタル VLSI回路において、前記演算処理に力かるデータが、 複数のマクロブロックデータを含み、一定の処理期間(フレーム処理期間)が定めら れて 、るフレームデータを備えた動画像データであり、前記演算処理が前記動画像 の符号化 '復号化処理であり、前記演算器が、前記パイプライン演算処理を前記マク ロブロックデータ単位で実行し、前記クロック供給制御手段が、前記フレームデータ 中の最後のマクロブロックデータにかかる演算処理においては、前記検知手段により すべての前記演算器における演算処理の終了が検知されてもすべての前記演算器 に対するクロック供給の停止を継続し、前記フレーム処理期間の経過後、次のフレー ムデータのパイプライン演算処理に向けてすべての前記演算器へのクロック供給を 再開することを特徴とする。
上記構成により、すべての演算器がアイドル状態に入っている期間を省略するこ とができ、演算器でのノ ィプライン演算処理を詰めて行なうことによりクロックゲーティ ングの開始回数、終了回数を減少させることができ、従来のクロックゲーティングより 一層の低消費電力化を図ることができる。なお、ノ ィプライン演算処理はマクロブロッ クデータ単位で行なうものである。
[0029] また、上記第 1のデジタル VLSI回路において、前記演算処理に力かるデータが、 複数のマクロブロックデータ(前記マクロブロックデータは複数個のブロックデータに より構成される)を含み、一定の処理期間(フレーム処理期間)が定められて ヽるフレ ームデータを備えた動画像データであり、前記演算処理が前記動画像の符号化 ·復 号化処理であり、前記演算器が、前記パイプライン演算処理を前記ブロックデータ単 位で実行し、前記クロック供給制御手段が、前記フレームデータ中の最後の前記マク ロブロックデータ中の最後の前記ブロックデータにかかる演算処理においては、前記 検知手段によりすベての前記演算器における演算処理の終了が検知されてもすべ ての前記演算器に対するクロック供給の停止を継続し、前記フレーム処理期間の経 過後、次のフレームデータのパイプライン演算処理に向けてすべての前記演算器へ のクロック供給を再開することを特徴とする。
なお、マクロブロックに含まれるブロック数は例えば 24個とする。
上記構成により、すべての演算器がアイドル状態に入っている期間を省略すること ができ、演算器でのパイプライン演算処理を詰めて行なうことによりクロックゲーティン グの開始回数、終了回数を減少させることができ、従来のクロックゲーティングより一 層の低消費電力化を図ることができる。なお、ノ ィプライン演算処理はブロックデータ 単位で行なうものである。
[0030] 次に、本発明の第 2のデジタル VLSI回路は、パイプライン演算処理の各ステージ を担い、クロックに同期して演算処理を実行する複数の演算器と、前記演算器におけ る担当ステージの演算処理の終了を検知する検知手段と、前記演算器ごとにクロック の供給 Z停止を制御するクロック供給制御手段とを備え、前記クロック供給制御手段 力 前記検知手段により、前記パイプライン演算処理において相前後する前段演算 器と次段演算器のうち次段演算器の演算処理の終了が先に検知された場合、当該 次段演算器に対してクロック供給を停止し、前記次段演算器へのクロック供給の停止 後、前記前段演算器の演算処理の終了が検知された場合、次のパイプライン演算処 理に向けて前記次段演算器に対してクロック供給を再開するように構成されたことを 特徴とする。
[0031] なお、上記の第 2のデジタル VLSI回路において、前記クロック供給制御手段が、 前記検知手段により、前記前段演算器の演算処理の終了が先に検知された場合、 前記前段演算器が前記次段演算器に対して処理済みの演算処理データを出力でき るまで、当該前段演算器に対してクロック供給を停止し、前記前段演算器へのクロッ ク供給の停止後、前記前段演算器が前記次段演算器に対して処理済みの演算処理 データを出力できる状態となれば、当該前段演算器に対してクロック供給を再開する ように構成されることが好まし 、。
上記構成により、ノ ィプライン制御において前後に並ぶ演算器間で処理済みデー タの受け渡しができる限り、シームレスにどんどんパイプライン処理を実行し、演算器 でのパイプライン演算処理を詰めて行なうことによりクロックゲーティングの開始回数、 終了回数を減少させることができ、低消費電力化を図ることができる。
[0032] 次に、上記の第 2のデジタル VLSI回路において、前記演算処理に力かるデータが 、複数のマクロブロックデータ力 構成され、処理を完了すべき制約時間(フレーム処 理時間)が定められているフレームデータ力も構成される動画像データであり、前記 演算処理が前記動画像の符号化 '復号化処理であり、前記演算器が、前記パイプラ イン演算処理を前記マクロブロックデータ単位で実行し、前記クロック供給制御手段 力 前記フレームデータ中の最後のマクロブロックデータにかかる演算処理において は、前記次段演算器へのクロック供給の停止後、前記前段演算器の演算処理の終 了が検知されても、前記次段演算器へのクロック供給の停止を継続し、前記フレーム 処理期間の経過後、次のフレームデータのパイプライン演算処理に向けてすべての 前記演算器へのクロック供給を再開することを特徴とする。
[0033] また、上記の第 2のデジタル VLSI回路にぉ 、て、前記クロック供給制御手段が、前 記フレームデータ中の最後のマクロブロックデータにかかる演算処理においては、前 記前段演算器へのクロック供給の停止後、前記前段演算器が前記次段演算器に対 して処理済みの演算処理データを出力できる状態となった場合でも、前記前段演算 器へのクロック供給の停止を継続し、前記フレーム処理期間の経過後、次のフレーム データのパイプライン演算処理に向けてすべての前記演算器へのクロック供給を再 開することを特徴とする。
上記構成により、ノ ィプライン制御において前後に並ぶ演算器間で処理済みデー タの受け渡しができる限り、シームレスにどんどんパイプライン処理を実行し、演算器 でのパイプライン演算処理を詰めて行なうことによりクロックゲーティングの開始回数、 終了回数を減少させることができ、低消費電力化を図ることができる。なお、パイプラ イン演算処理はマクロブロックデータ単位で行なうものである。
[0034] 次に、上記の第 2のデジタル VLSI回路において、前記演算処理に力かるデータが 、複数のマクロブロックデータ(前記マクロブロックデータは複数個のブロックデータに より構成される)を含み、一定の処理期間(フレーム処理期間)が定められて ヽるフレ ームデータを備えた動画像データであり、前記演算処理が前記動画像の符号化 ·復 号化処理であり、前記演算器が、前記パイプライン演算処理を前記ブロックデータ単 位で実行し、前記クロック供給制御手段が、前記フレームデータ中の最後の前記マク ロブロックデータ中の最後の前記ブロックデータにかかる演算処理においては、前記 次段演算器へのクロック供給の停止後、前記前段演算器の演算処理の終了が検知 されても前記次段演算器へのクロック供給の停止を継続し、前記フレーム処理期間 の経過後、次のフレームデータのパイプライン演算処理に向けてすべての前記演算 器へのクロック供給を再開することを特徴とすることが好ましい。
なお、マクロブロックに含まれるブロック数は例えば 24個とする。
[0035] また、上記の第 2のデジタル VLSI回路にぉ 、て、前記クロック供給制御手段が、前 記フレームデータ中の最後の前記マクロブロックデータの最後の前記ブロックデータ にかかる演算処理においては、前記前段演算器へのクロック供給の停止後、前記前 段演算器が前記次段演算器に対して処理済みの演算処理データを出力できる状態 となった場合でも、前記前段演算器へのクロック供給の停止を継続し、前記フレーム 処理期間の経過後、次のフレームデータのパイプライン演算処理に向けてすべての 前記演算器へのクロック供給を再開することを特徴とすることが好ましい。
上記構成により、ノ ィプライン制御において前後に並ぶ演算器間で処理済みデー タの受け渡しができる限り、シームレスにどんどんパイプライン処理を実行し、演算器 でのパイプライン演算処理を詰めて行なうことによりクロックゲーティングの開始回数、 終了回数を減少させることができ、低消費電力化を図ることができる。なお、パイプラ イン演算処理はブロックデータ単位で行なうものである。
[0036] なお、上記第 1または第 2のデジタル VLSI回路において、前記パイプライン演算処 理のデータ処理量を前記所定の制約時間ごとにカウントし、次の所定の制約時間の 処理時における前記演算器の動作電源電圧と基板バイアス電圧と動作周波数とを 決定するフィードバック制御部と、前記演算器の動作電源電圧と基板バイアス電圧と 動作周波数とを調整する演算器調整部とを備え、前記演算器の動作電源電圧と基 板バイアス電圧と動作周波数に関してフィードバック制御による動的制御を行なうこと を特徴とする。
上記構成によれば、フィードバック制御により、適切な演算器の動作電源電圧と基 板バイアス電圧と動作周波数の調整を行なうことができる。
[0037] また、上記第 1または第 2のデジタル VLSI回路において、前記パイプライン演算処 理に供される前に、前記パイプライン演算処理に供される前記所定の制約時間に含 まれるデータ量を検知し、前記ノ ィプライン演算処理に力かる処理負荷を予測する 処理負荷予測部と、前記処理負荷予測部による予測に基づき、前記演算器の動作 電源電圧と基板バイアス電圧と動作周波数とを決定するフィードフォワード制御部と 、前記演算器の動作電源電圧と基板バイアス電圧と動作周波数とを調整する演算器 調整部とを備え、前記演算器の動作電源電圧と基板バイアス電圧と動作周波数に関 してフィードフォワード制御による動的制御を行なうことを特徴とする。
上記構成によれば、フィードフォワード制御により、適切な演算器の動作電源電圧 と基板バイアス電圧と動作周波数の調整を行なうことができる。
[0038] 上記の第 1または第 2のデジタル VLSI回路は、クロックゲーティング技術を利用す るものであつたが、演算器への電力供給のオンオフを制御することによる低消費電力 ィ匕も可能である。つまり、クロックゲーティングにおいて演算器へのクロック供給を停 止するタイミングで、演算器への電力供給を停止することとしても良い。
例えば、本発明の第 3のデジタル VLSI回路は、パイプライン演算処理の各ステー ジを担い、演算処理を実行する複数の演算器と、前記演算器における担当ステージ の演算処理の終了を検知する検知手段と、前記演算器ごとに電力の供給 Z停止を 制御する電力供給制御手段とを備え、前記電力供給制御手段が、前記検知手段に より演算処理終了が検知された前記演算器に対する電力供給を停止し、前記検知 手段によりすベての前記演算器における演算処理終了が検知されれば次のパイプ ライン演算処理に向けてすべての前記演算器への電力供給を再開するように構成さ れたことを特徴としている。
上記構成により、同様に、すべての演算器がアイドル状態に入っている期間を省略 することができ、低消費電力化を図ることができる。さらに、演算器でのパイプライン演 算処理を詰めて行なうことにより電力供給の開始回数、終了回数を減少させることが でき、より一層の低消費電力化を図ることができる。
[0039] また、本発明の第 3のデジタル VLSI回路において、前記演算処理にかかるデータ 力 複数のマクロブロックデータを含み、一定の処理期間(フレーム処理期間)が定め られているフレームデータを備えた動画像データであり、前記演算処理が前記動画 像の符号化 '復号化処理であり、前記演算器が、前記パイプライン演算処理を前記 マクロブロックデータ単位で実行し、前記電力供給制御部が、前記フレームデータ中 の最後のマクロブロックデータにかかる演算処理においては、前記検知手段によりす ベての前記演算器における演算処理の終了が検知されてもすべての前記演算器に 対する電力供給の停止を継続し、前記フレーム処理期間の経過後、次のフレームデ ータのパイプライン演算処理に向けてすべての前記演算器への電力供給を再開す ることを特徴とする。
[0040] 次に、本発明の第 4のデジタル VLSI回路は、パイプライン演算処理の各ステージ を担い、演算処理を実行する複数の演算器と、前記演算器における担当ステージの 演算処理の終了を検知する検知手段と、前記演算器ごとに電力の供給 Z停止を制 御する電力供給制御手段とを備え、前記電力供給制御手段が、前記検知手段により 、前記パイプライン演算処理において相前後する前段演算器と次段演算器のうち次 段演算器の演算処理の終了が先に検知された場合、当該次段演算器に対して電力 供給を停止し、前記次段演算器への電力供給の停止後、前記前段演算器の演算処 理の終了が検知された場合、次のパイプライン演算処理に向けて前記次段演算器 に対して電力供給を再開するように構成されたことを特徴とする。
[0041] また、本発明の第 4のデジタル VLSI回路において、前記電力供給制御手段が、前 記検知手段により、前記前段演算器の演算処理の終了が先に検知された場合、前 記前段演算器が前記次段演算器に対して処理済みの演算処理データを出力できる まで、当該前段演算器に対して電力供給を停止し、前記前段演算器への電力供給 の停止後、前記前段演算器が前記次段演算器に対して処理済みの演算処理データ を出力できる状態となれば、当該前段演算器に対して電力供給を再開するように構 成されたことを特徴とする。
上記構成により、ノ ィプライン制御において前後に並ぶ演算器間で処理済みデー タの受け渡しができる限り、シームレスにどんどんパイプライン処理を実行し、演算器 のアイドル状態の発生を極力抑えることができるので、低消費電力化を図ることがで きる。また、演算器でのパイプライン演算処理を詰めて行なうことにより電力供給の開 始回数、終了回数を減少させることができ、より一層の低消費電力化を図ることがで きる。
[0042] 上記第 3または第 4のデジタル VLSI回路において、上記第 1または第 2のデジタル VLSI回路と同様の種々の変形を行なうことができる。例えば、クロック供給を電力供 給と読み替え、クロック供給制御手段を電力供給制御手段とする。
発明の効果
[0043] 本発明に係るデジタル VLSI回路によれば、実際のパイプライン演算処理において 、クロックゲーティング技術を利用しつつ、制約時間内における演算器へのクロック供 給のオンオフ切り替えを少なくすることにより低消費電力化を達成することができる。 また、本発明に係るデジタル VLSI回路によれば、実際のパイプライン演算処理に おいて、演算器ごとの電力供給を制御しつつ、制約時間内における演算器への電力 供給オンオフの切り替えを少なくすることにより低消費電力化を達成すること
[0044] また、本発明に係るデジタル VLSI回路によれば、実際のパイプライン演算処理に おいて、制約時間内に演算処理を完了しなければならないデータの演算処理に必 要なサイクル数を最悪サイクル数から削減することができ、したがって制約時間内の デジタル VLSI回路の動作周波数を低下させても制約時間内に所定の演算処理を 完了することができ、したがってデジタル VLSI回路の動作周波数、動作電源電圧、 基板バイアス電圧を適切に制御することにより低消費電力化を達成することができる
[0045] 本発明の画像処理システムによれば、動画像のデータ処理において低消費電力 化が図られており、低消費電力が求められる様々なシステムに対して組み込むことが 容易となり、柔軟なシステム設計が可能となる。
[0046] 本発明の携帯端末によれば、動画像のデータ処理において低消費電力化が図ら れており、携帯電話のような小型の端末においても動画像の符号化復号化処理を行 なうことができ、携帯端末の用途が様々に広がる。
発明を実施するための最良の形態
[0047] 以下、本発明のデジタル VLSI回路の実施例について、図面を参照しながら詳細 に説明していく。本発明はパイプライン処理を実行するデジタル VLSI回路に広く適 用できるものであるが、ここでは一例として動画像の符号化'復号ィ匕を行なう用途に 使用するものを示す。
なお、信号のオンオフに関し、以下の実施例ではハイアクティブとし、論理レベルが ハイのときにアクティブになるように説明している力 ローアクティブとし、論理レベル 力 Sローのときにアクティブになる構成であつても良い。 実施例 1
[0048] 図 1は本発明の実施例 1にかかるデジタル VLSI回路の構成を模式的に示す図で ある。
演算器 10a〜10cが処理終了検知器 20と接続された構成となっており、各演算器 10と処理終了検知器 20は終了フラグライン 30と処理開始フラグライン 40により接続 されている。
演算器は図示の便宜上、演算器 1 (10a)、演算器 2 (10b)、演算器 3 (10c)の 3つ のみ示しているが、パイプラインの段数に応じて演算器の数を増減した設計とするこ とができることは言うまでもな!/、。
それぞれの演算器 10a〜: LOcは、パイプライン演算処理の各ステージを担い、クロ ックに同期して演算処理を実行するものである。動画像の符号化'復号化を行なうパ ィプライン処理を実行するのであれば、例えば、演算器 1 (10a)がエントロピーデコー ドステージを担当し、演算器 2 (10b)が逆 Q処理ステージを担当し、演算器 3 (10c) が逆 T処理ステージを担当するものとする。それ以降のノィプライン処理ステージは 図示を省略している。各演算器はシーケンシャルに接続され、前後の演算器の間で 処理済のデータを前段から次段へ次々と受け渡していく構成となっている。各演算 器で演算処理がなされたデータは各演算器に備えられるバッファに保存され、前後 の演算器はこのバッファを介してデータを受け渡す。例えば演算器 1で演算処理が なされたデータは演算器 1に備えられたバッファに保存され、このノ ッファカも次段の 演算器 2へデータが受け渡される。バッファはフリップフロップや RAM (Random A ccess Memory)などで構成される。
[0049] 図 1において、各演算器 10a〜 10cは処理終了検知器 20と接続され、各演算器が 担当する演算処理が終了すると演算器は終了フラグを立てる。つまり、処理終了検 知器 20と接続されている終了フラグライン 30にアクティブ信号を出力する。この例で はハイアクティブ論理としてハイ信号を出力して 、る(ローアクティブ論理の場合は口 一信号を出力すれば良い)。
[0050] この実施例 1の構成では、処理終了検知器 20が各演算器 10の終了フラグが立て られたことを検知することにより、演算器 10における担当ステージの演算処理の終了 を検知することができる。本実施例 1では、処理終了検知器 20は演算器 10における 演算処理の終了を終了フラグを介して検知する仕組みとなっている。なお、この処理 終了検知器 20は多入力の AND回路となって 、る。
[0051] 演算処理が終了した演算器 10に対してクロックゲーティングが行なわれる。この実 施例では、クロック供給制御手段は、各演算器 10は演算処理が終了すると終了フラ グを出力するとともにクロックの供給が一時停止されるように構成されている。
図 2はクロック供給制御手段の構成例を示す図である。図 2に示すように、状態マシ ンとフリップフロップと AND回路により自動的にクロック供給のオンオフが制御される 構成となっている。状態マシン 11からフリップフロップ 12を介して AND回路 13に接 続されている。演算器 10の出力ラインの一部は状態マシン 11に接続され、一部は終 了フラグライン 30に接続されている。図 2 (a)は演算器 10にクロック供給を開始する 際の動作の流れの一例を示す図、図 2 (b)は演算器 10のクロック供給を一時停止す る動作の流れの一例を示す図である。いま、図 2 (a)に示すように、フリップフロップ 1 2がオンとなり AND回路 13を通してクロック入力ライン 14からクロックが供給されてい る状態にあるとする。 図 2 (b)において、演算器 10はパイプラインの演算処理が終了すると終了信号を出 力する。当該終了信号は状態マシン 11を介してフリップフロップ 12に入力され、フリ ップフロップ 12を反転させてオフとなる。当該オフ信号により AND回路 13はオフとな る。そのため図 2 (a)の状態ではクロック入力ライン 14力も供給されていたクロックの 供給が停止する。
上記クロック供給の停止処理は演算器ごとに行なわれる。そのため、終了フラグを 出力した演算器 10から順々にクロックの供給が停止されて行くこととなる。
[0052] 図 1に戻って説明を続ける。処理終了検知器 20は、多入力の AND回路となってお り、すべての演算器 10の終了フラグを検知した場合 (すべての演算器において処理 が終了した場合)、次のパイプライン演算処理に向けてすべての演算器に対してクロ ック供給を再開するよう処理開始フラグライン 40に処理開始フラグ信号を出力する構 成となっている。処理開始フラグライン 40はすべての演算器 10に対して並列に接続 されており、処理開始フラグ信号はすべての演算器 10に対して一斉に通知される仕 組みとなっている。各演算器 10は処理開始フラグ信号を受け取ると一斉にクロック供 給の開始を受け、次のパイプライン処理の演算処理に移る。図 2 (a)に示すように、処 理開始フラグ信号が処理開始フラグライン 40から状態マシン 11を介してフリップフロ ップ 12に入力され、フリップフロップ 12は反転する(オフ力もオンへ)。図 2 (b)の状態 ではフリップフロップ 12がオフでありクロックの供給が停止されていた力 図 2 (a)の状 態ではフリップフロップ 12がオンとなり AND回路 13を介して演算器 10に対してクロッ ク供給が再開する。このクロック供給は演算器 10のすべてにおいて一斉に再開され る仕組みとなっている。
本実施例 1では、上記構成の制御に基づくクロック供給 Z停止制御の仕組みがクロ ック供給制御手段となっている。このように、本実施例 1の構成によれば、パイプライ ン処理において担当ステージの演算処理が終了した演算器 10から順にクロックの供 給が停止して行き、すべての演算器 10の演算処理が終了した場合、次のパイプライ ン処理の演算処理に向けてすべての演算器 10へのクロック供給が一斉に再開され る。
[0053] 図 3は実施例 1のデジタル VLSI回路によるパイプライン処理を示すタイミングチヤ ートである。このタイミングチャートでも演算器は、演算器 1 (10a)、演算器 2 (10b)、 演算器 3 (10c)の 3つのみ示して!/、る。
図 3のタイミングチャートにおいて、第 1クロックの時点では演算器 1 (10a)において データ (n+ 2)が処理され、演算器 2 (10b)においてデータ (n+ 1)が処理され、演 算器 3 (10c)においてデータ(n)が処理されている。デジタル VLSI回路がマクロブロ ックパイプラインで構成されて ヽればデータ(n)をマクロブロック(n)で、ブロックパイ プラインで構成されて ヽればデータ (n)をブロック (n)と置き換えればよ 、。
[0054] 演算器 1 (10a)におけるデータ (n+ 2)の処理は、図中の第 1クロックで完了してい る。演算器 1 (10a)はこの第 1クロックで終了フラグを立て終了フラグライン 30から処 理終了検知器 20に対して演算処理終了を通知するとともに、クロックゲーティングを 行なう。つまり、図 2 (b)に示したようにクロックの供給が停止される。
[0055] 演算器 2 (10b)におけるデータ(n+ 1)の処理は、図中の第 3クロックで完了してい る。演算器 2はこの第 3クロックで終了フラグを立てて終了フラグライン 30から処理終 了検知器 20に対して演算処理終了を通知する。この場合、後述するように処理終了 検知器の処理開始信号によりクロックゲーティングに遷移することなく次マクロブロッ クの処理〖こ移ることとなる。
[0056] 演算器 3 (10c)におけるマクロブロック(n)の処理も、図中の第 3クロックで完了して V、る。演算器 3はこの第 3クロックで終了フラグを立てて終了フラグライン 30から処理 終了検知器 20に対して演算処理終了を通知する。この場合も、後述するように処理 終了検知器の処理開始信号によりクロックゲーティングに遷移することなく次マクロブ ロックの処理に移ることとなる。
[0057] 処理終了検知器 20は、演算器 1 (10a)、演算器 2 (10b)、演算器 3 (10c)の終了フ ラグ信号について AND処理を行なう。この例では、第 3クロックで演算器 1 (10a)、演 算器 2 (10b)、演算器 3 (10c)からの終了フラグがすべて揃うこととなり AND条件が 成立する。処理終了検知器 20は次のパイプライン演算処理に向けてすべての演算 器 10へのクロック供給を再開するよう処理開始フラグライン 40に処理開始フラグ信号 を出力する。処理終了検知器 20の出力ラインは演算器 1 (10a)、演算器 2 (10b)、 演算器 3 (10c)のすべてに対して並列に接続されているので、処理開始フラグ信号 は演算器 1 (10a)、演算器 2 (10b)、演算器 3 (10c)のすべてに対して一斉に通知さ れる。演算器 1 (10a)、演算器 2 (10b)、演算器 3 (10c)は処理開始フラグ信号を受 け取ると一斉にクロックの供給が再開され、次のパイプライン処理の演算処理に移る
[0058] 第 4クロックにおいて、演算器 10はクロックの供給開始を受けると、まず、終了フラグ を非アクティブにする。この例ではハイアクティブであるのでローに切り替える。次に、 次のパイプライン処理が開始される。演算器 1 (10a)はデータ (n+ 3)に対して処理 の実行を開始し、演算器 2 (10b)はデータ (n+ 2)に対して処理の実行を開始し、演 算器 3 (10c)はデータ (n+ 1)に対して処理の実行を開始する。
[0059] 図 4は、実施例 1にかかるデジタル VLSI回路におけるパイプライン動作の進行を模 式的に示す図である。
縦軸に演算器 1、 2、 3など各演算器を模式的に並べている。演算器 1 (10a)による 処理が終了すれば、演算処理済みデータが演算器 2 (10b)に受け渡されて演算器 2 による処理が行なわれ、当該処理が終了すれば、演算処理済みデータが演算器 3 ( 10c)に受け渡されて演算器 3による処理が行なわれる。このように縦軸方向にパイプ ライン処理の流れが展開されて!、る。
横軸はタイミングである。第 1段目には、従来技術による最悪サイクル数を確保せし めつつパイプライン処理を実行する場合のタイミング (タイミング 0〜: L000)が表示さ れている。第 2段目には本発明の実施例 1にかかるデジタル VLSI回路のパイプライ ン処理を実行する場合のタイミング (タイミング 0〜: L000)が表示されている。図 4では ノ、ツチングが施されている部分は演算処理が実行されている期間を示しており、ハツ チングがない部分は演算処理が終了し、次のデータの処理の実行までの間のクロッ クゲーティング期間を示して 、る。
[0060] 図 4のタイミングチャートに示すように、各演算器のデータの処理について、その開 始タイミングが一斉に揃っていることがわかる。つまり、処理終了検知器 20から処理 開始フラグ信号が一斉に通知されると当該タイミングを持って各演算器が担当するパ ィプラインステージの処理を一斉に開始することが分かる。このタイミングは処理終了 検知器 20から処理開始フラグ信号が出力されたタイミング (第 2段目に表示したタイミ ング)となっている。図 4のタイミングチャートを見れば分力るように、演算器 10のうち、 担当するパイプライン処理が早く終了したものは、他の演算器 10におけるノ ィプライ ン処理が終了するまでの間、クロックゲーティングが行なわれている。例えば演算器 1 (10a)では、 275サイクノレと 300サイクノレの間、 425サイクノレと 450サイクノレの間、 57 5サイクルと 600サイクルの間、クロックゲーティングが行なわれる期間がある。
[0061] 図 4のタイミングチャートと図 23のタイミングチャートを比較すると明らかなように、従 来技術のデジタル VLSI回路におけるパイプライン動作の進行と、本発明の実施例 1 のデジタル VLSI回路におけるパイプライン動作の進行では、図 4と図 23のハツチン グを施した部分の面積の総合計は同じである。つまり、各演算器 10が動作している 期間の総合計サイクル数は同じである。同様に、ノ、ツチングを施していない部分の面 積の総合計は同じであり、クロックゲーティングの総合計時間が同じものであることが 分かる。つまり、クロックゲーティングを実行する時間を長くすることにより得られる低 消費電力効果は図 4の場合も図 23の場合も基本的に同じである。
[0062] しかし、本発明の実施例 1にかかる図 4では、各演算器 10での演算処理をできるだ け詰めてシームレスに連続処理として実行することにより、クロックゲーティング開始 によるクロックの供給停止の回数と、クロックゲーティング停止によるクロックの供給開 始の回数が減少している。例えば、図 4では、演算器 1がデータ 9の処理を完了する までのパイプライン処理までを見た場合、演算器 1におけるクロックゲーティング開始 回数が 3回(275サイクル、 425サイクル、 575サイクル)であり、クロックゲーティング 停止回数も 3回ある(300サイクル、 450サイクル、 600サイクル)。同様に演算器 2に 関してはクロックゲーティング開始回数が 3回、クロックゲーティング停止回数が 3回あ る。演算器 3に関してもクロックゲーティング開始回数が 3回、クロックゲーティング停 止回数が 3回ある。一方、図 23の場合は、 1回のパイプライン処理ごとにクロックゲー ティングが発生しているので、演算器 1がデータ 9の処理を完了するまでのパイプライ ン処理を見た場合、クロックゲーティング開始回数が 9回、クロックゲーティング停止 回数が 9回ある。同様に演算器 2についてはクロックゲーティング開始回数が 8回、ク ロックゲーティング停止回数が 8回、演算器 3についてもクロックゲーティング開始回 数が 7回、クロックゲーティング停止回数が 7回ある。 [0063] このように明らかに、クロックゲーティング開始回数およびクロックゲーティング停止 回数が減少して 、る。上記では 9回のパイプライン処理実行での回数で比較した力 ノ ィプライン処理実行回数が多くなるほどその差は広がり、本発明のデジタル VLSI 回路のクロックゲーティング開始回数、停止回数とも、従来のデジタル VLSI回路のク ロックゲーティング開始回数、停止回数がより少なくなることが理解されよう。例えば、 HDTV画像(1920 X 1088画素)の 1フレームは 8160個のマクロブロックで構成さ れるため、 HDTV画像の処理をマクロブロックパイプラインで実施する場合、ノィプラ イン処理の実行回数は 8160回となる。
[0064] 次に、フレームに含まれている最終のマクロブロックのパイプライン処理の終了時の 動作について説明する。
演算器 10は動画像の符号化'復号ィ匕処理をパイプライン演算処理によりマクロブロッ クデータ単位で実行する演算器である場合、上記のように演算器 10におけるパイプ ライン処理を詰めて!/、くので、 1フレームの処理時間中に次のフレームに含まれるマ クロブロックの処理を行える時間が存在する。例えば、図 4において 1フレームに含ま れるマクロブロックが 8個であるとした場合、演算器 1は 575サイクルで 1フレーム分の マクロブロックの処理が完了し、最終段の演算器 3が 8番目のマクロブロックの処理を 完了する 750サイクルまで次フレームに含まれるマクロブロックの処理を行えるサイク ルが存在する。しカゝし、フレームデータで構成されている動画像処理は、 1フレームの 処理期間が所定の制約時間内に定められているため、演算器 1はフレームに含まれ る最後のマクロブロックの処理を完了したのち、次のフレームに含まれるマクロブロッ クの処理は行わない。演算器 2についても同様に次のフレームに含まれるマクロブロ ックの処理は行わない(図 28)。
[0065] 従って、フレームデータ中の最後のマクロブロックデータに力かる演算処理におい ては、処理終了検知器 20によりすベての演算器 10における演算処理の終了が検知 されてもすべての演算器 10に対するクロック供給の停止を継続し、フレーム処理期 間の経過後、次のフレームデータのパイプライン演算処理に向けてすべての演算器 10へのクロック供給を再開する仕組みとする。
例えば、演算器 10が最終マクロブロックの演算処理が終了しても終了信号を出力 しないという仕組みが考えられる。この場合、処理終了検知器 20が別途、フレーム処 理期間の終了の通知を制御部(図示せず)力 受け、処理開始フラグラインに処理開 始フラグ信号を出力すれば良い。
他には例えば、演算器 10は最終マクロブロックの演算処理が終了すれば同様に終 了信号を出力するが、その際、最終マクロブロックの処理の終了という属性信号を付 して出力する仕組みが考えられる。この場合、処理終了検知器 20は、フレーム処理 期間の終了の通知を制御部(図示せず)から受けるまでは、処理開始フラグラインに 処理開始フラグ信号を出力するのを待ち、フレーム処理期間の終了の通知を受けた 後、処理開始フラグラインに処理開始フラグ信号を出力すれば良い。
[0066] 図 4に示すように、最終マクロブロックの演算処理を終了した後、クロックゲーティン グを継続する期間が設けられる。
上記のように、次のフレーム処理期間の開始まで、まとめてクロックゲーティングが 持続的に行なわれているので、クロックゲーティングの開始回数、停止回数としては 1 回とカウントされる。
[0067] 以上、クロックゲーティング開始回数およびクロックゲーティング停止回数が減少す ることにより、より低消費電力化を図ることができる。
なお、上記の実施例 1のデジタル VLSI回路は、ノ ィプライン演算処理をマクロプロ ック単位(マクロブロックデータは複数個のブロックデータにより構成される。なお 24 個で構成されることが多 ヽ)で実行する構成例であった力 ブロック単位で実行する 構成とすることも可能である。
ブロック単位でパイプライン処理を実行する場合は、演算器がパイプライン演算処 理をブロックデータ単位で実行し、クロック供給制御手段力 フレームデータ中の最 後のマクロブロックデータ中の最後のブロックデータに力かる演算処理においては、 検知手段によりすベての演算器における演算処理の終了が検知されてもすべての 演算器に対するクロック供給の停止を継続し、フレーム処理期間の経過後、次のフレ ームデータのパイプライン演算処理に向けてすべての演算器へのクロック供給を再 開するものとする。
以上、本実施例 1のデジタル VLSI回路によれば、図 4と図 23の比較から明らかな ように、クロックゲーティング開始回数およびクロックゲーティング停止回数が減少して おり、より一層の低消費電力化が図られていることが分力る。
実施例 2
実施例 2は、パイプライン処理の前後の演算器同士がハンドシェイク型の連携をも つてパイプライン処理の演算処理を詰めて行な 、、クロックゲーティング開始回数お よびクロックゲーティング停止回数を少なくし、低消費電力化を図ったデジタル VLSI 回路の例である。なお、本実施例 2では、パイプライン演算処理の単位をマクロブロッ ク単位とする構成として説明する力 ブロック単位とする構成も可能である。
図 6は本発明の実施例 2にかかるデジタル VLSI回路の構成を模式的に示す図で ある。
パイプライン演算処理の並びにおいて前後の演算器 10a〜 10cがシーケンシャル に接続された構成となっており、演算器同士がハンドシェイクすることにより連携する 構成となっている。
演算器 10は図示の便宜上、演算器 1 (10a)、演算器 2 (10b)、演算器 3 (10c)の 3 つのみ示しているが、パイプラインの段数に応じて演算器の数を増減した設計とする ことができることは言うまでもな!/、。
それぞれの演算器 10a〜: LOcは、パイプライン演算処理の各ステージを担い、クロ ックに同期して演算処理を実行するものである。ここでは、動画像の符号化'復号ィ匕 を行なうパイプライン処理を実行するので、例えば、演算器 1 (10a)がエントロピーデ コードステージを担当し、演算器 2 (10b)が逆 Q処理ステージを担当し、演算器 3 (10 c)が逆 T処理ステージを担当するものとする。それ以降のパイプライン処理ステージ は図示を省略している。各演算器はシーケンシャルに接続され、前後の演算器の間 で処理済のデータを前段から次段へ次々と受け渡していく構成となっている。各演算 器で演算処理がなされたデータは各演算器に備えられるバッファに保存され、前後 の演算器はこのバッファを介してデータを受け渡す。例えば演算器 1で演算処理が なされたデータは演算器 1に備えられたバッファに保存され、このノ ッファカも次段の 演算器 2へデータが受け渡される。バッファはフリップフロップや RAM (Random A ccess Memory)などで構成される。 [0069] 図 6に示すように、前段の演算器 10と次段の演算器 10は要求信号と受理信号を交 換し合 、、両者の信号の交換が成立した場合に前段の演算器 10から処理済みデー タが次段の演算器 10に受け渡される仕組みとなっている。
各演算器はパイプライン処理の並びにおいて前後の演算器と受理信号ライン 50と 要求信号ライン 60とデータライン 70の 3本ずつのラインを介して接続されている。
[0070] 本実施例 2のデジタル VLSI回路の演算器およびクロック供給制御手段は、例えば 、以下の 7つのルールに従って動作するように構成されて 、る。
(ルール 1)当段の演算器 10は自らの処理が終了すれば、受理信号ライン 50を介 して次段の演算器 10に対して受理信号を発し、自らが処理した処理済データを受け 渡し準備が完了した旨を伝える。
(ルール 2)次段の演算器 10は自らの処理が終了すれば、要求信号ライン 60を介 して当段の演算器 10に対して要求信号を発し、当段の演算器 10からデータの受け 入れ準備が完了した旨を伝える。
(ルール 3)当段演算器 10と次段演算器 10が相互にデータの受け渡しに関する状 況を交換し合い、両者ともデータの受け渡しの準備が出来ていることを確認できれば データの受け渡しを行なう。
(ルール 4)前段演算器 10は自らの処理が終了すれば、受理信号ライン 50を介して 当段の演算器 10に対して受理信号を発し、自らが処理した処理済データを受け渡し 準備が完了した旨を伝える。
(ルール 5)当段の演算器 10は自らの処理が終了すれば、要求信号ライン 60を介 して前段の演算器 10に対して要求信号を発し、前段の演算器 10からデータの受け 入れ準備が完了した旨を伝える。
(ルール 6)前段演算器 10と当段演算器 10が相互にデータの受け渡しに関する状 況を交換し合い、両者ともデータの受け渡しの準備が出来ていることを確認できれば データの受け渡しを行なう。
(ルール 7)演算器 10は処理済データを次段の演算器 10に出力するまで、次のデ ータの処理を開始しない。
上記 7つのルールに従って、パイプライン処理の前後の演算器同士がハンドシエイ ク型の連携をもってパイプライン処理の演算処理を詰めて行な 、、クロックゲーティン グ開始回数およびクロックゲーティング停止回数を少なくし、低消費電力化を図るも のである。
[0071] クロック供給制御手段は上記ルールを実現する回路構成であれば特に限定されな い。図 7は、本実施例 2にかかる演算器 10およびクロック供給制御手段の構成の一 例を詳しく示したものである。
図 7に示すように、 4つの状態マシン(111〜114)と 4つのフリップフロップ(121〜1 24)と 3つの AND回路(131〜133)により自動的にクロック供給のオンオフが制御さ れる構成となっている。
[0072] 当段演算器 10と次段演算器 10との間の接続関係は以下のようになつている。
受理信号ライン 50bは状態マシン 113からフリップフロップ 123を介して AND回路 132に接続されている。また、要求信号ライン 60bは状態マシン 114からフリップフロ ップ 124を介して AND回路 132に接続されている。
受理信号ライン 50bと要求信号ライン 60bの両者がともにアクティブ (ノヽィ)になるこ とにより、 AND回路 132がオンとなる。つまり、 AND回路 132がオンとなった場合、 当該演算器 10と次段演算器 10との間でハンドシェイクが成立し、両者間でデータの 受け渡しが行なわれ得る状態となっている。この状態で上記のルール 1、ルール 2、 ルール 3が成立する。次段演算器 10が既に処理済みデータを出力済みであればル ール 7も成立している。
ここで、また、 AND回路 132の出力が直接演算器 10に入力されており、当段演算 器 10は次段演算器 10との間でノヽンドシェイクが成立し、両者間でデータの受け渡し が行なわれ得る状態となったことを検知できる仕組みとなっており、演算器 10は次段 演算器 10に処理済みデータを受け渡すことができる構成例となっている。
[0073] 一方、当段演算器 10と前段演算器 10との間の接続関係は以下のようになつている 受理信号ライン 50aは状態マシン 111からフリップフロップ 121を介して AND回路 131に接続されている。また、要求信号ライン 60aは状態マシン 112からフリップフロ ップ 122を介して AND回路 131に接続されている。 受理信号ライン 50aと要求信号ライン 60aの両者がともにアクティブ (ノヽィ)になるこ とにより、 AND回路 131がオンとなる。つまり、 AND回路 131がオンとなった場合、 当該演算器 10と前段演算器 10との間でハンドシェイクが成立し、両者間でデータの 受け渡しが行なわれ得る状態となっている。この状態で上記のルール 4、ルール 5、 ルール 6が成立する。ただし、当段演算器 10においてルール 7が成立しないと実際 のデータの受け渡しは行なわれない。例えば、当段演算器 10の処理済みデータが 既に次段演算器 10に出力済みであればルール 7も満たされるので前段演算器と当 段演算器との間でデータの受け渡しが行なわれる。
[0074] 図 7に示した構成例の動作の例を 2つ示す。
まず、第 1の動作例は、当段演算器 10のパイプライン処理が先に終了し、次に次段 演算器 10のパイプライン処理が終了し、最後に前段演算器 10のパイプライン処理が 終了した場合の流れの場合の動作例である。その流れを図 8から図 10に分けて説明 する。
[0075] 図 8は、当段演算器 10のパイプライン処理が終了した場合の処理の流れを示す図 である。
当段演算器 10は処理が終了すると、当段演算器 10は状態マシン 113に対して終 了信号を発する。状態マシン 113は次段演算器 10への受理信号ラインに対して接 続されており、受理信号レベルをアクティブ (ハイ)とする。また、状態マシン 113はフ リップフロップ 123に信号を出力し、フリップフロップ 123を反転させる(オフ→オン)。 この状態マシン 113はこの状態遷移を維持し、出力状態を保つ。
[0076] 図 9は、図 8の状態力も次段演算器 10のパイプライン処理が終了し、次段演算器 1 0の要求信号を検知した場合の処理の流れを示す図である。
次段演算器 10はパイプライン処理が終了すると、当段演算器 10への要求信号ライ ンをアクティブ (ハイ)とし、当段演算器 10の状態マシン 114に対して要求信号を発 する。要求信号が出されたということは次段演算器 10は既に処理済みデータを次々 段以降の演算器に出力済みであることを意味する。当段演算器 10の状態マシン 11 4はフリップフロップ 124に信号を出力し、フリップフロップ 124を反転させる(オフ→ オン)。この状態マシン 114はこの状態遷移を維持し、出力状態を保つ。 [0077] この図 9の状態において、 AND回路 132は両入力ともアクティブ(ハイ)となってい るのでオンとなる。
AND回路 132の出力は当段演算器 10に入力されており、当段演算器 10は、当段 演算器 10と次段演算器 10の間でノ、ンドシェイクが成立し、当段演算器 10の処理済 みデータを次段に出力する状態となったことを検知するので、当段演算器 10の処理 済データを次段演算器 10に対して出力する。
当段演算器 10は処理済データを次段に出力したので、前段から処理済データを 受け入れられる状態となったので、状態マシン 112を介して前段演算器 10に対して 要求信号ラインを介して要求信号を出力する。
この図 9の後、当段演算器 10はクロック供給が停止され、クロックゲーティング状態 に入る。
[0078] 図 10は、図 9の状態力 前段演算器 10のデータ処理が終了した場合の処理の流 れを示す図である。
前段演算器 10は処理が終了すると、当段演算器 10への受理信号ラインをァクティ ブ (ハイ)とし、当段演算器 10の状態マシン 111に対して受理信号を発する。状態マ シン 111はフリップフロップ 121に信号を出力し、フリップフロップ 121を反転させる( オフ→オン)。この状態マシン 111はこの状態遷移を維持し、出力状態を保つ。 この図 10の状態にお!/、て、 AND回路 131は両入力ともアクティブ(ハイ)となって いるのでオンとなる。前段演算器 10と当段演算器 10との間でノヽンドシェイクが成立し ていることとなる。なお、前段演算器 10と当段演算器 10との間でノヽンドシェイクが成 立している場合は当段演算器 10は既に処理済みデータを次段に出力済みであるの で、前段演算器 10から処理済みデータが当段演算器 10に受け渡される。
さらに、図 10の状態において、 AND回路 133は、 AND回路 131からの入力およ び AND回路 132からの入力ともアクティブ (ノヽィ)となっているのでオンとなる。
ここで、 AND回路 133はクロック入力との間ではゲートとして動作し、クロックゲート がアクティブになったので、クロック供給が開始される。
[0079] 以上が、当段演算器 10のデータ処理が先に終了し、次に次段演算器 10のデータ 処理が終了し、最後に前段演算器 10の処理が終了した場合の流れの場合の動作 例である。このように、当段演算器 10、次段演算器 10、前段演算器 10の順序にてパ ィプライン処理が終了する場合、次段演算器 10のパイプライン処理終了後、前段演 算器 10のパイプライン処理が終了するまでの間、クロックゲーティングが行なわれる こととなる。
[0080] 次に、第 2の動作例は、当段演算器 10のデータ処理が先に終了し、次に前段演算 器 10のデータ処理が終了し、最後に次段演算器 10の処理が終了した場合の流れ の場合の動作例である。その流れを図 11から図 13に分けて説明する。
当段演算器 10のデータ処理が終了した場合の図 11に示す動作は、図 8に示した ものと同じであるので、ここでの説明は省略する。
[0081] 次に、図 12は、図 11の状態力も前段演算器 10のデータ処理が終了した場合の処 理の流れを示す図である。前段演算器 10は処理が終了すると、当段演算器 10への 受理信号ラインをアクティブ (ハイ)とし、当段演算器 10の状態マシン 111に対して受 理信号を発する。状態マシン 111はフリップフロップ 121に信号を出力し、フリップフ ロップ 121を反転させる(オフ→オン)。この状態マシン 111はこの状態遷移を維持し 、出力状態を保つ。この図 12の状態では、 AND回路 131、 AND回路 132とも一方 の信号のみがアクティブ (ノヽィ)であり、他方は非アクティブ(ロー)となっているのでォ フのままであり、前段演算器 10と当段演算器 10との間でもハンドシェイクが成立して おらず、次段演算器 10と当段演算器 10との間でもハンドシェイクが成立していないこ ととなる。
この図 12の後、当段演算器 10はクロック供給が停止され、クロックゲーティング状 態に入る。
[0082] 次に、図 13は、図 12の状態力も次段演算器 10のパイプライン処理が終了し、次段 演算器 10から要求信号が検知された場合の動作を示す図である。
次段演算器 10はパイプライン処理が終了すると、当段演算器 10への要求信号ライ ンをアクティブ (ハイ)とし、当段演算器 10の状態マシン 114に対して要求信号を発 する。要求信号が出されたということは次段演算器 10は既に処理済みデータを次々 段以降の演算器に出力済みであることを意味する。当段演算器 10の状態マシン 11 4はフリップフロップ 124に信号を出力し、フリップフロップ 124を反転させる(オフ→ オン)。この状態マシン 114はこの状態遷移を維持し、出力状態を保つ。
この図 13の状態において、 AND回路 132は両入力ともアクティブ(ハイ)となって いるので才ンとなる。
AND回路 132の出力は当段演算器 10に入力されており、当段演算器 10は、当段 演算器 10と次段演算器 10の間でノ、ンドシェイクが成立し、当段演算器 10の処理済 みデータを次段に出力する状態となったことを検知するので、当段演算器 10の処理 済データを次段演算器 10に対して出力する。
[0083] 当段演算器 10は処理済データを次段に出力したので、前段から処理済データを 受け入れられる状態となったので、状態マシン 112を介して前段演算器 10に対して 要求信号ラインを介して要求信号を出力する。
前段演算器 10は、当段演算器 10からの要求信号ラインがアクティブ (ハイ)になつ たことを受け、当段演算器 10に対して処理済みデータを受け渡す。
この図 13の状態にお!/、て、 AND回路 131は両入力ともアクティブ(ハイ)となって いるので才ンとなる。
さらに、図 13の状態において、 AND回路 133は、 AND回路 131からの入力およ び AND回路 132からの入力ともアクティブ (ノヽィ)となっているのでオンとなる。
ここで、 AND回路 133はクロック入力との間ではゲートとして動作し、クロックゲート がアクティブになったので、クロック供給が開始される。
[0084] 以上が、当段演算器 10のデータ処理が先に終了し、次に前段演算器 10のデータ 処理が終了し、最後に次段演算器 10の処理が終了した場合の流れの場合の動作 例である。このように、当段演算器 10、前段演算器 10、次段演算器 10の順序にてパ ィプライン処理が終了する場合、前段演算器 10のパイプライン処理終了後、次段演 算器 10のパイプライン処理が終了するまでの間、クロックゲーティングが行なわれる こととなる。
[0085] 図 14は実施例 2のデジタル VLSI回路によるパイプライン処理を示すタイミングチヤ ートである。このタイミングチャートでも演算器は、演算器 1 (10a)、演算器 2 (10b)、 演算器 3 (10c)の 3つのみ示して!/、る。
図 14のタイミングチャートにおいて、第 1クロックの時点では演算器 1 (10a)におい てデータ (n+ 2)が処理され、演算器 2 (10b)においてデータ (n+ 1)が処理され、演 算器 3 (10c)にお 、てデータ (n)が処理されて 、る。
[0086] 演算器 1 (10a)は、データ (n+ 2)のパイプライン処理を図中の第 1クロックで完了し ており要求信号を発している。なお、演算器 2 (10b)から受理信号を第 3クロックに受 けている。
演算器 2 (10b)は、データ (n+ 1)のノ ィプライン処理を図中の第 4クロックで完了し ており要求信号を発している。なお、演算器 3 (10c)から受理信号を第 1クロックに受 けている。
演算器 3 (10c)は、データ (n)のパイプライン処理を図中の第 2クロックで完了して おり要求信号を発している。なお、次段演算器から受理信号を第 4クロックに受けて いる。
[0087] 図 14のタイミングチャートでは、演算器 1 (10a)では上記の第 2の動作例(図 11から 図 13に示した動作例)〖こよりクロックゲーティングが第 2クロック力ゝら第 3クロックまで行 なわれて!、る。演算器 2 (10b)では演算器 2 (10b)での処理完了前に演算器 1 (10a )および演算器 3 (10c)が処理を完了しているため、クロックゲーティングされる期間 は存在せず、(n+ 1)番目のデータの処理が完了した次のクロックで (n+ 2)番目の データの処理が開始されている。演算器 3 (10a)では上記の第 1の動作例(図 11か ら図 13に示した動作例)によりクロックゲーティングが第 3クロック力ら第 4クロックまで 行なわれている。
[0088] 図 15は、実施例 2にかかるデジタル VLSI回路におけるパイプライン動作の進行を 模式的に示す図である。図 15に示した各図の要素の説明は図 4に示した各図の要 素の説明と同様でありここでの説明は省略する。
図 15のタイミングチャートに示すように、演算器 10のうち、担当するパイプライン処 理が早く終了したものは、前後の演算器 10におけるパイプライン処理が終了するま での間、クロックゲーティングが行なわれて 、る。
[0089] 例えば、演算器 2 (10b)の 500サイクル力ら 525サイクルの間、演算器 3 (10c)の 6 25サイクルから 650サイクルの間にクロックゲーティングが行なわれる期間がある。こ の例では、前段の演算器の処理が完了した後、当段のクロックゲーティングが解除さ れる上記動作例 1の場合(図 8から図 10に示した動作例)のクロックゲーティングであ る。
例えば、演算器 1 (10a)の 350サイクル力ら 375サイクルの間と 425サイクルと 450 サイクルの間、演算器 2 (10b)の 275サイクルから 300サイクルの間にクロックゲーテ イングが行なわれる期間がある。この例では、次段の演算器の処理が完了した後、当 段の演算器のクロックゲーティングが解除される上記の動作例 2の場合(図 11から図 13に示した動作例)のクロックゲーティングである。
実施例 1では、同じパイプラインステージにおいてすベての演算器の処理が終了す るまでクロックゲーティング期間が設けられた力 実施例 2では、上記のように、同じパ ィプラインステージにおいて当段演算器の前後の演算器の処理が終了するまでクロ ックゲ一ティング期間が設けられているので、実施例 2の方が最終マクロブロックのパ ィプライン処理がより早く終了し、また、クロックゲーティングの開始回数、停止回数が 低減される可能性があることが分かる。
なお、上記の実施例 2のデジタル VLSI回路は、ノ ィプライン演算処理をマクロプロ ック単位(マクロブロックデータは複数個のブロックデータにより構成される。なお 24 個で構成されることが多 ヽ)で実行する構成例であった力 ブロック単位で実行する 構成とすることも可能である。
ブロック単位でパイプライン処理を実行する場合は、演算器がパイプライン演算処 理をブロックデータ単位で実行し、クロック供給制御手段力 フレームデータ中の最 後のマクロブロックデータ中の最後のブロックデータに力かる演算処理においては、 次段演算器へのクロック供給の停止後、前段演算器の演算処理の終了が検知され ても次段演算器へのクロック供給の停止を継続し、フレーム処理期間の経過後、次 のフレームデータのノ ィプライン演算処理に向けてすべての演算器へのクロック供給 を再開する構成とする。
また、クロック供給制御手段力 フレームデータ中の最後のマクロブロックデータの 最後のブロックデータにかかる演算処理においては、前段演算器へのクロック供給の 停止後、前段演算器が次段演算器に対して処理済みの演算処理データを出力でき る状態となった場合でも、前段演算器へのクロック供給の停止を継続し、フレーム処 理期間の経過後、次のフレームデータのパイプライン演算処理に向けてすべての演 算器へのクロック供給を再開する構成とする。 実施例 3
[0091] 実施例 3のデジタル VLSI回路は、演算器の動作電源電圧と基板バイアス電圧と動 作周波数に関してフィードバック制御又はフィードフォワード制御による動的制御を 行なうものである。
実施例 1、実施例 2に示した本発明のパイプラインで構成されるデジタル VLSI回路 では従来のパイプラインで構成されるデジタル VLSI回路と比較して、復号化処理対 象ビットストリームに含まれる有効ブロック数の大小や有効係数の数の大小に依存し て、復号ィ匕処理に必要なサイクル数がフレーム単位に大きく変動する。また、符号ィ匕 処理にお 、ても、動き補償処理で実行されるブロックマッチング回数や発生する有効 ブロック数の大小や有効係数の大小に依存して、符号化処理に必要なサイクル数が フレーム単位に大きく変動する。したがって、実施例 実施例 2に示した本発明の専 用ハードウェア構成のデジタル VLSI回路であれば、フィードバック型の動的制御や フィードフォワード型の動的制御を適用して演算器の動作電源電圧と基板バイアス 電圧を適切な値に抑えて消費電力を削減することができる。また、演算器の動作周 波数を適切な値に抑えることも消費電力削減には効果的である。
[0092] 図 16は、専用ハードウェア構成のデジタル VLSI回路に対してフィードバック型の 動的制御を適用したブロック図である。
処理済マクロブロックカウンタ 80は、パイプライン演算処理のデータ処理量をフレー ムデータごとにカウントする部分である。
フィードバック制御部 81は、処理済マクロブロックカウンタ 80がカウントした処理済 マクロブロックのカウント数に応じて現在処理中のフレームに含まれている未処理の マクロブロックの個数と現在処理中のフレームの処理を完了しなければならない時刻 とから演算器の動作周波数を計算する部分である。
演算器調整部 82は、フィードバック制御部 80の決定した動作周波数に基づき、演 算器の動作電源電圧と基板バイアス電圧と動作周波数とを調整する部分である。
[0093] 図 16に示すように、実施例 1または実施例 2に示したデジタル VLSI回路 100に対 して、処理済マクロブロックカウンタ 80、フィードバック制御部 81、演算器調整部 82 によりフィードバックループを形成することにより、デジタル VLSI回路 100中の演算 器の動作電源電圧と基板バイアス電圧と動作周波数に関してフィードバック制御によ る動的制御を行なうことができる。フィードバックループにおいて、処理済マクロブロッ クカウンタ 80、フィードバック制御部 81、演算器調整部 82の協働によりフィードバック 制御する方法には多様な方法がある。処理済マクロブロックカウンタ 80のカウントによ り経過時間ごとの処理済マクロブロックの数が分かる。例えばパイプラインを実施例 1 や実施例 2で説明したマクロブロックパイプラインで構成した場合、実施例 1や実施例 2で説明したように処理時間の短!ヽマクロブロックがあればマクロブロックパイプライン 処理のサイクル余裕が生まれる。演算器調整部 81はこのサイクル余裕時間を利用し て演算器の動作電源電圧と基板電圧と動作周波数を下げるように調整する。例えば 、ノ ィプラインの各段のマクロブロックの最悪サイクル数を nとし、実際にパイプライン 処理に必要なサイクル数がすべてのステージで mサイクルであったとすると(n— m) サイクルの余裕が生まれる。したがって、次のパイプライン処理においては n+ (n- m) = 2n— mサイクル分の処理時間がある。次のパイプライン処理に必要なサイクル が最悪サイクル数であってもたかだ力 nサイクルであるので、次のパイプライン処理は 動作周波数を下げても制約時間内に最悪サイクル数を確保することができる。そこで 、フィードバック制御部 81は演算器調整部 82に対し、演算器の動作電源電圧と基板 電圧と動作周波数を nZ (2n—m)となるように調整させることができる。
図 17は、専用ハードウェア構成のデジタル VLSI回路に対してフィードフォワード型 の動的制御を適用したブロック図である。
処理負荷予測部 90は、パイプライン演算処理に供される前に、パイプライン演算処 理に供されるフレームデータに含まれるマクロブロックデータ量を検知し、パイプライ ン演算処理に力かる処理負荷を予測する部分である。
フィードフォワード制御部 91は、処理負荷予測部 90による予測に基づき、演算器 の動作電源電圧と基板バイアス電圧と動作周波数とを決定する部分である。
演算器調整部 92は、フィードフォワード制御部 91の決定に基づき、演算器の動作 電源電圧と基板バイアス電圧と動作周波数とを調整する部分である。 図 17に示すように、実施例 1または実施例 2に示したデジタル VLSI回路 100に対 して、処理負荷予測部 90、フィードフォワード制御部 91、演算器調整部 92によりフィ ードフォワードループを形成することにより、デジタル VLSI回路 100中の演算器の動 作電源電圧と基板バイアス電圧と動作周波数に関してフィードフォワード制御による 動的制御を行なうことができる。処理負荷予測部 90、フィードフォワード制御部 91、 演算器調整部 92の協働によりフィードフォワード制御する方法には多様な方法があ る。処理負荷予測部 90は過去における制約時間内に処理すべきデータの処理に必 要であった処理サイクル数を記憶しておく。例えば、 MPEGx、 H. 26xによる動画像 処理では制約時間は 1フレームの時間であり、制約時間内に処理すべきデータとは 1 フレームに含まれるすべてのマクロブロックとなる。 MPEGxや H. 26xによる動画像 処理では、処理対象のフレームのフレームタイプとして Iフレーム、 Pフレーム、 Bフレ ームのタイプがあるので、処理負荷予測部 90は各フレームタイプごとに処理サイクル 数を記憶しておく。処理負荷予測部 90はこれから処理に供されるフレームのフレー ムタイプを調べ、フレームタイプに応じた過去の処理サイクル数を処理負荷サイクル として予測し、フィードフォワード部 91に対して予測サイクル数を表わす信号を出力 する。フィードフォワード部 91は処理負荷予測部 90の予測に基づき、演算器調整部 92に対して動作電源電圧と基板バイアス電圧と動作周波数を低下させるように制御 する。最悪サイクル数を nとし、予測サイクル数が mとすると、 mZnに低下するように 制御することができる。なお、 mZnに低下させると予測が外れ、実際の処理サイクル 数が mより大きい場合に制約時間内に処理が完了しないペナルティが発生してしまう ので、このリスクを低減させるためにも nよりも大きく見積もる工夫を施すことも可能で ある。例えば、処理負荷予測部 90の予測サイクル数が mのとき、フィードフォワード部 91が(1. l) mに調整したり、 (1. 2) mに調整した上で演算器調整部 92に(1. l) m Znや(1. 2) mZnに調整させる。
なお、過去の処理サイクル数に基づく予測サイクル数 mの予測の方法には幾通りか の方法がある。第 1には、同タイプのマクロブロックのうち時間的にもっとも近い過去の ものを用いる方法がある。動画では時間的に近いほどマクロブロック同士の処理サイ クル数が同程度になることが期待できる。第 2には、同タイプのマクロブロックのうち時 間的に近い数個のマクロブロックの処理サイクル数の平均値を用いる方法がある。 実施例 4
[0096] 実施例 4は、実施例 実施例 2、実施例 3に示したデジタル VLSI回路の構成のク ロック供給制御手段に替えて電力供給制御手段とした構成である。低消費電力化は クロックゲーティングの開始回数、停止回数を低減することにより図ることが可能であ る力 クロックゲーティングではなく電力供給自体のオンオフを制御する構成とし、そ の電力供給の開始回数、停止回数を低減することにより図ることによつても同様の効 果が得られる。
例えば、実施例 2、 3の説明においてクロック供給に関する部分を電力供給に関 する部分に替え、また、タイミングチャートの説明においてクロックゲーティング期間を 電力停止期間に替え、対応する図面も書き替えて読めば良い(例えば、クロック入力 14を電力供給ライン 14と読み替えれば良 、)。
実施例 5
[0097] 実施例 5は、上記実施例 1、 2、 3、 4に示した本発明のデジタル VLSI回路を組み 込んだ応用例である。
図 18は本発明のデジタル VLSI回路を組み込んだ画像処理システム 200の構成 例を示す図である。例えば、本発明のデジタル VLSI回路をマイクロプロセッサに採 用したパーソナルコンピュータとして構成したものでも良ぐまた、本発明のデジタル VLSI回路を画像処理チップとして画像処理ボード中に組み込んでも良 ヽ。
図 19は本発明のデジタル VLSI回路を組み込んだ携帯端末 300の構成例を示す 図である。この構成例では携帯電話に組み込んだ例である。携帯電話でも近年は動 画像を扱う能力を備えるものが投入されつつある一方、低消費電力化への要求は極 めて強ぐ本発明のデジタル VLSI回路を組み込んだ携帯電話とすれば、低消費電 力化を図るとともに動画像の処理速度の向上の両面を図ることが可能となる。
[0098] 以上、本発明のデジタル VLSI回路によれば、すべての演算器で処理が終了して アイドル状態に入っている期間を省略することができ、さらに、演算器でのパイプライ ン演算処理を詰めて行なうことによりクロックゲーティングまたは電力供給の開始回数 、終了回数を減少させることができ、より一層の低消費電力化を図ることができる。 また、本発明のデジタル VLSI回路によれば、パイプライン処理の前後の演算器の 処理が終了すれば演算器間で処理済データをやり取りすることによりデータ処理サイ クル数 (動画像処理の場合はマクロブロック処理やブロック処理のサイクル数)を小さ くすることができ、さらに、演算器でのパイプライン演算処理を詰めて行なうことにより クロックゲーティングまたは電力供給の開始回数、停止回数を減少させることができ、 より一層の低消費電力化を図ることができる。さらに、実際のデータに対する処理に 必要となるサイクル数に応じて制約時間内に必要となるサイクル数を大幅に削減する ことができ、削減されたサイクルによって生ずる時間余裕を用いて動作周波数 '電圧 の動的制御を行うことにより、よりいつそうの低消費電力化を図ることができる。
産業上の利用可能性
[0099] 本発明の画像処理システムによれば、動画像のデータ処理において低消費電力 化が図られており、低消費電力が求められる様々なシステムに対して組み込むことが 容易となり、柔軟なシステム設計が可能となる。
本発明の携帯端末によれば、動画像のデータ処理において低消費電力化が図ら れており、携帯電話のような小型の端末においても動画像の符号化復号化処理を行 なうことができ、携帯端末の用途が様々に広がる。
図面の簡単な説明
[0100] [図 1]本発明の実施例 1にかかるデジタル VLSI回路の構成を模式的に示す図
[図 2]クロック供給制御手段の構成例を示す図
[図 3]実施例 1のデジタル VLSI回路によるノ ィプライン処理を示すタイミングチャート [図 4]実施例 1にかかるデジタル VLSI回路におけるパイプライン動作の進行を模式 的に示す図
[図 5]動作周波数の制御例
[図 6]本発明の実施例 2にかかるデジタル VLSI回路の構成を模式的に示す図
[図 7]本実施例 2にかかる演算器 10およびクロック供給制御手段の構成の一例を詳 しく示す図
[図 8]当段演算器 10のパイプライン処理が終了した場合の処理の流れを示す図 [図 9]図 8の状態力 次段演算器 10のパイプライン処理が終了し、次段演算器 10の 要求信号を検知した場合の処理の流れを示す図
圆 10]図 9の状態力も前段演算器 10のデータ処理が終了した場合の処理の流れを 示す図
圆 11]当段演算器 10のパイプライン処理が終了した場合の処理の流れを示す図 圆 12]図 11の状態力も前段演算器 10のデータ処理が終了した場合の処理の流れを 示す図
圆 13]図 12の状態力も次段演算器 10のパイプライン処理が終了し、次段演算器 10 から要求信号が検知された場合の動作を示す図
[図 14]実施例 2のデジタル VLSI回路によるノ ィプライン処理を示すタイミングチヤ一 卜
[図 15]実施例 2にかかるデジタル VLSI回路におけるパイプライン動作の進行を模式 的に示す図
[図 16]専用ハードウェア構成のデジタル VLSI回路に対してフィードバック型の動的 制御を適用したブロック図
[図 17]専用ハードウェア構成のデジタル VLSI回路に対してフィードフォワード型の動 的制御を適用したブロック図
[図 18]本発明のデジタル VLSI回路を組み込んだ画像処理システムの構成例を示す 図
圆 19]本発明のデジタル VLSI回路を組み込んだ携帯端末の構成例を示す図
[図 20]H. 264デコード処理モジュールのブロック図
[図 21]図 20に示したブロック図に基づき、複数の演算器を用いた専用ハードウェアに より H. 264デコード処理モジュールを構築した例を示す図
圆 22]パイプライン分割を示す図
圆 23]従来のパイプライン処理における演算器の動作状況を模式的に示すタイムチ ヤート
[図 24]動画像処理ソフトウェアのフローチャート
[図 25]動画像処理ソフトウェアで、処理に必要なサイクルの実際例を示す図
[図 26]動画像処理ソフトウェアで、処理に必要な最悪サイクル数を示す図 [図 27]動作周波数の制御例を示すグラフ
[図 28]実施例 2にかかるデジタル VLSI回路におけるパイプライン動作の進行
[図 29]動作周波数の制御例を示すグラフ
符号の説明
演算器
11, 111, 112, 113, 114 状態マシン
12, 121, 122, 123, 124 フリップフロ
13, 131, 132, 133 AND回路
14 クロック人力ライン
20 処理終了検知器
30 終了フラグライン
40 処理開始フラグライン
50 受理信号ライン
60 要求信号ライン
70 データライン
80 処理済マクロブロックカウンタ
81 フィードバック制御部
82 演算器調整部
90 処理負荷予測部
91 フィードフォワード制御部
92 演算器調整部
100 デジタル VLSI回路
200 画像処理システム
300 携帯端末

Claims

請求の範囲
[1] ノ ィプライン演算処理の各ステージを担い、クロックに同期して演算処理を実行す る複数の演算器と、
前記演算器における担当ステージの演算処理の終了を検知する検知手段と、 前記演算器ごとにクロックの供給 Z停止を制御するクロック供給制御手段とを備え、 前記クロック供給制御手段が、
前記検知手段により演算処理の終了が検知された前記演算器に対するクロック供 給を停止し、
前記検知手段によりすベての前記演算器における演算処理の終了が検知されれ ば次のパイプライン演算処理に向けてすべての前記演算器へのクロック供給を再開 するように構成されたことを特徴としたデジタル VLSI回路。
[2] 前記演算処理に力かるデータ力 複数のマクロブロックデータを含み、一定の処理 期間(フレーム処理期間)が定められているフレームデータを備えた動画像データで あり、前記演算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記マクロブロックデータ単位で実行 し、
前記クロック供給制御手段力 前記フレームデータ中の最後の前記マクロブロック データにかかる演算処理においては、前記検知手段によりすベての前記演算器に おける演算処理の終了が検知されてもすべての前記演算器に対するクロック供給の 停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器へのクロック供給を再開することを特徴とする請求項 1に 記載のデジタル VLSI回路。
[3] 前記演算処理に力かるデータ力 複数のマクロブロックデータ(前記マクロブロック データは複数個のブロックデータにより構成される)を含み、一定の処理期間(フレー ム処理期間)が定められているフレームデータを備えた動画像データであり、前記演 算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記ブロックデータ単位で実行し、 前記クロック供給制御手段力 前記フレームデータ中の最後の前記マクロブロック データ中の最後の前記ブロックデータにかかる演算処理においては、前記検知手段 によりすベての前記演算器における演算処理の終了が検知されてもすべての前記 演算器に対するクロック供給の停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器へのクロック供給を再開することを特徴とする請求項 1に 記載のデジタル VLSI回路。
[4] ノ ィプライン演算処理の各ステージを担い、クロックに同期して演算処理を実行す る複数の演算器と、
前記演算器における担当ステージの演算処理の終了を検知する検知手段と、 前記演算器ごとにクロックの供給 Z停止を制御するクロック供給制御手段とを備え、 前記クロック供給制御手段が、
前記検知手段により、前記パイプライン演算処理において相前後する前段演算器 と次段演算器のうち次段演算器の演算処理の終了が先に検知された場合、当該次 段演算器に対してクロック供給を停止し、
前記次段演算器へのクロック供給の停止後、前記前段演算器の演算処理の終了 が検知された場合、次のパイプライン演算処理に向けて前記次段演算器に対してク ロック供給を再開するように構成されたことを特徴とするデジタル VLSI回路。
[5] 前記クロック供給制御手段が、
前記検知手段により、前記前段演算器の演算処理の終了が先に検知された場合、 前記前段演算器が前記次段演算器に対して処理済みの演算処理データを出力でき るまで、当該前段演算器に対してクロック供給を停止し、
前記前段演算器へのクロック供給の停止後、前記前段演算器が前記次段演算器 に対して処理済みの演算処理データを出力できる状態となれば、当該前段演算器に 対してクロック供給を再開するように構成されたことを特徴とする請求項 4に記載のデ ジタル VLSI回路。
[6] 前記演算処理に力かるデータ力 複数のマクロブロックデータを含み、一定の処理 期間(フレーム処理期間)が定められているフレームデータを備えた動画像データで あり、前記演算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記マクロブロックデータ単位で実行 し、
前記クロック供給制御手段が、
前記フレームデータ中の最後のマクロブロックデータにかかる演算処理においては 、前記次段演算器へのクロック供給の停止後、前記前段演算器の演算処理の終了 が検知されても、前記次段演算器へのクロック供給の停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器へのクロック供給を再開することを特徴とする請求項 4ま たは 5に記載のデジタル VLSI回路。
[7] 前記クロック供給制御手段が、
前記フレームデータ中の最後のマクロブロックデータにかかる演算処理においては 、前記前段演算器へのクロック供給の停止後、前記前段演算器が前記次段演算器 に対して処理済みの演算処理データを出力できる状態となった場合でも、前記前段 演算器へのクロック供給の停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器へのクロック供給を再開することを特徴とする請求項 6に 記載のデジタル VLSI回路。
[8] 前記演算処理に力かるデータ力 複数のマクロブロックデータ(前記マクロブロック データは複数個のブロックデータにより構成される)を含み、一定の処理期間(フレー ム処理期間)が定められているフレームデータを備えた動画像データであり、前記演 算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記ブロックデータ単位で実行し、 前記クロック供給制御手段が、
前記フレームデータ中の最後の前記マクロブロックデータ中の最後の前記ブロック データにかかる演算処理においては、前記次段演算器へのクロック供給の停止後、 前記前段演算器の演算処理の終了が検知されても前記次段演算器へのクロック供 給の停止を継続し、 前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器へのクロック供給を再開することを特徴とする請求項 4ま たは 5に記載のデジタル VLSI回路。
[9] 過去における制約時間内に演算処理すべきデータ量と演算処理に要したサイクル 数を記憶するサイクル数記憶手段と、
次の制約時間内に演算処理すべきデータ量と前記記憶手段に記憶されているサイ クル数から、次の制約時刻までに演算処理しなければならな 、データの演算処理に 必要なサイクル数を予測するサイクル数予測手段と、
前記サイクル数予測手段で計算されたサイクル数カゝら専用デジタル VLSI回路の 動作周波数、動作電源電圧、基板バイアス電圧を決定し決定した動作周波数のクロ ックとこのクロックに対応した動作電源電圧、基板バイアス電圧を供給する動作周波 数 ·電圧制御手段を備え、
前記サイクル数記憶手段が制約時間内に演算処理すべきデータ量と当該データ に対して演算処理を完了するために要したサイクル数を記憶し、次の制約時間の開 始直後において前記動作周波数 ·電圧制御手段がサイクル数記憶手段に記憶され ている過去のデータの演算処理に要したサイクル数と次の制約時刻までに演算処理 を完了しなければならないデータ量と力 次の制約時刻までの専用デジタル VLSI 回路の動作周波数、動作電源電圧、基板バイアス電圧を決定し専用デジタル VLSI 回路の動作周波数、動作電源電圧、基板バイアス電圧を決定した値に制御する請 求項 2、 3、 6、 7、 8のいずれ力 1項に記載のデジタル VLSI回路。
[10] 次の制約時間内に演算処理すべきデータ量とパイプラインを少なくとも 1回動作さ せることによって演算処理が完了するデータ量とから次のパイプライン動作時の専用 デジタル VLSI回路の動作周波数、動作電源電圧、基板バイアス電圧を決定し決定 した動作周波数のクロックとこのクロックに対応した動作電源電圧、基板バイアス電圧 を供給する動作周波数 ·電圧制御手段を備え、
前記動作周波数'電圧制御手段が次のパイプライン演算処理において演算処理を 行うデータの演算処理に必要なサイクルが想定される最悪サイクル数であっても当該 データの演算処理が完了すべき時刻に間に合うように動作周波数を決定し、次のパ ィプライン演算処理における専用デジタル VLSI回路のクロック、動作電源電圧、基 板バイアス電圧を決定した動作周波数とこの動作周波数に対応した動作電源電圧、 基板バイアス電圧に制御する請求項 2、 3、 6、 7、 8のいずれか 1項に記載のデジタ ノレ VLSI回路。
[11] ノ ィプライン演算処理の各ステージを担い、演算処理を実行する複数の演算器と、 前記演算器における担当ステージの演算処理の終了を検知する検知手段と、 前記演算器ごとに電力の供給 Z停止を制御する電力供給制御手段とを備え、 前記電力供給制御手段が、
前記検知手段により演算処理終了が検知された前記演算器に対する電力供給を 停止し、
前記検知手段によりすベての前記演算器における演算処理終了が検知されれば 次のパイプライン演算処理に向けてすべての前記演算器への電力供給を再開する ように構成されたことを特徴としたデジタル VLSI回路。
[12] 前記演算処理に力かるデータ力 複数のマクロブロックデータを含み、一定の処理 期間(フレーム処理期間)が定められているフレームデータを備えた動画像データで あり、前記演算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記マクロブロックデータ単位で実行 し、
前記電力供給制御部が、前記フレームデータ中の最後の前記マクロブロックデータ にかかる演算処理においては、前記検知手段によりすベての前記演算器における演 算処理の終了が検知されてもすべての前記演算器に対する電力供給の停止を継続 し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器への電力供給を再開することを特徴とする請求項 11に 記載のデジタル VLSI回路。
[13] 前記演算処理に力かるデータ力 複数のマクロブロックデータ(前記マクロブロック データは複数個のブロックデータにより構成される)を含み、一定の処理期間(フレー ム処理期間)が定められているフレームデータを備えた動画像データであり、前記演 算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記ブロックデータ単位で実行し、 前記電力供給制御部が、前記フレームデータ中の最後の前記マクロブロックデータ 中の最後の前記ブロックデータにかかる演算処理においては、前記検知手段により すべての前記演算器における演算処理の終了が検知されてもすべての前記演算器 に対する電力供給の停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器への電力供給を再開することを特徴とする請求項 11に 記載のデジタル VLSI回路。
[14] ノ ィプライン演算処理の各ステージを担 、、演算処理を実行する複数の演算器と、 前記演算器における担当ステージの演算処理の終了を検知する検知手段と、 前記演算器ごとに電力の供給 Z停止を制御する電力供給制御手段とを備え、 前記電力供給制御手段が、
前記検知手段により、前記パイプライン演算処理において相前後する前段演算器 と次段演算器のうち次段演算器の演算処理の終了が先に検知された場合、当該次 段演算器に対して電力供給を停止し、
前記次段演算器への電力供給の停止後、前記前段演算器の演算処理の終了が 検知された場合、次のパイプライン演算処理に向けて前記次段演算器に対して電力 供給を再開するように構成されたことを特徴とするデジタル VLSI回路。
[15] 前記電力供給制御手段が、
前記検知手段により、前記前段演算器の演算処理の終了が先に検知された場合、 前記前段演算器が前記次段演算器に対して処理済みの演算処理データを出力でき るまで、当該前段演算器に対して電力供給を停止し、
前記前段演算器への電力供給の停止後、前記前段演算器が前記次段演算器に 対して処理済みの演算処理データを出力できる状態となれば、当該前段演算器に 対して電力供給を再開するように構成されたことを特徴とする請求項 14に記載のデ ジタル VLSI回路。
[16] 前記演算処理に力かるデータ力 複数のマクロブロックデータを含み、一定の処理 期間(フレーム処理期間)が定められているフレームデータを備えた動画像データで あり、前記演算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記マクロブロックデータ単位で実行 し、
前記電力供給制御手段が、
前記フレームデータ中の最後のマクロブロックデータにかかる演算処理においては 、前記次段演算器への電力供給の停止後、前記前段演算器の演算処理の終了が 検知されても、前記次段演算器への電力供給の停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器への電力供給を再開することを特徴とする請求項 14ま たは 15に記載のデジタル VLSI回路。
[17] 前記電力供給制御手段が、
前記フレームデータ中の最後のマクロブロックデータにかかる演算処理においては 、前記前段演算器への電力供給の停止後、前記前段演算器が前記次段演算器に 対して処理済みの演算処理データを出力できる状態となった場合でも、前記前段演 算器への電力供給の停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器への電力供給を再開することを特徴とする請求項 16に 記載のデジタル VLSI回路。
[18] 前記演算処理に力かるデータ力 複数のマクロブロックデータ(前記マクロブロック データは複数個のブロックデータにより構成される)を含み、一定の処理期間(フレー ム処理期間)が定められているフレームデータを備えた動画像データであり、前記演 算処理が前記動画像の符号化 ·復号化処理であり、
前記演算器が、前記パイプライン演算処理を前記ブロックデータ単位で実行し、 前記電力供給制御手段が、
前記フレームデータ中の最後の前記マクロブロックデータ中の最後の前記ブロック データにかかる演算処理においては、前記次段演算器への電力供給の停止後、前 記前段演算器の演算処理の終了が検知されても、前記次段演算器への電力供給の 停止を継続し、
前記フレーム処理期間の経過後、次のフレームデータのパイプライン演算処理に 向けてすべての前記演算器への電力供給を再開することを特徴とする請求項 14ま たは 15に記載のデジタル VLSI回路。
[19] 過去における制約時間内に演算処理すべきデータ量と演算処理に要したサイクル 数を記憶するサイクル数記憶手段と、
次の制約時間内に演算処理すべきデータ量と前記記憶手段に記憶されているサイ クル数から、次の制約時刻までに演算処理しなければならな 、データの演算処理に 必要なサイクル数を予測するサイクル数予測手段と、
前記サイクル数予測手段で計算されたサイクル数カゝら専用デジタル VLSI回路の 動作周波数、動作電源電圧、基板バイアス電圧を決定し決定した動作周波数のクロ ックとこのクロックに対応した動作電源電圧、基板バイアス電圧を供給する動作周波 数 ·電圧制御手段を備え、
前記サイクル数記憶手段が制約時間内に演算処理すべきデータ量と当該データ に対して演算処理を完了するために要したサイクル数を記憶し、次の制約時間の開 始直後において前記動作周波数 ·電圧制御手段がサイクル数記憶手段に記憶され ている過去のデータの演算処理に要したサイクル数と次の制約時刻までに演算処理 を完了しなければならないデータ量と力 次の制約時刻までの専用デジタル VLSI 回路の動作周波数、動作電源電圧、基板バイアス電圧を決定し専用デジタル VLSI 回路の動作周波数、動作電源電圧、基板バイアス電圧を決定した値に制御する請 求項 12、 13、 16、 17、 18の!ヽずれ力 1項に記載のデジタノレ VLSI回路。
[20] 次の制約時間内に演算処理すべきデータ量とパイプラインを少なくとも 1回動作さ せることによって演算処理が完了するデータ量とから次のパイプライン動作時の専用 デジタル VLSI回路の動作周波数、動作電源電圧、基板バイアス電圧を決定し決定 した動作周波数のクロックとこのクロックに対応した動作電源電圧、基板バイアス電圧 を供給する動作周波数 ·電圧制御手段を備え、
前記動作周波数'電圧制御手段が次のパイプライン演算処理において演算処理を 行うデータの演算処理に必要なサイクルが想定される最悪サイクル数であっても当該 データの演算処理が完了すべき時刻に間に合うように動作周波数を決定し、次のパ ィプライン演算処理における専用デジタル VLSI回路のクロック、動作電源電圧、基 板バイアス電圧を決定した動作周波数とこの動作周波数に対応した動作電源電圧、 基板バイアス電圧に制御する請求項 13、 14、 17、 18のいずれ力 1項に記載のデジ タノレ VLSI回路。
PCT/JP2007/051927 2006-02-03 2007-02-05 デジタルvlsi回路およびそれを組み込んだ画像処理システム Ceased WO2007089014A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2007556947A JP4521508B2 (ja) 2006-02-03 2007-02-05 デジタルvlsi回路およびそれを組み込んだ画像処理システム
US12/278,015 US8291256B2 (en) 2006-02-03 2007-02-05 Clock stop and restart control to pipelined arithmetic processing units processing plurality of macroblock data in image frame per frame processing period

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2006027431 2006-02-03
JP2006-027431 2006-02-03

Publications (1)

Publication Number Publication Date
WO2007089014A1 true WO2007089014A1 (ja) 2007-08-09

Family

ID=38327571

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2007/051927 Ceased WO2007089014A1 (ja) 2006-02-03 2007-02-05 デジタルvlsi回路およびそれを組み込んだ画像処理システム

Country Status (3)

Country Link
US (1) US8291256B2 (ja)
JP (1) JP4521508B2 (ja)
WO (1) WO2007089014A1 (ja)

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008192040A (ja) * 2007-02-07 2008-08-21 Nec Corp 半導体集積回路及び動作条件制御方法
JP2010020598A (ja) * 2008-07-11 2010-01-28 Univ Of Tsukuba ネットワークシステムおよびネットワークシステムにおける電源制御方法
JP2010073854A (ja) * 2008-09-18 2010-04-02 Nuflare Technology Inc 荷電粒子ビーム描画装置および荷電粒子ビーム描画方法
JP2011503683A (ja) * 2007-10-11 2011-01-27 クゥアルコム・インコーポレイテッド グラフィック処理ユニットにおける需要ベースの電力制御
JP2011187045A (ja) * 2010-02-09 2011-09-22 Canon Inc データ処理装置及びその制御方法、プログラム
JP2012128738A (ja) * 2010-12-16 2012-07-05 Canon Inc データ処理装置、データ処理方法及びプログラム
JP2015508528A (ja) * 2011-12-28 2015-03-19 インテル・コーポレーション パイプライン化された画像処理シーケンサ
JPWO2017119123A1 (ja) * 2016-01-08 2018-03-22 三菱電機株式会社 プロセッサ合成装置、プロセッサ合成方法及びプロセッサ合成プログラム
JP2019521454A (ja) * 2016-07-21 2019-07-25 アドバンスト・マイクロ・ディバイシズ・インコーポレイテッドAdvanced Micro Devices Incorporated 非同期パイプラインのステージの動作速度の制御

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009075973A (ja) * 2007-09-21 2009-04-09 Canon Inc 電子機器及び当該電子機器の電力制御方法
JP5449791B2 (ja) * 2009-02-02 2014-03-19 オリンパス株式会社 データ処理装置および画像処理装置
US20130004071A1 (en) * 2011-07-01 2013-01-03 Chang Yuh-Lin E Image signal processor architecture optimized for low-power, processing flexibility, and user experience
CN113918222A (zh) * 2020-07-08 2022-01-11 上海寒武纪信息科技有限公司 流水线控制方法、运算模块及相关产品

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH04143819A (ja) * 1989-12-15 1992-05-18 Hitachi Ltd 消費電力制御方法、半導体集積回路装置およびマイクロプロセツサ
WO2001033351A1 (fr) * 1999-10-29 2001-05-10 Fujitsu Limited Architecture de processeur
JP2002268877A (ja) * 2001-03-08 2002-09-20 Matsushita Electric Ind Co Ltd クロック制御方法及び当該クロック制御方法を用いた情報処理装置
JP2003263311A (ja) * 2002-03-07 2003-09-19 Seiko Epson Corp 並列演算処理装置及び並列演算処理用の命令コードのデータ構造、並びに並列演算処理用の命令コードの生成方法
WO2004027528A2 (en) * 2002-09-20 2004-04-01 Koninklijke Philips Electronics N.V. Adaptive data processing scheme based on delay forecast

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1157641C (zh) * 1997-09-03 2004-07-14 松下电器产业株式会社 处理器
TWI230855B (en) * 2002-01-05 2005-04-11 Via Tech Inc Transmission line circuit structure saving power consumption and operating method thereof
KR100719360B1 (ko) * 2005-11-03 2007-05-17 삼성전자주식회사 디지털 로직 프로세싱 회로, 그것을 포함하는 데이터 처리 장치, 그것을 포함한 시스템-온 칩, 그것을 포함한 시스템, 그리고 클록 신호 게이팅 방법
US7773236B2 (en) * 2006-06-01 2010-08-10 Toshiba Tec Kabushiki Kaisha Image forming processing circuit and image forming apparatus

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH04143819A (ja) * 1989-12-15 1992-05-18 Hitachi Ltd 消費電力制御方法、半導体集積回路装置およびマイクロプロセツサ
WO2001033351A1 (fr) * 1999-10-29 2001-05-10 Fujitsu Limited Architecture de processeur
JP2002268877A (ja) * 2001-03-08 2002-09-20 Matsushita Electric Ind Co Ltd クロック制御方法及び当該クロック制御方法を用いた情報処理装置
JP2003263311A (ja) * 2002-03-07 2003-09-19 Seiko Epson Corp 並列演算処理装置及び並列演算処理用の命令コードのデータ構造、並びに並列演算処理用の命令コードの生成方法
WO2004027528A2 (en) * 2002-09-20 2004-04-01 Koninklijke Philips Electronics N.V. Adaptive data processing scheme based on delay forecast

Cited By (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008192040A (ja) * 2007-02-07 2008-08-21 Nec Corp 半導体集積回路及び動作条件制御方法
JP2011503683A (ja) * 2007-10-11 2011-01-27 クゥアルコム・インコーポレイテッド グラフィック処理ユニットにおける需要ベースの電力制御
US8458497B2 (en) 2007-10-11 2013-06-04 Qualcomm Incorporated Demand based power control in a graphics processing unit
JP2010020598A (ja) * 2008-07-11 2010-01-28 Univ Of Tsukuba ネットワークシステムおよびネットワークシステムにおける電源制御方法
JP2010073854A (ja) * 2008-09-18 2010-04-02 Nuflare Technology Inc 荷電粒子ビーム描画装置および荷電粒子ビーム描画方法
JP2011187045A (ja) * 2010-02-09 2011-09-22 Canon Inc データ処理装置及びその制御方法、プログラム
JP2012128738A (ja) * 2010-12-16 2012-07-05 Canon Inc データ処理装置、データ処理方法及びプログラム
JP2015508528A (ja) * 2011-12-28 2015-03-19 インテル・コーポレーション パイプライン化された画像処理シーケンサ
JPWO2017119123A1 (ja) * 2016-01-08 2018-03-22 三菱電機株式会社 プロセッサ合成装置、プロセッサ合成方法及びプロセッサ合成プログラム
JP2019521454A (ja) * 2016-07-21 2019-07-25 アドバンスト・マイクロ・ディバイシズ・インコーポレイテッドAdvanced Micro Devices Incorporated 非同期パイプラインのステージの動作速度の制御

Also Published As

Publication number Publication date
US20090024866A1 (en) 2009-01-22
US8291256B2 (en) 2012-10-16
JPWO2007089014A1 (ja) 2009-06-25
JP4521508B2 (ja) 2010-08-11

Similar Documents

Publication Publication Date Title
WO2007089014A1 (ja) デジタルvlsi回路およびそれを組み込んだ画像処理システム
US8106804B2 (en) Video decoder with reduced power consumption and method thereof
US9582060B2 (en) Battery-powered device with reduced power consumption based on an application profile data
US20220377322A1 (en) Intra/inter mode decision for predictive frame encoding
US8213511B2 (en) Video encoder software architecture for VLIW cores incorporating inter prediction and intra prediction
US9317103B2 (en) Method and system for selective power control for a multi-media processor
EP2163097A2 (en) Adaptive video encoding apparatus and methods
TW201630423A (zh) 具有上下文切換之視訊編碼器
KR100564010B1 (ko) 화상 처리 장치
JPWO2004093458A1 (ja) 動画像符号化又は復号化処理システム及び動画像符号化又は復号化処理方法
Hameed et al. Understanding sources of ineffciency in general-purpose chips
JP2007166192A (ja) 情報処理装置、制御方法およびプログラム
US8238429B2 (en) Statistically cycle optimized bounding box for high definition video decoding
Senn et al. Joint DVFS and parallelism for energy efficient and low latency software video decoding
EP2490102A2 (en) Video decodere and/or battery-powered device with reduced power consumption and methods thereof
JPH1155668A (ja) 画像符号化装置
Nguyen et al. Implementation of H. 264/AVC encoder on coarse-grained dynamically reconfigurable computing system
JP2006155223A (ja) データ処理装置
Kawakami et al. Power and memory bandwidth reduction of an H. 264/AVC HDTV decoder LSI with elastic pipeline architecture
Kawakami et al. A 50% power reduction in H. 264/AVC HDTV video decoder LSI by dynamic voltage scaling in elastic pipeline
Nguyen et al. An Efficient Implementation of H. 264/AVC Integer Motion Estimation Algorithm on Coarse-grained Reconfigurable Computing System.
Wei et al. Realization and optimization of DSP based H. 264 encoder
Rapaka et al. Scalable software architecture for high performance video codec's on parallel processing engines
De-Gregorio et al. Bringing streaming video to wireless handheld devices
Nakata et al. Development of full-HD multi-standard video CODEC IP based on heterogeneous multiprocessor architecture

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 12278015

Country of ref document: US

ENP Entry into the national phase

Ref document number: 2007556947

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 07708045

Country of ref document: EP

Kind code of ref document: A1