WO2012008616A1 - Video decoder for low resolution power reduction using low resolution data - Google Patents
Video decoder for low resolution power reduction using low resolution data Download PDFInfo
- Publication number
- WO2012008616A1 WO2012008616A1 PCT/JP2011/066638 JP2011066638W WO2012008616A1 WO 2012008616 A1 WO2012008616 A1 WO 2012008616A1 JP 2011066638 W JP2011066638 W JP 2011066638W WO 2012008616 A1 WO2012008616 A1 WO 2012008616A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- resolution
- data
- low
- low resolution
- decoder
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/59—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial sub-sampling or interpolation, e.g. alteration of picture size or resolution
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/156—Availability of hardware or computational resources, e.g. encoding based on power-saving criteria
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/44—Decoders specially adapted therefor, e.g. video decoders which are asymmetric with respect to the encoder
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/60—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
- H04N19/61—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
- H04N19/82—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation involving filtering within a prediction loop
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/85—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression
- H04N19/86—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression involving reduction of coding artifacts, e.g. of blockiness
Definitions
- the present invention relates to a video decoder with power reduction.
- Power reduction is generally achieved by using two primary techniques.
- the first technique for power reduction is opportunistic, where a video coding system reduces its processing capability when operating on a sequence that is easy to decode. This reduction in processing capability may be achieved by frequency scaling, voltage scaling, on-chip data pre-fetching (caching) , and/ or a systematic idling strategy. In many cases the resulting decoder operation conforms to the standard.
- the second technique for power reduction is to discard frame or image data during the decoding process. This typically allows for more significant power savings but generally at the expense of visible degradation in the image quality. In addition, in many cases the resulting decoder operation does not conform to the standard.
- One embodiment of the present invention discloses a video decoder that decodes video comprising: (a) an entropy decoder that decodes a bitstream defining said video; (b) an inverse transformation that transforms said decoded bitstream; (c) a predictor that selectively performs an intra- prediction and a motion compensated prediction based on said decoded bitstream; (d) a buffer comprising compressed image data used for said motion compensated prediction, including low-resolution data and high-resolution data, where said predictor predicts both a low-resolution data set and a high-resolution data set based upon said low resolution data using high-resolution prediction information decoded from said bitstream without using said high resolution data.
- a video decoder that decodes video comprising: (a) an entropy decoder that decodes a bitstream defining said video; (b) an inverse transformation that transforms said decoded bitstream; (c) a predictor that selectively performs an intra- prediction and a motion compensated prediction on said decoded bitstream; (d) a buffer comprising compressed image data used for said motion compensated prediction; (e) wherein said compressed image data includes a low-resolution data set and a high-resolution data set, where said low resolution data set is independent of said high resolution data set and used for decoding in a low resolution mode, and wherein said low resolution data set and said high resolution data set are both used when decoding in a high resolution mode.
- FIG. 1 illustrates a decoder
- FIG. 2 illustrates low resolution prediction
- FIGS . 3A and 3B illustrate a decoder and data flow for the decoder.
- FIG. 4 illustrates a sampling structure of the frame buffer.
- FIG. 5 illustrates integration of the frame buffer in the decoder.
- FIG. 6A and 6B illustrates representative pixel values of two blocks.
- the system may be used with minimal impact on coding efficiency.
- the system should operate alternatively on low resolution data and high resolution data.
- the combination of low resolution data and high resolution data may result in full resolution data.
- the use of low resolution data is particularly suitable when the display has a resolution lower than the resolution of the transmitted content.
- Power is a factor when designing higher resolution decoders.
- One major contributor to power usage is memory bandwidth. Memory bandwidth traditionally increases with higher resolutions and frame rates, and it is often a significant bottleneck and cost factor in system design.
- a second major contributor to power usage is high pixel counts . High pixel counts are directly determined by the resolution of the image frame and increase the amount of pixel processing and computation . The amount of power required for each pixel operation is determined by the complexity of the decoding process. Historically, the decoding complexity has increased in each "improved" video coding standard.
- the system may include an entropy decoding module 10, a transformation module (such as inverse transformation using a dequant IDCT) 20, an intra prediction module 30, a motion compensated prediction module 40 , an adder 80, a deblocking module 50, an adaptive loop filter module 60, and a memory compression / decompression module associated with a frame buffer 70.
- the arrangement and selection of the different modules for the video system may be modified, as desired.
- the system in one aspect, preferably reduces the power requirements of both memory bandwidth and high pixel counts of the frame buffer.
- the memory bandwidth is reduced by incorporating a frame buffer compression technique within a video coder design.
- the purpose of the frame buffer compression technique is to reduce the memory bandwidth (and power) required to access data in the reference picture buffer. Given that the reference picture buffer is itself a compressed version of the original image data, compressing the reference frames can be achieved without significant coding loss for many applications.
- the video codec should support a low resolution processing mode without drift.
- the decoder may switch between low-resolution and full-resolution operating points and be compliant with the standard . This may be accomplished by performing prediction of both the low-resolution and high-resolution data using the full-resolution prediction information but only the low-resolution image data. Additionally, this may be improved using a de-blocking process that makes de-blocking decisions using only the low-resolution data. De-blocking is applied to the low-resolution data and, also if desired, the high- resolution data. The de-blocking of the low-resolution pixels does not depend on the high-resolution pixels.
- the low resolution deblocking and high resolution deblocking may be performed serially and/ or in parallel. However, the deblocking of the high resolution pixels may depend on the low- resolution pixels. In this manner the low resolution process is independent of the high resolution process, thus enabling a power savings mode, while the high resolution process may depend on the low resolution process, thus enabling greater image quality when desired .
- a decoder when operating in the low-resolution mode (S 10) , may exploit the properties of low- resolution prediction and modified de-blocking to significantly reduce the number of pixels to be processed. This may be accomplished by predicting only the low-resolution data (S 12) . Then after predicting the low resolution data, computing the residual data for only the low-resolution pixels (i.e. , pixel locations) and not the high resolution pixels (i.e . , pixel locations) (S 14) . The residual data is typically transmitted in a bit-stream. The residual data computed for the low- resolution pixel values has the same pixel values as the full resolution residual data at the low-resolution pixel locations.
- the residual data needs to only be calculated at the position of the low-resolution pixels.
- the low-resolution residual is added to the low-resolution prediction (S 16) , to provide the low resolution pixel values.
- the resulting signal is then de-blocked. Again, the de-blocking is preferably performed at only the low-resolution sample locations (S I 8) to reduce power consumption.
- the result may be stored in the reference picture frame buffer for future prediction.
- the result may be processed with an adaptive loop filter.
- the adaptive loop filter may be related to the adaptive loop filter for the full resolution data, or it may be signaled independently, or it may be omitted.
- FIGS . 3A and 3B An exemplary depiction of the system operating in low- resolution mode is shown in FIGS . 3A and 3B .
- the system may likewise include a mode that operates in full resolution mode .
- entropy decoding 100 may be performed at full resolution, while the inverse transform (Dequant IDCT) 200 and prediction (Intra Prediction 300 ; Motion Compensated Prediction (MCP) 400) are preferably performed at low resolution.
- Dequant IDCT inverse transform
- MCP Motion Compensated Prediction
- a frame buffer that includes memory compression stores the low-resolution data used for future prediction.
- the entropy decoding 100 shown in FIG. 3A entropy decodes the residual data for full-resolution pixels ( 101 ) .
- the shaded pixels in the residual 101 represent low resolution positions, while the un-shaded pixels represent high resolution positions .
- the Dequant IDCT 200 inverse transforms only the low resolution pixel data in the residual 101 , so as to produce a residual-after-Dequant-and-IDCT 201 .
- the Intra Prediction 300 produces a prediction 30 1 only for the low resolution positions (depicted by the shaded pixels) .
- Adder 800 adds the low resolution pixel data in the residual-after-Dequant-and- IDCT 20 1 to the low resolution pixel data in the prediction 30 1 , so as to produce a reconstruction 80 1 only for the low resolution positions (depicted by the shaded pixels) .
- the MCP 400 shown in FIG. 3B reads out the low resolution pixel data of the reference picture (depicted by the shaded pixels in the reference picture data 702) from the Memory 700, and produces by interpolation the high resolution pixel data which have been removed .
- the MCP 400 produces by interpolation the high resolution pixel data C from low resolution pixel data of the neighboring pixels. Taking an average of low resolution pixel data of pixels located in the upper and bottom side of C, taking an average of low resolution pixel data of pixels located in the left and right side of C, taking an average of low resolution pixel data of pixels located in the upper, bottom, left and right side of C may be employed as the interpolation.
- Deblocking 500 is performed in a cascade fashion .
- the Deblocking 500 filters the low resolution data in the first time (50 1 ) , while it filters the high resolution data in the second time (502) . More specifically, Deblocking 500 is performed in the following manner.
- Deblocking 500 applies only to the low resolution data using the low resolution data and the high resolution data by interpolation.
- Deblocking 500 applies only to the high resolution data using the low resolution data and the high resolution data by interpolation.
- a picture 502 after the Deblocking 500 is a full resolution picture, which may be referred to as a picture 70 1.
- the picture 70 1 after the Deblocking 500 is decimated (702) in a checker-board pattern such that only the low resolution positions are remained and stored in the Memory 700.
- the decimated high resolution pixel data (depicted by the unshaded pixels of 702) is interpolated and the interpolated picture is used for producing a predicted picture.
- the frame buffer compression technique is preferably a component of the low resolution functionality.
- the frame buffer compression technique preferably divides the image pixel data into multiple sets, and that a first set of the pixel data does not depend on other sets.
- the system employs a checker-board pattern as shown in FIG. 4.
- the shaded pixel locations belong to the first set and the un-shaded pixels belong to the second set.
- Other sampling structures may be used, as desired . For example, every other column of pixels may be assigned to the first set. Alternatively, every other row of pixels may be assigned to the first set. Similarly, every other column and row of pixels may be assigned to the first set. Any suitable partition into multiple sets of pixels may be used .
- the frame buffer compression technique preferably has the pixels in a second set of pixels be linearly predicted from pixels in the first set of pixels.
- the prediction may be pre-defined. Alternatively, it may be spatially varying or determined using any other suitable technique.
- the pixels in the first set of pixels are coded.
- This coding may use any suitable technique, such as for example, block truncation coding (BTC) , such as described by Healy, D . ; Mitchell, O . , " Digital Video Bandwidth Compression Using Block Truncation Coding, " IEEE Transactions on Communications [legacy, pre - 1988] , vol.29 , no. 12 pp. 1809- 1817 , Dec 198 1 , absolute moment block truncation coding (AMBTC) , such as described by Lema, M . ; Mitchell, O . , "Absolute Moment Block Truncation Coding and Its A pp.
- BTC block truncation coding
- AMBTC absolute moment block truncation coding
- the pixels in the second set of pixels may be coded and predicted using any suitable technique, such as for example being predicted using a linear process known to the frame buffer compression encoder and frame buffer compression decoder. Then the difference between the prediction and the pixel value may be computed. Finally, the difference may be compressed.
- the system may use block truncation coding (BTC) to compress the first set of pixels.
- AMBTC absolute moment block truncation coding
- the system may use quantization to compress the first set of pixels.
- the system may use bi-linear interpolation to predict the pixel values in the second set of pixels.
- the system may use bi-cubic interpolation to predict the pixel values in the second set of pixels.
- the system may use bi-linear interpolation to predict the pixel values in the second set of pixels and absolute moment block truncation coding (AMBTC) to compress the residual difference between the predicted pixel values in the second set and the pixel value in the second set.
- AMBTC absolute moment block truncation coding
- a property of the frame buffer compression technique is that it is controlled with a flag to signal low resolution processing capability.
- this flag does not signal low resolution processing capability, then the frame buffer decoder produces output frames that contain the first set of pixel values (i. e . low resolution pixel data) , possibly compressed, and the second set of pixel values (i.e . high resolution pixel data) that are predicted from the first set of pixel values and refined with optional residual data.
- this flag does signal low resolution processing capability
- the frame buffer decoder produces output frames that contain the first set of pixel values, possibly compressed, and the second set of pixel values that are predicted from the first set of pixel values but not refined with optional residual data. Accordingly, the flag indicates whether or not to use the optional residual data.
- the residual data may represent the differences between the predicted pixel values and the actual pixel values.
- the encoder when the flag does not signal low resolution processing capability, then the encoder stores the first set of pixel values, possibly in compressed form. Then, the encoder predicts the second set of pixel values from the first set of pixel values. In some embodiments, the encoder determines the residual difference between the prediction and actual pixel value and stores the residual difference, possibly in compressed form. In some embodiments, the encoder selects from multiple prediction mechanisms a preferred prediction mechanism for the second set pixels. The encoder then stores the selected prediction mechanism in the frame buffer. In one embodiment, the multiple prediction mechanisms consist of multiple linear filters and the encoder selects the prediction mechanism by computing the predicted pixel value for each linear filter and selecting the linear filter that computes a predicted pixel value that is closest to the pixel value .
- the multiple prediction mechanisms consist of multiple linear filters and the encoder selects the prediction mechanism by computing the predicted pixel values for each linear filter for a block of pixel locations and selecting the linear filter that computes a block of predicted pixel value that are closest to the block of pixel values.
- a block of pixels is a set of pixels within an image. The determination of the block of pixel predicted pixel values that are closest to the block of pixel values may be determined by selecting the block of predicted pixel values that result in the smallest sum of absolute differences between the block of predicted pixels values and block of pixels values. Alternatively, the sum of squared differences may be used to select the block. In other embodiments, the residual difference is compressed with block truncation coding (BTC) .
- BTC block truncation coding
- the residual difference is compressed with the absolute moment block truncation coding (AMBTC) .
- the parameters used for the compression of the second set pixels are determined from the parameters used for the compression of the first set of pixels.
- the first set of pixels and second set of pixels use AMBTC, and a first parameter used for the AMBTC method of the first set of pixels is related to a first parameter used for the AMBTC method for the second set of pixels.
- said first parameter used for the second set of pixels is equal to said first parameter used for the first set of pixels and not stored.
- said first parameter used for the second set of pixels is related to said first parameter used for the first set of pixels .
- the relationship may be defined as a scale factor, and the scale factor stored in place of said first parameter used for the second set of pixels. In other embodiments, the relationship may be defined as an index into a look-up-table of scale factors, the index stored in place of said first parameter used for the second set of pixels. In other embodiments, the relationship may be predefined.
- the encoder combines the selected prediction mechanism and residual difference determination step. By comparison, when the flag signals low resolution processing capability, then the encoder stores the first set of pixel values, possibly in compressed form. However, the encoder does not store residual information. In embodiments described above that determine a selected prediction mechanism, the encoder does not compute the selected prediction mechanism from the reconstructed data. Instead, any selected prediction mechanism is signaled from the encoder to the decoder.
- the signaling of a flag enables low resolution decoding capability.
- the decoder is not required to decode a low resolution sequence even when the flag signals a low resolution decoding capability. Instead, it may decode either a full resolution or low resolution sequence . These sequences will have the same decoded pixel values for pixel locations on the low resolution grid. The sequences may or may not have the same decoded pixel values for pixel locations on the high resolution grid.
- the signaling the flag may be on a frame-by- frame basis, on a sequence-by-sequence basis, or any other basis.
- the decoder When the flag appears in the bit-stream, the decoder preferably performs the following steps:
- (a) Disables the residual calculation in the frame buffer compression technique. This includes disabling the calculation of residual data during the loading of reference frames as well as disabling the calculation of residual data during the storage of reference frames, as illustrated in FIG. 5.
- the decoder may continue to operate in full resolution mode . Specifically, for future frames, it can retrieve the full resolution frame from the compressed reference buffer, perform motion compensation, residual addition, de-blocking, and loop filter. The result will be a full resolution frame . This frame can still contain frequency content that occupies the entire range of the full resolution pixel grid.
- the decoder may choose to operate only on the low-resolution data. This is possible due to the independence of the lower resolution grid on the higher resolution grid in the buffer compression structure.
- the interpolation process is modified to exploit the fact that high resolution pixels are linearly related to the low-resolution data.
- the motion estimation process may be performed at low resolution with modified interpolation filters.
- the system may exploit the fact that the low resolution data does not rely on the high resolution data in subsequent steps of the decoder.
- the system uses a reduced inverse transformation process that only computes the low resolution pixels from the full resolution transform coefficients.
- the system employs a de-blocking filter that de-blocks the low-resolution data independent from the high-resolution pixels (the high-resolution may be dependent on the low- resolution) . This is again due to the linear relationship between the high-resolution and lower-resolution data.
- d is a variable and pij and qij are pixel values.
- the location of the pixel values are depicted in FIG. 6.
- two 4x4 coding units are shown.
- the pixel values may be determined from any block size by considering the location of the pixels relative to the block boundary.
- the value computed for d is compared to a threshold. If the value d is less than the threshold, the deblocking filter is engaged . If the value d is greater than or equal to the threshold, then no filtering is applied and the deblocked pixels have the same values as the input pixel values.
- the threshold may be a function of a quantization parameter, and it may be described as beta(QP) .
- the deblocking decision is made independently for horizontal and vertical boundaries .
- the process continues to determine the type of filter to apply.
- the de-blocking operation uses either strong or weak filter types.
- the choice of filtering strength is based on the previously computed d, beta(QP) and also additional local differences. This is computed for each line (row or column) of the de-blocked boundary. For example, for the first row of the pixel locations shown in FIG. 6, the calculation is computed as
- StrongFilterFlag ((d ⁇ beta(QP)) && ((
- the filtering process may be described as follows. Here, this is described by the filtering process for the boundary between block A and block B in FIG. 6. The process is:
- ⁇ is an offset
- Clipo-255() is an operator that maps the input value to the range [0,255].
- the operator may map the input values to alternative ranges, such as [16,235], [0,1023] or other ranges.
- the filtering process may be described as follows. Here, this is described by the filtering process for the boundary between block A and block B in FIG. 6. The process is:
- Clipo-255() is an operator that maps the input value to the range [0,255].
- the operator may map the input values to alternative ranges, such as [16,235], [0,1023] or other ranges.
- ⁇ also referred to as delta in this specification, is an offset
- Clipo-255 () is an operator that maps the input value to the range [0,255] .
- the operator may map the input values to alternative ranges, such as [ 16, 235] , [0, 1023] or other ranges.
- the pixel locations within an image frame may be partitioned into two or more sets.
- a flag is signaled in the bit-stream, or communicated in any manner, the system enables the processing of the first set of pixel locations without the pixel values at the second set of pixel locations.
- FIG. 4 An example of this partitioning is shown in FIG. 4.
- a block is divided into two sets of pixels. The first set corresponds to the shaded locations; the second set corresponds to the unshaded locations.
- the system may modify the previous de-blocking operations as follows:
- the system uses the previously described equations, or other suitable equations. However, for the pixel values corresponding to pixel locations that are not in the first set of pixels, the system may use pixel values that are derived from the first set of pixel locations.
- pOl, p03, p05, p07, qOO, q02, q04, q06 in FIG. 6A and 6B are first set of pixels which are calculated by entropy decoding, inverse transformation and prediction.
- pOO, p02, p04, p06, qOl, q03, q05, q07 are second set of pixels which are calculated by such an equation shown in FIG.3B or FIG.5,
- Eq.l, Eq 2, Eqs 3, Eqs 4, and Eqs 5 are calculated using these pixel values.
- the system derives the pixel values as a linear summation of neighboring pixel values located in the first set of pixels.
- the system uses bi-linear interpolation of the pixel values located in the first set of pixels.
- the system computes the linear average of the pixel value located in the first set of pixels that is above the current pixel location and the pixel value located in the first set of pixels that is below the current pixel location.
- the system is operating on a vertical block boundary (and applying horizontal de-blocking). For the case that the system is operating on a horizontal block boundary (and applying vertical de-blocking) , then the system computes the average of the pixel to the left and right of the current location.
- the system may restrict the average calculation to pixel values within the same block. For example, if the pixel value located above a current pixel is not in the same block but the pixel value located below the current pixel is in the same block, then the current pixel is set equal to the pixel value below the current pixel.
- the system may use the same approach as described above. Namely, the pixels values that do not correspond to the first set of pixels are derived from the first set of pixels. After computing the above decision, the system may use the decision for the processing of the first set of pixels. Decoders processing subsequent sets of pixels use the same decision to process the subsequent sets of pixels.
- the system may use the weak filtering process described above .
- the system does not use the pixel values that correspond to the set of pixels subsequent to the first set. Instead, the system may derive the pixel values as discussed above .
- the value for ⁇ is then applied to the actual pixel values in the first set and the delta value is applied to the actual pixel values in the second set.
- the system may do the following:
- the system may use the equations for the luma strong filter described above. However, for the pixel values not located in the first set of pixel locations, the system may derive the pixel values from the first set of pixel locations as described above. The system then store the results of the filter process for the first set of pixel locations. Subsequently, for decoders generating the subsequent pixel locations as output, the system uses the equations for the luma strong filter described above with the previously computed strong filtered results for the first pixel locations and the reconstructed (not filtered) results for the subsequent pixel locations . The system then applies the filter at the subsequent pixel locations only. The output are filtered first pixel locations corresponding to the first filter operation and filtered subsequent pixel locations corresponding to the additional filter passes.
- the system takes the first pixel values and interpolates the missing pixel vales, computes the strong filter result for the first pixel values, updates the missing pixel values to be the actual reconstructed values, and computes the strong filter result for the missing pixel locations.
- the system uses the equations for the strong luma filter described above. For the pixel values not located in the first set of pixel locations, the system derives the pixel values from the first set of pixel locations as described above. The system then computes the strong filter result for both the first and subsequent sets of pixel locations using the derived values. Finally, the system computes a weighted average of the reconstructed pixel values at the subsequent locations and the output of the strong filter at the subsequent locations. In one embodiment, the weight is transmitted from the encoder to the decoder. In an alternative embodiment, the weight is fixed.
- the system uses the weak filtering process for chroma as described above .
- the system does not use the pixel values that correspond to the set of pixels subsequent to the first set. Instead, the system preferably derives the pixel values as in the previously described.
- the value for ⁇ is then applied to the actual pixel values in the first set and the delta value is applied to the actual pixel values in the second set.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computing Systems (AREA)
- Theoretical Computer Science (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
A video decoder that uses power reduction techniques is disclosed. The video decoder comprising; (a) an entropy decoder that decodes a bitstream defining said video having data representative of a low-resolution data set and a high-resolution data set; combiner forming a reconstructed image data based upon a predicted data and a transformed bitstream; (f) a de-blocking module that selectively de-blocks said reconstructed image data by making de-blocking decisions based upon data at said low-resolution pixel positions and not at said high-resolution pixel positions; (g) said de-blocking module using only the reconstructed image data at said low-resolution pixel positions for de-blocking said reconstructed image data at said low-resolution pixel position.
Description
DESCRIPTION
TITLE OF INVENTION: VIDEO DECODER FOR LOW RESOLUTION POWER REDUCTION USING LOW RESOLUTION DATA
TECHNICAL FIELD
The present invention relates to a video decoder with power reduction.
BACKGROUND ART
Existing video coding standards, such as H .264 /AVC, generally provide relatively high coding efficiency at the expense of increased computational complexity. The relatively high computational complexity has resulted in significant power consumption, which is especially problematic for low power devices such as cellular phones.
Power reduction is generally achieved by using two primary techniques. The first technique for power reduction is opportunistic, where a video coding system reduces its processing capability when operating on a sequence that is easy to decode. This reduction in processing capability may be achieved by frequency scaling, voltage scaling, on-chip
data pre-fetching (caching) , and/ or a systematic idling strategy. In many cases the resulting decoder operation conforms to the standard. The second technique for power reduction is to discard frame or image data during the decoding process. This typically allows for more significant power savings but generally at the expense of visible degradation in the image quality. In addition, in many cases the resulting decoder operation does not conform to the standard.
SUMMARY OF INVENTION
One embodiment of the present invention discloses a video decoder that decodes video comprising: (a) an entropy decoder that decodes a bitstream defining said video; (b) an inverse transformation that transforms said decoded bitstream; (c) a predictor that selectively performs an intra- prediction and a motion compensated prediction based on said decoded bitstream; (d) a buffer comprising compressed image data used for said motion compensated prediction, including low-resolution data and high-resolution data, where said predictor predicts both a low-resolution data set and a high-resolution data set based upon said low resolution data using high-resolution prediction information decoded from said bitstream without using said high resolution data.
Another embodiment of the present invention discloses a video decoder that decodes video comprising: (a) an entropy decoder that decodes a bitstream defining said video; (b) an inverse transformation that transforms said decoded bitstream; (c) a predictor that selectively performs an intra- prediction and a motion compensated prediction on said decoded bitstream; (d) a buffer comprising compressed image data used for said motion compensated prediction; (e) wherein said compressed image data includes a low-resolution data set and a high-resolution data set, where said low resolution data set is independent of said high resolution data set and used for decoding in a low resolution mode, and wherein said low resolution data set and said high resolution data set are both used when decoding in a high resolution mode.
The foregoing and other objectives, features, and advantages of the invention will be more readily understood upon consideration of the following detailed description of the invention, taken in conjunction with the accompanying drawings.
BRIEF DESCRIPTION OF DRAWINGS
FIG. 1 illustrates a decoder.
FIG. 2 illustrates low resolution prediction.
FIGS . 3A and 3B illustrate a decoder and data flow for the decoder.
FIG. 4 illustrates a sampling structure of the frame buffer.
FIG. 5 illustrates integration of the frame buffer in the decoder.
FIG. 6A and 6B illustrates representative pixel values of two blocks.
DESCRIPTION OF EMBODIMENTS
It is desirable to enable significant power savings typically associated with discarding frame data without visible degradation in the resulting image quality and standard nonconformance. Suitably implemented the system may be used with minimal impact on coding efficiency. In order to facilitate such power savings with minimal image derogation and loss of coding efficiency, the system should operate alternatively on low resolution data and high resolution data. The combination of low resolution data and high resolution data may result in full resolution data. The use of low resolution data is particularly suitable when the display has a resolution lower than the resolution of the transmitted content.
Power is a factor when designing higher resolution decoders. One major contributor to power usage is memory bandwidth. Memory bandwidth traditionally increases with higher resolutions and frame rates, and it is often a significant bottleneck and cost factor in system design. A
second major contributor to power usage is high pixel counts . High pixel counts are directly determined by the resolution of the image frame and increase the amount of pixel processing and computation . The amount of power required for each pixel operation is determined by the complexity of the decoding process. Historically, the decoding complexity has increased in each "improved" video coding standard.
Referring to FIG. 1 , the system may include an entropy decoding module 10, a transformation module (such as inverse transformation using a dequant IDCT) 20, an intra prediction module 30, a motion compensated prediction module 40 , an adder 80, a deblocking module 50, an adaptive loop filter module 60, and a memory compression / decompression module associated with a frame buffer 70. The arrangement and selection of the different modules for the video system may be modified, as desired. The system, in one aspect, preferably reduces the power requirements of both memory bandwidth and high pixel counts of the frame buffer. The memory bandwidth is reduced by incorporating a frame buffer compression technique within a video coder design. The purpose of the frame buffer compression technique is to reduce the memory bandwidth (and power) required to access data in the reference picture buffer. Given that the reference picture buffer is itself a compressed version of the original image data, compressing the reference frames can be achieved
without significant coding loss for many applications.
To address the high pixel counts, the video codec should support a low resolution processing mode without drift. This means that the decoder may switch between low-resolution and full-resolution operating points and be compliant with the standard . This may be accomplished by performing prediction of both the low-resolution and high-resolution data using the full-resolution prediction information but only the low-resolution image data. Additionally, this may be improved using a de-blocking process that makes de-blocking decisions using only the low-resolution data. De-blocking is applied to the low-resolution data and, also if desired, the high- resolution data. The de-blocking of the low-resolution pixels does not depend on the high-resolution pixels. The low resolution deblocking and high resolution deblocking may be performed serially and/ or in parallel. However, the deblocking of the high resolution pixels may depend on the low- resolution pixels. In this manner the low resolution process is independent of the high resolution process, thus enabling a power savings mode, while the high resolution process may depend on the low resolution process, thus enabling greater image quality when desired .
Referring to FIG. 2 , when operating in the low-resolution mode (S 10) , a decoder may exploit the properties of low- resolution prediction and modified de-blocking to significantly
reduce the number of pixels to be processed. This may be accomplished by predicting only the low-resolution data (S 12) . Then after predicting the low resolution data, computing the residual data for only the low-resolution pixels (i.e. , pixel locations) and not the high resolution pixels (i.e . , pixel locations) (S 14) . The residual data is typically transmitted in a bit-stream. The residual data computed for the low- resolution pixel values has the same pixel values as the full resolution residual data at the low-resolution pixel locations. The principal difference is that the residual data needs to only be calculated at the position of the low-resolution pixels. Following calculation of the residual, the low-resolution residual is added to the low-resolution prediction (S 16) , to provide the low resolution pixel values. The resulting signal is then de-blocked. Again, the de-blocking is preferably performed at only the low-resolution sample locations (S I 8) to reduce power consumption. Finally, the result may be stored in the reference picture frame buffer for future prediction. Optionally, the result may be processed with an adaptive loop filter. The adaptive loop filter may be related to the adaptive loop filter for the full resolution data, or it may be signaled independently, or it may be omitted.
An exemplary depiction of the system operating in low- resolution mode is shown in FIGS . 3A and 3B . The system may likewise include a mode that operates in full resolution
mode . As shown in FIGS . 3A and 3B, entropy decoding 100 may be performed at full resolution, while the inverse transform (Dequant IDCT) 200 and prediction (Intra Prediction 300 ; Motion Compensated Prediction (MCP) 400) are preferably performed at low resolution. The de-blocking
500 is preferably performed in a cascade fashion so that the de-blocking of the low resolution pixels do not depend on the additional, high resolution data. Finally, a frame buffer that includes memory compression stores the low-resolution data used for future prediction.
The entropy decoding 100 shown in FIG. 3A entropy decodes the residual data for full-resolution pixels ( 101 ) . The shaded pixels in the residual 101 represent low resolution positions, while the un-shaded pixels represent high resolution positions . The Dequant IDCT 200 inverse transforms only the low resolution pixel data in the residual 101 , so as to produce a residual-after-Dequant-and-IDCT 201 .
In the case of intra pictures, the Intra Prediction 300 produces a prediction 30 1 only for the low resolution positions (depicted by the shaded pixels) . Adder 800 adds the low resolution pixel data in the residual-after-Dequant-and- IDCT 20 1 to the low resolution pixel data in the prediction 30 1 , so as to produce a reconstruction 80 1 only for the low resolution positions (depicted by the shaded pixels) .
In the case of inter pictures, the MCP 400 shown in FIG.
3B reads out the low resolution pixel data of the reference picture (depicted by the shaded pixels in the reference picture data 702) from the Memory 700, and produces by interpolation the high resolution pixel data which have been removed . For example, as indicated in the interpolation 40 1 , the MCP 400 produces by interpolation the high resolution pixel data C from low resolution pixel data of the neighboring pixels. Taking an average of low resolution pixel data of pixels located in the upper and bottom side of C, taking an average of low resolution pixel data of pixels located in the left and right side of C, taking an average of low resolution pixel data of pixels located in the upper, bottom, left and right side of C may be employed as the interpolation.
Deblocking 500 is performed in a cascade fashion . The Deblocking 500 filters the low resolution data in the first time (50 1 ) , while it filters the high resolution data in the second time (502) . More specifically, Deblocking 500 is performed in the following manner.
STEP 1 ) (50 1 )
Deblocking 500 applies only to the low resolution data using the low resolution data and the high resolution data by interpolation.
STEP 2) (502)
Deblocking 500 applies only to the high resolution data using the low resolution data and the high resolution data by
interpolation.
Pictures after the Deblocking 500 are stored in the Memory 700. The followings explain pictures (701 , 702 , 703) which have been once stored in the Memory 700 after the Deblocking 500, and read out for the MCP 400. A picture 502 after the Deblocking 500 is a full resolution picture, which may be referred to as a picture 70 1. The picture 70 1 after the Deblocking 500 is decimated (702) in a checker-board pattern such that only the low resolution positions are remained and stored in the Memory 700. When used in prediction, the decimated high resolution pixel data (depicted by the unshaded pixels of 702) is interpolated and the interpolated picture is used for producing a predicted picture.
The frame buffer compression technique is preferably a component of the low resolution functionality. The frame buffer compression technique preferably divides the image pixel data into multiple sets, and that a first set of the pixel data does not depend on other sets. In one embodiment, the system employs a checker-board pattern as shown in FIG. 4. In FIG. 4 , the shaded pixel locations belong to the first set and the un-shaded pixels belong to the second set. Other sampling structures may be used, as desired . For example, every other column of pixels may be assigned to the first set. Alternatively, every other row of pixels may be assigned to the first set. Similarly, every other column and row of pixels may
be assigned to the first set. Any suitable partition into multiple sets of pixels may be used .
For memory compression / decompression the frame buffer compression technique preferably has the pixels in a second set of pixels be linearly predicted from pixels in the first set of pixels. The prediction may be pre-defined. Alternatively, it may be spatially varying or determined using any other suitable technique.
In one embodiment, the pixels in the first set of pixels are coded. This coding may use any suitable technique, such as for example, block truncation coding (BTC) , such as described by Healy, D . ; Mitchell, O . , " Digital Video Bandwidth Compression Using Block Truncation Coding, " IEEE Transactions on Communications [legacy, pre - 1988] , vol.29 , no. 12 pp. 1809- 1817 , Dec 198 1 , absolute moment block truncation coding (AMBTC) , such as described by Lema, M . ; Mitchell, O . , "Absolute Moment Block Truncation Coding and Its A pp. ication to Color Images, " IEEE Transactions on Communications [legacy, pre - 1988] , vol.32 , no . 10 pp. 1 148- 1 157, Oct 1984 , or scalar quantization. Similarly, the pixels in the second set of pixels may be coded and predicted using any suitable technique, such as for example being predicted using a linear process known to the frame buffer compression encoder and frame buffer compression decoder. Then the difference between the prediction and the pixel value may be
computed. Finally, the difference may be compressed. In one embodiment, the system may use block truncation coding (BTC) to compress the first set of pixels. In another embodiment, the system may use absolute moment block truncation coding (AMBTC) to compress the first set of pixels. In another embodiment, the system may use quantization to compress the first set of pixels. In yet another embodiment, the system may use bi-linear interpolation to predict the pixel values in the second set of pixels. In a further embodiment, the system may use bi-cubic interpolation to predict the pixel values in the second set of pixels. In another embodiment, the system may use bi-linear interpolation to predict the pixel values in the second set of pixels and absolute moment block truncation coding (AMBTC) to compress the residual difference between the predicted pixel values in the second set and the pixel value in the second set.
A property of the frame buffer compression technique is that it is controlled with a flag to signal low resolution processing capability. In one configuration when this flag does not signal low resolution processing capability, then the frame buffer decoder produces output frames that contain the first set of pixel values (i. e . low resolution pixel data) , possibly compressed, and the second set of pixel values (i.e . high resolution pixel data) that are predicted from the first set of pixel values and refined with optional residual data. In
another configuration when this flag does signal low resolution processing capability, then the frame buffer decoder produces output frames that contain the first set of pixel values, possibly compressed, and the second set of pixel values that are predicted from the first set of pixel values but not refined with optional residual data. Accordingly, the flag indicates whether or not to use the optional residual data. The residual data may represent the differences between the predicted pixel values and the actual pixel values.
For the frame buffer compression encoder, when the flag does not signal low resolution processing capability, then the encoder stores the first set of pixel values, possibly in compressed form. Then, the encoder predicts the second set of pixel values from the first set of pixel values. In some embodiments, the encoder determines the residual difference between the prediction and actual pixel value and stores the residual difference, possibly in compressed form. In some embodiments, the encoder selects from multiple prediction mechanisms a preferred prediction mechanism for the second set pixels. The encoder then stores the selected prediction mechanism in the frame buffer. In one embodiment, the multiple prediction mechanisms consist of multiple linear filters and the encoder selects the prediction mechanism by computing the predicted pixel value for each linear filter and selecting the linear filter that computes a predicted pixel
value that is closest to the pixel value . In one embodiment, the multiple prediction mechanisms consist of multiple linear filters and the encoder selects the prediction mechanism by computing the predicted pixel values for each linear filter for a block of pixel locations and selecting the linear filter that computes a block of predicted pixel value that are closest to the block of pixel values. A block of pixels is a set of pixels within an image. The determination of the block of pixel predicted pixel values that are closest to the block of pixel values may be determined by selecting the block of predicted pixel values that result in the smallest sum of absolute differences between the block of predicted pixels values and block of pixels values. Alternatively, the sum of squared differences may be used to select the block. In other embodiments, the residual difference is compressed with block truncation coding (BTC) . In one embodiment, the residual difference is compressed with the absolute moment block truncation coding (AMBTC) . In one embodiment, the parameters used for the compression of the second set pixels are determined from the parameters used for the compression of the first set of pixels. In one embodiment, the first set of pixels and second set of pixels use AMBTC, and a first parameter used for the AMBTC method of the first set of pixels is related to a first parameter used for the AMBTC method for the second set of pixels. In one embodiment, said
first parameter used for the second set of pixels is equal to said first parameter used for the first set of pixels and not stored. In another embodiment, said first parameter used for the second set of pixels is related to said first parameter used for the first set of pixels . In one embodiment, the relationship may be defined as a scale factor, and the scale factor stored in place of said first parameter used for the second set of pixels. In other embodiments, the relationship may be defined as an index into a look-up-table of scale factors, the index stored in place of said first parameter used for the second set of pixels. In other embodiments, the relationship may be predefined. In other embodiments, the encoder combines the selected prediction mechanism and residual difference determination step. By comparison, when the flag signals low resolution processing capability, then the encoder stores the first set of pixel values, possibly in compressed form. However, the encoder does not store residual information. In embodiments described above that determine a selected prediction mechanism, the encoder does not compute the selected prediction mechanism from the reconstructed data. Instead, any selected prediction mechanism is signaled from the encoder to the decoder.
The signaling of a flag enables low resolution decoding capability. The decoder is not required to decode a low resolution sequence even when the flag signals a low
resolution decoding capability. Instead, it may decode either a full resolution or low resolution sequence . These sequences will have the same decoded pixel values for pixel locations on the low resolution grid. The sequences may or may not have the same decoded pixel values for pixel locations on the high resolution grid. The signaling the flag may be on a frame-by- frame basis, on a sequence-by-sequence basis, or any other basis.
When the flag appears in the bit-stream, the decoder preferably performs the following steps:
(a) Disables the residual calculation in the frame buffer compression technique. This includes disabling the calculation of residual data during the loading of reference frames as well as disabling the calculation of residual data during the storage of reference frames, as illustrated in FIG. 5.
(b) Uses low resolution pixel values for low resolution deblocking, as previously described . Uses an alternative deblocking operation for the sample locations in the higher resolution locations, as previously described.
(c) Stores reference frames prior to applying the adaptive loop filter.
With these changes, the decoder may continue to operate in full resolution mode . Specifically, for future frames, it can retrieve the full resolution frame from the compressed reference buffer, perform motion compensation, residual
addition, de-blocking, and loop filter. The result will be a full resolution frame . This frame can still contain frequency content that occupies the entire range of the full resolution pixel grid.
Alternatively though, the decoder may choose to operate only on the low-resolution data. This is possible due to the independence of the lower resolution grid on the higher resolution grid in the buffer compression structure. For motion estimation, the interpolation process is modified to exploit the fact that high resolution pixels are linearly related to the low-resolution data. Thus, the motion estimation process may be performed at low resolution with modified interpolation filters. Similarly, for residual calculation, the system may exploit the fact that the low resolution data does not rely on the high resolution data in subsequent steps of the decoder. Thus, the system uses a reduced inverse transformation process that only computes the low resolution pixels from the full resolution transform coefficients. Finally, the system employs a de-blocking filter that de-blocks the low-resolution data independent from the high-resolution pixels (the high-resolution may be dependent on the low- resolution) . This is again due to the linear relationship between the high-resolution and lower-resolution data.
An existing deblocking filter in the JCT-VC Test Model under Consideration JCTVC-A 1 19 is desired in the context of
8x8 block sizes . For luma deblocking filtering, the process begins by determining if a block boundary should be deblocked . This is accomplished by computing the following d = | p22 - 2*p l 2 + p02 | + | q22 - 2*q l 2 + q02 | + | p25 - 2*p l 5 + p05 | + | q25 - 2*q l 5 + qOs | , ... (Eq. 1 )
where d is a variable and pij and qij are pixel values. The location of the pixel values are depicted in FIG. 6. In FIG. 6, two 4x4 coding units are shown. However, the pixel values may be determined from any block size by considering the location of the pixels relative to the block boundary.
Next, the value computed for d is compared to a threshold. If the value d is less than the threshold, the deblocking filter is engaged . If the value d is greater than or equal to the threshold, then no filtering is applied and the deblocked pixels have the same values as the input pixel values. Note that the threshold may be a function of a quantization parameter, and it may be described as beta(QP) . The deblocking decision is made independently for horizontal and vertical boundaries .
If the d value for a boundary results in a decision to deblock, then the process continues to determine the type of filter to apply. The de-blocking operation uses either strong or weak filter types. The choice of filtering strength is based on the previously computed d, beta(QP) and also additional local differences. This is computed for each line (row or column) of
the de-blocked boundary. For example, for the first row of the pixel locations shown in FIG. 6, the calculation is computed as
StrongFilterFlag = ((d< beta(QP)) && (( | p3_ - pOi | + | q0_ - q3i|) < (β>>3) && I pOi - qOi| < ((5*tc + 1)»1) ). ... (Eq.2) where tc is a threshold that is typically a function of the quantization parameter, QP.
For the case of luminance samples, if the previously described process results in the decision to de-block a boundary and subsequently to de-block a line (row or column) with a weak filter, then the filtering process may be described as follows. Here, this is described by the filtering process for the boundary between block A and block B in FIG. 6. The process is:
Δ = Clip(-tc,tc, (13*(q0_ - p0_) + 4*( qli - pli) - 5*( q2i - p2i)+16)>>5)) i = 0,7
pOi = Clip0-255(p0i + Δ) i = 0,7
qOi = Clipo-255(qOi - Δ) i = 0,7
pli = Clipo-255(pli + Δ/2) i = 0,7
qli = Clipo-255(qli - Δ/2) i = 0,7
... (Eqs.3) where Δ is an offset and Clipo-255() is an operator that maps the input value to the range [0,255]. In alternative embodiments, the operator may map the input values to alternative ranges, such as [16,235], [0,1023] or other ranges.
For the case of luminance samples, if the previously described process results in the decision to de-block a boundary and subsequently to de-block a line (row or column) with a strong filter, then the filtering process may be described as follows. Here, this is described by the filtering process for the boundary between block A and block B in FIG. 6. The process is:
p0i=Clip0-255((p2i + 2*pli + 2*p0i +2*q0i + qli + 4)>>3); i = 0,7 q0i=Clip0-255((pli + 2*p0i + 2*q0j + 2*qli + q2i + 4)>>3); i = 0,7 pli=Clip0-255((p2i + pli + pOi + qOi +2)>>2); i = 0,7
qli=Clipo-255((pOi + qOi + qli + q2i +2)>>2); i = 0,7
p2i=Clip0-255((2*p3i + 3*p2i + pli + pOi + qOi + 4)>>3); i = 0,7 q2i=Clip0-255((p0i + qOi + qli + 3*q2i + 2*q3i + 4)>>3); i = 0,7
... (Eqs.4)
where Clipo-255() is an operator that maps the input value to the range [0,255]. In alternative embodiments, the operator may map the input values to alternative ranges, such as [16,235], [0,1023] or other ranges.
For the case of chrominance samples, if the previously described process results in the decision to de-block a boundary, then all lines (row or column) or the chroma component is processed with a weak filtering operation. Here, this is described by the filtering process for the boundary between block A and block B in FIG. 6, where the blocks are now assumed to contain chroma pixel values. The process is:
Δ = Clip(-tc,tc, ((((qOi - pOi) < < 2) + p l i - q li + 4) > > 3)) i = 0, 7 pOi = Clip0 -255(p0i + Δ) i = 0, 7 qOi = Clipo-25 s(qOi - Δ) i = 0, 7
... (Eqs . 5)
where Δ, also referred to as delta in this specification, is an offset and Clipo-255 () is an operator that maps the input value to the range [0,255] . In alternative embodiments, the operator may map the input values to alternative ranges, such as [ 16, 235] , [0, 1023] or other ranges.
The pixel locations within an image frame may be partitioned into two or more sets. When a flag is signaled in the bit-stream, or communicated in any manner, the system enables the processing of the first set of pixel locations without the pixel values at the second set of pixel locations. An example of this partitioning is shown in FIG. 4. In FIG. 4 , a block is divided into two sets of pixels. The first set corresponds to the shaded locations; the second set corresponds to the unshaded locations.
When this alternative mode is enabled, the system may modify the previous de-blocking operations as follows:
First in calculating if a boundary should be de-blocked, the system uses the previously described equations, or other suitable equations. However, for the pixel values corresponding to pixel locations that are not in the first set of pixels, the system may use pixel values that are derived from
the first set of pixel locations. pOl, p03, p05, p07, qOO, q02, q04, q06 in FIG. 6A and 6B are first set of pixels which are calculated by entropy decoding, inverse transformation and prediction. pOO, p02, p04, p06, qOl, q03, q05, q07 are second set of pixels which are calculated by such an equation shown in FIG.3B or FIG.5,
pOO = (plO + qOO) >> 1
p02 = (pOl + p03) >> 1
p04 = (p03 + p05) >> 1 q07 = (p07 + ql7) >> 1. ... (Eqs.6)
The Eq.l, Eq 2, Eqs 3, Eqs 4, and Eqs 5 are calculated using these pixel values.
In one embodiment, the system derives the pixel values as a linear summation of neighboring pixel values located in the first set of pixels. In a second embodiment, the system uses bi-linear interpolation of the pixel values located in the first set of pixels. In a preferred embodiment, the system computes the linear average of the pixel value located in the first set of pixels that is above the current pixel location and the pixel value located in the first set of pixels that is below the current pixel location. Please note that the above description assumes that the system is operating on a vertical block boundary (and applying horizontal de-blocking). For the case that the system is operating on a horizontal block
boundary (and applying vertical de-blocking) , then the system computes the average of the pixel to the left and right of the current location. In an alternative embodiment, the system may restrict the average calculation to pixel values within the same block. For example, if the pixel value located above a current pixel is not in the same block but the pixel value located below the current pixel is in the same block, then the current pixel is set equal to the pixel value below the current pixel.
Second, in calculating if a boundary should use the strong or weak filter, the system may use the same approach as described above. Namely, the pixels values that do not correspond to the first set of pixels are derived from the first set of pixels. After computing the above decision, the system may use the decision for the processing of the first set of pixels. Decoders processing subsequent sets of pixels use the same decision to process the subsequent sets of pixels.
If the previously described process results in the decision to de-block a boundary and subsequently to de-block a line (row or column) with a weak filter, then the system may use the weak filtering process described above . However, when computing the value for Δ, the system does not use the pixel values that correspond to the set of pixels subsequent to the first set. Instead, the system may derive the pixel values as discussed above . By way of example , the value for Δ is then
applied to the actual pixel values in the first set and the delta value is applied to the actual pixel values in the second set.
If the previously described process results in the decision to de-block a boundary and subsequently to de-block a line (row or column) with a strong filter, then the system may do the following:
In one embodiment, the system may use the equations for the luma strong filter described above. However, for the pixel values not located in the first set of pixel locations, the system may derive the pixel values from the first set of pixel locations as described above. The system then store the results of the filter process for the first set of pixel locations. Subsequently, for decoders generating the subsequent pixel locations as output, the system uses the equations for the luma strong filter described above with the previously computed strong filtered results for the first pixel locations and the reconstructed (not filtered) results for the subsequent pixel locations . The system then applies the filter at the subsequent pixel locations only. The output are filtered first pixel locations corresponding to the first filter operation and filtered subsequent pixel locations corresponding to the additional filter passes.
To summarize, as previously described, the system takes the first pixel values and interpolates the missing pixel vales,
computes the strong filter result for the first pixel values, updates the missing pixel values to be the actual reconstructed values, and computes the strong filter result for the missing pixel locations.
In a second embodiment, the system uses the equations for the strong luma filter described above. For the pixel values not located in the first set of pixel locations, the system derives the pixel values from the first set of pixel locations as described above. The system then computes the strong filter result for both the first and subsequent sets of pixel locations using the derived values. Finally, the system computes a weighted average of the reconstructed pixel values at the subsequent locations and the output of the strong filter at the subsequent locations. In one embodiment, the weight is transmitted from the encoder to the decoder. In an alternative embodiment, the weight is fixed.
If the previously described process results in the decision to de-block a boundary, then the system uses the weak filtering process for chroma as described above . However, when computing the value for Δ, the system does not use the pixel values that correspond to the set of pixels subsequent to the first set. Instead, the system preferably derives the pixel values as in the previously described. By way of example, the value for Δ is then applied to the actual pixel values in the first set and the delta value is applied to the actual pixel
values in the second set.
The terms and expressions which have been employed in the foregoing specification are used therein as terms of description and not of limitation, and there is no intention, in the use of such terms and expressions, of excluding equivalents of the features shown and described or portions thereof, it being recognized that the scope of the invention is defined and limited only by the claims which follow.
Claims
1 . A video decoder that decodes video comprising:
(a) an entropy decoder that decodes a bitstream defining said video having data representative of a low resolution data set and a high resolution data set;
(b) an inverse transformation that transforms said decoded bitstream;
(c) a predictor that selectively performs an intra- prediction and a motion compensated prediction based on said decoded bitstream for only said low resolution data set when operating in a low resolution mode;
(d) a buffer comprising compressed image data used for said motion compensated prediction;
(e) a combiner forming a reconstructed image data based upon said predicted data and said transformed bitstream;
(f) a de-blocking module that selectively de-blocks said reconstructed image data by making de-blocking decisions based upon data at said low resolution pixel positions and not at said high resolution pixel positions;
(g) said de-blocking module using only the reconstructed image data at said low-resolution pixel positions for deblocking said reconstructed image data at said low resolution pixel position.
2. The decoder of claim 1 wherein compressed image data includes low resolution data used for predicting a low resolution data set.
3. The decoder of claim 1 wherein said inverse transformation operates only at low resolution pixel positions from said bit stream when operating in said low resolution mode .
4. The decoder of claim 1 wherein said predicted data and said transformed bitstream are added together.
5. The decoder of claim 1 wherein said predictor operates only at low resolution pixel positions when operating in said low resolution mode.
6. The decoder of claim 1 wherein said de-blocking module is enabled when a flag is signed in said bitstream.
7. The decoder of claim 1 wherein said buffer comprising compressed image data used for said motion compensated prediction, including low-resolution data and high-resolution data, where said predictor predicts both a low-resolution data set and a high-resolution data set based upon said low resolution data using high-resolution prediction information decoded from said bitstream without using said high resolution data.
8. A video decoder that decodes video comprising:
(a) an entropy decoder that decodes a bitstream defining said video having data representative of a low resolution data set and a high resolution data set;
(b) an inverse transformation that transforms said decoded bitstream;
(c) a predictor that selectively performs an intra- prediction and a motion compensated prediction based on said decoded bitstream for only said low resolution data set when operating in a high resolution mode;
(d) a buffer comprising compressed image data used for said motion compensated prediction;
(e) a combiner forming a reconstructed image data based upon said predicted data and said transformed bitstream;
(f) a de-blocking module that selectively de-blocks said reconstructed image data by making de-blocking decisions based upon data at said low resolution pixel positions and not at said high resolution pixel positions;
(g) a de-blocking module that selectively de-blocks said reconstructed image data by making de-blocking decisions based upon data at said low resolution pixel positions and not at said high resolution pixel positions;
(h) said deblocking module using the reconstructed image data at said high resolution pixel position for deblocking said reconstructed image data at said high resolution pixel positions in a manner dependent on said reconstructed image data at said low resolution pixel positions.
9. The decoder of claim 8 wherein said buffer comprising compressed image data used for said motion compensated prediction, including low-resolution data and high-resolution data, where said predictor predicts both a low-resolution data set and a high-resolution data set based upon said low resolution data using high-resolution prediction information decoded from said bitstream without using said high resolution data.
10. The decoder of claim 8 wherein said deblocking is enabled based upon receiving a flag in said bistream.
1 1 . A video decoder that decodes video from a bit- stream comprising:
(a) an entropy decoder that decodes a bitstream of said video; (b) an inverse transformation that transforms said decoded bitstream;
(c) a predictor that selectively performs an intra- prediction and a motion compensated prediction based on said decoded bitstream;
(d) a buffer comprising compressed image data used for said motion compensated prediction having a low resolution data set and a high resolution data set that is determined based upon said low resolution data set;
(e) wherein said decoder receives a flag in said bit- stream indicating a option to decode said bitstream at a low resolution mode or a high resolution mode, from which said decoder selects one of said low resolution mode and said high resolution mode;
(f) wherein said decoder does not receive said flag in said bitstream and as a result said decoder selects said high resolution mode.
Applications Claiming Priority (6)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US12/838,389 US8548062B2 (en) | 2010-07-16 | 2010-07-16 | System for low resolution power reduction with deblocking flag |
| US12/838,389 | 2010-07-16 | ||
| US12/838,354 | 2010-07-16 | ||
| US12/838,367 | 2010-07-16 | ||
| US12/838,367 US20120014447A1 (en) | 2010-07-16 | 2010-07-16 | System for low resolution power reduction with high resolution deblocking |
| US12/838,354 US9313523B2 (en) | 2010-07-16 | 2010-07-16 | System for low resolution power reduction using deblocking |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2012008616A1 true WO2012008616A1 (en) | 2012-01-19 |
Family
ID=45469598
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2011/066638 Ceased WO2012008616A1 (en) | 2010-07-16 | 2011-07-14 | Video decoder for low resolution power reduction using low resolution data |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2012008616A1 (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2012161345A1 (en) * | 2011-05-26 | 2012-11-29 | Sharp Kabushiki Kaisha | Video decoder |
| WO2020252745A1 (en) * | 2019-06-20 | 2020-12-24 | Alibaba Group Holding Limited | Loop filter design for adaptive resolution video coding |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0794675A2 (en) * | 1996-03-04 | 1997-09-10 | Kokusai Denshin Denwa Kabushiki Kaisha | Apparatus for decoding coded video data |
| JPH11146399A (en) * | 1997-11-05 | 1999-05-28 | Sanyo Electric Co Ltd | Image decoder |
| WO2005088983A2 (en) * | 2004-03-08 | 2005-09-22 | Koninklijke Philips Electronics N.V. | Video decoder with scalable compression and buffer for storing and retrieving reference frame data |
| JP2007221697A (en) * | 2006-02-20 | 2007-08-30 | Hitachi Ltd | Image decoding apparatus and image decoding method |
-
2011
- 2011-07-14 WO PCT/JP2011/066638 patent/WO2012008616A1/en not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0794675A2 (en) * | 1996-03-04 | 1997-09-10 | Kokusai Denshin Denwa Kabushiki Kaisha | Apparatus for decoding coded video data |
| JPH11146399A (en) * | 1997-11-05 | 1999-05-28 | Sanyo Electric Co Ltd | Image decoder |
| WO2005088983A2 (en) * | 2004-03-08 | 2005-09-22 | Koninklijke Philips Electronics N.V. | Video decoder with scalable compression and buffer for storing and retrieving reference frame data |
| JP2007221697A (en) * | 2006-02-20 | 2007-08-30 | Hitachi Ltd | Image decoding apparatus and image decoding method |
Non-Patent Citations (1)
| Title |
|---|
| ZHAN MA ET AL.: "System for Graceful Power Degradation", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC) OF ITU-T SG16 WP3 AND ISO/IEC JTC1/SC29/WG11, JCTVC-B114, 2ND MEETING: GENEVA, CH, July 2010 (2010-07-01), GENEVA, CH, pages 1 - 7 * |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2012161345A1 (en) * | 2011-05-26 | 2012-11-29 | Sharp Kabushiki Kaisha | Video decoder |
| JP2014519212A (en) * | 2011-05-26 | 2014-08-07 | シャープ株式会社 | Video decoder |
| WO2020252745A1 (en) * | 2019-06-20 | 2020-12-24 | Alibaba Group Holding Limited | Loop filter design for adaptive resolution video coding |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8548062B2 (en) | System for low resolution power reduction with deblocking flag | |
| CA2755889C (en) | Image processing device and method | |
| CN104396249B (en) | Method and apparatus for inter-layer prediction for scalable video coding | |
| US8767828B2 (en) | System for low resolution power reduction with compressed image | |
| KR101530832B1 (en) | Method and device for optimizing encoding/decoding of compensation offsets for a set of reconstructed samples of an image | |
| EP2437499A1 (en) | Video encoder, video decoder, video encoding method, and video decoding method | |
| US20120300850A1 (en) | Image encoding/decoding apparatus and method | |
| WO2010144408A9 (en) | Digital image compression by adaptive macroblock resolution coding | |
| WO2010144406A1 (en) | Digital image compression by residual decimation | |
| US20250260821A1 (en) | Method and device for processing video signal by using cross-component linear model | |
| EP4289139A1 (en) | Metadata for signaling information representative of an energy consumption of a decoding process | |
| EP2594074A1 (en) | Video decoder for low resolution power reduction using low resolution data | |
| US9313523B2 (en) | System for low resolution power reduction using deblocking | |
| US20120014445A1 (en) | System for low resolution power reduction using low resolution data | |
| US20120300844A1 (en) | Cascaded motion compensation | |
| WO2012008616A1 (en) | Video decoder for low resolution power reduction using low resolution data | |
| US20120300838A1 (en) | Low resolution intra prediction | |
| GB2498225A (en) | Encoding and Decoding Information Representing Prediction Modes | |
| US20120014447A1 (en) | System for low resolution power reduction with high resolution deblocking | |
| WO2012161345A1 (en) | Video decoder | |
| WO2025163017A1 (en) | Internal chroma format increase | |
| KR20090076019A (en) | Interpolation filter, multi codec decoder and decoding method using the interpolation filter |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 11806943 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 11806943 Country of ref document: EP Kind code of ref document: A1 |