EP4717030A1 - Merge based cross-component prediction and inter prediction with filtering - Google Patents
Merge based cross-component prediction and inter prediction with filteringInfo
- Publication number
- EP4717030A1 EP4717030A1 EP24745869.8A EP24745869A EP4717030A1 EP 4717030 A1 EP4717030 A1 EP 4717030A1 EP 24745869 A EP24745869 A EP 24745869A EP 4717030 A1 EP4717030 A1 EP 4717030A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- block
- prediction
- parameters
- blocks
- pixels
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/186—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being a colour or a chrominance component
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Techniques are described for improving merge mode for cross-component prediction and for introducing a merge mode for inter prediction with filtering. Previously coded blocks used to generate (e.g., filter) parameters for each technique can be used in merge mode even where those blocks were not encoded using cross-component prediction or inter prediction with filtering. Techniques are also described that improve cross-component prediction, whether merge mode is used to not, by allowing sub-block determination of parameters, which can improve compression efficiency with minimal increases in signaling cost.
Description
MERGE BASED CROSS-COMPONENT PREDICTION AND INTER PREDICTION
WITH FILTERING
CROSS-REFERENCES TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application Nos. 63/511,882 and 63/511,883, each filed July 4, 2023, and each of which is incorporated herein in its entirety by reference.
BACKGROUND
[0002] Digital video streams may represent video using a sequence of frames or still images. Digital video can be used for various applications including, for example, video conferencing, high-definition video entertainment, video advertisements, or sharing of usergenerated videos. A digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data. Various approaches have been proposed to reduce the amount of data in video streams, including compression and other encoding techniques.
[0003] Encoding and decoding may be performed by breaking frames or images into blocks that are predicted using one or more prediction blocks. Differences (i.e., residual errors) between a block and its prediction block are compressed and encoded in a bitstream. A decoder uses the differences and the prediction block to reconstruct the original data.
SUMMARY
[0004] Disclosed herein are aspects, features, elements, and implementations for encoding and decoding blocks using merge based cross -component prediction and inter prediction with filtering and further allowing sub-block based cross-component prediction, whether merge mode is used or not.
[0005] An aspect of the teachings herein is a method for decoding that includes determining that a current block was encoded using cross-component prediction, wherein the current block has blocks in each of a first color plane of data and a second color plane of data, and cross-component prediction determines a prediction block for a block in the second color plane using pixels from a corresponding block in the first color plane. The method also
includes determining parameters for cross-component prediction from a preceding block to the current block in a coding order, wherein the preceding block was encoded using a coding technique other than cross-component prediction, performing cross-component prediction for the current block using the parameters from the preceding block to obtain a prediction block, and reconstructing the current block using the prediction block.
[0006] In some implementations of this method, the parameters are determined using a prediction signal comprising reconstructed pixels of the preceding block from each of the first color plane and the second color plane. In some implementations, the parameters are determined using a prediction signal comprising prediction pixels of the preceding block from each of the first color plane and the second color plane. In some implementations, the parameters are determined using a prediction signal comprising residual pixels of the preceding block from each of the first color plane and the second color plane. In some implementations, the parameters are determined using any combination of these prediction signals.
[0007] In some implementations of this method, the preceding block is determined from a merge list of (e.g., candidate) blocks of the current block. The merge list may be one of multiple available merge lists, where each merge list is based on the prediction signal used to determine the parameters. Alternatively, or additionally, each merge list may be based on a prediction mode used for encoding candidate blocks for the multiple available merge lists. [0008] In some implementations of this method, blocks of the current frame preceding the current block in the coding order that have a block size smaller than a defined block size may be omitted from the merge list.
[0009] In some implementations of this method, the parameters may be determined using fewer than all pixels of a prediction block for the preceding block in the first color plane and fewer than all pixels of a prediction block for the preceding block in the second color plane. For example, the fewer than all pixels can comprise one of a right half of each of the prediction block for the preceding block in the first color plane and the prediction block for the preceding block in the second color plane, or a bottom half of each of the prediction block for the preceding block in the first color plane and the prediction block for the preceding block in the second color plane. In other variations of these implementations, the fewer than all pixels can comprise every other pixel in x and y directions of each of the prediction block for the preceding block in the first color plane and the prediction block for the preceding block in the second color plane.
[0010] In some implementations of this method, responsive to a block size of the
preceding block being larger than a defined block size, the parameters are determined from one or more sub-blocks of the preceding block having the defined sub-block size.
[0011] An aspect of the teachings herein is a method for decoding that includes determining that a current block was encoded using cross-component prediction, determining that the current block should be split into sub-blocks, determining respective parameters for applying cross -component prediction to the sub-blocks of the current block, reconstructing the sub-blocks using the respective parameters, and reconstructing the current block using the sub-blocks as reconstructed.
[0012] In some implementations of this method, determining that the current block should be split into sub-blocks includes determining that a block size of the current block is larger than a defined sub-block size. The defined sub-block size can be a fixed block size or a defined fraction of the block resulting in sub-blocks no smaller than the fixed block size. [0013] In some implementations of this method, the respective parameters may be determined as parameters for each sub-block based on one of a block coded previously to the current block within the current frame or a coded sub-block of the current block.
[0014] In some implementations of this method, multiple cross-component prediction models are available, cross -component prediction used by the sub-blocks is the same crosscomponent prediction model, and parameters for each of the sub-blocks are allowed to be different.
[0015] In some implementations of this method, the current block can be a subsequent block after the current block in the coding order that was decoded according to the first method described above.
[0016] An aspect of the teachings herein is a method for decoding that includes determining that a current block was encoded using inter prediction with filtering, wherein inter prediction with filtering determines a prediction block for a block by identifying an intermediate prediction block for the block using a motion vector and a reference frame and filtering the intermediate prediction block from pixel positions peripheral to the block and pixel positions peripheral to the intermediate prediction block. The method also includes determining parameters for inter prediction with filtering from a preceding block to the current block in a coding order, wherein the preceding block was encoded using a coding technique other than inter prediction with filtering, performing inter prediction with filtering using the parameters from the preceding block to obtain a prediction block, and reconstructing the current block using the prediction block.
[0017] In some implementations of this method, the parameters are determined using a
prediction signal including reconstructed pixels peripheral to the preceding block and reconstructed pixels peripheral to a prediction block for the preceding block. In some implementations of this method, the parameters are determined using a prediction signal including prediction pixels peripheral to the preceding block and prediction pixels peripheral to a prediction block for the preceding block. In some implementations of this method, the parameters are determined using a prediction signal including residual pixels peripheral to the preceding block and residual pixels peripheral to a prediction block for the preceding block. Any combination of these prediction signals may be used to determine the parameters.
[0018] In some implementations of this method, the preceding block may be determined a merge list of blocks of the current frame. The merge list may be one of multiple available merge lists. Each merge list may be based on the prediction signal used to determine the parameters, a prediction mode used for encoding candidate blocks for the multiple available merge lists, or some combination thereof.
[0019] In some implementations of this method, blocks of the current frame preceding the current block in the coding order that have a block size smaller than a defined block size may be omitted from the merge list.
[0020] An aspect of the teachings herein includes an apparatus for decoding including a processor configured to perform any of the methods described above.
[0021] An aspect of the teachings herein includes a computer-readable storage medium storing instructions that, when executed, cause a processor to perform any of the methods described above.
[0022] As aspect of the teachings herein includes a computer-readable storage medium storing a bitstream that, when provided to a decoder, results in the decoder executing anyh of the methods described above.
[0023] Variations in and further details of these methods, techniques, and apparatuses are found in the drawings, description, and claims that follow.
BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The description herein makes reference to the accompanying drawings described below, wherein like reference numerals refer to like parts throughout the several views. [0025] FIG. 1 is a schematic of a video encoding and decoding system.
[0026] FIG. 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
[0027] FIG. 3 is a diagram of a video stream to be encoded and subsequently decoded.
[0028] FIG. 4 is a block diagram of an encoder according to implementations of this disclosure.
[0029] FIG. 5 is a block diagram of a decoder according to implementations of this disclosure.
[0030] FIG. 6 is a diagram illustrating an example of a cross-component linear model prediction mode.
[0031] FIG. 7 is an example of a filter used with the teachings herein.
[0032] FIG. 8 is a diagram of an example of a template for explaining prediction using the filter according to FIG. 7.
[0033] FIG. 9 is a diagram used to explain a cross-component residual model prediction mode.
[0034] FIG. 10 is a diagram used to explain the determination of filter coefficients of a cross-component residual model according to FIG. 9.
[0035] FIG. 11 is a flowchart diagram of a technique for decoding a current block using merge based cross-component prediction or inter prediction with filtering.
[0036] FIG. 12 is a flowchart diagram of a technique for decoding a current block using sub-block based cross-component prediction.
DETAILED DESCRIPTION
[0037] As mentioned above, compression schemes related to coding video streams may include breaking images into blocks and generating a digital video output bitstream (i.e., an encoded bitstream) using one or more techniques to limit the information included in the output bitstream. A received bitstream can be decoded to re-create the blocks and the source images from the limited information. Encoding a video stream, or a portion thereof, such as a frame or a block, can include using temporal similarities in the video stream to improve coding efficiency. For example, a current block of a video stream may be encoded based on identifying a difference (residual) between the previously coded pixel values, or between a combination of previously coded pixel values, and those in the current block.
[0038] Coding using temporal similarities is known as inter prediction. Inter prediction attempts to predict the pixel values of a block using a possibly displaced block or blocks from a temporally nearby frame (i.e., reference frame) or frames. A temporally nearby frame is a frame that appears earlier or later in time in the video stream than the frame of the block being encoded. Inter prediction can be performed using a motion vector that represents translational motion, i.e., pixel shifts of a prediction block in a reference frame in the x- and
y-axes as compared to the block being predicted.
[0039] Coding using spatial similarities is known as intra prediction. Intra prediction attempts to predict the pixel values of a block using nearby pixels. An intra-prediction mode defines how the nearby pixels are used to form a prediction block.
[0040] A variation of intra prediction used to reduce redundancies among different components generates filter parameters that, when applied to reconstructed luma samples, are used to predict chroma samples within the same coding unit. Techniques described herein allow this intra prediction to use sub-blocks of a coding unit. Further, filters generated according to the teachings herein can be used in a merge mode to improve coding efficiency in both inter prediction and intra prediction. Further details of sub-block based crosscomponent prediction and merge based techniques to improve cross-component prediction and inter prediction with filtering are described herein with initial reference to a system in which they can be implemented.
[0041] FIG. 1 is a schematic of a video encoding and decoding system 100. A transmitting station 102 can be, for example, a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.
[0042] A network 104 can connect the transmitting station 102 and a receiving station 106 for encoding and decoding of the video stream. Specifically, the video stream can be encoded in the transmitting station 102, and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), wide area network (WAN), virtual private network (VPN), cellular telephone network, or any other means of transferring the video stream from the transmitting station 102 to, in this example, the receiving station 106.
[0043] The receiving station 106, in one example, can be a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
[0044] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit the network 104. In another implementation, a video stream can be encoded and then stored for transmission at a later time to the receiving station 106 or any other device having memory. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and/or some
communication pathway) the encoded video stream and stores the video stream for later decoding. In an example implementation, a real-time transport protocol (RTP) is used for transmission of the encoded video over the network 104. In another implementation, a transport protocol other than RTP may be used, e.g., a video streaming protocol based on the Hypertext Transfer Protocol (HTTP).
[0045] When used in a video conferencing system, for example, the transmitting station 102 and/or the receiving station 106 may include the ability to both encode and decode a video stream as described below. For example, the receiving station 106 could be a video conference participant who receives an encoded video bitstream from a video conference server (e.g., the transmitting station 102) to decode and view and further encodes and transmits his or her own video bitstream to the video conference server for decoding and viewing by other participants.
[0046] FIG. 2 is a block diagram of an example of a computing device 200 that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement one or both of the transmitting station 102 and the receiving station 106 of FIG. 1. The computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of one computing device, for example, a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.
[0047] A CPU 202 in the computing device 200 can be a conventional central processing unit. Alternatively, the CPU 202 can be any other type of device, or multiple devices, capable of manipulating or processing information now existing or hereafter developed. Although the disclosed implementations can be practiced with one processor as shown (e.g., the CPU 202), advantages in speed and efficiency can be achieved by using more than one processor.
[0048] A memory 204 in computing device 200 can be a read only memory (ROM) device or a random-access memory (RAM) device in an implementation. Any other suitable type of storage device can be used as the memory 204. The memory 204 can include code and data 206 that is accessed by the CPU 202 using a bus 212. The memory 204 can further include an operating system 208 and application programs 210, the application programs 210 including at least one program that permits the CPU 202 to perform the methods described herein. For example, the application programs 210 can include applications 1 through N, which further include a video coding application that performs the techniques described here, such as the techniques for performing inter-prediction of a current block with filtering.
Computing device 200 can also include a secondary storage 214, which can, for example, be
a memory card used with a mobile computing device. Because the video communication sessions may contain a significant amount of information, they can be stored in whole or in part in the secondary storage 214 and loaded into the memory 204 as needed for processing. [0049] The computing device 200 can also include one or more output devices, such as a display 218. The display 218 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 218 can be coupled to the CPU 202 via the bus 212. Other output devices that permit a user to program or otherwise use the computing device 200 can be provided in addition to or as an alternative to the display 218. When the output device is or includes a display, the display can be implemented in various ways, including by a liquid crystal display (LCD), a cathode-ray tube (CRT) display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.
[0050] The computing device 200 can also include or be in communication with an image-sensing device 220, for example, a camera, or any other image-sensing device 220 now existing or hereafter developed that can sense an image such as the image of a user operating the computing device 200. The image-sensing device 220 can be positioned such that it is directed toward the user operating the computing device 200. In an example, the position and optical axis of the image-sensing device 220 can be configured such that the field of vision includes an area that is directly adjacent to the display 218 and from which the display 218 is visible.
[0051] The computing device 200 can also include or be in communication with a soundsensing device 222, for example, a microphone, or any other sound-sensing device now existing or hereafter developed that can sense sounds near the computing device 200. The sound-sensing device 222 can be positioned such that it is directed toward the user operating the computing device 200 and can be configured to receive sounds, for example, speech or other utterances, made by the user while the user operates the computing device 200.
[0052] Although FIG. 2 depicts the CPU 202 and the memory 204 of the computing device 200 as being integrated into one unit, other configurations can be utilized. The operations of the CPU 202 can be distributed across multiple machines (wherein individual machines can have one or more processors) that can be coupled directly or across a local area or other network. The memory 204 can be distributed across multiple machines such as a network-based memory or memory in multiple machines performing the operations of the computing device 200. Although depicted here as one bus, the bus 212 of the computing device 200 can be composed of multiple buses. Further, the secondary storage 214 can be
directly coupled to the other components of the computing device 200 or can be accessed via a network and can comprise an integrated unit such as a memory card or multiple units such as multiple memory cards. The computing device 200 can thus be implemented in a wide variety of configurations.
[0053] FIG. 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes multiple adjacent frames 304. While three frames are depicted as the adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. The adjacent frames 304 can then be further subdivided into individual frames, for example, a frame 306. At the next level, the frame 306 can be divided into a series of planes or segments 308. The segments 308 can be subsets of frames that permit parallel processing, for example. The segments 308 can also be subsets of frames that can separate the video data into separate colors. For example, a frame 306 of color video data can include a luminance plane and two chrominance planes. The segments 308 may be sampled at different resolutions.
[0054] Whether or not the frame 306 is divided into segments 308, the frame 306 may be further subdivided into blocks 310, which can contain data corresponding to, for example, 16x16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more segments 308 of pixel data. The blocks 310 can also be of any other suitable size such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger. Unless otherwise noted, the terms block and macroblock are used interchangeably herein.
[0055] FIG. 4 is a block diagram of an encoder 400 according to implementations of this disclosure. The encoder 400 can be implemented, as described above, in the transmitting station 102, such as by providing a computer software program stored in memory, for example, the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the transmitting station 102 to encode video data in the manner described in FIG. 4. The encoder 400 can also be implemented as specialized hardware included in, for example, the transmitting station 102. In one particularly desirable implementation, the encoder 400 is a hardware encoder.
[0056] The encoder 400 has the following stages to perform the various functions in a forward path (shown by the solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra/inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408. The encoder 400 may also include a reconstruction path (shown by the dotted connection lines) to
reconstruct a frame for encoding of future blocks. In FIG. 4, the encoder 400 has the following stages to perform the various functions in the reconstruction path: a dequantization stage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.
[0057] When the video stream 300 is presented for encoding, respective adjacent frames 304, such as the frame 306, can be processed in units of blocks. At the intra/inter prediction stage 402, respective blocks can be encoded using intra-frame prediction (also called intraprediction) or inter- frame prediction (also called inter-prediction). In any case, a prediction block can be formed. In the case of intra-prediction, a prediction block may be formed from samples in the current frame that have been previously encoded and reconstructed. In the case of inter-prediction, a prediction block may be formed from samples in one or more previously constructed reference frames.
[0058] Next, still referring to FIG. 4, the prediction block can be subtracted from the current block at the intra/inter prediction stage 402 to produce a residual block (also called a residual). The transform stage 404 transforms the residual into transform coefficients in, for example, the frequency domain using block-based transforms. The quantization stage 406 converts the transform coefficients into discrete quantum values, which are referred to as quantized transform coefficients, using a quantizer value or a quantization level. For example, the transform coefficients may be divided by the quantizer value and truncated. The quantized transform coefficients are then entropy encoded by the entropy encoding stage 408. The entropy-encoded coefficients, together with other information used to decode the block (which may include, for example, the type of prediction used, transform type, motion vectors and quantizer value), are then output to the compressed bitstream 420. The compressed bitstream 420 can be formatted using various techniques, such as variable length coding (VLC) or arithmetic coding. The compressed bitstream 420 can also be referred to as an encoded video stream or encoded video bitstream, and the terms will be used interchangeably herein.
[0059] The reconstruction path in FIG. 4 (shown by the dotted connection lines) can be used to ensure that the encoder 400 and a decoder 500 (described below) use the same reference frames to decode the compressed bitstream 420. The reconstruction path performs similar functions to those that take place during the decoding process (described below), including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage
412 to produce a derivative residual block (also called a derivative residual). At the reconstruction stage 414, the prediction block that was predicted at the intra/inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filtering stage 416 can be applied to the reconstructed block to reduce distortion such as blocking artifacts.
[0060] Other variations of the encoder 400 can be used to encode the compressed bitstream 420. For example, a non-transform-based encoder can quantize the residual signal directly without the transform stage 404 for certain blocks or frames. In another implementation, an encoder can have the quantization stage 406 and the dequantization stage 410 combined in a common stage.
[0061] FIG. 5 is a block diagram of a decoder 500 according to implementations of this disclosure. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described in FIG. 5. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.
[0062] The decoder 500, similar to the reconstruction path of the encoder 400 discussed above, includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra/inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512, and a post filtering stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.
[0063] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by the quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derivative residual that can be identical to that created by the inverse transform stage 412 in the encoder 400. Using header information decoded from the compressed bitstream 420, the decoder 500 can use the intra/inter prediction stage 508 to create the same prediction block as was created in the encoder 400, e.g., at the intra/inter prediction stage 402. At the reconstruction stage 510, the prediction
block can be added to the derivative residual to create a reconstructed block. The loop filtering stage 512 can be applied to the reconstructed block to reduce blocking artifacts. [0064] Other filtering can be applied to the reconstructed block. In this example, the post filtering stage 514 is applied to the reconstructed block to reduce blocking distortion or perform other post-processing on a frame, and the result is output as the output video stream 516. The output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein. Other variations of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can produce the output video stream 516 without the post filtering stage 514.
[0065] One intra-prediction mode that may be used in encoding and decoding is a crosscomponent linear model (CCLM) prediction mode. This prediction mode is described in detail in U.S. Patent Publication No. 2022/0272351, which is incorporated herein by reference. In brief, a reconstructed luma block is used to obtain the prediction block for a corresponding chroma block by a linear model. More specifically, chroma samples are predicted using reconstructed luma samples of the same coding unit (which may be referred to as a largest coding unit, a macroblock, a prediction block, or other such nomenclature), herein referred to as the current block. The linear model may be represented by equation (1) below.
Predc(i,j) = a ■ RecL' (i,j) + ft (1)
[0066] In equation (1), Predc(i, f) represents the predicted chroma sample at respective pixel positions (i,j) of an NxN chroma block of the current block corresponding to the 2Nx2N luma block of the current block. Rec^' i ) represents down-sampled reconstructed luma predictions of the luma block. As shown in the example of FIG. 6, down-sampling is used when the case that the chroma samples and the luma samples do not have the same resolution. For example, down-sampling may be performed where a 4:2:2 or a 4:2:0 format is used. The down-sampling aligns the resolution of luma and chroma blocks. If the image or frame is not down-sampled, the reconstructed luma samples may be used without downsampling.
[0067] The cross-component parameters (a and ?) can be derived with at most four neighboring chroma samples and their corresponding down- sampled luma samples. The division operation to calculate parameter a may be implemented with a look-up table. In FIG. 6, the locations of left and above samples and the sample of the current block involved in the cross-component prediction are shown. With respect to a luma block 602, the locations of left
and above samples are shown as filled circles, such as a filled circle 612. With respect to a chroma block 604, the locations of left and above samples are shown as filled circles, such as a filled circle 614.
[0068] Another example of a cross -component prediction uses a convolutional crosscomponent model (CCCM). A 7-tap convolutional filter may be used to obtain the chroma prediction block from a luminance prediction block. The convolutional filter may include a 5- tap plus sign shape spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of a center (C) luma sample that is collocated with the chroma sample (e.g., pixel value) to be predicted and its above/north (N), below/south (S), left/west (W) and right/east (E) neighboring luma samples (e.g., pixel values). FIG. 7 is a filter illustrating positions of these pixels, which are used for predicting chroma pixels according to a convolutional cross-component model prediction mode. The locations indicate the related downsampled luma samples in the event the color format is other than 4:4:4.
[0069] A chroma prediction pixel value can be obtained using equation (2) below, where C, N, S, E, and W correspond to the values of the luma prediction values, such as shown in FIG. 7. predChromaVal = c0C + c N + c2S + c3E + c4W + c5P + c6B (2) [0070] In equation (2), P is the non-linear term. In an example, the non-linear term P is represented as power of two of the center luma sample C and is scaled to the sample value range of the content (e.g., bitDepth) according to equation (3) below, where midVai is the middle chroma value given the bit depth.
P = (C2 + midVal) » bitDepth (3)
[0071] For example, if the content is 10-bit content, then P is calculated as P = (C2 + 512 ) » 10. Other non-linear terms are possible.
[0072] The bias term B in equation (2) can represent a scalar offset between the input and output, which is like the offset term used in the cross -component linear model prediction mode described above. The output of the filter is calculated as a convolution between the filter coefficients ct and the input values and is clipped to the range of valid chroma samples. [0073] The example 800 shown in FIG. 8 may be used to explain the determination of the filter coefficients of the convolutional cross-component model represented by equation (2). In general, the coefficients ct (i.e., c0, c1? c2, c3, c4, c5, c6) can be obtained by minimizing the mean-squared-error (MSE) between predicted and reconstructed chroma samples in a
reference area. In FIG. 8, the block 808 corresponds to a location of a current prediction unit (e.g., a chroma block of pixels to be encoded or decoded), whose pixels are identified with a first pattern 802. While the block 808 is shown as being of size 8x4, the disclosure is not so limited. The block 808 can be of any other size.
[0074] Pixels filled with a second pattern 806 are pixels that define the reference area. The reference area includes 6 lines of chroma samples above and left of the block 808. The reference area extends one block width to the right and one block height below the boundaries of the block 808. The reference area is adjusted to include only available samples. The extensions to the pixels of the reference area identified with a third pattern 804 support the “side samples” of the plus-shaped spatial filter (in this example, pixels identified by N, S, W, and E in the filter of FIG. 7). When a sample is unavailable, the pixel is padded (i.e., is set to a padding value). Depending on the neighborhood used for a filter, for example, one or more pixels needed by the filter may not be available (such as because the pixels are outside the frame boundary or are outside a largest coding unit that includes the block 808). A padding value may be used (e.g., assumed) for such pixels. For other filter shapes and other sizes for the block 808, the reference area may be a different size and/or a different shape.
[0075] The MSE minimization may be performed by calculating an autocorrelation matrix for the luma input and a cross -correlation vector between the luma input and chroma output. The autocorrelation matrix may be decomposed using, e.g., square-root-free Cholesky (LDL) decomposition, and the final filter coefficients are computed using back-substitution. Although the proposed approach uses only integer arithmetic, the LDL decomposition and back-substitution process have a long latency, which is relatively unfriendly for a hardware decoder.
[0076] In some implementations, the convolutional cross-component model prediction mode is considered a sub-mode to the cross-component linear model prediction mode. In such an implementation, a flag (e.g., at the block level) signaling that the convolutional crosscomponent model prediction mode is used can only be signaled if the block has been signaled as coded with a cross-component linear model prediction mode.
[0077] The cross-component linear model prediction mode and the convolutional crosscomponent model prediction mode may be considered intra prediction modes. Crosscomponent prediction for inter predicted blocks is not considered in these models. A crosscomponent residual model prediction mode is described with regards to FIGS. 9 and 10. [0078] Particularly, FIG. 9 shows a reconstruction path 900 for an inter predicted block encoded using the cross-component residual model prediction mode. A residual luma block
resY (e.g., decoded from an encoded bitstream) is added to a luma prediction block predY generated using inter prediction at adder 902. For example, the luma prediction block predY may be generated at the intra/inter prediction stage 508 of the decoder 500. The result is a reconstructed luma block Y.
[0079] Each of the chroma blocks Cb and Cr may be reconstructed using a crosscomponent residual model. In general, chroma samples are predicted from reconstructed luma samples when the block uses inter prediction or intra block copy (IBC). At decoder, the cross-component filters are derived at stage 904 using the prediction signals of the luma and chroma components. The derived filters are applied to the reconstructed luma signal at stage 906 to produce the final chroma predictions. The reconstructed chroma block Cb is then generated by adding the chroma residual block Cb (e.g., decoded from an encoded bitstream) to the chroma prediction for block Cb at adder 908. Similarly, the reconstructed chroma block Cr is generated by adding the chroma residual block Cr (e.g., decoded from an encoded bitstream) to the chroma prediction for block Cr at adder 910.
[0080] The filter derived at stage 904 may use 6 spatial luma samples as shown in FIG. 10, a non-linear term, and a bias term. According to an implementation of the teachings herein, a predicted chroma value predChromaVal may be determined according to equation (4) below. predChromaVal = coL0 + c Ll + c2L2 + c3L3 + c4L4 + c5L5 +
[0081] In equation (4), the filter coefficients c0 - c7 may be determined using a similar technique to those described above or using any known technique. The luma spatial samples L0 - L5 are shown in FIG. 10. The bias value B is similar to the bias terms previously described.
[0082] In some implementations, a cross-component merge mode (also called a merge mode for cross -component prediction) may be used. Instead of deriving the cross-component prediction parameters, including the parameters of the cross-component linear model prediction mode or the convolutional cross-component model prediction mode, the parameters are merged from the adjacent and non-adjacent spatial neighbors of the current block and from blocks that use cross-component prediction. That is, the parameters from another block that uses cross -component prediction can be reused. A merge flag may be encoded into the bitstream to indicate that the current block was encoded using merge mode, and a merge index may be encoded into the bitstream to indicate from which block (e.g., in a
merge list) to obtain the parameters. Merging from temporal neighbors is also possible. [0083] Techniques according to this disclosure may be used also to support merge mode for inter prediction with filtering. Inter prediction with filtering addresses situations where, for example, when source data is noisy, when there are illumination differences between a current block being predicted and a reference frame used to obtain a prediction block, or when the source data includes motion blur, inter-prediction accuracy may suffer.
Consequently, compression performance may also suffer. To address this, a filtered inter prediction block may be used as the final prediction block for encoding or decoding a current block, where the filter coefficients are derived based on neighboring pixel samples according to one of the previously discussed techniques. In an example, the filter coefficients for refining the prediction block identified by a motion vector for the current block (whether a luma block or a chroma block) are obtained using first reconstructed pixels that are peripheral to the current block and second reconstructed pixels that are peripheral to the prediction block in one of the filtering techniques previously described.
[0084] More specifically, FIGS. 7 and 8 may be used also to explain inter prediction with filtering. That is, as described above, FIG. 8 is a diagram of an example 800 of a template for explaining prediction using the filter according to FIG. 7. How FIG. 8 relates to crosscomponent prediction is also described above. How FIG. 8 relates to inter prediction with filtering using the filter according to FIG. 7 is next described.
[0085] In inter prediction with filtering, an intermediate prediction block is first identified for the current block. The current block may be a luminance block (e.g., a Y block) or a chrominance block (e.g., a Cb block, a Cr block, a U block, or a V block). The intermediate prediction block may also be referred to as a reference block. In an example, a motion vector and a reference frame may be identified for the current block. In an example, the reference block and the motion vector may be identified using data obtained from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5. The motion vector and reference may be identified as described with respect to FIG. 5.
[0086] The data obtained from the compressed bitstream can indicate that a motion vector and/or a reference of another block, which may be a temporal or spatial neighboring block to the current block, are to be used for the current block. In such a situation, the current block may be said to be merged with the neighboring block. The intermediate prediction block is the block in the reference frame that is pointed to by the motion vector. As is known, the intermediate prediction block (i.e., the reference block) may be at integer pixel locations or at sub-pixel locations. In the case that the intermediate prediction block is at sub-pixel locations,
and as is known, interpolation filtering may be performed to obtain values at the sub-pixels. [0087] The example 800 is used to illustrate both a current block template and a reference block template.
[0088] When used to describe a current block template, a block 808 illustrates a current block (i.e., the block being decoded). While the block 808 is shown as being of size 8x4, the disclosure is not so limited. The block 808 can be of any other size. Pixels filled with a pattern 802 are pixels of the current block of a current frame. Pixels filled with a pattern 806 (or a subset thereof, as further described herein) illustrate (e.g., reconstructed) pixels of the current frame. Pixels filled with a pattern 804 are pixels that are not available and may contain a padding value (i.e., are set to a padding value). Depending on the neighborhood used for a filter, one or more pixels used by the filter may not be available (such as because these pixels are outside the frame boundary or are outside a largest coding unit that includes the block 808). As such, a padding value may be used (e.g., assumed) for such pixels if needed in the processing described herein.
[0089] When used to describe a reference block template, the block 808 illustrates a reference block in a reference frame. Again, while the block 808 is shown as being of size 8x4, the disclosure is not so limited. The block 808 can be of a size corresponding to the size of the current block. As such, pixels filled with the pattern 802 are pixels of the reference block of a reference frame. Pixels filled with the pattern 806 (or a subset thereof, as further described herein) illustrate (e.g., reconstructed) pixels of the reference frame. Pixels filled with the pattern 804 are pixels that are not available and may contain a padding value.
[0090] The template may include a top region 810 that may include 1 to N (where N>1) rows of pixels. The template may include a top-right region 812 that includes 1 to N rows. The template may include a left region 814 of 1 to M (where M>1) columns of pixels. The template may include a bottom- left region 816 of 1 to M (where M>1) columns of pixels. [0091] In an example, N=M. In an example, if the current block is a luma block, then the template can be 4-sample wide. If the current block is a chroma block, the template (i.e., a chroma template) may be based on the chroma color format. For example, for 4:4:4 content, the chroma template can also be 4-sample wide; and for 4:2:0 or 4:2:2 color formats, the chroma template can be 2-sample wide. In an example, when the top-right region 812 is available, only a 4x4 luma block at the top-right is included in the template. Similarly, if the bottom-left region 816 is available, only a 4x4 luma block at bottom-right is included in the template. The chroma template can be adjusted accordingly based on the chroma color format. In another example, the top template may always be 1 -sample wide for both luma and
chroma, while the left template may be 4-sample wide for luma.
[0092] Like the parameters described above, the parameters in the form of filter coefficients can include at least two coefficients. In an example, the filter coefficients include more than two coefficients for at least one of the color components (i.e., at least of the luma or the chroma component). In an example, the number (i.e., cardinality) of the filter coefficients can be decoded from the compressed bitstream. For example, an indicator of the number of filter coefficients can be decoded from the compressed bitstream. For example, in response to the indicator of the number of the filter coefficients being a first value (e.g., 0), inter prediction with filtering may be omitted for the current block. That is, if the indicator of the number of coefficients is the first value, then no filtering may be performed on the prediction block. If the indicator of the number of the filter coefficients is a second value (e.g., 1), then two fdter coefficients are derived; and if the indicator of the number of fdter coefficients is a third value (e.g., 2), then more than two filter coefficients are derived. This signaling is by example only, and other examples are possible.
[0093] When the indicator of the number of filter coefficients is two (e.g., when the number of the filter coefficients is the second value), the two filter coefficients may be obtained using a technique known as Local Illumination Compensation (LIC). LIC is described in U.S. Patent Publication No. 2021/0352309, which is incorporated herein by reference. Briefly, LIC is an inter prediction technique to model local illumination variations between a current block and its prediction block as a function of illumination between a current block template and a reference block template. The parameters of the function can be denoted by a scale a and an offset p , therewith forming a linear equation: a x p[x] + p to compensate for illumination changes, where p[x] is a reference sample pointed to by a motion vector (MV) at a location x in the reference frame. Because a and p can be derived based on the current block template and the reference block template, no signaling overhead is required for them. That is, an encoder need not encode, and the decoder need not decode, values for the parameters a and p .
[0094] In an example, the filter can be a convolutional filter. The filter coefficients can be obtained by minimizing an error metric between the first pixels of a block and the second pixels of its reference block. The error metric can be a mean square error (MSE) between pixel values of the respective pixels. The error can be a sum of absolute differences (SAD) error between the pixel values of the pixels. Any other suitable error metric can be used. [0095] In an example, the number of coefficients to be obtained depends on which pixels
within the neighborhood of the intermediate prediction pixel to which the filter is to be applied are used in the filtering. The pixels within the neighborhood of an intermediate prediction pixel that are used for filtering are referred to herein as at least a subset of pixels of the neighborhood. FIG. 7 illustrates a filter for this process, particularly illustrating the pixel values to which the parameters (also referred to as the filter coefficients) are both derived and applied.
[0096] The filter defines a neighborhood for filtering about an intermediate pixel (here, center pixel C) of an intermediate prediction block. The neighborhood is a 3x3 neighborhood. However, the neighborhood can be larger or smaller, rectangular, or some other shape (e.g., diamond). The example illustrates that pixels N, S, E, and W are used in the filtering. Accordingly, the filter is a 5-tap filter, and the prediction pixel corresponding to the intermediate pixel C can be obtained using equation the equation below, where Ci (for z = 0, ... ,4) are the filter coefficients, pred is the filtered pixel of the final prediction block. A constant term (i.e., cs), which may also be considered a parameter for the inter prediction with filtering), can be derived and used in some implementations. pred = cQC + c N + c2S + c3E + c4W + c5
[0097] In an example, one or more but not all filter coefficients may be further refined after being derived. Thus, the filter coefficients obtained as described above may be considered predicted filter coefficients. The difference (i.e., a coefficient refinement value) between a predicted filter coefficient and the actual (e.g., refined) value of the filter coefficient may be signaled in the compressed bitstream. Accoredingly, obtaining the filter coefficients for the filter can include obtaining a predicted filter coefficient for a filter coefficient of the filter coefficients as described above, decoding, from the compressed bitstream, a coefficient refinement value, and adjusting the predicted filter coefficient using the coefficient refinement value to obtain the filter coefficient.
[0098] In an example, the coefficient refinement value corresponds (i.e., is used for) the intermediate prediction pixel itself. That is, for example, the coefficient refinement value may be used to refine the filter coefficient obtained for the intermediate pixel C of FIG. 7. In an example, the coefficient refinement value can be used to refine a coefficient corresponding to a non-linear term of the filter. For example, the filter may include a filter coefficient corresponding to the intermediate prediction pixel, one non-linear term, and a constant value. The non-linear term (i.e., a non-linear component) can be a square term of the intermediate prediction pixel. Accordingly, the filter can be given by a X p[x]2 + b X p[x] + c, where a
and b are the filter coefficients, c is a constant component, and p[ ] is the value of the intermediate prediction pixel at location x.
[0099] In an example, and as described with respect to FIG. 7, the filter coefficients can be applied to at least a subset of pixels in a 3x3 neighborhood of respective intermediate prediction pixels to obtain the prediction pixel of the final prediction block for the pixel location of the intermediate prediction pixel. A 3x3 neighborhood can be used whether the current block is a luma block or a chroma block. The subset of the pixels can form (e.g., can be of any) shape. In an example, the subset of the pixels in a 3x3 neighborhood can be those pixels that form a cross shape, such as shown in FIG. 7.
[0100] In an example, the filter can use at least a subset of pixels in a 3x3 neighborhood and may further include a constant term (also referred to as a DC value). In an example, one shape (i.e., a first filter shape) may be used for a current block that is a luma block, and a second filter shape may be used for a current block that is a chroma block. The filter so developed can be used to filter the intermediate prediction block to determine or obtain a final prediction block. The current block is reconstructed using the final prediction block by, for example, decoding a residual block from a compressed bitstream and adding the final prediction block to obtain the current block.
[0101] Referring back to merge-based cross -component prediction, the process may be efficient. However, there may be an insufficient number of blocks to merge to. In other words, there may be too few blocks encoded using one of the cross -component prediction techniques described above to find a good match for the current block. Further, there is currently no merge mechanism for the inter prediction with filtering.
[0102] The teachings herein derive cross-component prediction parameters and/or inter prediction with filtering parameters for a block after coding the block using a technique other than cross -component prediction and/or inter prediction with filtering. Then, the parameters can be stored for use by later blocks, such as via merge mode. While cross-component prediction and inter prediction with filtering are used as examples, the teachings herein can be extended to other techniques that would benefit from deriving parameters based on reconstructed or prediction samples from previously coded blocks.
[0103] FIG. 11 is a flowchart diagram of a method or technique 1100 for decoding a current block using merge based cross -component prediction and inter prediction with filtering. The technique 1100 can be implemented in a decoder such as the decoder 500 of FIG. 5 or in the reconstruction path of FIG. 4. The technique 1100 can be implemented, for example, as a software program that can be executed by computing devices such as
transmitting station 102 or the receiving station 106 of FIG. 1. The software program can include machine-readable instructions (e.g., executable instructions) that can be stored in a memory such as the memory 204 or the secondary storage 214, and that can be executed by a processor, such as CPU 202, to cause the computing device to perform the technique 1100. In at least some implementations, the technique 1100 can be performed in whole or in part by the intra/inter prediction stage 508 of the decoder 500 of FIG. 5.
[0104] The technique 1100 can be implemented using specialized hardware or firmware. Some computing devices can have multiple memories, multiple processors, or both. The steps or operations of the technique 1100 can be distributed using different processors, memories, or both. Use of the term “processor” or “memory” in the singular encompasses computing devices that have one processor or one memory as well as devices that have multiple processors or multiple memories that can be used in the performance of some or all of the recited steps.
[0105] At 1102, the technique 1100 determines whether the current block was encoded using cross-component prediction or inter prediction with filtering. The cross -component prediction may be any cross -component prediction mode such as CCLM or CCCM described herein. The inter prediction with filtering may be one of the variations of the inter prediction filtering described herein or a similar technique. Whether the current block was so encoded may be determined from one or more flags decoded from the bitstream. In some implementations where the cross-component prediction was used, a first flag indicates that one of CCLM or CCCM is used, and a subsequent flag indicates that the CCCM mode is used. Other signaling techniques may be used depending upon how many cross-component prediction modes are available for cross -component prediction. In some implementations where inter prediction with filtering is used, and as described above, one or more flags may indicate use of the inter prediction with filtering and, optionally, whether motion refinement was or should be applied to the prediction block. If the current block was not encoded using cross-component prediction or inter prediction with filtering, the technique 1100 ends.
[0106] At 1104, whether the current block was encoded using merge mode is determined. Whether the current block was encoded using merge mode is indicated by, for example, a merge flag decoded from the bitstream, such as in a header associated with the current block. If the current block was not encoded using merge mode, the technique 1100 ends.
[0107] If the current block was encoded using merge mode, the technique 1100 identifies which preceding block to use for the merge mode at 1106. For example, a merge index may be included (and then decoded) within the bitstream that identifies one of the preceding
blocks from a merge list for the current block, where the merge list comprises a list of the preceding blocks available for coding the current block using merge mode.
[0108] According to the teachings herein, the preceding block was encoded using a coding technique other than cross-component prediction or inter prediction with filtering. That is, for example, where the current block was encoded using a cross-component prediction mode, the preceding block was encoded using another intra prediction mode or an inter prediction mode. Similarly, where the current block was encoded using inter prediction with filtering, the preceding block was encoded using another inter prediction mode or an intra prediction mode.
[0109] At 1108, the parameters for the cross-component prediction or inter prediction with filtering, whichever is applicable to the current block, are determined from the preceding block. In practice, the parameters are generally determined after reconstruction of the preceding block and are stored for usage in association with the merge list for the preceding block. That is, determining the preceding block (e.g., from a merge list) at 1106 also determines the parameters at 1108.
[0110] Accordingly, multiple candidate blocks preceding the current block may be used to form the merge list. For each of the candidate blocks, the parameters may be determined. Where a candidate block was predicted using cross-component prediction or inter prediction with filtering, for example, the (e.g., filter) parameters for the candidate block are already known from the reconstruction of that block. For candidate blocks encoded using other coding techniques, the parameters for a candidate block for the merge list may be determined or derived according to any of the techniques described herein. In some implementations, a candidate block may be used to derive parameters for both inter prediction with filtering and one or more cross-component prediction modes.
[0111] In the examples described above, the parameters are derived based a prediction signal comprising fully reconstructed luma and/or chroma samples of a preceding block and/or its prediction block. This is not required. For example, in some implementations, the parameters are derived from luma and/or chroma prediction samples, such as from an intra prediction mode, and inter prediction mode, an IBC prediction, etc., that was used for predicting the preceding block. In some implementations, the parameters are derived from luma and/or chroma residuals used to reconstruct the preceding block. The parameters can be determined from any combination of these variables.
[0112] In some implementations, different merge lists may be constructed and used. For example, the different merge lists may be constructed based on the type of prediction signal,
e.g., reconstructed pixels, prediction pixels, residual pixels, etc., from which the parameters are derived. As another example, different merge lists may be constructed based on the coding mode of a candidate block. For example, a candidate block preceding the current block may be included in one merge list if the block was encoded using an intra prediction mode or may be included in another merge list if the block was encoded using an inter prediction mode. Merge lists may be determined based on particular intra prediction modes and/or particular inter prediction modes.
[0113] In some implementations, some but not all the samples of a preceding block are used to derive the parameters. For example, the right half or the bottom half of the preceding block is used. In another example, every other sample in the x- and y-directions of the preceding block are used to derive (determine, identify, etc.) the parameters for a candidate block.
[0114] In another embodiment, the parameter derivation is at a defined block size. For a preceding block larger than the defined size, the block may be split into sub-blocks based on the defined sub-block size, and the parameters may be derived for each sub-block. For a preceding block with a smaller size, the parameters may be derived based on that block.
Alternatively, no parameters may be derived for a preceding block with the smaller size (e.g., a size below the defined sub-block size).
[0115] The defined sub-block size, when used, may be predefined at both encoder and decoder. Alternatively, the defined sub-block size may be signaled in the bitstreams, such as in sequence parameter set, picture parameter set, picture header, slice header, etc.
[0116] Finally, at 1110, the current block is reconstructed using the parameters from the preceding block. For example, where the parameters are filter parameters (e.g., filter coefficients and any non-linear or offset terms), the filter is applied to refine an inter prediction or to generate a cross-component prediction as described previously, and the block is reconstructed by adding the prediction to a residual decoded from the bitstream.
[0117] Referring specifically to cross -component prediction, the technique is normally more accurate for small blocks than large block. However, the signaling cost for small blocks is higher than for large blocks. For example, if cross -component prediction is used for small blocks, the prediction mode may be signaled for each of the blocks. As a result, crosscomponent prediction has been generally limited to a coding unit size. Stated another way, cross-component prediction has not been used with sub-blocks of the prediction block.
[0118] The teachings herein describe a sub-block based cross-component prediction. With this mode, cross-component prediction may be performed at the sub-block level. If a
current block is larger than the defined sub-block size, the block is split into sub-blocks (e.g., with the defined sub-block size). Then cross-component prediction is made for each subblock in a predefined order or signaled order within the block, such as in a raster scan order. The cross-component prediction parameters described above may be derived for each subblock based on the reconstruction of previously coded blocks, optionally including the coded sub-blocks in the current block. The sub-blocks may use the merge based cross-component prediction described above.
[0119] FIG. 12 is a flowchart diagram of a method or technique 1200 for decoding a current block using sub-block based cross -component prediction. The technique 1200 can be implemented in a decoder such as the decoder 500 of FIG. 5 or in the reconstruction path of FIG. 4. The technique 1200 can be implemented, for example, as a software program that can be executed by computing devices such as transmitting station 102 or the receiving station 106 of FIG. 1. The software program can include machine-readable instructions (e.g., executable instructions) that can be stored in a memory such as the memory 204 or the secondary storage 214, and that can be executed by a processor, such as CPU 202, to cause the computing device to perform the technique 1200. In at least some implementations, the technique 1200 can be performed in whole or in part by the intra/inter prediction stage 508 of the decoder 500 of FIG. 5.
[0120] The technique 1200 can be implemented using specialized hardware or firmware. Some computing devices can have multiple memories, multiple processors, or both. The steps or operations of the technique 1200 can be distributed using different processors, memories, or both. Use of the term “processor” or “memory” in the singular encompasses computing devices that have one processor or one memory as well as devices that have multiple processors or multiple memories that can be used in the performance of some or all of the recited steps.
[0121] At 1202, the technique 1200 determines whether the current block was encoded using cross-component prediction. The cross-component prediction may be any crosscomponent prediction mode such as CCLM or CCCM described herein. This may be determined from one or more flags decoded from the bitstream. In some implementations, a first flag indicates that one of CCLM or CCCM is used, and a subsequent flag indicates that the CCCM mode is used. Other signaling techniques may be used depending upon how many cross-component prediction modes are available for cross-component prediction. If the current block was not encoded using cross-component prediction, the technique 1200 ends. [0122] At 1204, the technique 1200 determines whether the current block is to be split
into sub-blocks. If not, the technique 1200 ends.
[0123] In some implementations, whether the current block is to be split into sub-blocks is determined by the decoder by examining the block size of the current block as compared to a defined sub-block size.
[0124] In some implementations, the defined sub-block size is predefined at both an encoder and a decoder. Alternatively, the defined sub-block size is signaled in the bitstream, such as in sequence parameter set, picture parameter set, picture header, slice header, etc. [0125] In an example, the defined sub-block size is fixed to 8x8 chroma samples. In this case, a 16x16 chroma block may be used to determine up to 4 sets of parameters for the cross-component prediction.
[0126] In some implementations, the desired sub-block size is always set to a certain ratio of the current block unless the current block is too small to be further split. For example, the resulting sub-block size may be a quarter (half in both vertical and horizontal directions) of the current block.
[0127] In some implementations, a threshold (e.g., different from the defined sub-block size) is predefined or signaled in the bitstream to indicate whether a block is split into subblocks when using cross-component prediction. In examples where the width of a block is larger than the threshold while the height is smaller than the threshold, the block may be split only in one direction.
[0128] At 1206, the parameters for cross-component prediction are determined at the subblock level. For example, filter parameters may be derived for a sub-block, and the filter is applied to generate a prediction block (e.g., a chroma prediction block) for the chroma block using the luma block of the current block as described previously. In some implementations, all sub-blocks of a block share the same model for cross-component prediction though the parameters of the model may be different for each sub-block. In other implementations, different cross-component prediction models may be used for the sub-blocks.
[0129] Finally, at 1208, the current block is reconstructed using the sub-blocks reconstructed using the parameters. That is, each of the sub-blocks is reconstructed and combined to reconstruct the current block. For example, a sub-block may be reconstructed by adding the prediction pixels (e.g., a prediction block resulting from the cross-component prediction mode using the parameters) to residual pixels (e.g., a residual block) decoded (e.g., by dequantization and inverse transformation). In some implementations, a transform applies to each sub-block, i.e., each sub-block is a transform unit. In some implementations, all subblocks of a block share the same transform type. In some implementations, residuals of all
sub-blocks of this mode are enforced to be zero such that the chroma block is the same as that derived from the cross-component prediction.
[0130] According to the teachings herein, cross-component prediction can be signaled once for prediction units above a defined sub-block size. Thereafter, the block can be coded at the sub-block level using cross-component prediction. As a result, a more accurate prediction can be achieved with minimal or no increase in signaled bits.
[0131] For simplicity of explanation, the techniques herein are depicted and described as respective series of steps or operations. However, the steps or operations in accordance with this disclosure can occur in various orders and/or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, not all illustrated steps or operations may be required to implement a method in accordance with the disclosed subject matter.
[0132] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it is to be understood that encoding and decoding, as those terms are used in the claims, could mean compression, decompression, transformation, or any other processing or change of data.
[0133] The word “example” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” is not necessarily to be construed as being preferred or advantageous over other aspects or designs. Rather, use of the word “example” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise or clearly indicated otherwise by the context, the statement “X includes A or B” is intended to mean any of the natural inclusive permutations thereof. That is, if X includes A; X includes B; or X includes both A and B, then “X includes A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more,” unless specified otherwise or clearly indicated by the context to be directed to a singular form. Moreover, use of the term “an implementation” or the term “one implementation” throughout this disclosure is not intended to mean the same embodiment or implementation unless described as such.
[0134] Implementations of the transmitting station 102 and/or the receiving station 106 (and the algorithms, methods, instructions, etc., stored thereon and/or executed thereby, including by the encoder 400 and the decoder 500) can be realized in hardware, software, or any combination thereof. The hardware can include, for example, computers, intellectual
property (IP) cores, application- specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, microcontrollers, servers, microprocessors, digital signal processors, or any other suitable circuit. In the claims, the term “processor” should be understood as encompassing any of the foregoing hardware, either singly or in combination. The terms “signal” and “data” are used interchangeably. Further, portions of the transmitting station 102 and the receiving station 106 do not necessarily have to be implemented in the same manner.
[0135] Further, in one aspect, for example, the transmitting station 102 or the receiving station 106 can be implemented using a general-purpose computer or a general-purpose processor with a computer program that, when executed, carries out any of the respective methods, algorithms, and/or instructions described herein. In addition, or alternatively, for example, a special purpose computer/processor can be utilized which can contain other hardware for carrying out any of the methods, algorithms, or instructions described herein. [0136] The transmitting station 102 and the receiving station 106 can, for example, be implemented on computers in a video conferencing system. Alternatively, the transmitting station 102 can be implemented on a server, and the receiving station 106 can be implemented on a device separate from the server, such as a handheld communications device. In this instance, the transmitting station 102, using an encoder 400, can encode content into an encoded video signal and transmit the encoded video signal to the communications device. In turn, the communications device can then decode the encoded video signal using a decoder 500. Alternatively, the communications device can decode content stored locally on the communications device, for example, content that was not transmitted by the transmitting station 102. Other suitable transmitting and receiving implementation schemes are available. For example, the receiving station 106 can be a generally stationary personal computer rather than a portable communications device, and/or a device including an encoder 400 may also include a decoder 500.
[0137] Further, all or a portion of implementations of the present disclosure can take the form of a computer program product accessible from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can, for example, tangibly contain, store, communicate, or transport the program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable mediums are also available.
[0138] The above-described embodiments, implementations, and aspects have been
- l-
described to facilitate easy understanding of this disclosure and do not limit this disclosure. On the contrary, this disclosure is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope is to be accorded the broadest interpretation as is permitted under the law to encompass all such modifications and equivalent arrangements.
Claims
1. A method for decoding, comprising: determining that a current block was encoded using cross-component prediction, wherein the current block comprises blocks in each of a first color plane of data and a second color plane of data, and cross -component prediction determines a prediction block for a block in the second color plane using pixels from a corresponding block in the first color plane; determining parameters for cross-component prediction from a preceding block to the current block in a coding order, wherein the preceding block was encoded using a coding technique other than cross-component prediction; performing cross-component prediction for the current block using the parameters from the preceding block to obtain a prediction block; and reconstructing the current block using the prediction block.
2. The method of claim 1, wherein the parameters are determined using a prediction signal comprising reconstructed pixels of the preceding block from each of the first color plane and the second color plane.
3. The method of one of claim 1 or claim 2, wherein the parameters are determined using a prediction signal comprising prediction pixels of the preceding block from each of the first color plane and the second color plane.
4. The method of any one of claims 1 to 3, wherein the parameters are determined using a prediction signal comprising residual pixels of the preceding block from each of the first color plane and the second color plane.
5. The method of any one of claims 1 to 4, comprising: determining the preceding block from a merge list of blocks of the current frame.
6. The method of claim 5, wherein the merge list is one of multiple available merge lists, and each merge list is based on the prediction signal used to determine the parameters.
7. The method of claim 5, wherein the merge list is one of multiple available
merge lists, and each merge list is based on a prediction mode used for encoding candidate blocks for the multiple available merge lists.
8. The method of any one of claims 5 to 7, wherein blocks of the current frame preceding the current block in the coding order that have a block size smaller than a defined block size are omitted from the merge list.
9. The method of any one of claims 1 to 8, wherein determining the parameters comprises determining the parameters using fewer than all pixels of a prediction block for the preceding block in the first color plane and fewer than all pixels of a prediction block for the preceding block in the second color plane.
10. The method of claim 9, wherein the fewer than all pixels comprise one of a right half of each of the prediction block for the preceding block in the first color plane and the prediction block for the preceding block in the second color plane, or a bottom half of each of the prediction block for the preceding block in the first color plane and the prediction block for the preceding block in the second color plane.
11. The method of claim 9, wherein the fewer than all pixels comprise every other pixel in x and y directions of each of the prediction block for the preceding block in the first color plane and the prediction block for the preceding block in the second color plane.
12. The method of any one of claims 1 to 11, wherein, responsive to a block size of the preceding block being larger than a defined sub-block size, determining the parameters comprises determining the parameters from one or more sub-blocks of the preceding block having the defined sub-block size.
13. The method of claim 1, comprising, for a subsequent block after the current block in the coding order: determining that the subsequent block was encoded using cross -component prediction; determining that the subsequent block should be split into sub-blocks; determining respective parameters for applying cross -component prediction to the sub-blocks of the subsequent block;
reconstructing the sub-blocks using the respective parameters; and reconstructing the subsequent block using the sub-blocks as reconstructed.
14. The method of claim 13, wherein determining that the subsequent block should be split into sub-blocks comprises determining that a block size of the subsequent block is larger than a defined sub-block size.
15. The method of claim 14, wherein the defined sub-block size comprises one of a fixed block size or a defined fraction of the block resulting in sub-blocks no smaller than the fixed block size.
16. The method of any one of claims 13 to 15, wherein determining the respective parameters comprises determining parameters for each sub-block based on one of a block coded previously to the subsequent block within the current frame or a coded sub-block of the subsequent block.
17. The method of any one of claims 13 to 16, wherein multiple cross -component prediction models are available, cross-component prediction used by the sub-blocks comprises a same cross-component prediction model, and parameters for each of the subblocks are allowed to be different.
18. A method for decoding, comprising: determining that a current block was encoded using inter prediction with filtering, wherein inter prediction with filtering determines a prediction block for a block by identifying an intermediate prediction block for the block using a motion vector and a reference frame and filtering the intermediate prediction block from pixel positions peripheral to the block and pixel positions peripheral to the intermediate prediction block; determining parameters for inter prediction with filtering from a preceding block to the current block in a coding order, wherein the preceding block was encoded using a coding technique other than inter prediction with filtering; performing inter prediction with filtering using the parameters from the preceding block to obtain a prediction block; and reconstructing the current block using the prediction block.
19. The method of claim 18, wherein the parameters are determined using a prediction signal comprising reconstructed pixels peripheral to the preceding block and reconstructed pixels peripheral to a prediction block for the preceding block.
20. The method of one of claim 18 or claim 19, wherein the parameters are determined using a prediction signal comprising prediction pixels peripheral to the preceding block and prediction pixels peripheral to a prediction block for the preceding block.
21. The method of any one of claims 18 to 20, wherein the parameters are determined using a prediction signal comprising residual pixels peripheral to the preceding block and residual pixels peripheral to a prediction block for the preceding block.
22. The method of any one of claims 18 to 21, comprising: determining the preceding block from a merge list of blocks of the current frame.
23. The method of claim 22, wherein the merge list is one of multiple available merge lists, and each merge list is based on the prediction signal used to determine the parameters.
24. The method of claim 22, wherein the merge list is one of multiple available merge lists, and each merge list is based on a prediction mode used for encoding candidate blocks for the multiple available merge lists.
25. The method of any one of claims 22 to 24, wherein blocks of the current frame preceding the current block in the coding order that have a block size smaller than a defined block size are omitted from the merge list.
26. An apparatus for decoding, comprising: a processor configured to perform the method of any one of claims 1 to 25.
27. A computer-readable storage medium storing instructions that, when executed, cause a processor to perform the method of any one of claims 1 to 25.
28. A computer-readable storage medium storing a bitstream that, when provided
to a decoder, results in the decoder executing the method of any one of claims 1 to 25.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363511883P | 2023-07-04 | 2023-07-04 | |
| US202363511882P | 2023-07-04 | 2023-07-04 | |
| PCT/US2024/036862 WO2025010397A1 (en) | 2023-07-04 | 2024-07-05 | Merge based cross-component prediction and inter prediction with filtering |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4717030A1 true EP4717030A1 (en) | 2026-04-01 |
Family
ID=91960684
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24745869.8A Pending EP4717030A1 (en) | 2023-07-04 | 2024-07-05 | Merge based cross-component prediction and inter prediction with filtering |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4717030A1 (en) |
| WO (1) | WO2025010397A1 (en) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116708833A (en) | 2018-07-02 | 2023-09-05 | Lg电子株式会社 | Codec and sending method and storage medium |
| WO2020151764A1 (en) | 2019-01-27 | 2020-07-30 | Beijing Bytedance Network Technology Co., Ltd. | Improved method for local illumination compensation |
| CN110446044B (en) * | 2019-08-21 | 2022-08-09 | 浙江大华技术股份有限公司 | Linear model prediction method, device, encoder and storage device |
-
2024
- 2024-07-05 EP EP24745869.8A patent/EP4717030A1/en active Pending
- 2024-07-05 WO PCT/US2024/036862 patent/WO2025010397A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025010397A1 (en) | 2025-01-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10798408B2 (en) | Last frame motion vector partitioning | |
| US12425636B2 (en) | Segmentation-based parameterized motion models | |
| US12034963B2 (en) | Compound prediction for video coding | |
| US12273533B2 (en) | Video stream adaptive filtering for bitrate reduction | |
| US20250343926A1 (en) | Chroma-From-Luma Prediction With Derived Scaling Factor | |
| WO2024081011A1 (en) | Filter coefficient derivation simplification for cross-component prediction | |
| WO2023219616A1 (en) | Local motion extension in video coding | |
| EP4717030A1 (en) | Merge based cross-component prediction and inter prediction with filtering | |
| WO2024081012A1 (en) | Inter-prediction with filtering | |
| US20250142050A1 (en) | Overlapped Filtering For Temporally Interpolated Prediction Blocks | |
| US20260067454A1 (en) | Post-Reconstruction Filtering Syntax Prediction | |
| US12610046B2 (en) | Tap-constrained convolutional cross-component model prediction | |
| US20240380924A1 (en) | Geometric transformations for video compression | |
| WO2025212281A1 (en) | Adaptive blending for cross-component prediction | |
| WO2026090559A1 (en) | Intra block copy with extended reference area | |
| WO2024173325A1 (en) | Wiener filter design for video coding | |
| WO2026011182A1 (en) | Reference frame motion field selection for wedge mode blocks | |
| WO2026055438A1 (en) | Scaling in cross-component sample offset | |
| WO2024145086A1 (en) | Content derivation for geometric partitioning mode video coding | |
| CN121533014A (en) | Frame-level nonlinear motion offset in video bitmap processing | |
| WO2026050279A1 (en) | Block-level control flag coding in cross-component sample offset |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251223 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |