WO2022032028A1 - Methods and apparatuses for affine motion-compensated prediction refinement - Google Patents
Methods and apparatuses for affine motion-compensated prediction refinement Download PDFInfo
- Publication number
- WO2022032028A1 WO2022032028A1 PCT/US2021/044840 US2021044840W WO2022032028A1 WO 2022032028 A1 WO2022032028 A1 WO 2022032028A1 US 2021044840 W US2021044840 W US 2021044840W WO 2022032028 A1 WO2022032028 A1 WO 2022032028A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sample
- prediction
- block
- sub
- determining
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/132—Sampling, masking or truncation of coding units, e.g. adaptive resampling, frame skipping, frame interpolation or high-frequency transform coefficient masking
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
- H04N19/517—Processing of motion vectors by encoding
- H04N19/52—Processing of motion vectors by encoding by predictive encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/527—Global motion vector estimation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/563—Motion estimation with padding, i.e. with filling of non-object values in an arbitrarily shaped picture block or region for estimation purposes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
- H04N19/82—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation involving filtering within a prediction loop
Definitions
- the present disclosure relates to video coding and compression, and in particular but not limited to, methods and apparatuses for affine motion-compensated prediction refinement (AMPR) in video coding.
- AMPR affine motion-compensated prediction refinement
- Video coding is performed according to one or more video coding standards.
- video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part2) and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO/IEC MPEG and ITU-T VECG.
- AV1 AOMedia Video 1
- AOM Alliance for Open Media
- Audio Video Coding which refers to digital audio and digital video compression standard
- AVS Audio Video Coding
- Most of the existing video coding standards are built upon the famous hybrid video coding framework i.e., using block-based prediction methods, e.g., inter-prediction, intra- prediction, to reduce redundancy present in video images or sequences and using transform coding to compact the energy of the prediction errors.
- An important goal of video coding techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradations to video quality.
- the first generation AVS standard includes Chinese national standard “Information Technology, Advanced Audio Video Coding, Part 2: Video” (known as AVS1) and “Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video” (known as AVS+). It can offer around 50% bit-rate saving at the same perceptual quality compared to MPEG-2 standard.
- the AVS1 standard video part was promulgated as the Chinese national standard in February 2006.
- the second generation AVS standard includes the series of Chinese national standard “Information Technology, Efficient Multimedia Coding” (knows as AVS2), which is mainly targeted at the transmission of extra HD TV programs.
- the coding efficiency of the AVS2 is double of that of the AVS+. In May 2016, the AVS2 was issued as the Chinese national standard.
- the AVS2 standard video part was submitted by Institute of Electrical and Electronics Engineers (IEEE) as one international standard for applications.
- the AVS3 standard is one new generation video coding standard for UHD video application aiming at surpassing the coding efficiency of the latest interational standard HEVC.
- March 2019, at the 68-th AVS meeting the AVS3-P2 baseline was finished, which provides approximately 30% bit-rate savings over the HEVC standard.
- HPM high performance model
- the present disclosure provides examples of techniques relating to AMPR for the
- a method for AMPR includes determining whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copying a padding sample to the neighboring sample location for a filter based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
- an apparatus for AMPR includes one or more processors and a memory configured to store instructions executable by the one or more processors.
- the one or more processors are configured to determine whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copy a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
- a non-transitory computer-readable storage medium for AMPR storing computer-executable instructions that, when executed by one or more computer processors, causing the one or more computer processors to perform acts including: determining whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copying a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
- FIG. 1 is a block diagram illustrating an exemplary video encoder in accordance with some implementations of the present disclosure.
- FIG. 2 is a block diagram illustrating an exemplary video decoder in accordance with some implementations of the present disclosure.
- FIGS. 3A-3E are schematic diagrams illustrating multi-type tree splitting modes in accordance with some implementations of the present disclosure.
- FIG. 4 is a schematic diagram illustrating an example of a bi-directional optical flow (BIO) model in accordance with some implementations of the present disclosure.
- FIGS. 5A-5B are schematic diagrams illustrating examples of 4-parameter affine model in accordance with some implementations of the present disclosure.
- FIG. 6 is a schematic diagram illustrating an example of 6-parameter affine model in accordance with some implementations of the present disclosure.
- FIG. 7 illustrates a prediction refinement with optical flow (PROF) process for affine mode in accordance with some implementations of the present disclosure.
- FIG. 8 illustrates an example of calculation of a horizontal offset and a vertical offset from a sample location to a specific position of a sub-block where a sub-block MV is derived in accordance with some implementations of the present disclosure.
- FIG. 9 illustrates an example of sub-blocks inside one affine CU in accordance with some implementations of the present disclosure.
- FIGS . 10 A- 10B illustrate examples of diamond-shape filters in accordance with some implementations of the present disclosure.
- FIGS . 11 A- 11 B illustrate examples of diamond-shape filters scaled by MV difference at horizontal and vertical directions in accordance with some implementations of the present disclosure.
- FIG. 12 illustrates a spatial relationship of one fractional position and its four neighboring integer positions in accordance with some implementations of the present disclosure.
- FIG. 13 illustrates an example of extending the boundary prediction samples of one sub-block to its extended area in accordance with some implementations of the present disclosure.
- FIG. 14 is a block diagram illustrating an exemplary apparatus for AMPR in accordance with some implementations of the present disclosure.
- FIG. 15 is a flowchart illustrating an exemplary process of AMPR in accordance with some implementations of the present disclosure.
- references throughout this specification to “one embodiment,” “an embodiment,” “an example,” “some embodiments,” “some examples,” or similar language means that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are also applicable to other embodiments, unless expressly specified otherwise.
- the terms “first,” “second,” “third,” etc. are all used as nomenclature only for references to relevant elements, e.g., devices, components, compositions, steps, etc., without implying any spatial or chronological orders, unless expressly specified otherwise.
- a “first device” and a “second device” may refer to two separately formed devices, or two parts, components, or operational states of a same device, and may be named arbitrarily.
- module may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors.
- a module may include one or more circuits with or without stored code or instructions.
- the module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to, or located adjacent to, one another.
- a method may comprise steps of: i) when or if condition X is present, function or action X’ is performed, and ii) when or if condition Y is present, function or action Y’ is performed.
- the method may be implemented with both the capability of performing function or action X’, and the capability of performing function or action Y’.
- the functions X’ and Y’ may both be performed, at different times, on multiple executions of the method.
- a unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software.
- the unit or module may include functionally related code blocks or software components, that are directly or indirectly linked together, so as to perform a particular function.
- FIG. 1 shows a block diagram of illustrating an exemplary block-based hybrid video encoder 100 which may be used in conjunction with many video coding standards using block- based processing.
- a video frame is partitioned into a plurality of video blocks for processing.
- a prediction is formed based on either an inter prediction approach or an intra prediction approach.
- inter prediction one or more predictors are formed through motion estimation and motion compensation, based on pixels from previously reconstructed frames.
- intra prediction predictors are formed based on reconstructed pixels in a current frame. Through mode decision, a best predictor may be chosen to predict a current block.
- Intra prediction uses pixels from the samples of already coded neighboring blocks (which are called reference samples) in the same video picture and/or slice to predict the current video block. Spatial prediction reduces spatial redundancy inherent in the video signal.
- Inter prediction uses reconstructed pixels from already-coded video pictures to predict the current video block.
- Temporal prediction reduces temporal redundancy inherent in the video signal.
- Temporal prediction signal for a given coding unit (CU) or coding block is usually signaled by one or more motion vectors (MVs) which indicate the amount and the direction of motion between the current CU and its temporal reference. Further, if multiple reference pictures are supported, one reference picture index is additionally sent, which is used to identify from which reference picture in the reference picture store the temporal prediction signal comes.
- MVs motion vectors
- an intra/inter mode decision circuitry 121 in the encoder 100 chooses the best prediction mode, for example based on the rate-distortion optimization method.
- the block predictor 120 is then subtracted from the current video block; and the resulting prediction residual is de-correlated using the transform circuitry 102 and the quantization circuitry 104.
- the resulting quantized residual coefficients are inverse quantized by the inverse quantization circuitry 116 and inverse transformed by the inverse transform circuitry 118 to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU.
- in-loop filtering 115 such as a deblocking filter, a sample adaptive offset (SAO), and/or an adaptive in-loop filter (ALF) may be applied on the reconstructed CU before it is put in the reference picture store of the picture buffer 117 and used to code future video blocks.
- coding mode inter or intra
- prediction mode information motion information
- quantized residual coefficients are all sent to the entropy coding unit 106 to be further compressed and packed to form the bit-stream
- a deblocking filter is available in AVC, HEVC as well as the now- current version of WC.
- SAO sample adaptive offset
- SAO sample adaptive offset
- ALF adaptive loop filter
- FIG. 2 is a block diagram illustrating an exemplary block-based video decoder 200 which may be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section residing in the encoder 100 of FIG. 1.
- an incoming video bitstream 201 is first decoded through an Entropy Decoding 202 to derive quantized coefficient levels and prediction-related information.
- the quantized coefficient levels are then processed through an Inverse Quantization 204 and an Inverse Transform 206 to obtain a reconstructed prediction residual.
- a block predictor mechanism implemented in an Intra/inter Mode Selector 212, is configured to perform either an Intra Prediction 208, or a Motion Compensation 210, based on decoded prediction information.
- a set of unfiltered reconstructed pixels are obtained by summing up the reconstructed prediction residual from the Inverse Transform 206 and a predictive output generated by the block predictor mechanism, using a summer 214.
- the reconstructed block may further go through an In-Loop Filter 209 before it is stored in a Picture Buffer 213 which functions as a reference picture store.
- the reconstructed video in the Picture Buffer 213 may be sent to drive a display device, as well as used to predict future video blocks.
- a filtering operation is performed on these reconstructed pixels to derive a final reconstructed Video Output 222.
- Video coding/decoding standards mentioned above, such as HE VC and AV3 are conceptually similar. For example, they all use block-based hybrid video coding framework. Block partitioning schemes in some standards are elaborated below.
- HE VC partitions blocks only based on quad-trees.
- the basic unit for compression is termed coding tree unit (CTU).
- CTU coding tree unit
- Each CTU may contain one coding unit (CU) or recursively split into four smaller CUs until the predefined minimum CU size is readied.
- CU also named leaf CU
- PUs prediction units
- one coding tree unit is split into CUs to adapt to varying local characteristics based on quad/binary/extended-quad-tree. Additionally, the concept of multiple partition unit type in the HEVC is removed, i.e., the separation of CU, PU and TU does not exist in the AVS3. Instead, each CU is always used as the basic unit for both prediction and transform without further partitions.
- one CTU is firstly partitioned based on a quad-tree structure. Then, each quad-tree leaf node may be further partitioned based on a binary and extended-quad-tree structure.
- FIGS. 3A-3E are schematic diagrams illustrating multi-type tree splitting modes in accordance with some implementations of the present disclosure. As shown in FIGS. 3A-3E, there are five splitting types in multi-type tree structure: quaternary partitioning 301, vertical binary partitioning 302, horizontal binary partitioning 303, vertical extended quaternary partitioning 304, and horizontal extended quaternary partitioning 305.
- block-based motion compensation may be applied to achieve a trade-off among coding efficiency, complexity, and memory access bandwidth.
- the average prediction accuracy is inferior to pixel-based prediction because all pixels within each block or sub-block share the same block level motion vector.
- PROF for affine mode is adopted as a coding tool in the current WC standards.
- AVS3 there is no similar tool.
- Some examples of the present disclosure provide alternative optical flow based methods to improve affine mode prediction accuracy as elaborated below.
- affine motion compensated prediction is applied by signaling one flag for each inter coding block to indicate whether the translation motion model or the affine motion model is applied for inter prediction.
- two affine modes including 4-par amter affine mode and 6-parameter affine mode, are supported for one affine coding block.
- the 4-parameter affine model may have following parameters: two parameters for translation movement in horizontal and vertical directions respectively, one parameter for zoom motion and one parameter for rotational motion for both directions.
- horizontal zoom parameter may be equal to vertical zoom parameter
- horizontal rotation parameter may be equal to vertical rotation parameter.
- those affine parameters may be derived from two MVs, which are also called control point motion vector (CPMV), located at the top-left comer and top-right comer of a current block.
- CPMV control point motion vector
- FIGS. 5A-5B are schematic diagrams illustrating examples of 4-parameter affine model in accordance with some implementations of the present disclosure.
- the affine motion field of the block is described by two control point MVs (Vo, Vi).
- the motion field (v x , v y ) of one affine coded block is described as equation (1) below:
- the 6-parameter affine mode may have following parameters: two parameters for translation movement in horizontal and vertical directions respectively, two parameters for zoom motion and rotation motion respectively in horizontal direction, and another two parameters for zoom motion and rotation motion respectively in vertical direction.
- the 6- parameter affine motion model is coded with three CPMVs.
- FIG. 6 is a schematic diagram illustrating an example of 6-parameter affine model in accordance with some implementations of the present disclosure.
- three control points of one 6-par amter affine block 601 are located at the top-left, top-right and bottom left comer of the block.
- the motion at top-left control point is related to translation motion
- the motion at top-right control point is related to rotation and zoom motion in horizontal direction
- the motion at bottom-left control point is related to rotation and zoom motion in vertical direction.
- the rotation and zoom motion in the horizontal direction of the 6-paramter may not be same as those motion in vertical direction.
- the PROF is adopted in the WC which refines the sub-block based affine motion compensation based on the optical flow model. Specifically, after performing the sub-block-based affine motion compensation, each luma prediction sample of one affine block is modified by one sample refinement value derived based on the optical flow equation.
- the operations of the PROF may be summarized as following four steps.
- the sub-block-based affine motion compensation is performed to generate sub-block prediction I(i, j) using the sub-block MVs as derived according to the equation (1) above for 4-parameter affine model, or equation (2) above for 6-parameter affine model.
- one additional row and / or column of prediction samples need to be generated on each of the four sides of one sub-block, which extends a 4x4 subblock into a 6x6 sub-block.
- the samples on the extended borders are copied from the nearest integer pixel position in the reference picture to avoid additional interpolation processes.
- FIG. 7 illustrates a PROF process for the affine mode in accordance with some implementations of the present disclosure. In the PROF, after adding the prediction refinement to the original prediction sample, one clipping operation “clip3” is performed to clip the value of the refined prediction sample to be within 15-bit, as shown in equations below:
- I r (i,j) I (i,j) + ⁇ I(i,j)
- Function clip3(min, max, val) confines a given value “val” in the range of [min, max].
- ⁇ v(i,j) can be calculated for the first sub-block, and reused for other sub-blocks in the same CU.
- ⁇ x and ⁇ y are the horizontal and vertical offset from the sample location (t,y) to the center of the sub-block that the sample belongs to, ⁇ v(i,j) can be derived as shown in equation (5) below:
- parameters c, d, e, and f in equation (5) above can be derived.
- parameters c, d, e, and f may be derived as shown in equation below:
- parameters c, d, e, and f may be derived as shown in equation below: where (v 0x , v 0y ), (y 1x , v 1y ), (y 2x , v 2y ) are the top-left, top-right and bottom-left control point MVs of the current coding block, w and h are the width and height of the block.
- the MV difference ⁇ v x and ⁇ v y are always derived at the precision of 1/32-pel.
- the luma prediction refinement ⁇ I(i,j) is added to the sub-block prediction /(i,y) .
- the final prediction /'(i,j) for sample at location (i,j) is generated as shown in equation (6) below:
- I'(i,j) I (i,j) + ⁇ I(i,j) (6)
- Bi-prediction in video coding is a simple combination of two temporal prediction blocks obtained from the reference pictures.
- the motion vectors received at a decoder end may not be so accurate.
- the BIO tool is adopted in both WC and AVS3 standards to compensate such motion for every sample inside one block.
- the BIO is sample-wise motion refinement that is performed on top of the block-based motion-compensated predictions when bi-prediction is used.
- the derivation of the refined motion vector for each sample in one block is based on the classical optical flow model.
- Let l (k) (x,y) be the sample value at the coordinate (x, y) of the prediction block derived from the reference picture list k(k 0, 1), and ⁇ l (k) (x,y)/ ⁇ x and ⁇ l (k) (x,y)/ ⁇ y are the horizontal and vertical gradients of the sample.
- the motion refinement ( v x , v y ) at (x, y) can be derived by optical flow equation (7) below: [0062] With the combination of the optical flow equation (7) and the interpolation of the prediction blocks along the motion trajectory, as shown in FIG. 4, the BIO prediction may be obtained as shown in equation (8) below:
- FIG. 4 is a schematic diagram illustrating an example of a BIO model in accordance with some implementations of the present disclosure.
- (MV x 0, MV y 0) and (MV x 1, MVy1) indicate the block-level motion vectors that are used to generate the two prediction blocks and .
- the motion refinement ( v x , v y ) at the sample location (x, y) is calculated by minimizing the difference ⁇ between the values of the samples after motion refinement compensation (i.e., A and B in FIG. 4), as shown in equation (9) below:
- Equation (8) and (10) in addition to the block-level MC, gradients need to be derived in the BIO for every sample of each motion compensated block, i.e., and in order to derive the local motion refinement and generate the final predication at that sample location.
- the gradients are calculated by a two-dimensional (2D) separable finite impulse response (FIR) filtering process which defines a set of 8-tap filters and applies different filters to derive the horizontal and vertical gradients according to the precision of the block-level motion vector, e.g., ( MV x 0, MV y 0) and (MV x 1, MV y 1) in FIG. 4.
- Table 1 illustrates coefficients of the gradient filters that are used by the BIO.
- the BIO is only applied to bi-prediction blocks which are predicted by two reference blocks from temporal neighboring pictures. Additionally, the BIO is enabled without sending additional information from encoder to decoder. Specifically, the BIO is applied to all the bi-directional predicted blocks which have both the forward and backward prediction signals.
- the UMVE mode in AVS3 standards is the same tool named merge mode with motion vector differences (MMVD) in the VVC standards.
- MMVD motion vector differences
- the MMVD/UMVE mode is introduced in both the VVC and AVS standards as one special merge mode.
- MMVD flag at coding block level.
- two base merge candidates are firstly generated as the first two candidates of the regular merge mode.
- additional syntax elements are signaled to indicate the MVDs that are added to the motion of the selected merge candidate.
- the MMVD syntax elements include a merge candidate flag to select the base merge candidate, a distance index to specify the MVD magnitude and a direction index to indicate the MVD direction.
- sub-block based affine motion compensation (affine mode), similar as the VVC standards, is used to generate inter-predicted pixel values.
- This sub-block- based prediction is a trade-off among coding efficiency, complexity, and memory access bandwidth.
- the average prediction accuracy is inferior to pixel-based prediction because all pixels within each sub-block share the same motion vector.
- the AVS3 standards does not have pixel level refinement after sub-block-based motion compensation.
- the present disclosure provides a new method to improve affine mode prediction accuracy.
- prediction value of each pixel is refined by adding a differential value derived by the optical flow equation.
- the proposed method may be referred as AMPR.
- the AMPR may achieve pixel level prediction accuracy without significantly increasing the complexity and also keeps the worst-case memory access bandwidth comparable to the regular sub-block-based motion compensation in the affine mode.
- the AMPR is also built upon optical flow, it is significantly different from the PROF in the VVC standards in the following aspects.
- AMPR utilizes an interpolation-based filtering for gradient calculation at each pixel, which allows a unified design between the AMPR and BIO workflow in AVS3.
- MV difference calculation is at each pixel. Unlike PROF where MV difference is always calculated based on the pixel location relative to the sub-block center, in AMPR, MV difference may be calculated based on the pixel location relative to different positions within a sub-block.
- AMPR may be adaptively skipped at the decoder side based on certain defined conditions indicating that applying AMPR is not a good performance and/or complexity trade-off.
- this early termination method may be also used to simplify encoder side operations.
- Some encoder side optimization methods for the AMPR process are also presented in the examples of the present disclosure to reduce its latency and energy consumption, sudi as skip AMPR for affine UMVE, check the best mode selection at the parent CU before applying AMPR, skip AMPR for motion estimation at certain block sizes, check the magnitude of pixel MV difference before applying AMPR, skip AMPR for some picture types, e.g., low-delay pictures or non-low-delay pictures, etc.
- the AMPR method may include five steps as explained below.
- regular sub-block-based affine motion compensation is performed to generate sub-block prediction I (i,j) at each pixel location (i,j).
- both the horizontal and vertical gradients of affine prediction samples are directly calculated from reference samples at integer sample positions in the temporal reference picture.
- One advantage of doing so is that for each affine sub-block, its gradient values may be generated at the same time when generating its prediction samples.
- Another design benefit of such a gradient calculation method is that it is also consistent with the gradient calculation process used by other coding tools in the AVS standard, such as BIO. Sharing the same process among different modules in a standard is friendly to the pipeline and/or parallelism design in practical hardware codec implementations.
- the input to the gradient derivation process is the same reference samples as that used for the motion compensation of the affine sub-block and the same fractional components (fracX, fracY) of the input motion (MV x, MV y ) of the sub-block.
- fracX, fracY fractional components of the input motion
- the order of applying the filters ho and hi is different.
- the gradient filter ho is firstly applied in the horizontal direction to derive the horizontal gradient values at horizontal fractional sample position fracX, then, the interpolation filter hi is applied vertically to interpolate the gradient values at vertical fractional sample position fracY.
- the interpolation filter h L is firstly applied horizontally to interpolate intermediate interpolation samples at horizontal sample position fracX, followed by the gradient filter ho being applied in the vertical direction to derive the vertical gradient values at vertical fraction sample position fracY from the intermediate interpolation samples.
- the gradient filters may be generated with different filter coefficient precisions and with different number of taps, which can provide various trade-off between gradient calculation precision and computational complexity. For instance, gradient filters with more filter taps and/or with a higher filter coefficient precision can generally lead to better coding efficiency, but at the expense of more computational operations, e.g., number of additions, multiplications, and shifts, due to the gradient calculation processes.
- following 8 -tap filters are proposed for the horizontal and/or vertical gradient calculations of the AMPR, as shown in Table 2.
- Table 2 is an exemplary table of predefined 8-tap interpolation filter coeffi cents fgra d [p] for generating spatial gradients based on 1/16-pel precision of input sample values.
- Table 3 is an exemplary table of predefined 4-tap interpolation filter coefficients f grad [p] for generating spatial gradients based on 1/16-pel precision of input sample values.
- MV difference ⁇ v(i,j) which is between each pixel’s MV and the MV of the sub-block to which the pixel belongs, is calculated at each pixel location (i,j).
- FIG. 8 illustrates an example of calculation of a horizontal offset and a vertical offset from a sample location to a specific position of a sub-block where a sub-block MV is derived.
- the horizontal ⁇ x and the vertical offset ⁇ y are calculated from the sample location (i,j) to the specific position (i',j') of the sub-block where the sub-block MV is derived.
- the specific position (i',j') may not be always the center of the sub-block.
- ⁇ v(i,j) is calculated based on the pixel location relative to a specific position within the sub-block by the equation (5).
- ⁇ x and ⁇ y may be calculated by equation shown below:
- ⁇ x and ⁇ y j Or if the sub-block MV is derived from the position at the top-right comer outside the sub- block, ⁇ x and ⁇ y can be calculated by equation shown below:
- ⁇ x and ⁇ y can be calculated by equation shown below:
- ⁇ x and ⁇ y can be calculated as
- ⁇ v(i,j) may be calculated by equation (5), ⁇ x and ⁇ y are horizontal and vertical offsets from the sample location (i,j) to the pilot sample location of the sub-block where the sample belongs to.
- the pilot sample location refers to the sample location inside one sub-block which is used to derive the MV for generating the sub-block- based prediction samples of the sub-block.
- the values of ⁇ x and ⁇ y are derived as follows.
- FIG. 9 illustrates an example of sub-blocks inside one affine CU in accordance with some implementations of the present disclosure.
- ⁇ x i
- ⁇ x (i — w + 1)
- sub-block C sub-block C in FIG.
- ⁇ v(i,j) may be derived by equation (11) below: where c, d, e and / are affine model parameters, which are already known since the current block is affine mode coded bock.
- equation (11) similar as the equation (5) of the
- a prediction refinement value is calculated by the equation (4).
- the prediction refinement is added to the sub-block prediction /(i,j).
- the final prediction /'(i,j) is generated as the equation (6).
- the proposed AMPR workflow may be applied to luma component and/or chroma components.
- the proposed AMPR is only applied to refine the affine prediction samples of luma component while the chroma prediction samples are still generated based on the existing sub-block-based affine motion compensation.
- both luma component and chroma components are refined by the proposed AMPR process.
- the sample-wise MV difference ⁇ v(i,j) may be derived in different manners.
- the sample-wise MV difference ⁇ v(i,j) may be always derived only once based on luma sub-blocks and then reused for chroma components when prediction refinement value is calculated in the above fourth step.
- the value of ⁇ v(i,j) used by chroma components may be scaled according to the sampling grid ratio between the collocated luma and chroma coding block. For example, for 4:2:0 video, the value of the reused ⁇ v(i,j) may be halved before used by chroma components, while for 4:4:4 video, the same value of the reused ⁇ v(i,j) may be used by chroma components. For 4:2:2 video, the horizontal offset of ⁇ v(i,j) may be halved while the vertical offset of ⁇ v(i,j) may not be changed before used by chroma components.
- the sample-wise MV difference ⁇ v(i,j) may be separately derived for luma and chroma components, where the derivation process may be the same as the above described third step.
- one flag is signaled to indicate whether the AMPR is applied to chroma components at various coding levels, e.g., sequence level, picture level, slice level and so forth. Further, if the above enabling/disabling flag is true, another flag may be signaled from encoder to decoder to indicate whether the chroma motion refinements are re-calculated from the corresponding control-point motion vectors or directly borrowed from the corresponding motion refinements of luma component.
- An alternative implementation of AMPR is to approximate the optical-flow equation- based multiplication by a filtering process. Specifically, it is proposed to replace the multiplication of the gradient value and motion vector difference at each sample location by performing a filtering process over regular sub-block-based affine motion prediction. This can be formulated by the following equation as where P AFF (x, y) is the sub-block-based motion compensation prediction samples, f k is the filter coefficients and P AMPR (i,j ) is the filtered affine prediction samples. In practice, different number of filter taps and filter shapes can be applied to achieve different trade-off between complexity and coding performance.
- the filtering process may be performed by a cross-shape filter, which may also be referred as a diamond-shape filter.
- the diamond-shape filter may be a combination of vertical shape and horizontal stupe of 3-tip filter [-1, 1, 1] or 5- tap filter [-1, -2, 4, 2, 1], as shown in FIGS. 10A-10B.
- the cross-shape filter may be an approximation of the gradient calculation process described in the previous optical-flow based refinement process.
- FIGS. 11 A-l IB illustrate examples of diamond-shape filters scaled by MV difference at horizontal and vertical directions in accordance with some implementations of the present disclosure.
- FIG. 11A illustrates a 5-tip scaled diamond shape filter.
- FIG. 1 IB illustrates a 9-tip scaled diamond filter.
- the filtering process may be performed by a square-shape filter.
- the square-shape filter may be 3x3 or 5x5 shape filter, where the significance of each coefficient of the square filter may depend on the distance between the location of each coefficient and center of the filter, which means the center coefficient may have the maximum value in the filter.
- the filter coefficients in the selected square-shape filter may be scaled by the value of ⁇ v(i,j ) at horizontal and vertical direction.
- the corresponding scaled filter is calculated at each sample location.
- the scale value is the motion vector (MV) difference ⁇ v(i,j) at each sample location, which is between each pixel’s MV and the MV of the sub-block to which the pixel belongs.
- the calculation process is the same as the optical-flow based implementation of AMPR, where the value of ⁇ v(i,j) is determined based on whether the associated sub-block MV is derived from the position at the sub-block center at integer position, or at the top-left comer of the sub-block, or at the top-right comer of the sub-block, or bottom-left comer of the sub-block.
- the corresponding filtered affine prediction samples can be calculated as: where M and N are initialized coefficients of constant values, and ⁇ and Ay are scaling factors which can be used to adjust the importance of the neighboring samples to the filter sample at one current position.
- M and N are initialized coefficients of constant values
- ⁇ and Ay are scaling factors which can be used to adjust the importance of the neighboring samples to the filter sample at one current position.
- the filtering process may be performed over the regular sub-block-based affine motion predicted samples.
- padding process may be needed.
- the padded sample values may be copied from nearest reference samples at integer position.
- the padded sample values may be duplicated values of the reference samples at integer position used by the current block/sub- block.
- the integer samples that are closest to the current boundary samples, which may be fractional, of the current CU are used to pad the extended region of the current CU.
- the integer samples whose positions are smaller than the corresponding boundary samples of the current CU, i.e., floor 0, are used to pad the samples in the extended region of the current CU.
- the integer reference samples that is left to the prediction sample in horizontal direction and below the prediction sample in vertical direction i.e., the integer sample LB in FIG. 12
- the integer reference sample that is closest to the prediction sample is used for affine prediction filtering.
- the average of the reference samples at multiple integer sample positions can also be used to fill the extended samples of one sub-block. For instance, according to the corresponding fractional position of the sub-block motion vector, two closest integer samples to the fractional position may be used and the average may be used to fill the extended sample. In another example, it is proposed to always average all four integer reference samples to calculate the corresponding extended sample, as indicated by the following equation:
- bilinear filter may be applied which interpolate the extended samples according to the distances of four integer reference samples to the fractional position, as shown as: where X frac and frac are the fractional sample position in x- and y-direction, which is floating number in the range [0, 1).
- the prediction refinement derived by applying AMPR may not be always beneficial or/and necessary.
- the significance of the derived ⁇ I(i,j) is determined by the precision and magnitude of the derived ⁇ 17(1,7) and g(i,j).
- AMPR operation may be conditionally applied based on certain conditions. This may be achieved by signaling a flag for each block to indicate if AMPR mode is applied or not. It may also be achieved by using the same conditions to enable AMPR operation at both encoder and decode sides, with no additional signaling required.
- AMPR operation may not help, or even hurt, coding performance, and therefore it is better to skip AMPR operation for the block.
- Another motivation of such conditional application of AMPR operation is that in some cases the benefit of applying AMPR may be marginal and from computation complexity point of view it is also better to turn the operation off.
- AMPR operation may be applied depending on whether the CMPVs are explicitly signaled or not.
- affine merge mode where the CPMVs are not explicitly signaled but implicitly derived from spatial neighbor CUs, AMPR may be skipped for the current CU because the CPMVs under this mode may not be accurate.
- AMPR may be skipped.
- threshold values may be determined based on various factors, e.g., CU aspect ratio and/or sub-block size, etc. Such an example may be implemented in different manners described below.
- AMPR may be skipped for this sub-block.
- This condition may have different implementation variations.
- the check of the absolute value of the derived ⁇ v(i,j) for all the pixels may be simplified by only checking the four comers of the current sub-block, where the maximum absolute value of the derived ⁇ v(i,j) for all the pixels within a sub-block can be found as shown in equation (14) below: where the pixel location (i,j) may be any pixel coordinate in the sub-block, or may be from four comers (0, 0), (w — 1, 0), (0, h — 1), (w — 1, h — 1).
- the calculation of the maximum absolute value of all ⁇ v(i,j) may be obtained by the equation below: where the sample locations (i,j) are the four comers of those sub-blocks in a CU except the top-left, i.e., sub-block A in FIG. 9, top-right, i.e., sub-block B in FIG. 9, and bottom-left sub- block, i.e., sub-block C in FIG. 9.
- the coordinates of the four comer pixels within a subblock are: (0, 0), (w — 1, 0), (0, h — 1), (w — 1, h — 1).
- ⁇ x ⁇ is the function to take absolute value of x.
- the check of the derived ⁇ v(i,j) may be combined with non- simplified AMPR operation.
- the equation (14) may be merged with equation (4), then the prediction refinement value is calculated by equation below: [0126]
- the threshv x or threshv y may be with different or the same value.
- the values of threshv x and threshv y may be determined depending on which position is used for deriving the sub- block MV. In other words, a different or same pair of values of threshv x and threshv y may be determined for two sub-blocks if their MVs are derived using different positions.
- its pair of values of threshv x and threshv y may be the same or different from that of a sub- block whose sub-block level MV is derived based on the position of the sub-block top-left comer.
- the values of threshv x and threshv y may be defined as in the range of [1/32, 1/16] in unit of pixels.
- a value of (1/16)*(10/16), (1/16)*(12/16), or (1/16)*(14/16) may be used as the thresholds.
- the threshold value is a floating point of 1/16-pel, e.g., 0.625 in unit of 1/16-pel, 0.75 in unit of 1/16-pel, or 0.875 in unit of 1/16-pel.
- the values of threshv x and threshv y may be defined based on picture types. For low-delay pictures, the derived affine model parameters may have smaller magnitude than other non-low-delay pictures, since low-delay pictures tends to have smaller and/or smoother motions and therefore smaller values may be preferred for those thresholds. [0130] In some examples, the values of threshv x and threshv y may be the same regardless of different picture types.
- AMPR may be skipped for this sub-block.
- a sub-block contains smooth surface which may consist of flat textures (e.g., with no or small number of high-frequency details).
- the significance of ⁇ v(i,j) and g(i,j) may be considered jointly or used in a hybrid manner to decide whether AMPR should be skipped for current sub- block or CU.
- the selective enabling of AMPR may be also performed. For example, based on the equation (15), if max ⁇ vx and/or max ⁇ vy is smaller than a predefined threshold threshvx and threshv y , the corresponding scaled coefficients in the selected filter may become 0, which means the filtering process may be simplified from 2D to one-dimensional (ID) filtering process. Taking the 5 -tap filters in FIG.
- affine UMVE mode is computation intensive for encoder because it involves choosing the best distance index for each merge mode candidate.
- SAD sum of absolute transformed difference
- AMPR operation is skipped during SATD based cost calculation for affine UMVE mode at the encoder side. It is found through experiments that while the best index is selected according to the best SATD cost, whether AMPR is applied during the SATD calculation or not usually does not change the ranking of the best SATD cost. Therefore, with the proposed method enabling AMPR mode wouldn’t incur obvious encoder complexity for affine UMVE mode.
- Motion estimation is another major overhead at the encoder side.
- AMPR process may be skipped depending on certain conditions. These conditions indicate that the best encoding mode of a CU is unlikely to be affine mode after mode selection process.
- One example of such a condition is whether a current CU has a parent CU which is already determined to be coded by explicit affine mode or affine merge mode. This is due to the strong correlation of coding mode selection between a CU and its parent CU, and it is more likely that the best coding mode for the current CU is also explicit affine mode if the condition above is true.
- Another exemplar condition used for enabling AMPR is whether the parent CU of the current CU is determined to be inter-predicted with explicit affine mode. If it is true, AMPR is applied during affine motion estimation of the current CU; otherwise, AMPR is skipped during affine motion estimation of the current CU.
- AMPR may be skipped for small size CUs.
- the size of a CU may be defined as the total number of pixels.
- a pixel number threshold such as 16x16 or 16x32 or 32*32 may be defined, and for a block with a size smaller than the defined threshold, AMPR may be skipped during affine motion estimation process for the block.
- the encoder side optimization may be also performed.
- the filtering process may not be performed for affine UMVE mode.
- Another encoder optimization used for enabling the filtering process based AMPR is whether the parent CU of the current CU is determined to be inter-predicted with explicit affine mode. If it is true, the filtering process of AMPR is applied during affine motion estimation of the current CU. Otherwise, the filtering process is skipped during affine motion estimation of the current CU.
- FIG. 14 is a block diagram illustrating an apparatus for predicting a sample at a pixel location in a sub-block by AMPR in accordance with some implementations of the present disclosure.
- the apparatus 1400 may be a terminal, such as a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.
- the apparatus 1400 may include one or more of the following components: a processing component 1402, a memory 1404, a power supply component 1406, a multimedia component 1408, an audio component 1410, an input/output (I/O) interface 1412, a sensor component 1414, and a communication component 1416.
- a processing component 1402 such as a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.
- the apparatus 1400 may include one or more of the following components: a processing component 1402, a memory 1404, a power supply component 1406, a multimedia component 1408, an audio component 1410, an input/output (I/
- the processing component 1402 usually controls overall operations of the apparatus 1400, such as operations relating to display, a telephone call, data communication, a camera operation, and a recording operation.
- the processing component 1402 may include one or more processors 1420 for executing instructions to complete all or a part of steps of the above method.
- the processing component 1402 may include one or more modules to facilitate interaction between the processing component 1402 and other components.
- the processing component 1402 may include a multimedia module to facilitate the interaction between the multimedia component 1408 and the processing component 1402.
- the memory 1404 is configured to store different types of data to support operations of the apparatus 1400. Examples of such data include instructions, contact data, phonebook data, messages, pictures, videos, and so on for any application or method that operates on the apparatus 1400.
- the memory 1404 may be implemented by any type of volatile or non-volatile storage devices or a combination thereof, and the memory 1404 may be a Static Random Access Memory (SRAM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), an Erasable Programmable Read-Only Memory (EPROM), a Programmable Read-Only Memory (PROM), a Read-Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disk or a compact disk.
- SRAM Static Random Access Memory
- EEPROM Electrically Erasable Programmable Read-Only Memory
- EPROM Erasable Programmable Read-Only Memory
- PROM Programmable Read-Only Memory
- ROM Read-Only Memory
- the power supply component 1406 supplies power for different components of the apparatus 1400.
- the power supply component 1406 may include a power supply management system, one or more power supplies, and other components associated with generating, managing and distributing power for the apparatus 1400.
- the multimedia component 1408 includes a screen providing an output interface between the apparatus 1400 and a user.
- the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen receiving an input signal from a user.
- the touch panel may include one or more touch sensors for sensing a touch, a slide and a gesture on the touch panel. The touch sensor may not only sense a boundary of a touching or sliding actions, but also detect duration and pressure related to the touching or sliding operation.
- the multimedia component 1408 may include a front camera and/or a rear camera.
- the audio component 1410 is configured to output and/or input an audio signal.
- the audio component 1410 includes a microphone (MIC).
- the microphone is configured to receive an external audio signal.
- the received audio signal may be further stored in the memory 1404 or sent via the communication component 1416.
- the audio component 1410 further includes a speaker for outputting an audio signal.
- the I/O interface 1412 provides an interface between the processing component 1402 and a peripheral interface module.
- the above peripheral interface module may be a keyboard, a click wheel, a button, or the like. These buttons may include but not limited to, a home button, a volume button, a start button, and a lock button.
- the sensor component 1414 includes one or more sensors for providing a state assessment in different aspects for the apparatus 1400.
- the sensor component 1414 may detect an on/off state of the apparatus 1400 and relative locations of components.
- the components are a display and a keypad of the apparatus 1400.
- the sensor component 1414 may also detect a position change of the apparatus 1400 or a component of the apparatus 1400, presence or absence of a contact of a user on the apparatus 1400, an orientation or acceleration/deceleration of the apparatus 1400, and a temperature change of apparatus 1400.
- the sensor component 1414 may include a proximity sensor configured to detect presence of a nearby object without any physical touch.
- the sensor component 1414 may further include an optical sensor, such as a CMOS or CCD image sensor used in an imaging application.
- the sensor component 1414 may further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
- the communication component 1416 is configured to facilitate wired or wireless communication between the apparatus 1400 and other devices.
- the apparatus 1400 may access a wireless network based on a communication standard, such as WiFi, 4G, or a combination thereof.
- the communication component 1416 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel.
- the communication component 1416 may further include a Near Field Communication (NFC) module for promoting short-range communication.
- the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, Ultra-Wide Band (UWB) technology, Bluetooth (BT) technology and other technology.
- RFID Radio Frequency Identification
- IrDA infrared data association
- UWB Ultra-Wide Band
- Bluetooth Bluetooth
- the apparatus 1400 may be implemented by one or more of Application Specific Integrated Circuits (ASIC), Digital Signal Processors (DSP), Digital Signal Processing Devices (DSPD), Programmable Logic Devices (PLD), Field Programmable Gate Arrays (FPGA), controllers, microcontrollers, microprocessors, or other electronic elements to perform the above method.
- ASIC Application Specific Integrated Circuits
- DSP Digital Signal Processors
- DSPD Digital Signal Processing Devices
- PLD Programmable Logic Devices
- FPGA Field Programmable Gate Arrays
- controllers microcontrollers, microprocessors, or other electronic elements to perform the above method.
- a non-transitoiy computer readable storage medium may be, for example, a Hard Disk Drive (HDD), a Solid-State Drive (SSD), Flash memory, a Hybrid Drive or Solid-State Hybrid Drive (SSHD), a Read-Only Memory (ROM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, etc.
- HDD Hard Disk Drive
- SSD Solid-State Drive
- SSHD Solid-State Hybrid Drive
- ROM Read-Only Memory
- CD-ROM Compact Disc Read-Only Memory
- magnetic tape a floppy disk, etc.
- FIG. 15 is a flowchart illustrating an exemplary process for AMPR in accordance with some implementations of the present disclosure.
- step 1502 the processor 1420 determines whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block.
- the sample location may include Y-sample location, U-sample location, V-sample location, or a corresponding pixel location.
- the corresponding pixel is the pixel at location (/, j) including the sample at the location (/, j) as described in Exemplary
- the neighboring sample location may be determined based on the shape of a filter, such as the cross-shape filter shown in FIGS. 10A-10B and 11A-11B or the square-shape filter.
- step 1504 the processor 1420 copies a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
- the current sample at the sample location is the current boundary sample, as shown in FIG. 13.
- the current boundary sample is located at one of the four boundaries of the sub-block.
- the padding sample may be a reference sample selected from reference samples at integer sample positions or locations in a reference picture, an average sample of the reference samples, a prediction sample, etc.
- the processor 1420 may obtain multiple affine motion- compensated predictions for a plurality of samples at the sample location and a plurality of neighboring sample locations including one or more neighboring sample locations outside of the sub-block; and obtain, at the sample location, a refined prediction for a sample at the sample location using a filter with a pre-determined shape based on the multiple affine motion- compensated predictions, where the pre-determined shape covers the sample location and the plurality of neighboring sample locations.
- the sample location may be at one boundary of the four boundaries of the sub-block, and the plurality of neighboring sample locations may include the one or more neighboring samples that are outside of the four boundaries of the sub-block, for example, in the extended region as shown in FIG. 13.
- the processor 1420 may obtain the multiple affine motion- compensated predictions by performing a sub-block-based affine motion compensation on a video picture that comprises a plurality of sub-blocks.
- the processor 1420 may determine a reference sample at an integer position in a reference picture and determine the reference sample as the padding sample. [0164] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining a neighboring reference sample in the reference picture, where the neighboring reference sample is a reference sample at an integer position nearest to a prediction sample that is corresponding to the sample location. In one case, the prediction sample is in a reference picture that is previously reconstructed and located at a corresponding position. The prediction sample at the corresponding position in the reference picture may have same or the most similar image area or contents as the current sample at the sample position in the current sub-block in the current picture.
- the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located left to a prediction sample and above the prediction sample, such as the integer sample LT in FIG. 12.
- the prediction sample may be corresponding to the sample position.
- the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located right to a prediction sample and above the prediction sample, such as the integer sample RT in FIG. 12.
- the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located left to die prediction sample and below the prediction sample, such as the integer sample LB in FIG. 12. The prediction sample is corresponding to the sample position. [0168] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located right to the prediction sample and below the prediction sample, such as the integer sample RB in FIG. 12. The prediction sample is corresponding to the sample position. [0169] In some examples, the processor 1420 may determine an average sample of a plurality of reference samples at a plurality of integer positions in the reference picture and determine the average sample as the padding example. In one example, the average sample may be an average of two closest integer samples to the sample location. In another example, the average sample may be an average of the four integer reference samples LT, RT, LB, and RB.
- the processor 1420 may determine the padding sample using an interpolation filter.
- the processor 1420 may determine the padding sample by interpolating according to distances between each of a plurality of reference samples at a plurality of integer positions in the reference picture and a prediction sample corresponding to the sample location. For example, the padding sample is determined by interpolating according to the distances of four integer reference samples LT, RT, LB, and RB to a fractional position of the prediction sample.
- the processor 1420 may determine a prediction sample at one boundary of the sub-block as the padding sample.
- FIG. 13 shows the sub-block of four boundaries. The four boundaries are four sides of the sub-block between the sub-block and the extended region.
- the processor 1420 may determine the prediction sample corresponding to the sample location as the padding sample when it is determined that the sample position is at one boundary.
- the prediction sample is in a reference picture previously reconstructed and located at a corresponding position.
- the prediction sample at the corresponding position may have same or the most similar image area or contents as the current sample at the sample position in the current picture.
- an apparatus for AMPR includes one or more processors 1420 and a memory 1404 configured to store instructions executable by the one or more processors; where the processor, upon execution of the instructions, is configured to perform any method as described in FIG. 15 and above.
- a non-transitory computer readable storage medium 1404 having instructions stored therein. When the instructions are executed by one or more processors 1420, the instructions cause the processor to perform any method as described in FIG.15 and above.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
A method and an apparatus for affine motion-compensated prediction refinement (AMPR) are provided. The method includes determining whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copying a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
Description
METHODS AND APPARATUSES FOR AFFINE MOTION-COMPENSATED
PREDICTION REFINEMENT
CROSS-REFERENCE TO RELATED APPLICATION
[0001] The present application claims priority to U.S. Provisional Application No. 63/062,380, entitled “Affine Motion-Compensated Prediction Refinement,” filed on August 6, 2020, the entirety of which is incorporated by reference for all purposes.
FIELD
[0002] The present disclosure relates to video coding and compression, and in particular but not limited to, methods and apparatuses for affine motion-compensated prediction refinement (AMPR) in video coding.
BACKGROUND
[0003] Various video coding techniques may be used to compress video data. Video coding is performed according to one or more video coding standards. For example, nowadays, some well-known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part2) and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO/IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by Alliance for Open Media (AOM) as a successor to its preceding standard VP9. Audio Video Coding (AVS), which refers to digital audio and digital video compression standard, is another video compression standard series developed by the Audio and Video Coding Standard Workgroup of China. Most of the existing video coding standards are built upon the famous hybrid video coding framework i.e., using block-based prediction methods, e.g., inter-prediction, intra- prediction, to reduce redundancy present in video images or sequences and using transform coding to compact the energy of the prediction errors. An important goal of video coding
techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradations to video quality.
[0004] The first generation AVS standard includes Chinese national standard “Information Technology, Advanced Audio Video Coding, Part 2: Video” (known as AVS1) and “Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video” (known as AVS+). It can offer around 50% bit-rate saving at the same perceptual quality compared to MPEG-2 standard. The AVS1 standard video part was promulgated as the Chinese national standard in February 2006. The second generation AVS standard includes the series of Chinese national standard “Information Technology, Efficient Multimedia Coding” (knows as AVS2), which is mainly targeted at the transmission of extra HD TV programs. The coding efficiency of the AVS2 is double of that of the AVS+. In May 2016, the AVS2 was issued as the Chinese national standard. Meanwhile, the AVS2 standard video part was submitted by Institute of Electrical and Electronics Engineers (IEEE) as one international standard for applications. The AVS3 standard is one new generation video coding standard for UHD video application aiming at surpassing the coding efficiency of the latest interational standard HEVC. In March 2019, at the 68-th AVS meeting, the AVS3-P2 baseline was finished, which provides approximately 30% bit-rate savings over the HEVC standard. Currently, there is one reference software, called high performance model (HPM), is maintained by the AVS group to demonstrate a reference implementation of the AVS3 standard.
SUMMARY
[0005] The present disclosure provides examples of techniques relating to AMPR for the
AVS3 standards.
[0006] According to a first aspect of the present disclosure, there is provided a method for AMPR. The method includes determining whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copying a padding sample to the neighboring sample location for a filter based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
[0007] According to a second aspect of the present disclosure, there is provided an apparatus for AMPR. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. Upon execution of the instructions, the one or more processors are configured to determine whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copy a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
[0008] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for AMPR storing computer-executable instructions that, when executed by one or more computer processors, causing the one or more computer processors to perform acts including: determining whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block and copying a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
BRIEF DESCRIPTION OF THE DRAWINGS
[0009] A more particular description of the examples of the present disclosure will be rendered by reference to specific examples illustrated in the appended drawings. Given that these drawings depict only some examples and are not therefore considered to be limiting in scope, the examples will be described and explained with additional specificity and details through the use of the accompanying drawings.
[0010] FIG. 1 is a block diagram illustrating an exemplary video encoder in accordance with some implementations of the present disclosure.
[0011] FIG. 2 is a block diagram illustrating an exemplary video decoder in accordance with some implementations of the present disclosure.
[0012] FIGS. 3A-3E are schematic diagrams illustrating multi-type tree splitting modes in accordance with some implementations of the present disclosure.
[0013] FIG. 4 is a schematic diagram illustrating an example of a bi-directional optical flow (BIO) model in accordance with some implementations of the present disclosure.
[0014] FIGS. 5A-5B are schematic diagrams illustrating examples of 4-parameter affine model in accordance with some implementations of the present disclosure.
[0015] FIG. 6 is a schematic diagram illustrating an example of 6-parameter affine model in accordance with some implementations of the present disclosure.
[0016] FIG. 7 illustrates a prediction refinement with optical flow (PROF) process for affine mode in accordance with some implementations of the present disclosure.
[0017] FIG. 8 illustrates an example of calculation of a horizontal offset and a vertical offset from a sample location to a specific position of a sub-block where a sub-block MV is derived in accordance with some implementations of the present disclosure.
[0018] FIG. 9 illustrates an example of sub-blocks inside one affine CU in accordance with some implementations of the present disclosure.
[0019] FIGS . 10 A- 10B illustrate examples of diamond-shape filters in accordance with some implementations of the present disclosure.
[0020] FIGS . 11 A- 11 B illustrate examples of diamond-shape filters scaled by MV difference at horizontal and vertical directions in accordance with some implementations of the present disclosure.
[0021] FIG. 12 illustrates a spatial relationship of one fractional position and its four neighboring integer positions in accordance with some implementations of the present disclosure.
[0022] FIG. 13 illustrates an example of extending the boundary prediction samples of one sub-block to its extended area in accordance with some implementations of the present disclosure.
[0023] FIG. 14 is a block diagram illustrating an exemplary apparatus for AMPR in accordance with some implementations of the present disclosure.
[0024] FIG. 15 is a flowchart illustrating an exemplary process of AMPR in accordance with some implementations of the present disclosure.
DETAILED DESCRIPTION
[0025] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. But it will be apparent to one of ordinary skill in the art that various alternatives may be used. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0026] Reference throughout this specification to “one embodiment,” “an embodiment,” “an example,” “some embodiments,” “some examples,” or similar language means that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are also applicable to other embodiments, unless expressly specified otherwise. [0027] Throughout the disclosure, the terms “first,” “second,” “third,” etc. are all used as nomenclature only for references to relevant elements, e.g., devices, components, compositions, steps, etc., without implying any spatial or chronological orders, unless expressly specified otherwise. For example, a “first device” and a “second device” may refer to two separately formed devices, or two parts, components, or operational states of a same device, and may be named arbitrarily.
[0028] The terms “module,” “sub-module,” “circuit,” “sub-circuit,” “circuitry,” “sub- circuitry,” “unit,” or “sub-unit” may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. The module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to, or located adjacent to, one another.
[0029] As used herein, the term “if’ or “when” may be understood to mean “upon” or “in response to” depending on the context. These terms, if appear in a claim, may not indicate that the relevant limitations or features are conditional or optional. For example, a method may
comprise steps of: i) when or if condition X is present, function or action X’ is performed, and ii) when or if condition Y is present, function or action Y’ is performed. The method may be implemented with both the capability of performing function or action X’, and the capability of performing function or action Y’. Thus, the functions X’ and Y’ may both be performed, at different times, on multiple executions of the method.
[0030] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, the unit or module may include functionally related code blocks or software components, that are directly or indirectly linked together, so as to perform a particular function.
[0031] FIG. 1 shows a block diagram of illustrating an exemplary block-based hybrid video encoder 100 which may be used in conjunction with many video coding standards using block- based processing. In the encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter prediction approach or an intra prediction approach. In inter prediction, one or more predictors are formed through motion estimation and motion compensation, based on pixels from previously reconstructed frames. In intra prediction, predictors are formed based on reconstructed pixels in a current frame. Through mode decision, a best predictor may be chosen to predict a current block.
[0032] Intra prediction (also referred to as “spatial prediction”) uses pixels from the samples of already coded neighboring blocks (which are called reference samples) in the same video picture and/or slice to predict the current video block. Spatial prediction reduces spatial redundancy inherent in the video signal.
[0033] Inter prediction (also referred to as “temporal prediction”) uses reconstructed pixels from already-coded video pictures to predict the current video block. Temporal prediction reduces temporal redundancy inherent in the video signal. Temporal prediction signal for a given coding unit (CU) or coding block is usually signaled by one or more motion vectors (MVs) which indicate the amount and the direction of motion between the current CU and its temporal reference. Further, if multiple reference pictures are supported, one reference picture
index is additionally sent, which is used to identify from which reference picture in the reference picture store the temporal prediction signal comes.
[0034] After spatial and/or temporal prediction is performed, an intra/inter mode decision circuitry 121 in the encoder 100 chooses the best prediction mode, for example based on the rate-distortion optimization method. The block predictor 120 is then subtracted from the current video block; and the resulting prediction residual is de-correlated using the transform circuitry 102 and the quantization circuitry 104. The resulting quantized residual coefficients are inverse quantized by the inverse quantization circuitry 116 and inverse transformed by the inverse transform circuitry 118 to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, in-loop filtering 115, such as a deblocking filter, a sample adaptive offset (SAO), and/or an adaptive in-loop filter (ALF) may be applied on the reconstructed CU before it is put in the reference picture store of the picture buffer 117 and used to code future video blocks. To form the output video bitstream 114, coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 106 to be further compressed and packed to form the bit-stream
[0035] For example, a deblocking filter is available in AVC, HEVC as well as the now- current version of WC. In HEVC, an additional in-loop filter called SAO (sample adaptive offset) is defined to further improve coding efficiency. In the now-current version of the WC standard, yet another in-loop filter called ALF (adaptive loop filter) is being actively investigated, and it has a good chance of being included in the final standard.
[0036] These in-loop filter operations are optional. Performing these operations helps to improve coding efficiency and visual qualify. They may also be turned off as a decision rendered by the encoder 100 to save computational complexify.
[0037] It should be noted that intra prediction is usually based on unfiltered reconstructed pixels, while inter prediction is based on filtered reconstructed pixels if these filter options are turned on by the encoder 100.
[0038] FIG. 2 is a block diagram illustrating an exemplary block-based video decoder 200 which may be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section residing in the encoder 100 of FIG. 1. In the decoder 200, an incoming video bitstream 201 is first decoded through an Entropy Decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through an Inverse Quantization 204 and an Inverse Transform 206 to obtain a reconstructed prediction residual. A block predictor mechanism, implemented in an Intra/inter Mode Selector 212, is configured to perform either an Intra Prediction 208, or a Motion Compensation 210, based on decoded prediction information. A set of unfiltered reconstructed pixels are obtained by summing up the reconstructed prediction residual from the Inverse Transform 206 and a predictive output generated by the block predictor mechanism, using a summer 214.
[0039] The reconstructed block may further go through an In-Loop Filter 209 before it is stored in a Picture Buffer 213 which functions as a reference picture store. The reconstructed video in the Picture Buffer 213 may be sent to drive a display device, as well as used to predict future video blocks. In situations where the In-Loop Filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive a final reconstructed Video Output 222. [0040] Video coding/decoding standards mentioned above, such as HE VC and AV3 are conceptually similar. For example, they all use block-based hybrid video coding framework. Block partitioning schemes in some standards are elaborated below.
[0041] HE VC partitions blocks only based on quad-trees. The basic unit for compression is termed coding tree unit (CTU). Each CTU may contain one coding unit (CU) or recursively split into four smaller CUs until the predefined minimum CU size is readied. Each CU (also named leaf CU) contains one or multiple prediction units (PUs) and a tree of transform units
(TUs).
[0042] In the AVS3, one coding tree unit (CTU) is split into CUs to adapt to varying local characteristics based on quad/binary/extended-quad-tree. Additionally, the concept of multiple partition unit type in the HEVC is removed, i.e., the separation of CU, PU and TU does not
exist in the AVS3. Instead, each CU is always used as the basic unit for both prediction and transform without further partitions. In the tree partition structure of the AVS3, one CTU is firstly partitioned based on a quad-tree structure. Then, each quad-tree leaf node may be further partitioned based on a binary and extended-quad-tree structure.
[0043] FIGS. 3A-3E are schematic diagrams illustrating multi-type tree splitting modes in accordance with some implementations of the present disclosure. As shown in FIGS. 3A-3E, there are five splitting types in multi-type tree structure: quaternary partitioning 301, vertical binary partitioning 302, horizontal binary partitioning 303, vertical extended quaternary partitioning 304, and horizontal extended quaternary partitioning 305.
[0044] In the current WC and AVS3 standards, block-based motion compensation may be applied to achieve a trade-off among coding efficiency, complexity, and memory access bandwidth. The average prediction accuracy is inferior to pixel-based prediction because all pixels within each block or sub-block share the same block level motion vector. To improve the prediction accuracy at each pixel, PROF for affine mode is adopted as a coding tool in the current WC standards. In AVS3, there is no similar tool.
[0045] Some examples of the present disclosure provide alternative optical flow based methods to improve affine mode prediction accuracy as elaborated below.
Affine Mode
[0046] In HEVC, only translation motion model is applied for motion compensated prediction. While in the real world, there are many kinds of motion, e.g., zoom in/out, rotation, perspective motions and other irregular motions. In the WC and AVS3, affine motion compensated prediction is applied by signaling one flag for each inter coding block to indicate whether the translation motion model or the affine motion model is applied for inter prediction. In the current WC and AVS3 design, two affine modes, including 4-par amter affine mode and 6-parameter affine mode, are supported for one affine coding block.
[0047] The 4-parameter affine model may have following parameters: two parameters for translation movement in horizontal and vertical directions respectively, one parameter for zoom motion and one parameter for rotational motion for both directions. In the 4-parameter
affine model, horizontal zoom parameter may be equal to vertical zoom parameter, and horizontal rotation parameter may be equal to vertical rotation parameter. To achieve a better accommodation of the motion vectors and affine parameter, those affine parameters may be derived from two MVs, which are also called control point motion vector (CPMV), located at the top-left comer and top-right comer of a current block.
[0048] FIGS. 5A-5B are schematic diagrams illustrating examples of 4-parameter affine model in accordance with some implementations of the present disclosure. As shown in FIGS. 5A-5B, the affine motion field of the block is described by two control point MVs (Vo, Vi). Based on the control point motion, the motion field (vx, vy) of one affine coded block is described as equation (1) below:
[0049] The 6-parameter affine mode may have following parameters: two parameters for translation movement in horizontal and vertical directions respectively, two parameters for zoom motion and rotation motion respectively in horizontal direction, and another two parameters for zoom motion and rotation motion respectively in vertical direction. The 6- parameter affine motion model is coded with three CPMVs.
[0050] FIG. 6 is a schematic diagram illustrating an example of 6-parameter affine model in accordance with some implementations of the present disclosure. As shown in FIG. 6, three control points of one 6-par amter affine block 601 are located at the top-left, top-right and bottom left comer of the block. The motion at top-left control point is related to translation motion, and the motion at top-right control point is related to rotation and zoom motion in horizontal direction, and the motion at bottom-left control point is related to rotation and zoom motion in vertical direction. Compared to the 4-parameter affine motion model, the rotation and zoom motion in the horizontal direction of the 6-paramter may not be same as those motion in vertical direction.
[0051] In some examples, when (V0, V1, V2) are the MVs of the top-left, top-right and bottom-left comers of the current block in FIG. 6, the motion vector of each sub-block (vx, vy) is derived using the three MVs at control points as equation (2) below:
Prediction Refinement with Optical Flow for Affine Mode (PROF)
[0052] To improve affine motion compensation precision, the PROF is adopted in the WC which refines the sub-block based affine motion compensation based on the optical flow model. Specifically, after performing the sub-block-based affine motion compensation, each luma prediction sample of one affine block is modified by one sample refinement value derived based on the optical flow equation. In some examples, the operations of the PROF may be summarized as following four steps.
[0053] In the first step, the sub-block-based affine motion compensation is performed to generate sub-block prediction I(i, j) using the sub-block MVs as derived according to the equation (1) above for 4-parameter affine model, or equation (2) above for 6-parameter affine model.
[0054] Additionally, in the second step, spatial gradients gx(i,j) and gy(i,j ) of each prediction samples are calculated as equation (3) below:
Accordingly, to calculate the gradients, one additional row and / or column of prediction samples need to be generated on each of the four sides of one sub-block, which extends a 4x4 subblock into a 6x6 sub-block. To reduce the memory bandwidth and complexity, the samples on the extended borders are copied from the nearest integer pixel position in the reference picture to avoid additional interpolation processes.
[0055] Further, in the third step, the luma prediction refinement value is calculated by equation (4) below:
ΔI(i,j) = gx(i,j) * Δvx(i,j) + gy(i,j) * Δvy(i,j) (4) where the Δv(i,j) is the difference between a pixel MV computed for sample location (i,j), denoted by v(i,j), and the sub-block MV of the sub-block where the pixel (i,j) locates at. [0056] FIG. 7 illustrates a PROF process for the affine mode in accordance with some implementations of the present disclosure. In the PROF, after adding the prediction refinement to the original prediction sample, one clipping operation “clip3” is performed to clip the value of the refined prediction sample to be within 15-bit, as shown in equations below:
Ir(i,j) = I (i,j) + ΔI(i,j)
Ir(i,j) = clip3(—dILimit, dlLimit — 1,Ir(i,j)) dILimit = (1 « max (13, BitDepth + 1)) where /(i,j) and Ir(i,j) are the original prediction and refined prediction sample values at location (i,j), respectively. Function clip3(min, max, val) confines a given value “val” in the range of [min, max].
[0057] Because the affine model parameters and the pixel location relative to the sub-block center are not changed from sub-block to sub-block, Δv(i,j) can be calculated for the first sub-block, and reused for other sub-blocks in the same CU. When Δx and Δy are the horizontal and vertical offset from the sample location (t,y) to the center of the sub-block that the sample belongs to, Δv(i,j) can be derived as shown in equation (5) below:
Avx(i,j) = c * Δx + d * Δy (5) bVy(i,j) = e * Δx + f * Δy
[0058] Based on the affine sub-block MV derivation equations (1) and (2), parameters c, d, e, and f in equation (5) above can be derived. Specifically, for 4-parameter affine model, parameters c, d, e, and f may be derived as shown in equation below:
Additionally, for 6-parameter affine model, parameters c, d, e, and f may be derived as shown in equation below:
where (v0x, v0y), (y1x, v1y), (y2x, v2y) are the top-left, top-right and bottom-left control point MVs of the current coding block, w and h are the width and height of the block. In the PROF, the MV difference Δvx and Δvy are always derived at the precision of 1/32-pel. [0059] Finally, in the fourth step, the luma prediction refinement ΔI(i,j) is added to the sub-block prediction /(i,y) . The final prediction /'(i,j) for sample at location (i,j) is generated as shown in equation (6) below:
I'(i,j) = I (i,j) + ΔI(i,j) (6)
Bi-Directional Optical Flow (BIO)
[0060] Bi-prediction in video coding is a simple combination of two temporal prediction blocks obtained from the reference pictures. However, due to signaling cost and accuracy trade- off of motion vectors, the motion vectors received at a decoder end may not be so accurate. As a result, there may be still remaining small motion that can be observed between the two prediction blocks, which could reduce the efficiency of motion compensated prediction. To solve this problem, the BIO tool is adopted in both WC and AVS3 standards to compensate such motion for every sample inside one block. Specifically, the BIO is sample-wise motion refinement that is performed on top of the block-based motion-compensated predictions when bi-prediction is used.
[0061] In the BIO design, the derivation of the refined motion vector for each sample in one block is based on the classical optical flow model. Let l(k)(x,y) be the sample value at the coordinate (x, y) of the prediction block derived from the reference picture list k(k = 0, 1), and ∂l(k)(x,y)/ ∂x and ∂l(k)(x,y)/ ∂y are the horizontal and vertical gradients of the sample. Assuming the optical flow model is valid, the motion refinement ( vx, vy) at (x, y) can be derived by optical flow equation (7) below:
[0062] With the combination of the optical flow equation (7) and the interpolation of the prediction blocks along the motion trajectory, as shown in FIG. 4, the BIO prediction may be obtained as shown in equation (8) below:
[0063] FIG. 4 is a schematic diagram illustrating an example of a BIO model in accordance with some implementations of the present disclosure. As shown in FIG. 4, (MVx0, MVy0) and (MVx1, MVy1) indicate the block-level motion vectors that are used to generate the two prediction blocks and . Further, the motion refinement ( vx, vy) at the sample location (x, y) is calculated by minimizing the difference Δ between the values of the samples after motion refinement compensation (i.e., A and B in FIG. 4), as shown in equation (9) below:
[0064] Additionally, in order to ensure the regularity of the derived motion refinement, it is assumed that the motion refinement is consistent within a local surrounding area centered at (x, y); therefore, in the BIO design in the AVS3, the values of (vx, vy) are derived by minimizing Δ inside the 4x4 window Ω around the current sample at (x, y) as shown in equation (10) below:
[0065] As shown in equations (8) and (10), in addition to the block-level MC, gradients need to be derived in the BIO for every sample of each motion compensated block, i.e., and in order to derive the local motion refinement and generate the final predication at that sample location. In the AVS3, the gradients are calculated by a two-dimensional (2D) separable
finite impulse response (FIR) filtering process which defines a set of 8-tap filters and applies different filters to derive the horizontal and vertical gradients according to the precision of the block-level motion vector, e.g., ( MVx0, MVy0) and (MVx1, MVy1) in FIG. 4. Table 1 illustrates coefficients of the gradient filters that are used by the BIO.
[0066] Finally, the BIO is only applied to bi-prediction blocks which are predicted by two reference blocks from temporal neighboring pictures. Additionally, the BIO is enabled without sending additional information from encoder to decoder. Specifically, the BIO is applied to all the bi-directional predicted blocks which have both the forward and backward prediction signals.
Ultimate Motion Vector Expression (UMVE)
[0067] The UMVE mode in AVS3 standards is the same tool named merge mode with motion vector differences (MMVD) in the VVC standards. In addition to conventional merge mode which derives the motion information of one current block from its spatial/temporal neighbors, the MMVD/UMVE mode is introduced in both the VVC and AVS standards as one special merge mode.
[0068] Specifically, in both the VVC and AVS3, it is signaled by one MMVD flag at coding block level. In the MMVD mode, two base merge candidates are firstly generated as the first two candidates of the regular merge mode. After one base merge candidate is selected and signaled, additional syntax elements are signaled to indicate the MVDs that are added to the motion of the selected merge candidate. The MMVD syntax elements include a merge candidate flag to select the base merge candidate, a distance index to specify the MVD magnitude and a direction index to indicate the MVD direction.
[0069] In the AVS3 standards, sub-block based affine motion compensation (affine mode), similar as the VVC standards, is used to generate inter-predicted pixel values. This sub-block-
based prediction is a trade-off among coding efficiency, complexity, and memory access bandwidth. The average prediction accuracy is inferior to pixel-based prediction because all pixels within each sub-block share the same motion vector. Unlike the VVC standards, in affine mode, the AVS3 standards does not have pixel level refinement after sub-block-based motion compensation.
[0070] The present disclosure provides a new method to improve affine mode prediction accuracy. After the regular sub-block based affine motion compensation, prediction value of each pixel is refined by adding a differential value derived by the optical flow equation. The proposed method may be referred as AMPR. The AMPR may achieve pixel level prediction accuracy without significantly increasing the complexity and also keeps the worst-case memory access bandwidth comparable to the regular sub-block-based motion compensation in the affine mode. Although the AMPR is also built upon optical flow, it is significantly different from the PROF in the VVC standards in the following aspects.
[0071] First, gradient calculation is at each pixel. Unlike PROF, which extends the sub-block prediction by one pixel at each side, AMPR utilizes an interpolation-based filtering for gradient calculation at each pixel, which allows a unified design between the AMPR and BIO workflow in AVS3.
[0072] Second, MV difference calculation is at each pixel. Unlike PROF where MV difference is always calculated based on the pixel location relative to the sub-block center, in AMPR, MV difference may be calculated based on the pixel location relative to different positions within a sub-block.
[0073] The third is early termination. Unlike PROF process, which is always invoked at the decoder side when a coding block is predicted by the affine mode, AMPR may be adaptively skipped at the decoder side based on certain defined conditions indicating that applying AMPR is not a good performance and/or complexity trade-off.
[0074] In some examples, this early termination method may be also used to simplify encoder side operations. Some encoder side optimization methods for the AMPR process are also presented in the examples of the present disclosure to reduce its latency and energy
consumption, sudi as skip AMPR for affine UMVE, check the best mode selection at the parent CU before applying AMPR, skip AMPR for motion estimation at certain block sizes, check the magnitude of pixel MV difference before applying AMPR, skip AMPR for some picture types, e.g., low-delay pictures or non-low-delay pictures, etc.
Exemplary Workflow of AMPR
[0075] The AMPR method may include five steps as explained below. In the first step, regular sub-block-based affine motion compensation is performed to generate sub-block prediction I (i,j) at each pixel location (i,j).
[0076] In the second step, horizontal and vertical spatial gradients gx(i,j ) and gy(i,j ) of the sub-block prediction are calculated at each pixel location using an interpolation-based filtering. In some examples, both the horizontal and vertical gradients of affine prediction samples are directly calculated from reference samples at integer sample positions in the temporal reference picture. One advantage of doing so is that for each affine sub-block, its gradient values may be generated at the same time when generating its prediction samples. Another design benefit of such a gradient calculation method is that it is also consistent with the gradient calculation process used by other coding tools in the AVS standard, such as BIO. Sharing the same process among different modules in a standard is friendly to the pipeline and/or parallelism design in practical hardware codec implementations.
[0077] Specifically, the input to the gradient derivation process is the same reference samples as that used for the motion compensation of the affine sub-block and the same fractional components (fracX, fracY) of the input motion (MVx, MVy) of the sub-block. To derive the gradient values at each sample position, in addition to the default 8-tap FIR filters hL used for affine prediction, another set of new FIR filters hG are introduced in the proposed method to calculate the gradient values.
[0078] Additionally, depending on the direction of the derived gradients, the order of applying the filters ho and hi is different. In case of deriving horizontal gradients gx(i,j), the gradient filter ho is firstly applied in the horizontal direction to derive the horizontal gradient
values at horizontal fractional sample position fracX, then, the interpolation filter hi is applied vertically to interpolate the gradient values at vertical fractional sample position fracY.
[0079] On the contrary, when vertical gradients gy(i,j ) are derived, the interpolation filter hL is firstly applied horizontally to interpolate intermediate interpolation samples at horizontal sample position fracX, followed by the gradient filter ho being applied in the vertical direction to derive the vertical gradient values at vertical fraction sample position fracY from the intermediate interpolation samples.
[0080] In some examples, the gradient filters may be generated with different filter coefficient precisions and with different number of taps, which can provide various trade-off between gradient calculation precision and computational complexity. For instance, gradient filters with more filter taps and/or with a higher filter coefficient precision can generally lead to better coding efficiency, but at the expense of more computational operations, e.g., number of additions, multiplications, and shifts, due to the gradient calculation processes. In one example, following 8 -tap filters are proposed for the horizontal and/or vertical gradient calculations of the AMPR, as shown in Table 2.
[0081] Table 2 is an exemplary table of predefined 8-tap interpolation filter coeffi cents fgrad[p] for generating spatial gradients based on 1/16-pel precision of input sample values.
[0082] In another example, to reduce the complexity of gradient calculation, following 4-tap FIR filters as shown in Table 3 are used for the gradient generation of the proposed AMPR method. Table 3 is an exemplary table of predefined 4-tap interpolation filter coefficients fgrad[p] for generating spatial gradients based on 1/16-pel precision of input sample values.
[0083] In the third step, MV difference Δv(i,j), which is between each pixel’s MV and the MV of the sub-block to which the pixel belongs, is calculated at each pixel location (i,j). FIG. 8 illustrates an example of calculation of a horizontal offset and a vertical offset from a sample location to a specific position of a sub-block where a sub-block MV is derived. As shown in FIG. 8, the horizontal Δx and the vertical offset Δy are calculated from the sample location (i,j) to the specific position (i',j') of the sub-block where the sub-block MV is derived. In some examples, the specific position (i',j') may not be always the center of the sub-block. As shown in FIG. 8, Δv(i,j) is calculated based on the pixel location relative to a specific position within the sub-block by the equation (5).
[0084] In one example, for affine mode, let (i,j) be the pixel location/coordinate within a sub-block to which the pixel belongs, w and h be the width and height of the sub-block (e.g., w = h = 4 for 4x4 sub-block, w = h = 8 for 8x8 sub-block), the horizontal offset Δx and the vertical offset Δy (Δx and Δy are defined in equation (5)) for each pixel (i,j) may be derived as follows for i = 0... ( w — 1) and j = 0...( h — 1).
[0085] In one example, if the sub-block MV is derived from the position at the sub-block center at an integer position, Δx and Δy may be calculated by equation shown below: Δx = (i — (w » 1)) Δy = (J - (h » 1))
Or if the the sub-block MV is derived from the position at the sub-block center at a fractional position, Δx and Δy may be calculated by equation shown below:
(Δx = (i — (w » 1) — 0.5)
(Δy = (j — (h » 1) - 0.5)
[0086] In another example, if the sub-block MV is derived from the position at the top-left comer of the sub-block, Δx and Δy can be calculated by equation shown below: Δx = i Δy = j
[0087] In another example, if the sub-block MV is derived from the position at the top-right comer within the sub-block, Δx and Δy can be calculated by equation shown below: Δx = (i - w + 1)
Δy =j
Or if the sub-block MV is derived from the position at the top-right comer outside the sub- block, Δx and Δy can be calculated by equation shown below:
[0088] In another example, If the sub-block MV is derived from the position at the bottom- left comer within the sub-block, Δx and Δy can be calculated by equation shown below:
[0089] Or, if the sub-block MV is derived from the position at the bottom-left comer outside the sub-block, Δx and Δy can be calculated as
[0090] In another example, Δv(i,j) may be calculated by equation (5), Δx and Δy are horizontal and vertical offsets from the sample location (i,j) to the pilot sample location of the sub-block where the sample belongs to. The pilot sample location refers to the sample location inside one sub-block which is used to derive the MV for generating the sub-block- based prediction samples of the sub-block. In one example, based on the position of the sub- block within the current CU, the values of Δx and Δy are derived as follows.
[0091] FIG. 9 illustrates an example of sub-blocks inside one affine CU in accordance with some implementations of the present disclosure. For the top-left sub-block, i.e., sub-block A in FIG. 9, Δx = i, Δy = j. For the top-right sub-block, i.e., sub-block B in FIG. 9, Δx = (i — w + 1), Δy = j. For the bottom-left sub-block, i.e., sub-block C in FIG. 9, Δx = i, Δy = (j — h + 1) when 6-parameter affine model is applied; and Δx = (i — (w » 1) — 0.5) , Δy = (j — (h » 1) — 0.5) when 4-parameter affine model is applied. For other sub- blocks, Δx = (i — (w » 1) — 0.5), Δy = (j — (h » 1) — 0.5).
[0092] Once the horizontal offset Δx and the vertical offset Δy are calculated, Δv(i,j) may be derived by equation (11) below:
where c, d, e and / are affine model parameters, which are already known since the current block is affine mode coded bock. The equation (11) similar as the equation (5) of the
PROF tool in VVC standards.
[0093] In the fourth step, a prediction refinement value is calculated by the equation (4). [0094] In the fifth step, the prediction refinement is added to the sub-block prediction /(i,j). The final prediction /'(i,j) is generated as the equation (6).
[0095] In the present disclosure, the proposed AMPR workflow may be applied to luma component and/or chroma components.
[0096] In one example, to achieve a good performance/complexity trade-off, the proposed AMPR is only applied to refine the affine prediction samples of luma component while the chroma prediction samples are still generated based on the existing sub-block-based affine motion compensation.
[0097] In another example, for refinement alignment, both luma component and chroma components are refined by the proposed AMPR process. In this case, the sample-wise MV difference Δv(i,j) may be derived in different manners.
[0098] In one example, the sample-wise MV difference Δv(i,j) may be always derived only once based on luma sub-blocks and then reused for chroma components when prediction refinement value is calculated in the above fourth step. In this case, the value of Δv(i,j) used by chroma components may be scaled according to the sampling grid ratio between the collocated luma and chroma coding block. For example, for 4:2:0 video, the value of the reused Δv(i,j) may be halved before used by chroma components, while for 4:4:4 video, the same value of the reused Δv(i,j) may be used by chroma components. For 4:2:2 video, the horizontal offset of Δv(i,j) may be halved while the vertical offset of Δv(i,j) may not be changed before used by chroma components.
[0099] In another example, the sample-wise MV difference Δv(i,j) may be separately derived for luma and chroma components, where the derivation process may be the same as the above described third step.
[0100] In another example, it is proposed to adaptively switch between the method of reusing luma motion refinement for chroma and the method of separately deriving luma and chroma motion refinement based on the chroma sampling format of the input video. For instance, for 4:2:0 and 4:2:2 videos, given that the sampling grids of the luma and chroma components are misaligned, separate derivation of the motion refinements for luma and chroma components may be applied. On the other hand, when the input video is in 4:4:4 chroma sampling format, it only needs to derive the motion refinement once, i.e., for luma, which is then reused for the other two-color components, because the sampling grids of three-color components are fully aligned.
[0101] In another example, one flag is signaled to indicate whether the AMPR is applied to chroma components at various coding levels, e.g., sequence level, picture level, slice level and so forth. Further, if the above enabling/disabling flag is true, another flag may be signaled from encoder to decoder to indicate whether the chroma motion refinements are re-calculated from the corresponding control-point motion vectors or directly borrowed from the corresponding motion refinements of luma component.
Alternative Workflow of AMPR
[0102] An alternative implementation of AMPR is to approximate the optical-flow equation- based multiplication by a filtering process. Specifically, it is proposed to replace the multiplication of the gradient value and motion vector difference at each sample location by performing a filtering process over regular sub-block-based affine motion prediction. This can be formulated by the following equation as
where PAFF(x, y) is the sub-block-based motion compensation prediction samples, fk is the filter coefficients and PAMPR(i,j ) is the filtered affine prediction samples. In practice, different number of filter taps and filter shapes can be applied to achieve different trade-off between complexity and coding performance.
[0103] In one or more examples, the filtering process may be performed by a cross-shape filter, which may also be referred as a diamond-shape filter. For example, the diamond-shape filter may be a combination of vertical shape and horizontal stupe of 3-tip filter [-1, 1, 1] or 5- tap filter [-1, -2, 4, 2, 1], as shown in FIGS. 10A-10B. The cross-shape filter may be an approximation of the gradient calculation process described in the previous optical-flow based refinement process.
[0104] In order to capture the motion vector (MV) difference Δv(i,j), which is between each pixel’s MV and the MV of the sub-block to which the pixel belongs, the filter coefficients in the selected diamond-shape filter may be calculated based on the value of Δv(i,j) at horizontal and vertical direction. In another word, a scaled diamond-shape filter may be used to compensate the motion vector difference at each sample location. FIGS. 11 A-l IB illustrate examples of diamond-shape filters scaled by MV difference at horizontal and vertical directions in accordance with some implementations of the present disclosure. FIG. 11A illustrates a 5-tip scaled diamond shape filter. FIG. 1 IB illustrates a 9-tip scaled diamond filter. [0105] In another example, the filtering process may be performed by a square-shape filter. For example, the square-shape filter may be 3x3 or 5x5 shape filter, where the significance of each coefficient of the square filter may depend on the distance between the location of each coefficient and center of the filter, which means the center coefficient may have the maximum value in the filter. Similar to the above-mentioned diamond-shape filter, the filter coefficients in the selected square-shape filter may be scaled by the value of Δv(i,j ) at horizontal and vertical direction.
[0106] Once a specific type of filter, e.g., diamond-shape or square-shape, is selected, the corresponding scaled filter is calculated at each sample location. The scale value is the motion vector (MV) difference Δv(i,j) at each sample location, which is between each pixel’s MV and the MV of the sub-block to which the pixel belongs. The calculation process is the same as the optical-flow based implementation of AMPR, where the value of Δv(i,j) is determined based on whether the associated sub-block MV is derived from the position at the sub-block
center at integer position, or at the top-left comer of the sub-block, or at the top-right comer of the sub-block, or bottom-left comer of the sub-block.
[0107] In one specific example, when the 3-tap cross-shape filter is applied in the proposed scheme, the corresponding filtered affine prediction samples can be calculated as:
where M and N are initialized coefficients of constant values, and Δχ and Ay are scaling factors which can be used to adjust the importance of the neighboring samples to the filter sample at one current position. In one specific embodiment, it is proposed to setM= 2N, for example M = 16 and N = 8.
[0108] Once the scaled filter coefficients are calculated, the filtering process may be performed over the regular sub-block-based affine motion predicted samples. For the neighboring sample locations which are outside of current block/sub-block, padding process may be needed. In one or more examples, the padded sample values may be copied from nearest reference samples at integer position. In another example, the padded sample values may be duplicated values of the reference samples at integer position used by the current block/sub- block. In one example, the integer samples that are closest to the current boundary samples, which may be fractional, of the current CU are used to pad the extended region of the current CU. In another example, the integer samples whose positions are smaller than the corresponding boundary samples of the current CU, i.e., floor 0, are used to pad the samples in the extended region of the current CU.
Reference Sample Padding
[0109] As described in “Alternative Workflow of AMPR” above, one method is proposed to implement the AMPR sample refinement by replacing the sample-wise optical-flow refinement, i.e., multiplication of gradient and local motion refinements, by one convolution filtering. However, due to the length of the applied filters, such filtering operation needs to access additional rows and columns of prediction samples of each sub-block, i.e., the basic unit
of regular affine motion compensation. Such design could severely complicate the pipeline design when implementing the AMPR in hardware due to the interdependency between the prediction samples in different sub-blocks. To explain such problem, assuming there are two neighboring sub-blocks A and B, where A is left to B, in the current affine CU and the filter size that is applied to filter the affine prediction samples is 3-by-3. In such case, the sample refinements of the prediction samples located at the right boundary of the sub-block B cannot be started until the prediction samples at the left boundary of sub-block B are fully reconstructed.
[0110] To address such latency problem, various methods are proposed in the following to remove the dependency between different sub-blocks at affine motion compensation stage. In one example, it is proposed to directly use the reference samples at integer sample positions in the reference picture to fill the prediction samples at the extended region around each sub- block. Because there are neighboring integer sample positions around one fractional sample position, there could be different ways to select the corresponding integer sample position. [0111] In one or more examples, it is proposed to always select the integer reference sample that is located left to the prediction sample in horizontal direction and above the prediction sample in vertical direction, i.e., the integer sample LT in FIG. 12, to fill the extended region of each sub-block. In some examples, it is proposed to always select the integer reference samples that is right to the prediction sample in horizontal direction and above the prediction sample in vertical direction, i.e., the integer sample RT in FIG. 12, to fill the extended region of each sub-block.
[0112] In one example, it is proposed to always select the integer reference samples that is left to the prediction sample in horizontal direction and below the prediction sample in vertical direction, i.e., the integer sample LB in FIG. 12, to fill the extended region of each sub-block. In yet another example, it is proposed to always select the integer reference samples that is right to the prediction sample in horizontal direction and below the prediction sample in vertical direction, i.e., the integer sample RB in FIG. 12, to fill the extended region of each sub-block.
In some examples, the integer reference sample that is closest to the prediction sample is used for affine prediction filtering.
[0113] Alternatively, or additionally, instead of only selecting one specific integer sample position, the average of the reference samples at multiple integer sample positions can also be used to fill the extended samples of one sub-block. For instance, according to the corresponding fractional position of the sub-block motion vector, two closest integer samples to the fractional position may be used and the average may be used to fill the extended sample. In another example, it is proposed to always average all four integer reference samples to calculate the corresponding extended sample, as indicated by the following equation:
Pext = (LT + RT + LB + RB + 2) » 2
[0114] In another example, to improve the precision of the extended samples, instead of averaging, it is proposed to use other interpolation filters to generate the prediction samples in the extended region. In one specific example, bilinear filter may be applied which interpolate the extended samples according to the distances of four integer reference samples to the fractional position, as shown as:
where X frac and frac are the fractional sample position in x- and y-direction, which is floating number in the range [0, 1).
[0115] In another example, instead of using the reference samples at integer sample positions, it is proposed to directly pad out the prediction samples at the four boundaries of one sub-block to its extended region, as shown in FIG. 13.
Selective Enabling of AMPR
[0116] The prediction refinement derived by applying AMPR may not be always beneficial or/and necessary. According to the equation (4), the significance of the derived ΔI(i,j) is determined by the precision and magnitude of the derived Δ 17(1,7) and g(i,j).
[0117] In some examples, AMPR operation may be conditionally applied based on certain conditions. This may be achieved by signaling a flag for each block to indicate if AMPR mode
is applied or not. It may also be achieved by using the same conditions to enable AMPR operation at both encoder and decode sides, with no additional signaling required.
[0118] The motivation of such conditional application of AMPR operation is that if the CPMVs of a CU is not accurate, or the derived affine model, e.g., 2-parameter, 4-parameter, or 6-parameter affine mode, is not accurate, the subsequently derived Δv(i,j) by the equation (11) may not be accurate as well. In this case, AMPR operation may not help, or even hurt, coding performance, and therefore it is better to skip AMPR operation for the block. Another motivation of such conditional application of AMPR operation is that in some cases the benefit of applying AMPR may be marginal and from computation complexity point of view it is also better to turn the operation off.
[0119] In one or more examples, AMPR operation may be applied depending on whether the CMPVs are explicitly signaled or not. In affine merge mode where the CPMVs are not explicitly signaled but implicitly derived from spatial neighbor CUs, AMPR may be skipped for the current CU because the CPMVs under this mode may not be accurate.
[0120] In another example, if the magnitude of the derived Δv(i,j) or/and g(i,j) are small, for example, when compared with certain predefined or dynamically determined threshold values, AMPR may be skipped. Such threshold values may be determined based on various factors, e.g., CU aspect ratio and/or sub-block size, etc. Such an example may be implemented in different manners described below.
[0121] In one example, if the absolute value of the derived Δv(i,j) for all the pixels within a sub-block are smaller than a threshold value, AMPR may be skipped for this sub-block. This condition may have different implementation variations. For example, the check of the absolute value of the derived Δv(i,j) for all the pixels may be simplified by only checking the four comers of the current sub-block, where the maximum absolute value of the derived Δv(i,j) for all the pixels within a sub-block can be found as shown in equation (14) below:
where the pixel location (i,j) may be any pixel coordinate in the sub-block, or may be from four comers (0, 0), (w — 1, 0), (0, h — 1), (w — 1, h — 1).
[0122] In another example, the calculation of the maximum absolute value of all Δv(i,j) may be obtained by the equation below:
where the sample locations (i,j) are the four comers of those sub-blocks in a CU except the top-left, i.e., sub-block A in FIG. 9, top-right, i.e., sub-block B in FIG. 9, and bottom-left sub- block, i.e., sub-block C in FIG. 9. The coordinates of the four comer pixels within a subblock are: (0, 0), (w — 1, 0), (0, h — 1), (w — 1, h — 1). \x\ is the function to take absolute value of x.
[0123] In another example, the check of the derived Δv(i,j) for all the pixels within a sub- block may be performed jointly as equation (15) below or separately as equation (16) below from horizontal and vertical directions. if maxΔvx < threshvx or if maxΔvy < threshvy, ΔI(i,j) = 0 (15)
[0124] In the equations (15) and (16) above, different close-form expressions of ΔI(i,j) represents different simplification methods. For example, in the case of ΔI(i,j) = gy(i,j) * Δ vy (i, j) , the calculation of gx and Δνχ can be skipped.
[0125] In another example, the check of the derived Δv(i,j) may be combined with non- simplified AMPR operation. In this case, the equation (14) may be merged with equation (4), then the prediction refinement value is calculated by equation below:
[0126] In some examples, the threshvx or threshvy may be with different or the same value.
[0127] In some examples, when Δν(i,j) is derived for a sub-block, the values of threshvx and threshvy may be determined depending on which position is used for deriving the sub- block MV. In other words, a different or same pair of values of threshvx and threshvy may be determined for two sub-blocks if their MVs are derived using different positions. For example, for a sub-block whose sub-block level MV is derived based on the sub-block center, its pair of values of threshvx and threshvy may be the same or different from that of a sub- block whose sub-block level MV is derived based on the position of the sub-block top-left comer.
[0128] In some examples, the values of threshvx and threshvy may be defined as in the range of [1/32, 1/16] in unit of pixels. For example, a value of (1/16)*(10/16), (1/16)*(12/16), or (1/16)*(14/16) may be used as the thresholds. In this case, the threshold value is a floating point of 1/16-pel, e.g., 0.625 in unit of 1/16-pel, 0.75 in unit of 1/16-pel, or 0.875 in unit of 1/16-pel.
[0129] In some examples, the values of threshvx and threshvy may be defined based on picture types. For low-delay pictures, the derived affine model parameters may have smaller magnitude than other non-low-delay pictures, since low-delay pictures tends to have smaller and/or smoother motions and therefore smaller values may be preferred for those thresholds. [0130] In some examples, the values of threshvx and threshvy may be the same regardless of different picture types.
[0131] In some examples, if the absolute value of the majority of derived g(i,j) for all the pixels within a sub-block are smaller than a threshold value, AMPR may be skipped for this sub-block. One example of this method is that a sub-block contains smooth surface which may consist of flat textures (e.g., with no or small number of high-frequency details).
[0132] In some examples, the significance of Δv(i,j) and g(i,j) may be considered jointly or used in a hybrid manner to decide whether AMPR should be skipped for current sub- block or CU.
[0133] In case the implementation of AMPR is approximated by using the above-described filtering process, the selective enabling of AMPR may be also performed. For example, based on the equation (15), if maxΔvx and/or maxΔvy is smaller than a predefined threshold threshvx and threshvy, the corresponding scaled coefficients in the selected filter may become 0, which means the filtering process may be simplified from 2D to one-dimensional (ID) filtering process. Taking the 5 -tap filters in FIG. lOA as example, when only maxAVx is smaller than threshvx, only the ID filter [1, 2, 1] is applied in the vertical direction. When only maxΔvy is smaller than threshvy, only the ID filter [1, 2, 1] is applied in the horizontal direction. When neither maxΔvx and maxAVy is smaller than the corresponding threshold, i. e. , threshvx and threshvy , the 2D filter is applied in both horizontal and vertical directions. When both maxΔvx and maxΔvy are smaller than the corresponding threshold, no filter is applied at all, i.e., no AMPR is applied.
Encoder Side Optimization
[0134] In the AVS3 standards, affine UMVE mode is computation intensive for encoder because it involves choosing the best distance index for each merge mode candidate. When sum of absolute transformed difference (SATD) cost is calculated for each candidate distance index, regular affine motion compensation is always applied. If AMPR is applied on top of affine motion compensation, the computation can increase dramatically.
[0135] In one example, AMPR operation is skipped during SATD based cost calculation for affine UMVE mode at the encoder side. It is found through experiments that while the best index is selected according to the best SATD cost, whether AMPR is applied during the SATD calculation or not usually does not change the ranking of the best SATD cost. Therefore, with the proposed method enabling AMPR mode wouldn’t incur obvious encoder complexity for affine UMVE mode.
[0136] Motion estimation is another major overhead at the encoder side. In another example, AMPR process may be skipped depending on certain conditions. These conditions indicate that the best encoding mode of a CU is unlikely to be affine mode after mode selection process.
[0137] One example of such a condition is whether a current CU has a parent CU which is already determined to be coded by explicit affine mode or affine merge mode. This is due to the strong correlation of coding mode selection between a CU and its parent CU, and it is more likely that the best coding mode for the current CU is also explicit affine mode if the condition above is true.
[0138] Another exemplar condition used for enabling AMPR is whether the parent CU of the current CU is determined to be inter-predicted with explicit affine mode. If it is true, AMPR is applied during affine motion estimation of the current CU; otherwise, AMPR is skipped during affine motion estimation of the current CU.
[0139] Compared to large block size CU such as 64x64 CUs, small size CUs such as 16x16 CU have much higher average per-pixel computation cost when AMPR is applied. To effectively save computation complexity, in another example of the present disclosure, during motion estimation process, AMPR may be skipped for small size CUs. The size of a CU may be defined as the total number of pixels. A pixel number threshold such as 16x16 or 16x32 or 32*32 may be defined, and for a block with a size smaller than the defined threshold, AMPR may be skipped during affine motion estimation process for the block.
[0140] In case the implementation of AMPR is approximated by using above-described filtering process, the encoder side optimization may be also performed. For example, the filtering process may not be performed for affine UMVE mode. Another encoder optimization used for enabling the filtering process based AMPR is whether the parent CU of the current CU is determined to be inter-predicted with explicit affine mode. If it is true, the filtering process of AMPR is applied during affine motion estimation of the current CU. Otherwise, the filtering process is skipped during affine motion estimation of the current CU.
[0141] FIG. 14 is a block diagram illustrating an apparatus for predicting a sample at a pixel location in a sub-block by AMPR in accordance with some implementations of the present disclosure. The apparatus 1400 may be a terminal, such as a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.
[0142] As shown in FIG. 14, the apparatus 1400 may include one or more of the following components: a processing component 1402, a memory 1404, a power supply component 1406, a multimedia component 1408, an audio component 1410, an input/output (I/O) interface 1412, a sensor component 1414, and a communication component 1416.
[0143] The processing component 1402 usually controls overall operations of the apparatus 1400, such as operations relating to display, a telephone call, data communication, a camera operation, and a recording operation. The processing component 1402 may include one or more processors 1420 for executing instructions to complete all or a part of steps of the above method. Further, the processing component 1402 may include one or more modules to facilitate interaction between the processing component 1402 and other components. For example, the processing component 1402 may include a multimedia module to facilitate the interaction between the multimedia component 1408 and the processing component 1402.
[0144] The memory 1404 is configured to store different types of data to support operations of the apparatus 1400. Examples of such data include instructions, contact data, phonebook data, messages, pictures, videos, and so on for any application or method that operates on the apparatus 1400. The memory 1404 may be implemented by any type of volatile or non-volatile storage devices or a combination thereof, and the memory 1404 may be a Static Random Access Memory (SRAM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), an Erasable Programmable Read-Only Memory (EPROM), a Programmable Read-Only Memory (PROM), a Read-Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disk or a compact disk.
[0145] The power supply component 1406 supplies power for different components of the apparatus 1400. The power supply component 1406 may include a power supply management system, one or more power supplies, and other components associated with generating, managing and distributing power for the apparatus 1400.
[0146] The multimedia component 1408 includes a screen providing an output interface between the apparatus 1400 and a user. In some examples, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen
may be implemented as a touch screen receiving an input signal from a user. The touch panel may include one or more touch sensors for sensing a touch, a slide and a gesture on the touch panel. The touch sensor may not only sense a boundary of a touching or sliding actions, but also detect duration and pressure related to the touching or sliding operation. In some examples, the multimedia component 1408 may include a front camera and/or a rear camera. When the apparatus 1400 is in an operation mode, such as a shooting mode or a video mode, the front camera and/or the rear camera may receive external multimedia data [0147] The audio component 1410 is configured to output and/or input an audio signal. For example, the audio component 1410 includes a microphone (MIC). When the apparatus 1400 is in an operating mode, such as a call mode, a recording mode and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal may be further stored in the memory 1404 or sent via the communication component 1416. In some examples, the audio component 1410 further includes a speaker for outputting an audio signal.
[0148] The I/O interface 1412 provides an interface between the processing component 1402 and a peripheral interface module. The above peripheral interface module may be a keyboard, a click wheel, a button, or the like. These buttons may include but not limited to, a home button, a volume button, a start button, and a lock button.
[0149] The sensor component 1414 includes one or more sensors for providing a state assessment in different aspects for the apparatus 1400. For example, the sensor component 1414 may detect an on/off state of the apparatus 1400 and relative locations of components. For example, the components are a display and a keypad of the apparatus 1400. The sensor component 1414 may also detect a position change of the apparatus 1400 or a component of the apparatus 1400, presence or absence of a contact of a user on the apparatus 1400, an orientation or acceleration/deceleration of the apparatus 1400, and a temperature change of apparatus 1400. The sensor component 1414 may include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 1414 may further include an optical sensor, such as a CMOS or CCD image sensor used in an
imaging application. In some examples, the sensor component 1414 may further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0150] The communication component 1416 is configured to facilitate wired or wireless communication between the apparatus 1400 and other devices. The apparatus 1400 may access a wireless network based on a communication standard, such as WiFi, 4G, or a combination thereof. In an example, the communication component 1416 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an example, the communication component 1416 may further include a Near Field Communication (NFC) module for promoting short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, Ultra-Wide Band (UWB) technology, Bluetooth (BT) technology and other technology.
[0151] In an example, the apparatus 1400 may be implemented by one or more of Application Specific Integrated Circuits (ASIC), Digital Signal Processors (DSP), Digital Signal Processing Devices (DSPD), Programmable Logic Devices (PLD), Field Programmable Gate Arrays (FPGA), controllers, microcontrollers, microprocessors, or other electronic elements to perform the above method.
[0152] A non-transitoiy computer readable storage medium may be, for example, a Hard Disk Drive (HDD), a Solid-State Drive (SSD), Flash memory, a Hybrid Drive or Solid-State Hybrid Drive (SSHD), a Read-Only Memory (ROM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, etc.
[0153] FIG. 15 is a flowchart illustrating an exemplary process for AMPR in accordance with some implementations of the present disclosure.
[0154] In step 1502, the processor 1420 determines whether a neighboring sample location to a sample location in a sub-block is outside of the sub-block.
[0155] In some examples, the sample location may include Y-sample location, U-sample location, V-sample location, or a corresponding pixel location. The corresponding pixel is the
pixel at location (/, j) including the sample at the location (/, j) as described in Exemplary
Workflow of AMPR.”
[0156] In some examples, the neighboring sample location may be determined based on the shape of a filter, such as the cross-shape filter shown in FIGS. 10A-10B and 11A-11B or the square-shape filter.
[0157] In step 1504, the processor 1420 copies a padding sample to the neighboring sample location for a filter-based AMPR implementation in response to determining that the neighboring sample location is outside of the sub-block.
[0158] In some examples, when it is determined that the neighboring sample location is located outside of the sub-block, the current sample at the sample location is the current boundary sample, as shown in FIG. 13. The current boundary sample is located at one of the four boundaries of the sub-block.
[0159] In some examples, the padding sample may be a reference sample selected from reference samples at integer sample positions or locations in a reference picture, an average sample of the reference samples, a prediction sample, etc.
[0160] In some examples, the processor 1420 may obtain multiple affine motion- compensated predictions for a plurality of samples at the sample location and a plurality of neighboring sample locations including one or more neighboring sample locations outside of the sub-block; and obtain, at the sample location, a refined prediction for a sample at the sample location using a filter with a pre-determined shape based on the multiple affine motion- compensated predictions, where the pre-determined shape covers the sample location and the plurality of neighboring sample locations.
[0161] In some examples, the sample location may be at one boundary of the four boundaries of the sub-block, and the plurality of neighboring sample locations may include the one or more neighboring samples that are outside of the four boundaries of the sub-block, for example, in the extended region as shown in FIG. 13.
[0162] In some examples, the processor 1420 may obtain the multiple affine motion- compensated predictions by performing a sub-block-based affine motion compensation on a video picture that comprises a plurality of sub-blocks.
[0163] In some examples, the processor 1420 may determine a reference sample at an integer position in a reference picture and determine the reference sample as the padding sample. [0164] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining a neighboring reference sample in the reference picture, where the neighboring reference sample is a reference sample at an integer position nearest to a prediction sample that is corresponding to the sample location. In one case, the prediction sample is in a reference picture that is previously reconstructed and located at a corresponding position. The prediction sample at the corresponding position in the reference picture may have same or the most similar image area or contents as the current sample at the sample position in the current sub-block in the current picture.
[0165] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located left to a prediction sample and above the prediction sample, such as the integer sample LT in FIG. 12. The prediction sample may be corresponding to the sample position. [0166] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located right to a prediction sample and above the prediction sample, such as the integer sample RT in FIG. 12.
[0167] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer position located left to die prediction sample and below the prediction sample, such as the integer sample LB in FIG. 12. The prediction sample is corresponding to the sample position. [0168] In some examples, the processor 1420 may determine the reference sample at the integer position in the reference picture by determining the reference sample at the integer
position located right to the prediction sample and below the prediction sample, such as the integer sample RB in FIG. 12. The prediction sample is corresponding to the sample position. [0169] In some examples, the processor 1420 may determine an average sample of a plurality of reference samples at a plurality of integer positions in the reference picture and determine the average sample as the padding example. In one example, the average sample may be an average of two closest integer samples to the sample location. In another example, the average sample may be an average of the four integer reference samples LT, RT, LB, and RB.
[0170] In some examples, the processor 1420 may determine the padding sample using an interpolation filter.
[0171] In some examples, the processor 1420 may determine the padding sample by interpolating according to distances between each of a plurality of reference samples at a plurality of integer positions in the reference picture and a prediction sample corresponding to the sample location. For example, the padding sample is determined by interpolating according to the distances of four integer reference samples LT, RT, LB, and RB to a fractional position of the prediction sample.
[0172] In some examples, the processor 1420 may determine a prediction sample at one boundary of the sub-block as the padding sample. FIG. 13 shows the sub-block of four boundaries. The four boundaries are four sides of the sub-block between the sub-block and the extended region.
[0173] Specifically, the processor 1420 may determine the prediction sample corresponding to the sample location as the padding sample when it is determined that the sample position is at one boundary. The prediction sample is in a reference picture previously reconstructed and located at a corresponding position. The prediction sample at the corresponding position may have same or the most similar image area or contents as the current sample at the sample position in the current picture.
[0174] In some examples, there is provided an apparatus for AMPR. The apparatus includes one or more processors 1420 and a memory 1404 configured to store instructions executable
by the one or more processors; where the processor, upon execution of the instructions, is configured to perform any method as described in FIG. 15 and above.
[0175] In some other examples, there is provided a non-transitory computer readable storage medium 1404, having instructions stored therein. When the instructions are executed by one or more processors 1420, the instructions cause the processor to perform any method as described in FIG.15 and above.
[0176] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those of ordinary skill in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.
[0177] The examples were chosen and described in order to explain the principles of the disclosure, and to enable others skilled in the art to understand the disclosure for various implementations and to best utilize the underlying principles and various implementations with various modifications as are suited to the particular use contemplated. Therefore, it is to be understood that the scope of the disclosure is not to be limited to the specific examples of the implementations disclosed and that modifications and other implementations are intended to be included within the scope of the present disclosure.
Claims
1. A method for affine motion-compensated prediction refinement (AMPR), comprising: determining whether a neighboring sample location to a sample location in a sub- block is outside of the sub-block; and in response to determining that the neighboring sample location is outside of the sub- block, copying a padding sample to the neighboring sample location for a filter-based AMPR implementation.
2. The method of claim 1, wherein the filter-based AMPR implementation comprises: obtaining multiple affine motion-compensated predictions for a plurality of samples at the sample location and a plurality of neighboring sample locations including one or more neighboring sample locations outside of the sub-block; and obtaining, at the sample location, a refined prediction for a sample at the sample location using a filter with a pre-determined shape based on the multiple affine motion- compensated predictions, wherein the pre-determined shape covers the sample location and the plurality of neighboring sample locations.
3. The method of claim 2, further comprising: obtaining the multiple affine motion-compensated predictions by performing a sub- block-based affine motion compensation on a video picture that comprises a plurality of sub- blocks.
4. The method of claim 1, further comprising: determining a reference sample at an integer position in a reference picture; and determining the reference sample as the padding sample.
5. The method of claim 4, wherein determining the reference sample at the integer position in the reference picture comprises:
determining a neighboring reference sample in the reference picture, wherein the neighboring reference sample is a reference sample at an integer position nearest to a prediction sample corresponding to the sample location.
6. The method of claim 4, wherein determining the reference sample at the integer position in the reference picture comprises: determining the reference sample at the integer position located left to a prediction sample and above the prediction sample, wherein the prediction sample is corresponding to the sample position.
7. The method of claim 4, wherein determining the reference sample at the integer position in the reference picture comprises: determining the reference sample at the integer position located right to a prediction sample and above the prediction sample, wherein the prediction sample is corresponding to the sample position.
8. The method of claim 4, wherein determining the reference sample at the integer position in the reference picture comprises: determining the reference sample at the integer position located left to the prediction sample and below the prediction sample, wherein the prediction sample is corresponding to the sample position.
9. The method of claim 4, wherein determining the reference sample at the integer position in the reference picture comprises: determining the reference sample at the integer position located right to the prediction sample and below the prediction sample, wherein the prediction sample is corresponding to the sample position.
10. The method of claim 1, further comprising: determining an average sample of a plurality of reference samples at a plurality of integer positions in a reference picture; and determining the average sample as the padding example.
11. The method of claim 1, further comprising: determining the padding sample using an interpolation filter.
12. The method of claim 11 , further comprising: determining the padding sample by interpolating according to distances between each of a plurality of reference samples at a plurality of integer positions in a reference picture and a prediction sample corresponding to the sample location.
13. The method of claim 1, further comprising: determining a prediction sample at a boundary of the sub-block as the padding sample.
14. The method of claim 13, further comprising: determining the prediction sample corresponding to the sample location at the boundary of the sub-block as the padding sample.
15. An apparatus for affine motion-compensated prediction refinement (AMPR), comprising: one or more processors; and a memory configured to store instructions executable by the one or more processors; wherein the one or more processors, upon execution of the instructions, are configured to perform the method in any of claims 1-14.
16. A non-transitory computer-readable storage medium for affine motion-compensated prediction refinement (AMPR) storing computer-executable instructions that, when executed by one or more processors, causing the one or more processors to perform the method in any of claims 1-14.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202180056990.9A CN116171576A (en) | 2020-08-06 | 2021-08-05 | Method and apparatus for affine motion compensated prediction refinement |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202063062380P | 2020-08-06 | 2020-08-06 | |
| US63/062,380 | 2020-08-06 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022032028A1 true WO2022032028A1 (en) | 2022-02-10 |
Family
ID=80118573
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2021/044840 Ceased WO2022032028A1 (en) | 2020-08-06 | 2021-08-05 | Methods and apparatuses for affine motion-compensated prediction refinement |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116171576A (en) |
| WO (1) | WO2022032028A1 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023220970A1 (en) * | 2022-05-18 | 2023-11-23 | Oppo广东移动通信有限公司 | Video coding method and apparatus, and device, system and storage medium |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090238276A1 (en) * | 2006-10-18 | 2009-09-24 | Shay Har-Noy | Method and apparatus for video coding using prediction data refinement |
| WO2019002171A1 (en) * | 2017-06-26 | 2019-01-03 | Interdigital Vc Holdings, Inc. | Method and apparatus for intra prediction with multiple weighted references |
| US20190082193A1 (en) * | 2017-09-08 | 2019-03-14 | Qualcomm Incorporated | Motion compensated boundary pixel padding |
| WO2020147747A1 (en) * | 2019-01-15 | 2020-07-23 | Beijing Bytedance Network Technology Co., Ltd. | Weighted prediction in video coding |
| US20200244965A1 (en) * | 2017-11-07 | 2020-07-30 | Huawei Technologies Co., Ltd. | Interpolation filter for an inter prediction apparatus and method for video coding |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20180199057A1 (en) * | 2017-01-12 | 2018-07-12 | Mediatek Inc. | Method and Apparatus of Candidate Skipping for Predictor Refinement in Video Coding |
| US20190116376A1 (en) * | 2017-10-12 | 2019-04-18 | Qualcomm Incorporated | Motion vector predictors using affine motion model in video coding |
-
2021
- 2021-08-05 CN CN202180056990.9A patent/CN116171576A/en active Pending
- 2021-08-05 WO PCT/US2021/044840 patent/WO2022032028A1/en not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090238276A1 (en) * | 2006-10-18 | 2009-09-24 | Shay Har-Noy | Method and apparatus for video coding using prediction data refinement |
| WO2019002171A1 (en) * | 2017-06-26 | 2019-01-03 | Interdigital Vc Holdings, Inc. | Method and apparatus for intra prediction with multiple weighted references |
| US20190082193A1 (en) * | 2017-09-08 | 2019-03-14 | Qualcomm Incorporated | Motion compensated boundary pixel padding |
| US20200244965A1 (en) * | 2017-11-07 | 2020-07-30 | Huawei Technologies Co., Ltd. | Interpolation filter for an inter prediction apparatus and method for video coding |
| WO2020147747A1 (en) * | 2019-01-15 | 2020-07-23 | Beijing Bytedance Network Technology Co., Ltd. | Weighted prediction in video coding |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023220970A1 (en) * | 2022-05-18 | 2023-11-23 | Oppo广东移动通信有限公司 | Video coding method and apparatus, and device, system and storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116171576A (en) | 2023-05-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12413770B2 (en) | Methods and apparatus of motion vector rounding, clipping and storage for inter prediction | |
| JP7717201B2 (en) | Video encoding and decoding method and apparatus for triangular prediction | |
| CN114128263B (en) | Method and apparatus for adaptive motion vector resolution in video coding | |
| WO2021188598A1 (en) | Methods and devices for affine motion-compensated prediction refinement | |
| JP2024016288A (en) | Method and apparatus for decoder side motion vector correction in video coding | |
| WO2022081878A1 (en) | Methods and apparatuses for affine motion-compensated prediction refinement | |
| CN114009017A (en) | Motion compensation using combined inter and intra prediction | |
| US20240098290A1 (en) | Methods and devices for overlapped block motion compensation for inter prediction | |
| US12470733B2 (en) | Overlapped block motion compensation for inter prediction | |
| WO2022032028A1 (en) | Methods and apparatuses for affine motion-compensated prediction refinement | |
| US12356001B2 (en) | Methods and apparatus of motion vector rounding, clipping and storage for inter prediction | |
| CN115567709B (en) | Method and apparatus for encoding samples at pixel locations in a sub-block | |
| WO2022026480A1 (en) | Weighted ac prediction for video coding | |
| CN115997382A (en) | Method and apparatus for prediction related residual scaling for video coding | |
| CN114051732A (en) | Method and apparatus for decoder-side motion vector refinement in video coding | |
| WO2021248135A1 (en) | Methods and apparatuses for video coding using satd based cost calculation | |
| WO2021007133A1 (en) | Methods and apparatuses for decoder-side motion vector refinement in video coding |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21853809 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21853809 Country of ref document: EP Kind code of ref document: A1 |






















