WO2025214829A1 - Motion vector refinement with downsampled blocks, or hierarchical sub-blocks - Google Patents
Motion vector refinement with downsampled blocks, or hierarchical sub-blocksInfo
- Publication number
- WO2025214829A1 WO2025214829A1 PCT/EP2025/058886 EP2025058886W WO2025214829A1 WO 2025214829 A1 WO2025214829 A1 WO 2025214829A1 EP 2025058886 W EP2025058886 W EP 2025058886W WO 2025214829 A1 WO2025214829 A1 WO 2025214829A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- block
- samples
- motion vector
- downscaled
- ovt
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/124—Quantisation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/523—Motion estimation or motion compensation with sub-pixel accuracy
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/56—Motion estimation with initialisation of the vector search, e.g. estimating a good candidate to initiate a search
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/577—Motion compensation with bidirectional frame interpolation, i.e. using B-pictures
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/59—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial sub-sampling or interpolation, e.g. alteration of picture size or resolution
Definitions
- This disclosure relates to methods and apparatus for motion vector refinement.
- VVC Versatile Video Coding
- ECM Enhanced Coding Model
- a video (a.k.a., “video sequence”) comprises of a series of pictures (a.k.a., frames or images).
- each picture is identified with a picture order count (POC) value.
- POC picture order count
- the POC value also represents the display order of the picture.
- a picture with a smaller POC value is displayed before another picture with a larger POC value.
- Each component can be described as a two-dimensional rectangular array of sample values. It is common that each picture consists of three components; one luma component Y where the sample values are luma values and two chroma components Cb and Cr, where the sample values are chroma values.
- the dimensions of the chroma components are smaller than the luma components by a factor of two in each dimension.
- the size of the luma component of an HD picture would be 1920x1080 and the chroma components would each have the dimension of 960x540.
- Components are sometimes referred to as color components.
- a block is a two-dimensional (2D) array of sample values (or “samples” for short).
- a block may be divided into two or more blocks (a.k.a., subblocks), where each subblock is a 2D array of samples, which is also known as a “matrix” of samples.
- each component of a picture is split into blocks and the coded video bitstream consists of a series of coded blocks. It is common in video coding that pictures are split into units that cover a specific area of the picture.
- Each unit consists of all blocks from all components that make up that specific area of the picture and each block belongs fully to one unit.
- the Coding Unit (CU) in VVC is an example of a unit.
- the CUs may be split recursively to smaller CUs.
- the CU at the top level is referred to as the coding tree unit (CTU).
- CTU coding tree unit
- a CU usually contains three coding blocks, i.e., one coding block for luma and two coding blocks for chroma.
- the size of luma coding block is same as the CU.
- the CUs can have size of 4x4 up to 128x128.
- the CUs can have size of 4x4 up to 256x256.
- VVC specifies three types of parameter sets: the picture parameter set (PPS), the sequence parameter set (SPS), and the video parameter set (VPS).
- PPS picture parameter set
- SPS sequence parameter set
- VPS video parameter set
- the PPS contains data that is common for a whole picture
- the SPS contains data that is common for a coded layer video sequence (CLVS)
- CLVS coded layer video sequence
- the VPS contains data that is common for multiple CLVSs, e.g., data for multiple layers in the bitstream.
- slices divides the picture into independently coded slices, where decoding of one slice in a picture is independent of other slices of the same picture.
- Each slice has a slice header comprising syntax elements. Decoded slice header values from these syntax elements are used when decoding the slice.
- intra prediction also known as spatial prediction
- a block is predicted using previous decoded blocks within the same picture.
- the samples from the previously decoded blocks within the same picture are used to predict the samples inside the current block.
- a picture consisting of only intra-predicted blocks is referred to as an intra picture.
- inter prediction also known as temporal prediction
- blocks of the current picture are predicted using blocks from previously decoded pictures (these blocks are referred to as reference blocks).
- the samples from the reference blocks in the previously decoded pictures are used to predict the samples inside the current block.
- a picture that comprises one or more inter-predicted blocks is referred to as an inter-picture.
- the previous decoded pictures used for inter prediction are referred to as reference pictures.
- MV motion vector
- x value a.k.a., x component
- y value a.k.a., y component
- the value of a component may have a resolution finer than an integer position. When that is the case, a filtering (typically interpolation) is done to calculate values used for prediction.
- FIG. 1 shows an example of an MV for the current block.
- the example MV consist of a horizontal value of 2 and a vertical value of 1.
- An inter picture may use several reference pictures.
- the reference pictures are usually put into two reference picture lists, L0 and LI.
- the reference pictures that are output before the current picture are typically the first pictures in L0.
- the reference pictures that are output after the current picture are typically the first pictures in LI.
- FIG. 2 shows an example of two different prediction types: uni -prediction and bi-prediction.
- Inter predicted blocks can use uni -prediction or bi-prediction.
- a uni-predicted block uses one MV (e.g., MV0 in FIG. 2) to predict from one reference picture, either using L0 or LL Bi-prediction uses a pair of motion vectors (denoted Bi-MV), to predict from a first reference picture, such as, for example, reference picture 0 in FIG. 2, and a second reference picture, such as, for example, reference picture 1 in FIG. 2.
- the Bi-MV consists of a first MV (MV1) and a second MV (MV2), where the first MV is used to predict from the first reference picture and the second MV is used to predict from the second reference picture.
- a low delay picture is a picture that has all its reference pictures displayed before the picture. In other words, for a low delay picture, all its reference pictures have smaller POC values than the current POC.
- a non-low delay picture is a picture that has at least one of its reference pictures displayed after the picture.
- a non-low delay picture has at least one reference picture with a larger POC value than the current POC.
- the value of the MV’s x or y component may correspond to a sample position which has finer granularity than integer (sample) position. Those positions are also referred to as fractional (sample) positions. In WC and current ECM, the MV can be at 1/16 sample position.
- FIG. 3 depicts several fractional positions in the horizontal (x-) dimension.
- the solid-square blocks represent integer positions.
- the circles represent 1/16-position.
- MV (4, 10) means the x component is at 4/16 position, the y component is at 10/16 position.
- an MV rounding process is sometimes used to convert an MV at one position to another target position.
- One example of rounding is to round a fractional MV position to the nearest integer position.
- a residual block is generated by computing the difference between samples of a source block (a.k.a., “input block”), which contains original samples, and samples of a prediction block (a.k.a., “residual block”) (e.g., an inter-prediction block or an intraprediction block).
- a source block a.k.a., “input block”
- samples of a prediction block a.k.a., “residual block”
- residual block e.g., an inter-prediction block or an intraprediction block.
- QP quantization parameter
- a coded block flag CBF is used to indicate if there are any non-zero quantized transform coefficients.
- All coding parameters are then entropy coded at the encoder and decoded at the decoder. If the coded block flag is one, a reconstructed block can then be derived by inverse quantization and inverse transformation of the quantized transform coefficients and then add that to the prediction block. If the coded block flag is zero, the reconstructed block is identical to the prediction block.
- inter prediction information For an inter block inside an inter picture in VVC, its inter prediction information consists of the following three elements:
- a reference picture list flag which signals which reference picture list is used for the block (when the value of the flag is equal to 0, it means only L0 is used for predicting the current block, when the value of the flag is equal to 1, it means only LI is used for predicting the current block, and when the value of the flag is equal to 2, it means both L0 and LI are used for predicting the current block);
- MV motion vector
- the inter prediction information is also referred to as motion information.
- the decoder stores the motion information for each inter block. In other words, an inter block maintains its own motion information.
- the RD cost is calculated as D + * Rate.
- D (Distortion) measures the difference between the reconstructed block and the corresponding source block.
- One commonly used metric for calculating D is the sum of squared error SSE ) —
- Rate is usually an estimation of the bits to be spent on encoding the mode, and is a trade-off parameter between Rate and D.
- VVC and ECM include several methods for implicit signaling of motion information for each block, including the merge method and the subblock merge method.
- a common motivation behind the implicit methods is to inherit or reuse motion information from neighboring coded blocks. This often works in practice due to spatial correlation of close-by blocks, i.e., the fact that nearby blocks often behave similarly.
- the merge method derives a set of motion information from previously decoded blocks and use the derived motion information for generating the samples of the entire block.
- the merge method is sometimes referred to as the block merge method.
- the merge method first generates a list of motion information candidates.
- the list is also referred to as the merge list.
- the candidates are derived from previously coded blocks. These previously coded blocks can be spatially adjacent neighboring blocks or temporal collocated blocks relative to the current block.
- FIG. 4 shows the spatial neighboring blocks: left (L), top (T), top-right (TR), left-bottom (LB) and top-left (TL).
- the merge list construction process usually checks the previously coded blocks in a predefined order, for example: T, L, TR, LB, then TL. For each previously coded block being checked, if this previously coded block is inter coded and its motion information has no duplicates in the list, then the motion information of this previously coded block is added to the merge list.
- the merge list is generated, one of the candidates inside the list is used to derive the motion information of the current block.
- the candidate selection process is done on the encoder side. An encoder would select a best candidate from the list and encode an index (merge index) in the bitstream to signal to a decoder. The decoder receives the index, it follows the same merge list derivation process as the encoder and uses the index to retrieve the correct candidate.
- the blocks that use the block merge method are sometimes referred to as blocks in merge mode.
- non-adjacent spatial blocks are also considered as sources of motion information during the merge list construction.
- FIG. 5 shows some examples (marked with NA1, NA2 and NA3) of those non-adjacent spatial blocks.
- VVC and ECM also include the subblock merge method. It splits a current block into a number of subblocks and allows each subblock to have its own motion information.
- FIG. 6 shows an example of a current block and its subblocks. Each subblock maintains its own motion information. It should be noted that the subblocks are all rectangular.
- VVC only one so called collocated reference picture can be indicated and the motion is temporally predicted from a corresponding position in the collocated picture (either from L0 or LI).
- ECM supports two collocated pictures to be indicated, one from L0 and one from LI.
- ECM also supports that the corresponding position in the collocated picture is offset vertically and/or horizontally with some variation.
- OBMC Overlapped Block Motion Compensation
- OBMC is a tool included in ECM which operates at the block boundaries or subblock boundaries of a current inter block.
- OBMC blends the current block or subblock’s prediction samples, P CUR, (generated using the current associated motion information) with another set of prediction samples, P_NB, which are generated using motion information from a neighboring block or a neighboring subblock at the current block or subblock boundaries.
- the OBMC blending process takes weighted average of P CUR and P_NB to produce a set of OBMC modified prediction samples, P OBMC.
- the set P CUR or P_NB each associate with a respective weighting factor
- the OBMC may give better prediction for samples that are close to the block or subblock boundary.
- BDOF is a tool included in VVC and the current ECM that can be used to refine prediction samples that are generated from a Bi-MV. BDOF relies on optical flow estimation to derive a pair of refinement parameter (Vx, Vy) which can be further used to refine the prediction samples.
- DMVR is a tool included in VVC and the current ECM to refine motion vectors for a Bi-MV.
- DMVR operates on subblock level, usually 16x16. Different from BDOF, which relies on optical flow estimation, DMVR relies on bilateral matching of two reference blocks to refine the Bi-MV. The matching is based on SAD (sum of absolute differences).
- the DMVR searches within a window around the Bi-MV to find whether there exists another Bi-MV (Bi-MV’) that gives a better match between the L0 reference block and the LI reference block. If so, the Bi-MV’ is further used instead for generating the prediction samples of the current block. After that a subpel adjustment is made based on the SAD costs around the motion with least SAD to determine subblock motion with sub pixel accuracy.
- DMVR has been further evolved to use multi-pass optimization, first bilateral block matching in a search area to find the refinement of the merge motion that gives best match, then bilateral 16x16 subblock matching, then bi-directional optical flow on 8x8 subblocks to refine motion further.
- the multi-pass DMVR also allows for modification of only one of the bi-predictive motions. It is also allowed to use DMVR when BCW (other weightings than just average) is used.
- ECM block or subblock motion can also be refined by matching a template outside the current block with a corresponding template on the reference picture.
- the search is limited to a small range to find a better motion without signaling additional motion information. This can also be used cascaded with the bilateral matching in multi-pass DMVR.
- Another issue is that when DMVR is used in bi-prediction with non-equal distance to reference pictures the derived refinement from DMVR based on symmetrical assumption is scaled so that the motion vector closer to a reference picture is reduced according to difference in the distances to the reference pictures. Which means the motion vector may become closer to the correct motion vector but likely with some inaccuracy since the bilateral match may be worse after the scaling of the motion vector.
- the method also includes producing a first refined MV, mv_refmedA, and a second refined MV, mv_refmedB, using downscaled blocks.
- the method also includes obtaining motion compensated predicted samples using MV_refmedA and MV_refmedB.
- Producing mv_refmedA and mv_refmedB using downscaled blocks comprises (1) producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector (e.g., associated with ov_l and a first reference point associated with the first motion vector) and a second downscaled block associated with ov_2 and the second motion vector (e.g., associated with ov_2 and a second reference point associated with the second motion vector); (2) producing a score for a second OVT comprising a third offset vector, ov_3, and a fourth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector (e.g.,
- the method also includes determining a first offset vector, ov_l, using a first motion vector for a second block of samples (e.g., a 32x32 block) and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples.
- the method also includes determining a second offset vector, ov_2, using a first reference motion vector for a third block of samples (e.g., a 16x16 block) and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples.
- the method also includes producing a first refined motion vector, mv_refmed_3_l, for the third block of samples based on mv_refA, ov_l, and ov_2.
- the method also includes obtaining motion compensated predicted samples using the first refined motion vector.
- the method also includes selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples (e.g., a 32x32 block), wherein the second block of samples is a subblock within the first block of samples.
- the method also includes producing a first refined motion vector, mv_refmedl, for the second block of samples based on mv_refA and the first OVT.
- the method also includes producing a second refined motion vector, mv_refmed2, for the second block of samples based on mv refB and the first OVT.
- the method also includes obtaining motion compensated predicted samples using the first and the second refined motion vectors.
- Selecting the first OVT comprises: (1) producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples (e.g., associated with the first offset vector of the first OVT and a reference point associated with the first motion vector for the second block of samples) and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples (e.g., associated with the second offset vector of the first OVT and a second reference point associated with the second motion vector for the second block of samples); (2) producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and (3) determining that the score for the first OVT is better than the
- a carrier containing the computer program of the above embodiment, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.
- An advantage of the embodiments disclosed herein is that the refined motion vectors are more accurate, thereby providing a bitrate savings while at the same time providing no reduction in objective quality. Furthermore, the more accurate motion vectors provide increased subjective quality.
- FIG. 1 shows an example of a motion vector (MV).
- FIG. 2 illustrates uni-inter prediction and bi-inter prediction.
- FIG. 3 depicts several fractional positions in the horizontal (x-) dimension.
- FIG. 4 shows spatial neighboring blocks.
- FIG. 5 shows some examples of non-adjacent spatial blocks.
- FIG. 6 shows an example of a current block and its subblocks.
- FIG. 7 illustrates a system according to some embodiments.
- FIG. 8 is a schematic block diagram of an encoder according to an embodiment.
- FIG. 9 is a schematic block diagram of a decoder according to an embodiment.
- FIG. 10 illustrates a bi-directional motion vector
- FIG. 11 illustrates downscaling a search area.
- FIGs. 12A and 12B illustrate a process of finding an optimal offset vector tuple.
- FIG. 13 illustrates Hierarchical Motion Estimation according to an embodiment.
- FIG. 14 is a flowchart illustrating a process according to some embodiments.
- FIG. 15 is a flowchart illustrating a process according to some embodiments.
- FIG. 16 is a flowchart illustrating a process according to some embodiments.
- FIG. 17 is a block diagram of an apparatus according to some embodiments.
- FIG. 7 illustrates a system 700 according to an embodiment.
- System 700 includes an encoder 702 and a decoder 704, wherein, in the example shown, encoder 702 is in communication with decoder 704 via a network 110 (e.g., the Internet or other network).
- Encoder 702 encodes a source video sequence 701 (e.g., encodes blocks of units of pictures of video sequence 701) into a bitstream comprising an encoded video sequence (e.g., encoded blocks) and transmits the bitstream to decoder 704 via network 708.
- a source video sequence 701 e.g., encodes blocks of units of pictures of video sequence 701
- bitstream comprising an encoded video sequence (e.g., encoded blocks)
- encoder 702 is not in communication with decoder 704, and, in such an embodiment, rather than transmitting bitstream to decoder 704, the bitstream is stored in a data storage unit 790 and decoder 703 can retrieve the bitstream from data storage unit 790.
- Decoder 704 decodes the pictures included in the encoded video sequence to produce video data for display and/or for further image processing (e.g., a machine vision task). Accordingly, decoder 704 may be part of a device 703 having an image processor 705 and/or a display 706. The image processor 705 may perform machine vision tasks on the decoded pictures.
- the device 703 may be a mobile device, a set-top device, a head-mounted display, or any other device.
- FIG. 8 illustrates functional components of encoder 702 according to some embodiments. It should be noted that encoders may be implemented differently so implementation other than this specific example can be used. Encoder 702 employs a subtractor 241 to produce a residual block which is the difference in sample values between an input block and a prediction block (i.e., the output of a selector 251, which is either an inter-prediction block output by an inter predictor 250 (a.k.a., motion compensator) or an intra-prediction block output by an intra predictor 249). Then a forward transform 242 is performed on the residual block to produce a transformed block comprising transform coefficients.
- a subtractor 241 to produce a residual block which is the difference in sample values between an input block and a prediction block (i.e., the output of a selector 251, which is either an inter-prediction block output by an inter predictor 250 (a.k.a., motion compensator) or an intra-prediction block output by an intra predictor 249). Then a
- a quantization unit 243 quantizes the transform coefficients based on a quantization parameter (QP) value (e.g., a QP value obtained based on a picture QP value for the picture in which the input block is a part and a block specific QP offset value for the input block), thereby producing quantized transform coefficients which are then encoded into the bitstream by encoder 244 (e.g., an entropy encoder) and the bitstream with the encoded transform coefficients is output from encoder 702.
- encoder 702 uses the quantized transform coefficients to produce a reconstructed block.
- LF stage 267 may include three sub-stages: i) a deblocking filter, ii) a sample adaptive offset (SAO) filter, and iii) an Adaptive Loop Filter (ALF).
- FIG. 9 illustrates functional components of decoder 704 according to some embodiments. It should be noted that decoder 704 may be implemented differently so implementations other than this specific example can be used. Decoder 704 includes a decoder module 361 (e.g., an entropy decoder) that decodes from the bitstream quantized transform coefficient values of a block.
- decoder module 361 e.g., an entropy decoder
- Decoder 704 also includes a reconstruction stage 398 in which the quantized transform coefficient values are subject to an inverse quantization process 362 and inverse transform process 363 to produce a residual block.
- This residual block is input to adder 364 that adds the residual block and a prediction block output from selector 390 to form a reconstructed block.
- Selector 390 either selects to output an inter-prediction block or an intra-prediction block.
- the reconstructed block is stored in a RPB 365.
- the inter-prediction block is generated by the inter-prediction module 350 and the intra-prediction block is generated by the intra prediction module 369.
- a loop filter stage 367 applies loop filtering and the final decoded picture may be stored in a decoded picture buffer (DPB) 368 and output to image processor 105.
- DPB decoded picture buffer
- Pictures are stored in the DPB for two primary reasons: 1) to wait for picture output and 2) to be used for reference when decoding future pictures.
- this disclosure describes, in one embodiment, applying the decoder side motion vector refinement, by for example DMVR, BDOF or TM, in a hierarchical manner.
- the number of samples used for the motion estimation can be fixed so that same hardware units/ software units can be used in each step of the hierarchy. For example, for DMVR using a basis of 16x16 and for BDOF using a basis of 8x8.
- downsampling is used before motion refinement to reduce complexity but to also increase robustness of the motion vector refinement.
- the embodiments described herein can be used at encoder 702 and/or decoder 704.
- the embodiments can be used for both blocks and subblocks. From here forward the term block should be interpreted broadly to encompass any subblock because a subblock is itself a block.
- the following description focuses on DMVR but can also be applicable to other kinds of motion estimation on the decoder side.
- the refined motion vector can be used directly or passed to further sub-pixel refinement (for example BDOF) and prediction sample refinement by OBMC and finally used to produce a block of predicted samples, as is well known in the art. This is performed at both encoder and decoder.
- FIG. 10 illustrates a reference bi-directional motion vector (Bi-MV) for a current block 1002 of a current picture 1091. More specifically the Bi-MV includes a backward reference MV 1004 and forward reference MV 1006. MV 1004 references a reference block 1010 of a first reference picture 1090 (refA). For example, MV 1004 together with the coordinates of the upper-left hand corner of block 1002 specifies the coordinates of the upper-left hand corner of reference block 1010.
- MV 1006 points to a reference block 1020 of a second reference picture 1092 (refB).
- MV 1004 is denoted mv refA
- MV 1006 is denoted mv refB.
- a search block 1081 within reference picture 1090 which is also known as “search area” 1081, is defined.
- a search area 1082 within reference picture 1092 is also defined.
- the downscaling feature described here is used only if a condition is satisfied.
- One criterion for enabling the downscaling feature is the size of the current block. For example, if any dimension of the block is equal to or greater than a size threshold N, then the downscaling feature is enabled. In one example N is 64. Accordingly, with this example value of N, if either the width (W) or height (H) of block 1002 is greater than 64, then the downscaling feature is enabled. In another embodiment, the downscaling feature is enabled if block 1002 is a subblock.
- FIG. 11 shows that when the downscaling feature is enabled, the search areas 1081 and 1082, which includes reference blocks 1010 and 1020, respectively, are downscaled by a factor of M in at least one dimension to obtain two downscaled search areas 1181 and 1182, respectively.
- search areas 1081 and 1082 are blocks of samples that have been interpolated according to sub-pixel accuracy of the motion vectors which typically are referred to as predicted samples or blocks of predicted samples. Then when downscaling is performed it is performed on the predicted samples or on blocks of predicted samples. The downscaling may be performed by either picking every M:th sample or picking the M:th sample after averaging M samples.
- the corresponding down-sampled search areas 1181 and 1182 are (16+2*R/4 x 16+2*R/4).
- FIGs. 12A and 12B show that, after downscaling, bilateral block matching (BBM) is then applied on the downscaled search areas 1181 and 1182 to determine an optimal candidate offset vector tuple (OVT) (i.e., a first candidate offset vector paired with a second candidate offset vector), which is then used to refine the block motion vectors 1004 and 1006.
- the optimal OVT is the OVT from a set of candidate OVTs having the best score. That is, each OVT in the set of candidate OVTs is given a score, and then the optimal OVT is the one with the best score.
- the set of candidate offset vectors consist of the following 25 offset vectors: ⁇ -2,-2 ⁇ , ⁇ -2,-1 ⁇ ⁇ -2,0 ⁇ , ⁇ -2,1 ⁇ , ⁇ -2,2 ⁇ , ⁇ -1,-2 ⁇ , ⁇ -1,-1 ⁇ ⁇ -1,0 ⁇ , ⁇ -1,1 ⁇ , ⁇ - 1,2 ⁇ , ⁇ 0,-2 ⁇ , ⁇ 0,-1 ⁇ , ⁇ 0,0 ⁇ , ⁇ 0,1 ⁇ , ⁇ 0,2 ⁇ , ⁇ 1,-2 ⁇ , ⁇ 1,-1 ⁇ ⁇ 1,0 ⁇ , ⁇ 1,1 ⁇ , ⁇ 1,2 ⁇ , ⁇ -2,-2 ⁇ ,-1 ⁇ ⁇ -2,0 ⁇ , ⁇ -2,1 ⁇ , ⁇ -2,2 ⁇ .
- each offset vector has a horizontal component (value) and a vertical component (value).
- 25 candidate offset vectors one can form 25*25 OVTs.
- Two example OVTs are ⁇ -2,-1 ⁇ , ⁇ 1,-1 ⁇ and ⁇ 0,-1 ⁇ , ⁇ 2, 2 ⁇ .
- FIG. 12A shows a first OVT that includes a first offset vector 1201 and a second offset vector 1202.
- FIG. 12B shows a second OVT that includes a third offset vector 1203 and a fourth offset vector 1204.
- a block 1211 of downscaled samples is referenced by a combination of the first offset vector (OV) 1201 and reference point 1115
- a block 1212 of downscaled samples is referenced by a combination of the second OV 1202 and reference point 1125.
- These downscaled blocks 1211 and 1212 have the same dimensions as the downscaled reference blocks 1110 and 1120.
- FIG. 12A shows a block 1211 of downscaled samples is referenced by a combination of the first offset vector (OV) 1201 and reference point 1115
- a block 1212 of downscaled samples is referenced by a combination of the second OV 1202 and reference point 1125.
- These downscaled blocks 1211 and 1212 have the same dimensions as the downscaled reference blocks 1110 and 1120.
- a score for the first OVT (which consists of OV 1201 and OV 1202) is calculated based on a difference between block 1211 and block 1212. For instance, the sum of absolute differences (SAD) between downscaled reference samples from block 1211 and block 1212 can be calculated and the score for the first OVT is this calculated sum.
- SAD sum of absolute differences
- a score for the second OVT (which consists of OV 1203 and OV 1204) can be calculated based on a difference between block 1213 and block 1214.
- Every OVT included in the set of candidate OVTs is likewise associated with a pair of downscaled blocks, one from each downscaled search area. Accordingly, each OVT included in the set of candidate OVTs is given a score based on the pair of downscaled blocks with which the OVT is associated in the manner described above for the first OVT and the second OVT.
- the candidate OVT with the best score is then chosen to refine the motion vectors 1004 and 1006. In the example given the OVT with the best score is the OVT with the lowest calculate sum of sample differences.
- the optimal offset is upscaled to the resolution of the current block.
- the refined motion vector is defined as follows: where ov_l_h is the horizontal component of the first candidate offset vector of the optimal OVT, ov_l_v is the vertical component of the first offset vector, ov_2_h is the horizontal component of the second candidate offset vector of the optimal OVT, ov_2_v is the vertical component of the second offset vector.
- a hierarchical motion estimation is performed by using the block motion vectors 1004, 1006 (mv refA and mv refB) and deriving refinements of the block motion vectors for large subblocks of the block and deriving refinements of the motion of the large subblocks for a smaller subblock size.
- the current block 1002 has a size of 64x64
- the current block is divided into four blocks of size 32x32, and each of these blocks correspond to two reference blocks of size 32x32.
- 32x32 block 1300 of block 1002 has two corresponding blocks of size 32x32, namely block 1301 and block 1302.
- an optimal offset vector from a first set of candidate offset vectors is determined. That is, for each of the 32x32 blocks, bilateral block matching (BBM) is applied on the corresponding reference blocks to determine an optimal offset vector.
- the optimal offset vector is the offset vector from the first set of candidate offset vectors that gives the least sum of absolute differences (SAD) between samples of the reference blocks, which may or may not be downscaled.
- the two-dimensional (2D) search area is 32+2*R?2 x 32+2*R?2, where 2*R.32 is the search range in one dimension.
- Each 32x32 block is then divided into four blocks of size 16x16, and each of these blocks correspond to two reference blocks of size 16x16.
- 16x16 block 1310 of block 1002 has two corresponding blocks of size 16x16, namely block 1311 and block 1312.
- BBM bilateral block matching
- the optimal offset vector is the offset vector from the second set of candidate offset vectors that gives the least sum of absolute differences (SAD) between samples of the reference blocks, which may or may not be downscaled.
- the two-dimensional (2D) search area is 16+2*Ri6 x 16+2*Ri6, where 2*Ri6 is the search range is one dimension.
- the final refined motion for a particular one of the 16x16 blocks can be formulated as: where offsetH32 and offsetV32 are the horizontal and vertical components, respectively, of the optimal offset vector determined for the 32x32 block in which the 16x16 block is found, and offsetH16 and offsetV16 are the horizontal and vertical components, respectively, of the optimal offset vector determined for the 16x16 block.
- offsetH32+offsetH16 should be less or equal to +/- R, where 2*R is the maximum search range in one dimension.
- the SAD for evaluating the match between the reference blocks is biased towards the motion determined at a larger subblock. This is a way to increase the robustness of the motion estimation of smaller subblocks.
- the factor can be greater for motions deviating more from the larger subblock motion than for motions deviating less from the larger subblock motion.
- the SAD for larger subblocks can be scaled with a factor smaller than 1.
- the SAD for a motion vector that corresponds to a larger subblock is A
- the SAD for a motion that is different from that motion has a SAD of B*2.
- B is the actual SAD
- 2 is the scaling factor.
- the search range is reduced for smaller subblocks compared to larger subblocks. This is a way to reduce complexity but may also increase robustness of motion estimation.
- the block size is 128x128 the largest subblocks are 64x64 and may have a search range of +- R samples around the block motion.
- the next level of subblocks are 32x32 and may then have a search range of +- R/2 samples around respective 64x64 subblock motion. If the smallest subblocks are 16x16 they may then have a search range of +-R/4 samples around corresponding 32x32 subblock motion.
- R is 8.
- the search range for the first stages of the hierarchy is +-R and the search range for the last stage of the hierarchy, is +- P, where P is smaller than R.
- the search range R is kept same for all subblocks. This makes sure that the size of the reference areas to be searched is kept same.
- the 64x64 subblocks are downsampled to 16x16 and thus the search range is also downscaled to maintain same search range to +-R/4, and the search range for the 32x32 subblocks is downscaled to +-R/2 while the search range for the 16x16 subblocks is kept as R.
- the enabling of the features described herein is controlled by the spatial activity of the block. If the spatial activity is below a threshold the features can be used.
- the spatial activity can for example be measured on respective uni -prediction blocks according to the respective uni-parts of the bi-predictive block motion vector. Alternatively, it can be performed on subblocks of the block.
- Rt j be the samples of the prediction for the current block that would be the result of using one uni-part of the unrefined bi-predictive block motion vector, i.e., the motion vector of the main block. Then, the activity Act t j for the current block can be calculated as Ac il)/2.
- the spatial activity for the block can be measured as Sa /MxN, where MxN is the size of the block or the subblock.
- the enabling of the features is controlled by the QP (quantization parameter) for the block. If the QP is equal to or greater than a threshold the feature is enabled. In one example, the QP threshold is 32.
- the enabling of the features is controlled by the absolute value of a component of the motion vector. If the magnitude of one component is greater than a threshold the method is enabled. In one example, the threshold is 7 or 112 in 16 :th pel accuracy (7*16).
- the enabling of the method is controlled by the size of the video picture. If the picture size is greater than a threshold the method is enabled. In one example, the threshold is 1920x1080.
- FIG. 14 is a flowchart illustrating process 1400, according to some embodiments, for obtaining motion compensated predicted samples (e.g., obtaining an interprediction subblock for use in obtaining a reconstructed block by adding the inter-prediction subblock and a reconstructed residual subblock).
- Process 1400 may begin in step sl402.
- mv refA ⁇ mv_refA_h, mv_refA_v ⁇ .
- mv refB ⁇ mv_refB_h, mv_refB_v ⁇ .
- Step sl406 comprises producing a first refined MV (mv refmedA) and a second refined MV (mv refmedB) using downscaled reference blocks.
- Step sl408 comprises obtaining motion compensated predicted samples using MV_refinedA and MV_refinedB.
- producing mv refmedA and mv refmedB using downscaled blocks comprises the following steps:
- [0139] (1) producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector (e.g., associated with ov_l and a first reference point associated with the first motion vector, such as reference point 1115) and a second downscaled block associated with ov_2 and the second motion vector (e.g., associated with ov_2 and a second reference point associated with the second motion vector, such as reference point 1125);
- a first downscaled block associated with ov_l and the first motion vector e.g., associated with ov_l and a first reference point associated with the first motion vector, such as reference point 1115
- a second downscaled block associated with ov_2 and the second motion vector e.g., associated with ov_2 and a second reference point associated with the second motion
- FIG. 15 is a flowchart illustrating process 1500, according to some embodiments, for obtaining motion compensated predicted samples (e.g., obtaining an interprediction subblock for use in obtaining a reconstructed block by adding the inter-prediction subblock and a reconstructed residual subblock).
- Process 1500 may begin in step si 502.
- Step si 504 comprises determining a first offset vector (ov_l) using a first reference motion vector for a second block of samples (e.g., a 32x32 block) and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples.
- a first offset vector ov_l
- Step si 506 comprises determining a second offset vector (ov_2) using a first reference motion vector for a third block of samples (e.g., a 16x16 block) and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples.
- a second offset vector ov_2
- Step si 508 comprises producing a first refined motion vector, mv_refined_3_l, for the third block of samples based on mv refA, ov_l, and ov_2.
- Step si 510 comprises obtaining motion compensated predicted samples using the first refined motion vector.
- FIG. 16 is a flowchart illustrating process 1600, according to some embodiments, for obtaining motion compensated predicted samples (e.g., obtaining an interprediction subblock for use in obtaining a reconstructed block by adding the inter-prediction subblock and a reconstructed residual subblock).
- Process 1600 may begin in step sl602.
- Step si 604 comprises selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples (e.g., a 32x32 block), wherein the second block of samples is a subblock within the first block of samples.
- Step sl606 comprises producing a first refined motion vector, mv refinedl, for the second block of samples based on mv refA and the first OVT.
- Step si 608 comprises producing a second refined motion vector, mv_refined2, for the second block of samples based on mv refB and the first OVT.
- Step s 1610 comprises obtaining motion compensated predicted samples using the first and the second refined motion vectors.
- Selecting the first OVT comprises the following steps:
- [0156] (1) producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples (e.g., associated with the first offset vector of the first OVT and a reference point associated with the first motion vector for the second block of samples) and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples (e.g., associated with the second offset vector of the first OVT and a second reference point associated with the second motion vector for the second block of samples);
- process 1600 further includes: defining a first search area within a first reference block using the first motion vector for the second block of samples; defining a second search area within a second reference block using the second motion vector for the second block of samples; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block is within the first downscaled search area, and the second downscaled block is within the second downscaled search area.
- FIG. 17 is a block diagram of an apparatus 1700 for implementing encoder 702 and/or decoder 704, according to some embodiments.
- apparatus 1700 implements encoder 702
- apparatus 1700 may be referred to as an encoder apparatus
- apparatus 1700 may be referred to as a decoder apparatus. As shown in FIG.
- apparatus 1700 may comprise: processing circuitry (PC) 1702, which may include one or more processors (P) 1755 (e.g., one or more general purpose microprocessors and/or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like), which processors may be co-located in a single housing or in a single data center or may be geographically distributed (i.e., apparatus 1700 may be a distributed computing apparatus); at least one network interface 1748 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 1745 and a receiver (Rx) 1747 for enabling apparatus 1700 to transmit data to and receive data from other nodes connected to a network 100 (e.g., an Internet Protocol (IP) network) to which network interface 1748 is connected (physically or wirelessly) (e.g., network interface 1748 may be coupled to an antenna arrangement comprising one or more antennas for enabling apparatus 1700 to wirelessly transmit/re
- a computer readable storage medium (CRSM) 1742 may be provided.
- CRSM 1742 may store a computer program (CP) 1743 comprising computer readable instructions (CRI) 1744.
- CP computer program
- CRSM 1742 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like.
- the CRI 1744 of computer program 1743 is configured such that when executed by PC 1702, the CRI causes apparatus 1700 to perform steps described herein (e.g., steps described herein with reference to the flow charts).
- apparatus 1700 may be configured to perform steps described herein without the need for code. That is, for example, PC 1702 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and/or software.
- DMVR is a powerful technology to reduce the signalling of motion information and refines a bi-directional block motion vector on block basis but also on subblocks of 16x16 samples.
- the large gap between the size of the block and the size of the subblock can make the refinements on subblocks incoherent. This contribution addresses that issue by deploying DMVR search in a hierarchical manner.
- Alt A When subblock DMVR is applicable and the width or height of the current block is equal to or greater than 64, hierarchical subblock-based DMVR is enabled.
- the reference areas and search range given by the respective block motion vector of the bidirectional motion vector are of the same size as in ECM.
- the current block e.g., 64x64
- the corresponding reference subblocks and search areas are down- sampled so that each subblock has a size of 16x16 samples.
- the integer DMVR search is then applied on the downscaled subblocks to derive a refinement vector for each subblock.
- the current subblock e.g., 32x32
- the current subblock e.g., 32x32
- the corresponding reference subblocks and search area are down-sampled so that each smaller subblock has a size of 16x16 samples.
- the integer DMVR search is then applied on the downscaled smaller subblocks to derive a refinement vector for each smaller subblock.
- Alt B To keep the number of search points not exceeding ECM- 12.0 a complexity reduced approach has also been implemented where the criterion to enable the approach is based of width being equal or greater than 64 when height is equal or greater than 32 or when height is equal or greater than 64 and width is equal or greater than 32.
- the integer search range is reduced to +-5. This reduces the number of integer search points for a 32x64 block to about half compared to 16x16 DMVR in ECM.
- Alt A The initial results are not complete since QP22 for ParkRunning is not complete so results from anchor have been copied for those points.
- the first downscaled block is a block of W2xH2 samples, and W2 ⁇ W1 and/or H2 ⁇ Hl.
- A7 The method of any one of embodiments A2-A6, wherein the method further comprises: prior to producing mv refmedA and mv refmedB using downscaled reference blocks, determining that a condition is satisfied, wherein the step of producing mv refmedA and mv refmedB is performed in response to determining that the condition is satisfied.
- ⁇ mv_refmed_3_2_h, mv_refmed_3_2_v ⁇ , and producing the second refined motion vector for the third block of samples based on mv refB, ov_3, and ov_4 comprises: setting mv_refmed_3_2_h equal to (mv_refB_h + ov_3_h + ov_4_h); and setting mv_refmed_3_2_v equal to (mv refB v + ov_3_v + ov_4_v).
- producing the score for the candidate offset vector using the first block of predicted samples comprises: producing the score for the candidate offset vector using the first block of predicted samples and a second block of predicted samples chosen based on a combination of a second reference motion vector for the second block and an offset vector corresponding to candidate offset vector.
- Bl 1. The method of any one of embodiments Bl -BIO, wherein the method further comprises: prior to producing the first refined motion vector determining that a condition is satisfied, wherein the step of producing the first refined motion vector is preformed in response to determining that the condition is satisfied.
- determining that the condition is satisfied comprises performing at least one of the following steps: determining that W1 is greater than or equal to a width threshold; determining that Hl is greater than or equal to a height threshold; determining that subblocks are used; determining that a spatial activity of the first block is less than a spatial activity threshold; determining that a quantization parameter, QP, for the first block is greater than a QP threshold; determining that the absolute value of a component of the first reference MV is greater than an MV magnitude threshold; determining that the absolute value of a component of the second reference MV is greater than the MV magnitude threshold; or determining that a size of the picture to which the first block belongs is greater than a picture size threshold.
- ⁇ mv_refmed_3_2_h, mv_refmed_3_2_v ⁇ , and producing the second refined motion vector for the third block of samples based on mv refB, ov_l, and ov_2 comprises: setting mv_refined_3_2_h equal to (mv refB h - ov_l_h - ov_2_h); and setting mv_refmed_3_2_v equal to (mv refB v - ov_l_v - ov_2_v).
- a computer program (1743) comprising instructions (1744) which when executed by processing circuitry (1702) cause the processing circuitry (1702) to perform the method of any one of the above embodiments.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
A method for obtaining motion compensated predicted samples. The method includes obtaining, for a current block, a first motion vector, mv_refA, associated with a first block, wherein mv_refA = {mv_refA_h, mv_refA_v}. The method also includes obtaining, for the current block, a second motion vector, mv_refB, associated with a second block, wherein mv_refB = {mv_refB_h, mv_refB_v}. The method also includes producing a first refined MV, mv_refinedA, and a second refined MV, mv_refinedB, using downscaled blocks. The method also includes obtaining motion compensated predicted samples using MV_refinedA and MV_refinedB.
Description
MOTION VECTOR REFINEMENT WITH DOWNSAMPLED BLOCKS, OR HIERARCHICAL SUB-BLOCKS
TECHNICAL FIELD
[0001] This disclosure relates to methods and apparatus for motion vector refinement.
BACKGROUND
[0002] VVC and ECM
[0003] Versatile Video Coding (VVC) is a block-based video codec standardized by ITU-T and MPEG. Enhanced Coding Model (ECM) is an exploratory codec which is currently under development. The aim of ECM is to demonstrate and try providing evidence of video coding capabilities beyond VVC.
[0004] Video and Picture
[0005] A video (a.k.a., “video sequence”) comprises of a series of pictures (a.k.a., frames or images). In VVC, each picture is identified with a picture order count (POC) value. The POC value also represents the display order of the picture. A picture with a smaller POC value is displayed before another picture with a larger POC value.
[0006] Components
[0007] Each component can be described as a two-dimensional rectangular array of sample values. It is common that each picture consists of three components; one luma component Y where the sample values are luma values and two chroma components Cb and Cr, where the sample values are chroma values.
[0008] It is also common that the dimensions of the chroma components are smaller than the luma components by a factor of two in each dimension. For example, the size of the luma component of an HD picture would be 1920x1080 and the chroma components would each have the dimension of 960x540. Components are sometimes referred to as color components.
[0009] Coding Unit and Coding Block
[0010] A block is a two-dimensional (2D) array of sample values (or “samples” for short). A block may be divided into two or more blocks (a.k.a., subblocks), where each
subblock is a 2D array of samples, which is also known as a “matrix” of samples. In video coding, each component of a picture is split into blocks and the coded video bitstream consists of a series of coded blocks. It is common in video coding that pictures are split into units that cover a specific area of the picture.
[0011] Each unit consists of all blocks from all components that make up that specific area of the picture and each block belongs fully to one unit. The Coding Unit (CU) in VVC is an example of a unit. In VVC, the CUs may be split recursively to smaller CUs. The CU at the top level is referred to as the coding tree unit (CTU).
[0012] A CU usually contains three coding blocks, i.e., one coding block for luma and two coding blocks for chroma. The size of luma coding block is same as the CU.
[0013] In VVC, the CUs can have size of 4x4 up to 128x128. In current ECM, the CUs can have size of 4x4 up to 256x256.
[0014] Parameter sets, slice headers, and picture headers
[0015] VVC specifies three types of parameter sets: the picture parameter set (PPS), the sequence parameter set (SPS), and the video parameter set (VPS). The PPS contains data that is common for a whole picture, the SPS contains data that is common for a coded layer video sequence (CLVS), and the VPS contains data that is common for multiple CLVSs, e.g., data for multiple layers in the bitstream.
[0016] The concept of slices divides the picture into independently coded slices, where decoding of one slice in a picture is independent of other slices of the same picture. Each slice has a slice header comprising syntax elements. Decoded slice header values from these syntax elements are used when decoding the slice.
[0017] In VVC, a coded picture contains a picture header. The picture header contains parameters that are common for all slices of the coded picture.
[0018] Intra prediction
[0019] In intra prediction, also known as spatial prediction, a block is predicted using previous decoded blocks within the same picture. The samples from the previously decoded blocks within the same picture are used to predict the samples inside the current block. A picture consisting of only intra-predicted blocks is referred to as an intra picture.
[0020] Inter prediction
[0021] In inter prediction, also known as temporal prediction, blocks of the current picture are predicted using blocks from previously decoded pictures (these blocks are referred to as reference blocks). The samples from the reference blocks in the previously decoded pictures are used to predict the samples inside the current block. A picture that comprises one or more inter-predicted blocks is referred to as an inter-picture. The previous decoded pictures used for inter prediction are referred to as reference pictures.
[0022] The location of a referenced block inside a reference picture is indicated using a vector (i.e., a set of values) (which is referred to as “motion vector (MV)”). Each MV consists of two values: an x value (a.k.a., x component) and y value (a.k.a., y component) which represents the displacements between current block and the referenced block in x or y dimension. The value of a component may have a resolution finer than an integer position. When that is the case, a filtering (typically interpolation) is done to calculate values used for prediction.
[0023] FIG. 1 shows an example of an MV for the current block. The example MV consist of a horizontal value of 2 and a vertical value of 1.
[0024] An inter picture may use several reference pictures. The reference pictures are usually put into two reference picture lists, L0 and LI. The reference pictures that are output before the current picture are typically the first pictures in L0. The reference pictures that are output after the current picture are typically the first pictures in LI.
[0025] FIG. 2 shows an example of two different prediction types: uni -prediction and bi-prediction. Inter predicted blocks can use uni -prediction or bi-prediction. A uni-predicted block uses one MV (e.g., MV0 in FIG. 2) to predict from one reference picture, either using L0 or LL Bi-prediction uses a pair of motion vectors (denoted Bi-MV), to predict from a first reference picture, such as, for example, reference picture 0 in FIG. 2, and a second reference picture, such as, for example, reference picture 1 in FIG. 2. More specifically, the Bi-MV consists of a first MV (MV1) and a second MV (MV2), where the first MV is used to predict from the first reference picture and the second MV is used to predict from the second reference picture.
[0026] Picture coding type (Low delay picture and non-low delay picture)
[0027] A low delay picture is a picture that has all its reference pictures displayed before the picture. In other words, for a low delay picture, all its reference pictures have smaller POC values than the current POC.
[0028] A non-low delay picture is a picture that has at least one of its reference pictures displayed after the picture. In other words, a non-low delay picture has at least one reference picture with a larger POC value than the current POC.
[0029] Fractional MVs, Interpolation filter and MV rounding
[0030] The value of the MV’s x or y component may correspond to a sample position which has finer granularity than integer (sample) position. Those positions are also referred to as fractional (sample) positions. In WC and current ECM, the MV can be at 1/16 sample position.
[0031] FIG. 3 depicts several fractional positions in the horizontal (x-) dimension.
The solid-square blocks represent integer positions. The circles represent 1/16-position. For example, MV = (4, 10) means the x component is at 4/16 position, the y component is at 10/16 position.
[0032] In video coding, an MV rounding process is sometimes used to convert an MV at one position to another target position. One example of rounding is to round a fractional MV position to the nearest integer position.
[0033] When an MV is at a fractional position, filtering (typically interpolation) is done to calculate the sample values at those positions. In VVC, the length (number of filter taps) of the interpolation filter for luma component is 8, as shown in table 1 below. In ECM, the length of the interpolation filter for luma component has been increased to 12.
TABLE 1
[0034] Residual, transform and quantization
[0035] A residual block is generated by computing the difference between samples of a source block (a.k.a., “input block”), which contains original samples, and samples of a prediction block (a.k.a., “residual block”) (e.g., an inter-prediction block or an intraprediction block). To remove further redundancy, the difference is then typically compressed by a spatial transform, thereby producing transform coefficients. The transform coefficients are then quantized based on a quantization parameter (QP) to control the fidelity of the residual block and thus also the bitrate required to compress the block. A coded block flag (CBF) is used to indicate if there are any non-zero quantized transform coefficients. All coding parameters are then entropy coded at the encoder and decoded at the decoder. If the coded block flag is one, a reconstructed block can then be derived by inverse quantization and inverse transformation of the quantized transform coefficients and then add that to the prediction block. If the coded block flag is zero, the reconstructed block is identical to the prediction block.
[0036] Inter prediction information / Motion information
[0037] For an inter block inside an inter picture in VVC, its inter prediction
information consists of the following three elements:
[0038] (1) a reference picture list flag (RefPicListFlag) which signals which reference picture list is used for the block (when the value of the flag is equal to 0, it means only L0 is used for predicting the current block, when the value of the flag is equal to 1, it means only LI is used for predicting the current block, and when the value of the flag is equal to 2, it means both L0 and LI are used for predicting the current block);
[0039] (2) a reference picture index (RefPicIdx) per reference picture list used (the index signals which reference picture inside the reference list to be used for predicting the current block), and
[0040] (3) a motion vector (MV) per reference picture used, which signals the position inside the reference picture that is used for predicting the current block.
[0041] The inter prediction information is also referred to as motion information. The decoder stores the motion information for each inter block. In other words, an inter block maintains its own motion information.
[0042] Encoder Decision and Rate Distortion (RD) Cost
[0043] In practice, for an encoder to decide the best prediction mode for a current block, the encoder would evaluate all the possible prediction modes for the current block and select the prediction mode that yields the smallest Rate-Distortion (RD) cost.
[0044] The RD cost is calculated as D + * Rate. D (Distortion) measures the difference between the reconstructed block and the corresponding source block. One commonly used metric for calculating D is the sum of squared error SSE
) —
PB(x,y))2 , where PA and PB are the sample values in the two blocks A and B respectively. Rate is usually an estimation of the bits to be spent on encoding the mode, and is a trade-off parameter between Rate and D.
[0045] Motion Information Signaling
[0046] VVC and ECM include several methods for implicit signaling of motion information for each block, including the merge method and the subblock merge method. A common motivation behind the implicit methods is to inherit or reuse motion information from neighboring coded blocks. This often works in practice due to spatial correlation of close-by blocks, i.e., the fact that nearby blocks often behave similarly.
[0047] Merge (block merge) Method and Merge Mode
[0048] The merge method derives a set of motion information from previously decoded blocks and use the derived motion information for generating the samples of the entire block. The merge method is sometimes referred to as the block merge method.
[0049] The merge method first generates a list of motion information candidates. The list is also referred to as the merge list. The candidates are derived from previously coded blocks. These previously coded blocks can be spatially adjacent neighboring blocks or temporal collocated blocks relative to the current block.
[0050] FIG. 4 shows the spatial neighboring blocks: left (L), top (T), top-right (TR), left-bottom (LB) and top-left (TL).
[0051] The merge list construction process usually checks the previously coded blocks in a predefined order, for example: T, L, TR, LB, then TL. For each previously coded block being checked, if this previously coded block is inter coded and its motion information has no duplicates in the list, then the motion information of this previously coded block is added to the merge list.
[0052] After the merge list is generated, one of the candidates inside the list is used to derive the motion information of the current block. The candidate selection process is done on the encoder side. An encoder would select a best candidate from the list and encode an index (merge index) in the bitstream to signal to a decoder. The decoder receives the index, it follows the same merge list derivation process as the encoder and uses the index to retrieve the correct candidate. The blocks that use the block merge method are sometimes referred to as blocks in merge mode.
[0053] In the current ECM, non-adjacent spatial blocks are also considered as sources of motion information during the merge list construction.
[0054] FIG. 5 shows some examples (marked with NA1, NA2 and NA3) of those non-adjacent spatial blocks.
[0055] Subblock Merge Method
[0056] VVC and ECM also include the subblock merge method. It splits a current block into a number of subblocks and allows each subblock to have its own motion information.
[0057] FIG. 6 shows an example of a current block and its subblocks. Each subblock maintains its own motion information. It should be noted that the subblocks are all rectangular. In VVC, only one so called collocated reference picture can be indicated and the motion is temporally predicted from a corresponding position in the collocated picture (either from L0 or LI). ECM supports two collocated pictures to be indicated, one from L0 and one from LI. ECM also supports that the corresponding position in the collocated picture is offset vertically and/or horizontally with some variation.
[0058] Overlapped Block Motion Compensation (OBMC)
[0059] OBMC is a tool included in ECM which operates at the block boundaries or subblock boundaries of a current inter block. OBMC blends the current block or subblock’s prediction samples, P CUR, (generated using the current associated motion information) with another set of prediction samples, P_NB, which are generated using motion information from a neighboring block or a neighboring subblock at the current block or subblock boundaries. The OBMC blending process takes weighted average of P CUR and P_NB to produce a set of OBMC modified prediction samples, P OBMC. In other words, the set P CUR or P_NB each associate with a respective weighting factor, and the set of OBMC modified prediction samples P_OBMC is derived as wl*P_CUR + w2*P_NB (where wl + w2 = 1). The OBMC may give better prediction for samples that are close to the block or subblock boundary.
[0060] Bi-Directional Optical Flow (BDOF)
[0061] BDOF is a tool included in VVC and the current ECM that can be used to refine prediction samples that are generated from a Bi-MV. BDOF relies on optical flow estimation to derive a pair of refinement parameter (Vx, Vy) which can be further used to refine the prediction samples.
[0062] Decoder-side Motion Vector Refinement (DMVR)
[0063] DMVR is a tool included in VVC and the current ECM to refine motion vectors for a Bi-MV. DMVR operates on subblock level, usually 16x16. Different from BDOF, which relies on optical flow estimation, DMVR relies on bilateral matching of two reference blocks to refine the Bi-MV. The matching is based on SAD (sum of absolute differences). The DMVR searches within a window around the Bi-MV to find whether there exists another Bi-MV (Bi-MV’) that gives a better match between the L0 reference block and the LI reference block. If so, the Bi-MV’ is further used instead for generating the prediction
samples of the current block. After that a subpel adjustment is made based on the SAD costs around the motion with least SAD to determine subblock motion with sub pixel accuracy.
[0064] Multi-pass DMVR
[0065] In ECM (enhanced compression beyond VVC), DMVR has been further evolved to use multi-pass optimization, first bilateral block matching in a search area to find the refinement of the merge motion that gives best match, then bilateral 16x16 subblock matching, then bi-directional optical flow on 8x8 subblocks to refine motion further. The multi-pass DMVR also allows for modification of only one of the bi-predictive motions. It is also allowed to use DMVR when BCW (other weightings than just average) is used.
[0066] Template matching (TM)
[0067] In ECM block or subblock motion can also be refined by matching a template outside the current block with a corresponding template on the reference picture. The search is limited to a small range to find a better motion without signaling additional motion information. This can also be used cascaded with the bilateral matching in multi-pass DMVR.
SUMMARY
[0068] Certain challenges presently exist. For example, with the multi-pass DMVR design in ECM there is a large gap between the motion vector derived on decoder side on the full block size (up to 256x256 in ECM12) and the motion vector derived on decoder side on 8x16 or 16x16 subblock size. This can result in inaccurate motion vectors in some cases, especially when spatial activity is low in a 16x16 subblock where several candidate motion vectors may give similar bilateral match. This issue can also appear for BDOF which either operates on 16x16 or 8x8 or 4x4 subblock sizes for large blocks (up to 256x256 in ECM12). Another issue is that when DMVR is used in bi-prediction with non-equal distance to reference pictures the derived refinement from DMVR based on symmetrical assumption is scaled so that the motion vector closer to a reference picture is reduced according to difference in the distances to the reference pictures. Which means the motion vector may become closer to the correct motion vector but likely with some inaccuracy since the bilateral match may be worse after the scaling of the motion vector.
[0069] Accordingly, in one aspect, there is provided a method for obtaining motion compensated predicted samples. The method includes obtaining, for a current block, a first motion
vector, mv_refA, associated with a first block, wherein mv_refA = {mv_refA_h, mv_refA_v}. The method also includes obtaining, for the current block, a second motion vector, mv refB, associated with a second block, wherein mv refB = {mv refB h, mv refB v}. The method also includes producing a first refined MV, mv_refmedA, and a second refined MV, mv_refmedB, using downscaled blocks. The method also includes obtaining motion compensated predicted samples using MV_refmedA and MV_refmedB. Producing mv_refmedA and mv_refmedB using downscaled blocks comprises (1) producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector (e.g., associated with ov_l and a first reference point associated with the first motion vector) and a second downscaled block associated with ov_2 and the second motion vector (e.g., associated with ov_2 and a second reference point associated with the second motion vector); (2) producing a score for a second OVT comprising a third offset vector, ov_3, and a fourth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector (e.g., associated with ov_3 and said first reference point) and a fourth downscaled block associated with ov_4 and the second motion vector (e.g., associated with ov_4 and said second reference point); (3) selecting the first OVT based on the score for the first OVT, wherein the selecting comprises comparing the score for the first OVT with the score for the second OVT; and (4) as a result of selecting the first OVT, producing mv_refmedA based on mv_refA and ov_l and producing mv refmedB based on mv refB and ov_2 or ov_l.
[0070] In another aspect, there is provided a method for obtaining motion compensated predicted samples. The method includes obtaining a first motion vector, mv_refA, for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv_refA_v}. The method also includes determining a first offset vector, ov_l, using a first motion vector for a second block of samples (e.g., a 32x32 block) and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples. The method also includes determining a second offset vector, ov_2, using a first reference motion vector for a third block of samples (e.g., a 16x16 block) and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples. The method also includes producing a first refined motion vector, mv_refmed_3_l, for the third block of samples based on mv_refA, ov_l, and ov_2. The method also includes obtaining motion compensated predicted samples using the first refined motion vector.
[0071] In another aspect, there is provided a method for obtaining motion compensated predicted samples. The method includes obtaining a first motion vector, mv_refA, and a second motion vector, mv_refB, for a first block of samples (e.g., a 64x64 block), wherein mv_refA =
{mv refA h, mv refA v} and mv refB = {mv refB h, mv refB v}. The method also includes selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples (e.g., a 32x32 block), wherein the second block of samples is a subblock within the first block of samples. The method also includes producing a first refined motion vector, mv_refmedl, for the second block of samples based on mv_refA and the first OVT. The method also includes producing a second refined motion vector, mv_refmed2, for the second block of samples based on mv refB and the first OVT. The method also includes obtaining motion compensated predicted samples using the first and the second refined motion vectors. Selecting the first OVT comprises: (1) producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples (e.g., associated with the first offset vector of the first OVT and a reference point associated with the first motion vector for the second block of samples) and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples (e.g., associated with the second offset vector of the first OVT and a second reference point associated with the second motion vector for the second block of samples); (2) producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and (3) determining that the score for the first OVT is better than the score for the second OVT.
[0072] In another aspect, there is provided a computer program comprising instructions which when executed by processing circuitry cause the processing circuitry to perform the method of any one of the above embodiments.
[0073] In a different aspect, there is provided a carrier containing the computer program of the above embodiment, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.
[0074] In another aspect, there is provided an apparatus for performing any of the methods disclosed herein.
[0075] An advantage of the embodiments disclosed herein is that the refined motion vectors are more accurate, thereby providing a bitrate savings while at the same time providing no reduction in objective quality. Furthermore, the more accurate motion vectors provide increased subjective quality.
BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.
[0077] FIG. 1 shows an example of a motion vector (MV).
[0078] FIG. 2 illustrates uni-inter prediction and bi-inter prediction.
[0079] FIG. 3 depicts several fractional positions in the horizontal (x-) dimension.
[0080] FIG. 4 shows spatial neighboring blocks.
[0081] FIG. 5 shows some examples of non-adjacent spatial blocks.
[0082] FIG. 6 shows an example of a current block and its subblocks.
[0083] FIG. 7 illustrates a system according to some embodiments.
[0084] FIG. 8 is a schematic block diagram of an encoder according to an embodiment.
[0085] FIG. 9 is a schematic block diagram of a decoder according to an embodiment.
[0086] FIG. 10 illustrates a bi-directional motion vector.
[0087] FIG. 11 illustrates downscaling a search area.
[0088] FIGs. 12A and 12B illustrate a process of finding an optimal offset vector tuple.
[0089] FIG. 13 illustrates Hierarchical Motion Estimation according to an embodiment.
[0090] FIG. 14 is a flowchart illustrating a process according to some embodiments.
[0091] FIG. 15 is a flowchart illustrating a process according to some embodiments.
[0092] FIG. 16 is a flowchart illustrating a process according to some embodiments.
[0093] FIG. 17 is a block diagram of an apparatus according to some embodiments.
DETAILED DESCRIPTION
[0094] FIG. 7 illustrates a system 700 according to an embodiment. System 700 includes an encoder 702 and a decoder 704, wherein, in the example shown, encoder 702 is in communication with decoder 704 via a network 110 (e.g., the Internet or other network). Encoder 702 encodes a source video sequence 701 (e.g., encodes blocks of units of pictures of video sequence 701) into a bitstream comprising an encoded video sequence (e.g., encoded
blocks) and transmits the bitstream to decoder 704 via network 708. In some embodiments, encoder 702 is not in communication with decoder 704, and, in such an embodiment, rather than transmitting bitstream to decoder 704, the bitstream is stored in a data storage unit 790 and decoder 703 can retrieve the bitstream from data storage unit 790. Decoder 704 decodes the pictures included in the encoded video sequence to produce video data for display and/or for further image processing (e.g., a machine vision task). Accordingly, decoder 704 may be part of a device 703 having an image processor 705 and/or a display 706. The image processor 705 may perform machine vision tasks on the decoded pictures. The device 703 may be a mobile device, a set-top device, a head-mounted display, or any other device.
[0095] FIG. 8 illustrates functional components of encoder 702 according to some embodiments. It should be noted that encoders may be implemented differently so implementation other than this specific example can be used. Encoder 702 employs a subtractor 241 to produce a residual block which is the difference in sample values between an input block and a prediction block (i.e., the output of a selector 251, which is either an inter-prediction block output by an inter predictor 250 (a.k.a., motion compensator) or an intra-prediction block output by an intra predictor 249). Then a forward transform 242 is performed on the residual block to produce a transformed block comprising transform coefficients. A quantization unit 243 quantizes the transform coefficients based on a quantization parameter (QP) value (e.g., a QP value obtained based on a picture QP value for the picture in which the input block is a part and a block specific QP offset value for the input block), thereby producing quantized transform coefficients which are then encoded into the bitstream by encoder 244 (e.g., an entropy encoder) and the bitstream with the encoded transform coefficients is output from encoder 702. Next, encoder 702 uses the quantized transform coefficients to produce a reconstructed block. This is done by first applying inverse quantization 245 and inverse transform 246 to the transform coefficients to produce a reconstructed residual block and using an adder 247 to add the prediction block to the reconstructed residual block, thereby producing the reconstructed block, which is stored in the reconstruction picture buffer (RPB) 266. Loop filtering by a loop filter (LF) stage 267 is applied and the final decoded picture is stored in a decoded picture buffer (DPB) 268, where it can then be used by the inter predictor 250 to produce an inter-prediction block for the next picture to be processed. LF stage 267 may include three sub-stages: i) a deblocking filter, ii) a sample adaptive offset (SAO) filter, and iii) an Adaptive Loop Filter (ALF).
[0096] FIG. 9 illustrates functional components of decoder 704 according to some embodiments. It should be noted that decoder 704 may be implemented differently so implementations other than this specific example can be used. Decoder 704 includes a decoder module 361 (e.g., an entropy decoder) that decodes from the bitstream quantized transform coefficient values of a block. Decoder 704 also includes a reconstruction stage 398 in which the quantized transform coefficient values are subject to an inverse quantization process 362 and inverse transform process 363 to produce a residual block. This residual block is input to adder 364 that adds the residual block and a prediction block output from selector 390 to form a reconstructed block. Selector 390 either selects to output an inter-prediction block or an intra-prediction block. The reconstructed block is stored in a RPB 365. The inter-prediction block is generated by the inter-prediction module 350 and the intra-prediction block is generated by the intra prediction module 369. Following the reconstruction stage 398, a loop filter stage 367 applies loop filtering and the final decoded picture may be stored in a decoded picture buffer (DPB) 368 and output to image processor 105. Pictures are stored in the DPB for two primary reasons: 1) to wait for picture output and 2) to be used for reference when decoding future pictures.
[0097] As noted above, in some situations, with the multi-pass DMVR design in ECM there is a large gap between the motion vector derived on decoder side on the full block size (up to 256x256 in ECM12) and the motion vector derived on decoder side on 8x16 or 16x16 subblock size, and this can result in inaccurate motion vectors.
[0098] Accordingly, to combat this issue with DMVR this disclosure describes, in one embodiment, applying the decoder side motion vector refinement, by for example DMVR, BDOF or TM, in a hierarchical manner. The number of samples used for the motion estimation can be fixed so that same hardware units/ software units can be used in each step of the hierarchy. For example, for DMVR using a basis of 16x16 and for BDOF using a basis of 8x8. In another embodiment, downsampling is used before motion refinement to reduce complexity but to also increase robustness of the motion vector refinement.
[0099] The embodiments described herein can be used at encoder 702 and/or decoder 704. The embodiments can be used for both blocks and subblocks. From here forward the term block should be interpreted broadly to encompass any subblock because a subblock is itself a block.
[0100] The following description focuses on DMVR but can also be applicable to other kinds of motion estimation on the decoder side. The refined motion vector can be used directly or passed to further sub-pixel refinement (for example BDOF) and prediction sample refinement by OBMC and finally used to produce a block of predicted samples, as is well known in the art. This is performed at both encoder and decoder.
[0101] Use of Downscaling
[0102] FIG. 10 illustrates a reference bi-directional motion vector (Bi-MV) for a current block 1002 of a current picture 1091. More specifically the Bi-MV includes a backward reference MV 1004 and forward reference MV 1006. MV 1004 references a reference block 1010 of a first reference picture 1090 (refA). For example, MV 1004 together with the coordinates of the upper-left hand corner of block 1002 specifies the coordinates of the upper-left hand corner of reference block 1010. For instance, if the upper-left hand corner of block 1002 is positioned at coordinates x,y within current picture 1091 and MV 1004 is equal to {-1,1 }, then the upper-left hand corner of reference block 1010 is positioned at coordinates x-l,y+l within first reference picture 1090. Similarly, MV 1006 points to a reference block 1020 of a second reference picture 1092 (refB). MV 1004 is denoted mv refA and MV 1006 is denoted mv refB. Both motion vectors have two components, a horizontal component and a vertical component. That is, mv refA = {mv refA h, mv refA v} and mv refB = {mv refB h, mv refB v}.
[0103] As shown in FIG. 10, a search block 1081 within reference picture 1090, which is also known as “search area” 1081, is defined. Likewise, a search area 1082 within reference picture 1092 is also defined. The position and size of search area 1081 is based on the position and size of reference block 1010, and the position and size of search area 1082 is based on the position and size of reference block 1020. More specifically, if block 1010 has a size of WxH, then search area 1081 has a size of (W+2*R)x(H+2*R), where R is an integer greater than 0. In one example R=8 (+/- 8). Likewise, if block 1020 has a size of WxH, then search area 1082 has a size of (W+2*R)x(H+2*R). And if the upper-left hand corner of block 1010 is located within picture 1090 at the coordinates xl,yl, then the upper-left hand corner of search area 1081 is located within picture 1090 at the coordinates xl-R,yl-R (assuming the coordinates of the upper-left hand corner of picture 1090 is 0,0). Similarly, if the upper-left hand corner of block 1020 is located within picture 1092 at the coordinates
x2,y2, then the upper-left hand corner of search area 1082 is located within picture 1092 at the coordinates x2-R,y2-R (assuming the coordinates of the upper-left hand corner of picture 1092 is 0,0).
[0104] In one embodiment, the downscaling feature described here is used only if a condition is satisfied. One criterion for enabling the downscaling feature is the size of the current block. For example, if any dimension of the block is equal to or greater than a size threshold N, then the downscaling feature is enabled. In one example N is 64. Accordingly, with this example value of N, if either the width (W) or height (H) of block 1002 is greater than 64, then the downscaling feature is enabled. In another embodiment, the downscaling feature is enabled if block 1002 is a subblock.
[0105] FIG. 11 shows that when the downscaling feature is enabled, the search areas 1081 and 1082, which includes reference blocks 1010 and 1020, respectively, are downscaled by a factor of M in at least one dimension to obtain two downscaled search areas 1181 and 1182, respectively. In one embodiment, search areas 1081 and 1082 are blocks of samples that have been interpolated according to sub-pixel accuracy of the motion vectors which typically are referred to as predicted samples or blocks of predicted samples. Then when downscaling is performed it is performed on the predicted samples or on blocks of predicted samples. The downscaling may be performed by either picking every M:th sample or picking the M:th sample after averaging M samples. If downscaling is made both horizontally and vertically, the downscaling applies to respective dimension subHorth sample horizontally and subVerth sample vertically or after averaging subHor*subVer samples, where subHor is W1/W2 and subVer=Hl/H2, where W1 is the width of the current block and Hl is the height of the current block, and W2 and H2 are the width and height of the downscaled reference blocks 1110 and 1120. If the size of the current block 1002 is 64x64 and the size of the downscaled reference blocks are 16x16 then subHor=64/16=4 and subVer=64/16=4. The corresponding down-sampled search areas 1181 and 1182 are (16+2*R/4 x 16+2*R/4). In this embodiment, the upper-left hand comer 1115 of downscaled reference block 1110 servers a reference point within downscaled search area 1181 and the upper-left hand comer 1125 of downscaled reference block 1120 servers a reference point within downscaled search area 1182.
[0106] FIGs. 12A and 12B show that, after downscaling, bilateral block matching (BBM) is then applied on the downscaled search areas 1181 and 1182 to determine an optimal candidate offset vector tuple (OVT) (i.e., a first candidate offset vector paired with a second candidate offset vector), which is then used to refine the block motion vectors 1004 and 1006. In one embodiment, the optimal OVT is the OVT from a set of candidate OVTs having the best score. That is, each OVT in the set of candidate OVTs is given a score, and then the optimal OVT is the one with the best score.
[0107] If |R|/4 is 2 (i.e., |R|=8), then the number of candidate offset vectors is 5*5=25. More specifically, with R=2, then the set of candidate offset vectors consist of the following 25 offset vectors: {-2,-2}, {-2,-1 } {-2,0}, {-2,1 }, {-2,2}, {-1,-2}, {-1,-1 } {-1,0}, {-1,1 }, {- 1,2}, {0,-2}, {0,-1 }, {0,0}, {0,1 }, {0,2}, { 1,-2}, { 1,-1 } { 1,0}, { 1,1 }, { 1,2}, {-2,-2}, {-2,-1 } {-2,0}, {-2,1}, {-2,2}. As illustrated, each offset vector has a horizontal component (value) and a vertical component (value). With 25 candidate offset vectors, one can form 25*25 OVTs. Two example OVTs are {{-2,-1 }, { 1,-1 }} and {{0,-1 }, {2, 2}}. In one embodiment there is constraint that the second offset vector of an OVT must be the inverse of the first offset vector of the OVT. For example, if the first offset vector of an OVT is {-1,1 }, then the second offset vector of the OVT must be { 1,-1 }. With this constraint there will be only 25 candidate OVTs when R=2.
[0108] FIG. 12A shows a first OVT that includes a first offset vector 1201 and a second offset vector 1202. FIG. 12B shows a second OVT that includes a third offset vector 1203 and a fourth offset vector 1204. As further shown in FIG. 12A, a block 1211 of downscaled samples is referenced by a combination of the first offset vector (OV) 1201 and reference point 1115 and a block 1212 of downscaled samples is referenced by a combination of the second OV 1202 and reference point 1125. These downscaled blocks 1211 and 1212 have the same dimensions as the downscaled reference blocks 1110 and 1120. And, as further shown in FIG. 12B, a block 1213 of downscaled samples is referenced by a combination of the third OV 1203 and reference point 1115 and a block 1214 of downscaled samples is referenced by a combination of the fourth OV 1204 and reference point 1125. These downscaled blocks 1213 and 1214 have the same dimensions as the downscaled reference blocks 1110 and 1120.
[0109] A score for the first OVT (which consists of OV 1201 and OV 1202) is calculated based on a difference between block 1211 and block 1212. For instance, the sum of absolute differences (SAD) between downscaled reference samples from block 1211 and block 1212 can be calculated and the score for the first OVT is this calculated sum.
Likewise, a score for the second OVT (which consists of OV 1203 and OV 1204) can be calculated based on a difference between block 1213 and block 1214. Every OVT included in the set of candidate OVTs is likewise associated with a pair of downscaled blocks, one from each downscaled search area. Accordingly, each OVT included in the set of candidate OVTs is given a score based on the pair of downscaled blocks with which the OVT is associated in the manner described above for the first OVT and the second OVT. The candidate OVT with the best score is then chosen to refine the motion vectors 1004 and 1006. In the example given the OVT with the best score is the OVT with the lowest calculate sum of sample differences.
[0110] Because the motion estimation is derived using downscaled blocks, in one embodiment, the optimal offset is upscaled to the resolution of the current block. The refined motion vector is defined as follows:
where ov_l_h is the horizontal component of the first candidate offset vector of the optimal OVT, ov_l_v is the vertical component of the first offset vector, ov_2_h is the horizontal component of the second candidate offset vector of the optimal OVT, ov_2_v is the vertical component of the second offset vector. As noted above, in one embodiment, the second candidate offset vector is the inverse of the first candidate offset vector - - i.e., ov_2_h = -l*ov_l_h, and ov_2_v = -l*ov_l_v.
[OHl] Hierarchical Motion Estimation
[0112] In one embodiment, a hierarchical motion estimation is performed by using the block motion vectors 1004, 1006 (mv refA and mv refB) and deriving refinements of the block motion vectors for large subblocks of the block and deriving refinements of the motion of the large subblocks for a smaller subblock size.
[0113] For example, as illustrated in FIG. 13, if the current block 1002 has a size of 64x64, the current block is divided into four blocks of size 32x32, and each of these blocks correspond to two reference blocks of size 32x32. For instance, 32x32 block 1300 of block 1002 has two corresponding blocks of size 32x32, namely block 1301 and block 1302.
[0114] For each of the 32x32 blocks, an optimal offset vector from a first set of candidate offset vectors is determined. That is, for each of the 32x32 blocks, bilateral block matching (BBM) is applied on the corresponding reference blocks to determine an optimal offset vector. In one embodiment, the optimal offset vector is the offset vector from the first set of candidate offset vectors that gives the least sum of absolute differences (SAD) between samples of the reference blocks, which may or may not be downscaled. In one embodiment, the two-dimensional (2D) search area is 32+2*R?2 x 32+2*R?2, where 2*R.32 is the search range in one dimension.
[0115] Each 32x32 block is then divided into four blocks of size 16x16, and each of these blocks correspond to two reference blocks of size 16x16. For instance, 16x16 block 1310 of block 1002 has two corresponding blocks of size 16x16, namely block 1311 and block 1312. For each of the 16x16 blocks, an optimal offset vector from a second set of candidate offset vectors is determined. That is, for each of the 16x16 blocks, bilateral block matching (BBM) is applied on the corresponding reference blocks to determine an optimal offset vector. In one embodiment, the optimal offset vector is the offset vector from the second set of candidate offset vectors that gives the least sum of absolute differences (SAD) between samples of the reference blocks, which may or may not be downscaled. In one embodiment, the two-dimensional (2D) search area is 16+2*Ri6 x 16+2*Ri6, where 2*Ri6 is the search range is one dimension.
[0116] The final refined motion for a particular one of the 16x16 blocks can be formulated as:
where offsetH32 and offsetV32 are the horizontal and vertical components, respectively, of the optimal offset vector determined for the 32x32 block in which the 16x16 block is found, and offsetH16 and offsetV16 are the horizontal and vertical components, respectively, of the optimal offset vector determined for the 16x16 block. offsetH32+offsetH16 should be less or equal to +/- R, where 2*R is the maximum search range in one dimension.
[0117] In one embodiment, the SAD for evaluating the match between the reference blocks is biased towards the motion determined at a larger subblock. This is a way to increase the robustness of the motion estimation of smaller subblocks.
[0118] This can for example be done by scaling the SAD for motions deviating from the larger subblock motion with a factor greater than 1. The factor can be greater for motions deviating more from the larger subblock motion than for motions deviating less from the larger subblock motion. Alternatively, or in addition the SAD for larger subblocks can be scaled with a factor smaller than 1.
[0119] For example, if the SAD for a motion vector that corresponds to a larger subblock is A, the SAD for a motion that is different from that motion has a SAD of B*2. Where B is the actual SAD and 2 is the scaling factor.
[0120] In one embodiment, the search range is reduced for smaller subblocks compared to larger subblocks. This is a way to reduce complexity but may also increase robustness of motion estimation.
[0121] For example, if the block size is 128x128 the largest subblocks are 64x64 and may have a search range of +- R samples around the block motion. The next level of subblocks are 32x32 and may then have a search range of +- R/2 samples around respective 64x64 subblock motion. If the smallest subblocks are 16x16 they may then have a search
range of +-R/4 samples around corresponding 32x32 subblock motion. One example of R is 8.
[0122] In an alternative embodiment the search range for the first stages of the hierarchy is +-R and the search range for the last stage of the hierarchy, is +- P, where P is smaller than R.
[0123] In one embodiment, after the best integer refinement is found it is further refined by sub-pixel refinement. This is preferably only done in the last stage (smallest subblocks).
[0124] In one embodiment, the search range R is kept same for all subblocks. This makes sure that the size of the reference areas to be searched is kept same.
[0125] For example, when downscaling is used and the downscaled size is 16x16 samples, the 64x64 subblocks are downsampled to 16x16 and thus the search range is also downscaled to maintain same search range to +-R/4, and the search range for the 32x32 subblocks is downscaled to +-R/2 while the search range for the 16x16 subblocks is kept as R.
[0126] In one embodiment, after the best integer refinement is found it is further refined by sub-pixel refinement. This is preferably only done in the last stage (smallest subblocks).
[0127] In one embodiment, the enabling of the features described herein is controlled by the spatial activity of the block. If the spatial activity is below a threshold the features can be used.
[0128] The spatial activity can for example be measured on respective uni -prediction blocks according to the respective uni-parts of the bi-predictive block motion vector. Alternatively, it can be performed on subblocks of the block.
[0129] Let Rt j be the samples of the prediction for the current block that would be the result of using one uni-part of the unrefined bi-predictive block motion vector, i.e., the motion vector of the main block. Then, the activity Actt j for the current block can be calculated as Ac
il)/2. The spatial activity for the block can be measured as Sa
/MxN, where MxN is the size of the block or the subblock.
[0130] In one embodiment the enabling of the features is controlled by the QP (quantization parameter) for the block. If the QP is equal to or greater than a threshold the feature is enabled. In one example, the QP threshold is 32.
[0131] In one embodiment, the enabling of the features is controlled by the absolute value of a component of the motion vector. If the magnitude of one component is greater than a threshold the method is enabled. In one example, the threshold is 7 or 112 in 16 :th pel accuracy (7*16).
[0132] In this embodiment, the enabling of the method is controlled by the size of the video picture. If the picture size is greater than a threshold the method is enabled. In one example, the threshold is 1920x1080.
[0133] FIG. 14 is a flowchart illustrating process 1400, according to some embodiments, for obtaining motion compensated predicted samples (e.g., obtaining an interprediction subblock for use in obtaining a reconstructed block by adding the inter-prediction subblock and a reconstructed residual subblock). Process 1400 may begin in step sl402.
[0134] Step sl402 comprises obtaining, for a current block (e.g., block 1002), a first reference motion vector (mv refA) (e.g., motion vector 1004) associated with a first reference block (e.g., block 1010), wherein mv_refA = {mv_refA_h, mv_refA_v}.
[0135] Step sl404 comprises obtaining, for the current block, a second reference motion vector (mv refB) (e.g., motion vector 1006) associated with a second reference block (e.g., block 1020), wherein mv_refB = {mv_refB_h, mv_refB_v}.
[0136] Step sl406 comprises producing a first refined MV (mv refmedA) and a second refined MV (mv refmedB) using downscaled reference blocks.
[0137] Step sl408 comprises obtaining motion compensated predicted samples using MV_refinedA and MV_refinedB.
[0138] In one embodiment, producing mv refmedA and mv refmedB using downscaled blocks comprises the following steps:
[0139] (1) producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector (e.g., associated with ov_l and a first reference point associated with the first motion vector, such as reference point 1115) and a second downscaled block associated with
ov_2 and the second motion vector (e.g., associated with ov_2 and a second reference point associated with the second motion vector, such as reference point 1125);
[0140] (2) producing a score for a second OVT comprising a third offset vector, ov_3, and a foruth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector (e.g., associated with ov_3 and said first reference point) and a fourth downscaled block associated with ov_4 and the second motion vector (e.g., associated with ov_4 and said second reference point);
[0141] (3) selecting the first OVT based on the score for the first OVT, wherein the selecting comprises comparing the score for the first OVT with the score for the second OVT; and
[0142] (4) as a result of selecting the first OVT, producing mv_refmedA based on mv_refA and ov_l and producing mv refmedB based on mv refB and ov_2 or ov_l.
[0143] FIG. 15 is a flowchart illustrating process 1500, according to some embodiments, for obtaining motion compensated predicted samples (e.g., obtaining an interprediction subblock for use in obtaining a reconstructed block by adding the inter-prediction subblock and a reconstructed residual subblock). Process 1500 may begin in step si 502.
[0144] Step si 502 comprises obtaining a first reference motion vector (mv refA) for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv_refA_v}.
[0145] Step si 504 comprises determining a first offset vector (ov_l) using a first reference motion vector for a second block of samples (e.g., a 32x32 block) and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples.
[0146] Step si 506 comprises determining a second offset vector (ov_2) using a first reference motion vector for a third block of samples (e.g., a 16x16 block) and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples.
[0147] Step si 508 comprises producing a first refined motion vector, mv_refined_3_l, for the third block of samples based on mv refA, ov_l, and ov_2.
[0148] Step si 510 comprises obtaining motion compensated predicted samples using the first refined motion vector.
[0149] FIG. 16 is a flowchart illustrating process 1600, according to some
embodiments, for obtaining motion compensated predicted samples (e.g., obtaining an interprediction subblock for use in obtaining a reconstructed block by adding the inter-prediction subblock and a reconstructed residual subblock). Process 1600 may begin in step sl602.
[0150] Step si 602 comprises obtaining a first motion vector, mv_refA, and a second motion vector, mv_refB, for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv_refA_v} and mv_refB = {mv_refB_h, mv_refB_v}.
[0151] Step si 604 comprises selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples (e.g., a 32x32 block), wherein the second block of samples is a subblock within the first block of samples.
[0152] Step sl606 comprises producing a first refined motion vector, mv refinedl, for the second block of samples based on mv refA and the first OVT.
[0153] Step si 608 comprises producing a second refined motion vector, mv_refined2, for the second block of samples based on mv refB and the first OVT.
[0154] Step s 1610 comprises obtaining motion compensated predicted samples using the first and the second refined motion vectors.
[0155] Selecting the first OVT comprises the following steps:
[0156] (1) producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples (e.g., associated with the first offset vector of the first OVT and a reference point associated with the first motion vector for the second block of samples) and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples (e.g., associated with the second offset vector of the first OVT and a second reference point associated with the second motion vector for the second block of samples);
[0157] (2) producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and
[0158] (3) determining that the score for the first OVT is better than the score for the second
OVT.
[0159] In some embodiments, process 1600 further includes: defining a first search area
within a first reference block using the first motion vector for the second block of samples; defining a second search area within a second reference block using the second motion vector for the second block of samples; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block is within the first downscaled search area, and the second downscaled block is within the second downscaled search area.
[0160] FIG. 17 is a block diagram of an apparatus 1700 for implementing encoder 702 and/or decoder 704, according to some embodiments. When apparatus 1700 implements encoder 702, apparatus 1700 may be referred to as an encoder apparatus, and when apparatus 1700 implements decoder 704, apparatus 1700 may be referred to as a decoder apparatus. As shown in FIG. 17, apparatus 1700 may comprise: processing circuitry (PC) 1702, which may include one or more processors (P) 1755 (e.g., one or more general purpose microprocessors and/or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like), which processors may be co-located in a single housing or in a single data center or may be geographically distributed (i.e., apparatus 1700 may be a distributed computing apparatus); at least one network interface 1748 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 1745 and a receiver (Rx) 1747 for enabling apparatus 1700 to transmit data to and receive data from other nodes connected to a network 100 (e.g., an Internet Protocol (IP) network) to which network interface 1748 is connected (physically or wirelessly) (e.g., network interface 1748 may be coupled to an antenna arrangement comprising one or more antennas for enabling apparatus 1700 to wirelessly transmit/receive data); and a storage unit (a.k.a., “data storage system”) 1708, which may include one or more non-volatile storage devices and/or one or more volatile storage devices. In embodiments where PC 1702 includes a programmable processor, a computer readable storage medium (CRSM) 1742 may be provided. CRSM 1742 may store a computer program (CP) 1743 comprising computer readable instructions (CRI) 1744. CRSM 1742 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 1744 of computer program 1743 is configured such that when executed by PC 1702, the CRI causes apparatus 1700 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, apparatus 1700 may be configured to perform steps described herein without the need for code. That is, for example, PC 1702 may consist
merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and/or software.
[0161] Additional Disclosure
[0162] As noted above, DMVR can refine a bi-directional block motion vector on subblocks of size 16x16. Considering that the largest block size in ECM is 256x256, subblocks of size 16x16 are very small and can in some cases get difficulties to find correct refinements of the block motion vectors. This disclosure combats the large size difference between block and subblocks by introducing a hierarchical subblock based DMVR. The reference blocks are divided into subblocks of reduced size in each dimension compared to the block size at each level of the hierarchy until the subblock size reach the ECM subblock size. Integer DMVR search is performed on each level of the hierarchy on downscaled subblocks of size 16x16. In the last level besides integer DMVR also subpixel DMVR search as in ECM-12.0 is performed. The search range is kept same as in ECM-12.0. The final subblock refinement is biased towards the larger subblock refinements by penalizing refinements from the corresponding larger subblock refinement. The objective benefit is asserted to be: -0.02% in random access configuration with most impact at 4K resolution. More importantly it is also asserted that subjective quality is improved by a more coherent motion field.
[0163] Introduction
[0164] In ECM-12.0 the largest block size is 256x256. DMVR is a powerful technology to reduce the signalling of motion information and refines a bi-directional block motion vector on block basis but also on subblocks of 16x16 samples. The large gap between the size of the block and the size of the subblock can make the refinements on subblocks incoherent. This contribution addresses that issue by deploying DMVR search in a hierarchical manner.
[0165] Proposal
[0166] Alt A: When subblock DMVR is applicable and the width or height of the current block is equal to or greater than 64, hierarchical subblock-based DMVR is enabled. The reference areas and search range given by the respective block motion vector of the bidirectional motion vector are of the same size as in ECM.
[0167] The current block (e.g., 64x64) is divided into subblocks of half size in each dimension when possible. The corresponding reference subblocks and search areas are down- sampled so that each subblock has a size of 16x16 samples. The integer DMVR search is then applied on the downscaled subblocks to derive a refinement vector for each subblock.
[0168] Then the current subblock (e.g., 32x32) is divided into smaller subblocks of half size in each dimension when possible. The corresponding reference subblocks and search area are down-sampled so that each smaller subblock has a size of 16x16 samples. The integer DMVR search is then applied on the downscaled smaller subblocks to derive a refinement vector for each smaller subblock.
[0169] When a subblock reaches size 16x16 no further downscaling is done and integer DMVR is applied, followed by sub-pixel DMVR as in ECM. The final subblock refinement is biased towards the higher level subblock refinements by penalizing refinements that differ from the corresponding larger subblocks motion. This is similar to the current ECM which penalizes large refinements from the block motion vector. One example of hierarchical subblock based DMVR is shown for a block size of 64x64 based on bidirectional motion vector (Mv refA, Mv refB) in FIG. 12.
[0170] Alt B: To keep the number of search points not exceeding ECM- 12.0 a complexity reduced approach has also been implemented where the criterion to enable the approach is based of width being equal or greater than 64 when height is equal or greater than 32 or when height is equal or greater than 64 and width is equal or greater than 32. At the last stage of the hierarchy the integer search range is reduced to +-5. This reduces the number of integer search points for a 32x64 block to about half compared to 16x16 DMVR in ECM.
[0171] Results
[0172] Initial results have been compared to ECM-12.0 on random access configuration where DMVR is applied. Encoding and decoding time are not accurate.
[0173] Alt A: The initial results are not complete since QP22 for ParkRunning is not complete so results from anchor have been copied for those points.
[0174] AltB:
[0175] Summary of Various Embodiments
[0176] Al. A method 1400 for obtaining motion compensated predicted samples, the method comprising: obtaining, for a current block, a first motion vector, mv refA, associated with a first block, wherein mv_refA = {mv_refA_h, mv_refA_v}; obtaining, for the current block, a second motion vector, mv refB, associated with a second block, wherein mv refB = {mv refB h, mv refB v}; producing a first refined MV, mv refmedA, and a second refined MV, mv refmedB, using downscaled blocks; and obtaining motion compensated predicted samples using MV refinedA and MV refinedB, wherein producing mv refmedA and mv refmedB using downscaled blocks comprises: producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector (e.g., associated with ov_l and a first reference point associated with the first motion vector, such as reference point 1115) and a second downscaled block associated with ov_2 and the second motion vector (e.g., associated with ov_2 and a second reference point associated with the second motion vector, such as reference point 1125); producing a score for a second OVT comprising a third offset vector, ov_3, and a fourth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector (e.g., associated with ov_3 and said first reference point) and a fourth downscaled block associated with ov_4 and the second motion vector (e.g., associated with ov_4 and said second reference point); selecting the first OVT based on the score for the first OVT, wherein the selecting comprises comparing the score for the first OVT with the score for the second OVT; and as a result of selecting the first OVT, producing mv refmedA based on mv refA and ov_l and producing mv refmedB based on mv refB and ov_2 or ov_l .
[0177] A2. The method of embodiment Al, wherein the first block is a block of
WlxHl samples, the first downscaled block is a block of W2xH2 samples, and W2 < W1 and/or H2 < Hl.
[0178] A3. The method of embodiment A2, wherein ov_l = {ov_l_h, ov_l_v}, mv_refmedA = {mv_refmedA_h, mv_refmedA_v}, ov_2 = {ov_2_h, ov_2_v}, mv refmedB = {mv refmedB h, mv refmedB v}, producing mv refmedA based on mv refA and ov_l comprises calculating: mv refmedA h = mv refA h + subH*ov_l_h, and mv refmedA v = mv refA v + subV*ov_l_v, producing mv refmedB based on mv refB and ov_2 comprising calculating: mv refmedB h = mv refB h + subH*ov_2_h, and mv refmedB v = mv refB v + subV*ov_2_v, subH is a first integer factor, and subV is a second integer
factor.
[0179] A4. The method of embodiment A3, wherein subH = W1/W2, and subV =
H1/H2.
[0180] A5. The method of any one of embodiments A2-A4, wherein the method further comprises: defining a first search area using the first motion vector, wherein the first block associated with the first motion vector is within the first search area; defining a second search area using the second motion vector, wherein the second block associated with the second motion vector is within the second search area; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block associated with the first offset vector is within the first downscaled search area, and the second downscaled block associated with the second offset vector is within the second downscaled search area.
[0181] A6. The method of embodiment A5, wherein the first search area has a size of
W3xH3, where W3=(W1+2*R), H3=(H1+2*R), and R > 2, and the first downscaled search area has a size of W4xH4, where W4=W3/(W1/W2) and or H4=H3/(H1/H2).
[0182] A7. The method of any one of embodiments A2-A6, wherein the method further comprises: prior to producing mv refmedA and mv refmedB using downscaled reference blocks, determining that a condition is satisfied, wherein the step of producing mv refmedA and mv refmedB is performed in response to determining that the condition is satisfied.
[0183] A8. The method of embodiment A7, wherein determining that the condition is satisfied comprises performing at least one of the following steps: determining that W1 is greater than or equal to a width threshold; determining that Hl is greater than or equal to a height threshold; determining that subblocks are used; determining that a spatial activity of the first block is less than a spatial activity threshold; determining that a quantization parameter, QP, for the first block is greater than a QP threshold; determining that the absolute value of a component of the first reference MV is greater than an MV magnitude threshold; determining that the absolute value of a component of the second reference MV is greater than the MV magnitude threshold; or determining that a size of the picture to which the first block belongs is greater than a picture size threshold.
[0184] A9. The method of any one of embodiments A1-A8, wherein producing the
score for the first OVT comprises computing a sum of absolute differences between samples of the first downscaled block and samples of the second downscaled block.
[0185] A10. The method of any one of embodiments A1-A9, wherein ov_2_h = - l*ov_l_h, ov_2_v = -l*ov_l_v, ov_4_h = -l*ov_3_h, and ov_4_v = -l*ov_3_v
[0186] Bl. A method 1500 for obtaining motion compensated predicted samples, the method comprising: obtaining a first motion vector, mv refA, for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv_refA_v}; determining a first offset vector, ov_l, using a first motion vector for a second block of samples (e.g., a 32x32 block) and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples; determining a second offset vector, ov_2, using a first reference motion vector for a third block of samples (e.g., a 16x16 block) and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples; producing a first refined motion vector, mv_refmed_3_l, for the third block of samples based on mv refA, ov_l, and ov_2; and obtaining motion compensated predicted samples using the first refined motion vector.
[0187] B2. The method of embodiment Bl, wherein mv_refmed_3_l =
{mv_refmed_3_l_h, mv_refmed_3_l_v}, ov_l = {ov_l_h, ov_l_v}, ov_2 = {ov_2_h, ov_2_v}, and producing the first refined motion vector for the third block of samples based on mv_refA, ov_l, and ov_2 comprises: setting mv_refmed_3_l_h equal to (mv_refA_h + ov_l_h + ov_2_h); and setting mv_refmed_3_l_v equal to (mv_refA_v + ov_l_v + ov_2_v).
[0188] B3. The method of embodiment B2, wherein the method further comprises: obtaining a second reference motion vector, mv refB, for the first block of samples, wherein mv refB = {mv refB h, mv reffi v}; determining a third offset vector, ov_3, using a second motion vector for the second block of samples; determining a fourth offset vector, ov_4, using a second reference motion vector for the third block of samples; producing a second refined motion vector, mv_refmed_3_2, for the third block of samples based on mv refB, ov_3, and ov_4.
[0189] B4. The method of embodiment B3, wherein mv_refmed_3_2 =
{mv_refmed_3_2_h, mv_refmed_3_2_v}, and producing the second refined motion vector for the third block of samples based on mv refB, ov_3, and ov_4 comprises: setting mv_refmed_3_2_h equal to (mv_refB_h + ov_3_h + ov_4_h); and setting mv_refmed_3_2_v
equal to (mv refB v + ov_3_v + ov_4_v).
[0190] B5. The method of any one of embodiments B3-B4, wherein the method further comprises: determining a fifth offset vector, ov_5, using a first reference motion vector for a fourth block of samples, wherein the fourth block of samples is a subblock within the second block of samples and does not overlap with the third block of samples; and producing a first refined motion vector, mv_refmed_4_l, for the fourth block of samples based on mv_refA, ov_l, and ov_5.
[0191] B6. The method of embodiment B5, wherein mv_refmed_4_l =
{mv_refmed_4_l_h, mv_refmed_4_l_v}, ov_5 = {ov_5_h, ov_5_v], and producing the first refined motion vector for the fourth block of samples based on mv refA, ov_l, and ov_5 comprises: setting mv_refmed_4_l_h equal to (mv_refA_h + ov_l_h + ov_5_h); and setting mv_refmed_4_l_v equal to (mv_refA_v + ov_l_v + ov_5_v).
[0192] B7. The method of any one of embodiments B1-B6, wherein determining the first offset vector, ov_l, using the first reference motion vector for the second block of samples and the first set of candidate offset vectors comprises: for each candidate offset vector included in the first set of candidate offset vectors, producing a score for the candidate offset vector using a first block of predicted samples chosen based on a combination of the first reference motion vector for the second block of samples and the candidate offset vector; and selecting a candidate offset vector from the first set of candidate offset vectors having the best score.
[0193] B8. The method of B7, wherein producing the score for the candidate offset vector using the first block of predicted samples comprises: producing the score for the candidate offset vector using the first block of predicted samples and a second block of predicted samples chosen based on a combination of a second reference motion vector for the second block and an offset vector corresponding to candidate offset vector.
[0194] B9. The method of embodiment B8, wherein producing the score for the candidate offset vector using the first block of predicted samples and the second block of predicted samples comprises computing a sum of absolute differences between samples of the first block and samples of the second block.
[0195] B10. The method of any one of embodiments B1-B9, wherein determining the second offset vector, ov_2, using the first reference motion vector for the third block of
samples and the second set of candidate offset vectors comprises: for each candidate offset vector included in the second set of candidate offset vectors, producing a score for the candidate offset vector using a block of predicted samples chosen based on a combination of the first reference motion vector for the third block of samples and the candidate offset vector; and selecting a candidate offset vector from the second set of candidate offset vectors having the best score.
[0196] Bl 1. The method of any one of embodiments Bl -BIO, wherein the method further comprises: prior to producing the first refined motion vector determining that a condition is satisfied, wherein the step of producing the first refined motion vector is preformed in response to determining that the condition is satisfied.
[0197] B12. The method of embodiment Bl 1, wherein determining that the condition is satisfied comprises performing at least one of the following steps: determining that W1 is greater than or equal to a width threshold; determining that Hl is greater than or equal to a height threshold; determining that subblocks are used; determining that a spatial activity of the first block is less than a spatial activity threshold; determining that a quantization parameter, QP, for the first block is greater than a QP threshold; determining that the absolute value of a component of the first reference MV is greater than an MV magnitude threshold; determining that the absolute value of a component of the second reference MV is greater than the MV magnitude threshold; or determining that a size of the picture to which the first block belongs is greater than a picture size threshold.
[0198] B13. The method of B12, wherein determining that W1 is greater than or equal to a width threshold also determining that Hl is greater than or equal to a second height threshold or wherein determining that Hl is greater than or equal to a height threshold also determining that W1 is greater than or equal to a second width threshold.
[0199] B14. The method of embodiment B2, wherein the method further comprises: obtaining a second reference motion vector, mv refB, for the first block of samples, wherein mv refB = {mv refB h, mv reffi v}; and producing a second refined motion vector, mv_refmed_3_2, for the third block of samples based on mv refB, ov_l, and ov_2.
[0200] Bl 5. The method of embodiment Bl 4, wherein mv_refmed_3_2 =
{mv_refmed_3_2_h, mv_refmed_3_2_v}, and producing the second refined motion vector for the third block of samples based on mv refB, ov_l, and ov_2 comprises: setting
mv_refined_3_2_h equal to (mv refB h - ov_l_h - ov_2_h); and setting mv_refmed_3_2_v equal to (mv refB v - ov_l_v - ov_2_v).
[0201] Cl. A method 1600 for obtaining motion compensated predicted samples, the method comprising: obtaining a first motion vector, mv refA, and a second motion vector, mv_refB, for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv refA v] and mv refB = {mv refB h, mv refB v}; selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples (e.g., a 32x32 block), wherein the second block of samples is a subblock within the first block of samples; producing a first refined motion vector, mv refmedl, for the second block of samples based on mv refA and the first OVT; producing a second refined motion vector, mv_refmed2, for the second block of samples based on mv refB and the first OVT; and obtaining motion compensated predicted samples using the first and the second refined motion vectors, wherein selecting the first OVT comprises: producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples (e.g., associated with the first offset vector of the first OVT and a reference point associated with the first motion vector for the second block of samples) and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples (e.g., associated with the second offset vector of the first OVT and a second reference point associated with the second motion vector for the second block of samples); producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and determining that the score for the first OVT is better than the score for the second OVT.
[0202] C2. The method of embodiment Cl, wherein the method further comprises: defining a first search area within a first reference block using the first motion vector for the second block of samples; defining a second search area within a second reference block using the second motion vector for the second block of samples; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block is within the first
downscaled search area, and the second downscaled block is within the second downscaled search area.
[0203] DI. A computer program (1743) comprising instructions (1744) which when executed by processing circuitry (1702) cause the processing circuitry (1702) to perform the method of any one of the above embodiments.
[0204] D2. A carrier containing the computer program of embodiment Cl, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1742).
[0205] El. An apparatus (1700) for obtaining motion compensated predicted samples, the apparatus being configured to perform a method comprising: obtaining, for a current block, a first motion vector, mv refA, associated with a first block, wherein mv refA = {mv refA h, mv refA v}; obtaining, for the current block, a second motion vector, mv refB, associated with a second block, wherein mv refB = {mv refB h, mv refB v}; producing a first refined MV, mv refmedA, and a second refined MV, mv refmedB, using downscaled blocks; and obtaining motion compensated predicted samples using MV refinedA and MV refinedB, wherein producing mv refmedA and mv refmedB using downscaled blocks comprises: producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector (e.g., associated with ov_l and a first reference point associated with the first motion vector, such as reference point 1115) and a second downscaled block associated with ov_2 and the second motion vector (e.g., associated with ov_2 and a second reference point associated with the second motion vector, such as reference point 1125); producing a score for a second OVT comprising a third offset vector, ov_3, and a fourth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector (e.g., associated with ov_3 and said first reference point) and a fourth downscaled block associated with ov_4 and the second motion vector (e.g., associated with ov_4 and said second reference point); selecting the first OVT based on the score for the first OVT, wherein the selecting comprises comparing the score for the first OVT with the score for the second OVT; and as a result of selecting the first OVT, producing mv refmedA based on mv refA and ov_l and producing mv refmedB based on mv refB and ov_2 or ov_l .
[0206] E2. The apparatus of embodiment El, wherein the apparatus is further configured to perform the method of any one of embodiments A2-A10.
[0207] Fl. An apparatus (1700) for obtaining motion compensated predicted samples, the apparatus being configured to perform a method comprising: obtaining a first motion vector, mv_refA, for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv_refA_v}; determining a first offset vector, ov_l, using a first motion vector for a second block of samples (e.g., a 32x32 block) and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples; determining a second offset vector, ov_2, using a first reference motion vector for a third block of samples (e.g., a 16x16 block) and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples; producing a first refined motion vector, mv_refmed_3_l, for the third block of samples based on mv refA, ov_l, and ov_2; and obtaining motion compensated predicted samples using the first refined motion vector.
[0208] F2. The apparatus of embodiment Fl, wherein the apparatus is further configured to perform the method of any one of embodiments B2-B15.
[0209] Gl. An apparatus (1700) for obtaining motion compensated predicted samples, the apparatus being configured to perform a method comprising: obtaining a first motion vector, mv refA, and a second motion vector, mv refB, for a first block of samples (e.g., a 64x64 block), wherein mv_refA = {mv_refA_h, mv_refA_v} and mv_refB = {mv refB h, mv refB v}; selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples (e.g., a 32x32 block), wherein the second block of samples is a subblock within the first block of samples; producing a first refined motion vector, mv refmedl, for the second block of samples based on mv refA and the first OVT; producing a second refined motion vector, mv_refmed2, for the second block of samples based on mv refB and the first OVT; and obtaining motion compensated predicted samples using the first and the second refined motion vectors, wherein selecting the first OVT comprises: producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples (e.g., associated with the first offset vector of the first OVT and a reference point associated with the first motion vector for
the second block of samples) and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples (e.g., associated with the second offset vector of the first OVT and a second reference point associated with the second motion vector for the second block of samples); producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and determining that the score for the first OVT is better than the score for the second OVT.
[0210] G2. The apparatus of embodiment Gl, wherein the apparatus is further configured to perform the method of embodiments C2.
[0211] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above described exemplary embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
[0212] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.
Claims
1. A method (1400) for obtaining motion compensated predicted samples, the method comprising: obtaining (si 402), for a current block, a first motion vector, mv refA, associated with a first block, wherein mv_refA = {mv_refA_h, mv_refA_v}; obtaining (sl405), for the current block, a second motion vector, mv refB, associated with a second block, wherein mv refB = {mv refB h, mv reffi v}; producing (si 406) a first refined MV, mv_refmedA, and a second refined MV, mv refmedB, using downscaled blocks; and obtaining (si 408) motion compensated predicted samples using MV refinedA and MV_refinedB, wherein producing mv refmedA and mv refmedB using downscaled blocks comprises: producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector and a second downscaled block associated with ov_2 and the second motion vector; producing a score for a second OVT comprising a third offset vector, ov_3, and a fourth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector and a fourth downscaled block associated with ov_4 and the second motion vector; selecting the first OVT based on the score for the first OVT, wherein the selecting comprises comparing the score for the first OVT with the score for the second OVT; and as a result of selecting the first OVT, producing mv refmedA based on mv refA and ov_l and producing mv refmedB based on mv refB and ov_2 or ov_l .
2. The method of claim 1, wherein the first block is a block of WlxHl samples, the first downscaled block is a block of W2xH2 samples, and
W2 < W1 and/or H2 < Hl.
3. The method of claim 2, wherein ov_l = {ov_l_h, ov_l_v}, mv_refinedA = {mv_refinedA_h, mv_refinedA_v}, ov_2 = {ov_2_h, ov_2_v}, mv refinedB = {mv refinedB h, mv refinedB v}, producing mv refinedA based on mv refA and ov_l comprises calculating: mv refinedA h = mv refA h + subH*ov_l_h, and mv_refinedA_v = mv_refA_v + subV*ov_l_v, producing mv refinedB based on mv refB and ov_2 comprising calculating: mv refinedB h = mv refB h + subH*ov_2_h, and mv_refmedB_v = mv_refB_v + subV*ov_2_v, subH is a first integer factor, and subV is a second integer factor.
4. The method of claim 3, wherein subH = W1/W2, and subV = H1/H2.
5. The method of any one of claims 2-4, wherein the method further comprises: defining a first search area using the first motion vector, wherein the first block associated with the first motion vector is within the first search area; defining a second search area using the second motion vector, wherein the second block associated with the second motion vector is within the second search area; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block associated with the first offset vector is within the first downscaled search area, and the second downscaled block associated with the second offset vector is within the second downscaled search area.
6. The method of claim 5, wherein
the first search area has a size of W3xH3, where W3=(W1+2*R), H3=(H1+2*R), and R > 2, and the first downscaled search area has a size of W4xH4, where W4=W3/(W1/W2) and or H4=H3/(H1/H2).
7. The method of any one of claims 2-6, wherein the method further comprises: prior to producing mv refinedA and mv refinedB using downscaled reference blocks, determining that a condition is satisfied, wherein the step of producing mv refinedA and mv refinedB is performed in response to determining that the condition is satisfied.
8. The method of claim 7, wherein determining that the condition is satisfied comprises performing at least one of the following steps: determining that W1 is greater than or equal to a width threshold; determining that Hl is greater than or equal to a height threshold; determining that subblocks are used; determining that a spatial activity of the first block is less than a spatial activity threshold; determining that a quantization parameter, QP, for the first block is greater than a QP threshold; determining that the absolute value of a component of the first reference MV is greater than an MV magnitude threshold; determining that the absolute value of a component of the second reference MV is greater than the MV magnitude threshold; or determining that a size of the picture to which the first block belongs is greater than a picture size threshold.
9. The method of any one of claims 1-8, wherein producing the score for the first OVT comprises computing a sum of absolute differences between samples of the first downscaled block and samples of the second downscaled block.
10. The method of any one of claims 1-9, wherein
ov_2_h = -ov_l_h, ov_2_v = -ov_l_v, ov_4_h = -ov_3_h, and ov_4_v = -ov_3_v.
11. A method (1500) for obtaining motion compensated predicted samples, the method comprising: obtaining (si 502) a first motion vector, mv refA, for a first block of samples, wherein mv_refA = {mv_refA_h, mv_refA_v}; determining (sl504) a first offset vector, ov_l, using a first motion vector for a second block of samples and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples; determining (si 506) a second offset vector, ov_2, using a first reference motion vector for a third block of samples and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples; producing (sl508) a first refined motion vector, mv_refmed_3_l, for the third block of samples based on mv refA, ov_l, and ov_2; and obtaining (si 510) motion compensated predicted samples using the first refined motion vector.
12. The method of claim 11, wherein mv_refmed_3_l = {mv_refmed_3_l_h, mv_refmed_3_l_v}, ov_l = {ov_l_h, ov_l_v}, ov_2 = {ov_2_h, ov_2_v}, and producing the first refined motion vector for the third block of samples based on mv_refA, ov_l, and ov_2 comprises: setting mv_refmed_3_l_h equal to (mv_refA_h + ov_l_h + ov_2_h); and setting mv_refmed_3_l_v equal to (mv_refA_v + ov_l_v + ov_2_v).
13. The method of claim 12, wherein the method further comprises: obtaining a second reference motion vector, mv refB, for the first block of samples, wherein mv refB = {mv refB h, mv refB v};
determining a third offset vector, ov_3, using a second motion vector for the second block of samples; determining a fourth offset vector, ov_4, using a second reference motion vector for the third block of samples; producing a second refined motion vector, mv_refmed_3_2, for the third block of samples based on mv refB, ov_3, and ov_4.
14. The method of claim 13, wherein mv_refmed_3_2 = {mv_refmed_3_2_h, mv_refmed_3_2_v}, and producing the second refined motion vector for the third block of samples based on mv refB, ov_3, and ov_4 comprises: setting mv_refmed_3_2_h equal to (mv_refB_h + ov_3_h + ov_4_h); and setting mv_refmed_3_2_v equal to (mv_refB_v + ov_3_v + ov_4_v).
15. The method of any one of claims 13-14, wherein the method further comprises: determining a fifth offset vector, ov_5, using a first reference motion vector for a fourth block of samples, wherein the fourth block of samples is a subblock within the second block of samples and does not overlap with the third block of samples; and producing a first refined motion vector, mv_refmed_4_l, for the fourth block of samples based on mv refA, ov_l, and ov_5.
16. The method of claim 15, wherein mv_refmed_4_l = {mv_refmed_4_l_h, mv_refmed_4_l_v}, ov_5 = {ov_5_h, ov_5_v}, and producing the first refined motion vector for the fourth block of samples based on mv_refA, ov_l, and ov_5 comprises: setting mv_refmed_4_l_h equal to (mv_refA_h + ov_l_h + ov_5_h); and setting mv_refmed_4_l_v equal to (mv_refA_v + ov_l_v + ov_5_v).
17. The method of any one of claims 11-16, wherein determining the first offset vector, ov_l, using the first reference motion vector for the second block of samples and the first set of candidate offset vectors comprises:
for each candidate offset vector included in the first set of candidate offset vectors, producing a score for the candidate offset vector using a first block of predicted samples chosen based on a combination of the first reference motion vector for the second block of samples and the candidate offset vector; and selecting a candidate offset vector from the first set of candidate offset vectors having the best score.
18. The method of claim 17, wherein producing the score for the candidate offset vector using the first block of predicted samples comprises: producing the score for the candidate offset vector using the first block of predicted samples and a second block of predicted samples chosen based on a combination of a second reference motion vector for the second block and an offset vector corresponding to candidate offset vector.
19. The method of claim 18, wherein producing the score for the candidate offset vector using the first block of predicted samples and the second block of predicted samples comprises computing a sum of absolute differences between samples of the first block and samples of the second block.
20. The method of any one of claims 11-19, wherein determining the second offset vector, ov_2, using the first reference motion vector for the third block of samples and the second set of candidate offset vectors comprises: for each candidate offset vector included in the second set of candidate offset vectors, producing a score for the candidate offset vector using a block of predicted samples chosen based on a combination of the first reference motion vector for the third block of samples and the candidate offset vector; and selecting a candidate offset vector from the second set of candidate offset vectors having the best score.
21. The method of any one of claims 11-20, wherein the method further comprises: prior to producing the first refined motion vector determining that a condition is satisfied, wherein
the step of producing the first refined motion vector is preformed in response to determining that the condition is satisfied.
22. The method of claim 21, wherein determining that the condition is satisfied comprises performing at least one of the following steps: determining that W1 is greater than or equal to a width threshold; determining that Hl is greater than or equal to a height threshold; determining that subblocks are used; determining that a spatial activity of the first block is less than a spatial activity threshold; determining that a quantization parameter, QP, for the first block is greater than a QP threshold; determining that the absolute value of a component of the first reference MV is greater than an MV magnitude threshold; determining that the absolute value of a component of the second reference MV is greater than the MV magnitude threshold; or determining that a size of the picture to which the first block belongs is greater than a picture size threshold.
23. The method of claim 22, wherein determining that W1 is greater than or equal to a width threshold also determining that Hl is greater than or equal to a second height threshold or wherein determining that Hl is greater than or equal to a height threshold also determining that W1 is greater than or equal to a second width threshold.
24. The method of claim 11, wherein the method further comprises: obtaining a second reference motion vector, mv refB, for the first block of samples, wherein mv refB = {mv refB h, mv refB v}; and producing a second refined motion vector, mv_refmed_3_2, for the third block of samples based on mv refB, ov_l, and ov_2.
25. The method of claim 24, wherein mv_refmed_3_2 = {mv_refmed_3_2_h, mv_refmed_3_2_v}, and
producing the second refined motion vector for the third block of samples based on mv refB, ov_l, and ov_2 comprises: setting mv_refined_3_2_h equal to (mv refB h - ov_l_h - ov_2_h); and setting mv_refined_3_2_v equal to (mv_refB_v - ov_l_v - ov_2_v).
26. A method (1600) for obtaining motion compensated predicted samples, the method comprising: obtaining (si 602) a first motion vector, mv refA, and a second motion vector, mv_refB, for a first block of samples, wherein mv_refA = {mv_refA_h, mv_refA_v} and mv refB = {mv refB h, mv reffi v}; selecting (si 604) a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples, wherein the second block of samples is a subblock within the first block of samples; producing (si 606) a first refined motion vector, mv refinedl, for the second block of samples based on mv refA and the first OVT; producing (si 608) a second refined motion vector, mv_refined2, for the second block of samples based on mv refB and the first OVT; and obtaining (si 610) motion compensated predicted samples using the first and the second refined motion vectors, wherein selecting the first OVT comprises: producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples; producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and determining that the score for the first OVT is better than the score for the second OVT.
27. The method of claim 26, wherein the method further comprises: defining a first search area within a first reference block using the first motion vector for the second block of samples; defining a second search area within a second reference block using the second motion vector for the second block of samples; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block is within the first downscaled search area, and the second downscaled block is within the second downscaled search area.
28. A computer program (1743) comprising instructions (1744) which when executed by processing circuitry (1702) cause the processing circuitry (1702) to perform the method of any one of the above claims.
29. A carrier containing the computer program of claim 28, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1742).
30. An apparatus (1700) for obtaining motion compensated predicted samples, the apparatus being configured to perform a method comprising: obtaining, for a current block, a first motion vector, mv refA, associated with a first block, wherein mv_refA = {mv_refA_h, mv_refA_v}; obtaining, for the current block, a second motion vector, mv refB, associated with a second block, wherein mv refB = {mv refB h, mv reffi v}; producing a first refined MV, mv refmedA, and a second refined MV, mv refmedB, using downscaled blocks; and obtaining motion compensated predicted samples using MV refinedA and MV_refinedB, wherein producing mv refmedA and mv refmedB using downscaled blocks comprises:
producing a score for a first offset vector tuple, OVT, comprising a first offset vector, ov_l, and a second offset vector, ov_2, using a first downscaled block associated with ov_l and the first motion vector and a second downscaled block associated with ov_2 and the second motion vector; producing a score for a second OVT comprising a third offset vector, ov_3, and a fourth offset vector, ov_4, using a third downscaled block associated with ov_3 and the first motion vector and a fourth downscaled block associated with ov_4 and the second motion vector; selecting the first OVT based on the score for the first OVT, wherein the selecting comprises comparing the score for the first OVT with the score for the second OVT; and as a result of selecting the first OVT, producing mv refmedA based on mv refA and ov_l and producing mv refmedB based on mv refB and ov_2 or ov_l .
31. The apparatus of claim 30, wherein the apparatus is further configured to perform the method of any one of claims 2-10.
32. An apparatus (1700) for obtaining motion compensated predicted samples, the apparatus being configured to perform a method comprising: obtaining a first motion vector, mv refA, for a first block of samples, wherein mv_refA = {mv_refA_h, mv_refA_v}; determining a first offset vector, ov_l, using a first motion vector for a second block of samples and a first set of candidate offset vectors, wherein the second block of samples is a subblock within the first block of samples; determining a second offset vector, ov_2, using a first reference motion vector for a third block of samples and a second set of candidate offset vectors, wherein the third block of samples is a subblock within the second block of samples; producing a first refined motion vector, mv_refmed_3_l, for the third block of samples based on mv refA, ov_l, and ov_2; and obtaining motion compensated predicted samples using the first refined motion vector.
33. The apparatus of claim 32, wherein the apparatus is further configured to perform the method of any one of claims 12-25.
34. An apparatus (1700) for obtaining motion compensated predicted samples, the apparatus being configured to perform a method comprising: obtaining a first motion vector, mv refA, and a second motion vector, mv refB, for a first block of samples, wherein mv_refA = {mv_refA_h, mv_refA_v} and mv_refB = {mv_refB_h, mv_refB_v}; selecting a first offset vector tuple, OVT, from a set of candidate OVTs using a first motion vector and a second motion vector for a second block of samples, wherein the second block of samples is a subblock within the first block of samples; producing a first refined motion vector, mv refmedl, for the second block of samples based on mv refA and the first OVT; producing a second refined motion vector, mv_refmed2, for the second block of samples based on mv refB and the first OVT; and obtaining motion compensated predicted samples using the first and the second refined motion vectors, wherein selecting the first OVT comprises: producing a score for the first OVT using i) a first downscaled block associated with a first offset vector of the first OVT and the first motion vector for the second block of samples and ii) a second downscaled block associated with a second offset vector of the first OVT and the second motion vector for the second block of samples; producing a score for a second OVT using i) a third downscaled block associated with a first offset vector of the second OVT and the first motion vector for the second block of samples ii) a fourth downscaled block associated with a second offset vector of the second OVT and the second motion vector for the second block of samples; and determining that the score for the first OVT is better than the score for the second OVT.
35. The apparatus of claim 34, wherein the method further comprises:
defining a first search area within a first reference block using the first motion vector for the second block of samples; defining a second search area within a second reference block using the second motion vector for the second block of samples; downscaling the first search area to produce a first downscaled search area; downscaling the second search area to produce a second downscaled search area, wherein the first downscaled block is within the first downscaled search area, and the second downscaled block is within the second downscaled search area.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463631525P | 2024-04-09 | 2024-04-09 | |
| US63/631,525 | 2024-04-09 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025214829A1 true WO2025214829A1 (en) | 2025-10-16 |
Family
ID=95290275
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2025/058886 Pending WO2025214829A1 (en) | 2024-04-09 | 2025-04-01 | Motion vector refinement with downsampled blocks, or hierarchical sub-blocks |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025214829A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023277755A1 (en) * | 2021-06-30 | 2023-01-05 | Telefonaktiebolaget Lm Ericsson (Publ) | Selective subblock-based motion refinement |
| WO2023198080A1 (en) * | 2022-04-12 | 2023-10-19 | Beijing Bytedance Network Technology Co., Ltd. | Method, apparatus, and medium for video processing |
| US11895302B2 (en) * | 2021-06-29 | 2024-02-06 | Qualcomm Incorporated | Adaptive bilateral matching for decoder side motion vector refinement |
-
2025
- 2025-04-01 WO PCT/EP2025/058886 patent/WO2025214829A1/en active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11895302B2 (en) * | 2021-06-29 | 2024-02-06 | Qualcomm Incorporated | Adaptive bilateral matching for decoder side motion vector refinement |
| WO2023277755A1 (en) * | 2021-06-30 | 2023-01-05 | Telefonaktiebolaget Lm Ericsson (Publ) | Selective subblock-based motion refinement |
| WO2023198080A1 (en) * | 2022-04-12 | 2023-10-19 | Beijing Bytedance Network Technology Co., Ltd. | Method, apparatus, and medium for video processing |
Non-Patent Citations (2)
| Title |
|---|
| ANDERSSON (ERICSSON) K ET AL: "AHG12: Hierarchical Subblock based DMVR", no. JVET-AH0144, 10 April 2024 (2024-04-10), XP030317474, Retrieved from the Internet <URL:https://jvet-experts.org/doc_end_user/documents/34_Rennes/wg11/JVET-AH0144-v1.zip JVET-AH0144.docx> [retrieved on 20240410] * |
| Y-J CHANG ET AL: "Compression efficiency methods beyond VVC", no. JVET-U0100, 31 December 2020 (2020-12-31), XP030293237, Retrieved from the Internet <URL:https://jvet-experts.org/doc_end_user/documents/21_Teleconference/wg11/JVET-U0100-v1.zip JVET-U0100.docx> [retrieved on 20201231] * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11973973B2 (en) | Prediction refinement based on optical flow | |
| KR102797628B1 (en) | Partial cost calculation | |
| JP7319386B2 (en) | Gradient calculation for different motion vector refinements | |
| CN110620923B (en) | Generalized motion vector difference resolution | |
| CN110662041B (en) | Method and apparatus for video bitstream processing, method of storing video bitstream, and non-transitory computer-readable recording medium | |
| KR102873107B1 (en) | Inter prediction method and device in a video coding system | |
| JP7366901B2 (en) | Video decoding method and device using inter prediction in video coding system | |
| WO2020164582A1 (en) | Video processing method and apparatus | |
| WO2021249375A1 (en) | Affine prediction improvements for video coding | |
| WO2020156515A1 (en) | Refined quantization steps in video coding | |
| WO2020182187A1 (en) | Adaptive weight in multi-hypothesis prediction in video coding | |
| US12439076B2 (en) | Selective subblock-based motion refinement | |
| US20260019583A1 (en) | A method, an apparatus and a computer program product for video encoding and decoding | |
| WO2020063598A1 (en) | A video encoder, a video decoder and corresponding methods | |
| WO2025214829A1 (en) | Motion vector refinement with downsampled blocks, or hierarchical sub-blocks | |
| WO2026005668A1 (en) | Obtaining motion compensated prediction samples | |
| US12621481B2 (en) | Selective application of decoder side refining tools | |
| US20250267273A1 (en) | A method, an apparatus and a computer program product for video encoding and decoding | |
| WO2025075541A1 (en) | Motion vector derivation based on block boundary distortion | |
| EP4690797A1 (en) | Encoder for encoding pictures of a video | |
| WO2025005846A1 (en) | Motion vector derivation | |
| WO2024175831A1 (en) | A method, an apparatus and a computer program product for video encoding and decoding | |
| EP4677853A1 (en) | A method, an apparatus and a computer program product for video encoding and decoding | |
| AU2024360744A1 (en) | A method, an apparatus and a computer program product for video encoding and decoding |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25716990 Country of ref document: EP Kind code of ref document: A1 |