2023PF00496 NEW HYPOTHESIS FOR MULTI-HYPOTHESIS INTER PREDICTION MODE 1. CROSS REFERENCE TO RELATED APPLICATIONS This application claims priority to European Application No.23306028.4, filed June 27, 2023, which is incorporated herein by reference in its entirety. 2. TECHNICAL FIELD At least one of the present embodiments generally relates to a method and a device for improving a multi-hypothesis inter prediction mode in video compression methods. 3. BACKGROUND To achieve high compression efficiency, video coding schemes usually employ predictions and transforms to leverage spatial and temporal redundancies in a video content. During an encoding, pictures of the video content are divided into blocks of samples (i.e., Pixels), these blocks being then partitioned into one or more sub-blocks, called original sub-blocks in the following. An intra or inter prediction is then applied to each sub-block to exploit intra or inter image correlations. Whatever the prediction method used (intra or inter), a predictor sub-block is determined for each original sub- block. Then, a sub-block representing a difference between the original sub-block and the predictor sub-block, often denoted as a prediction error sub-block, a prediction residual sub-block or simply a residual sub-block, is transformed, quantized and entropy coded to generate an encoded video stream. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to the transform, quantization and entropic coding. Inter prediction consists in predicting a current block of a current picture from at least one predictor block from a reference picture preceding or following the current picture. A predictor block is identified in the reference picture by motion information. Recently, a new inter prediction mode called multi-hypothesis inter prediction (MHP) mode had been proposed (see document JVET-M0425, CE10: Multi-hypothesis inter prediction (Test 10.1.2), Martin Winken, Heiko Schwarz, Detlev Marpe, Thomas
2023PF00496 Wiegand, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 1113th Meeting: Marrakech, MA, 9–18 Jan.2019). In the MHP mode, in addition to a traditional mono-prediction or bi-prediction predictor, one or more additional inter predictors (also called additional prediction hypothesis) are signaled. A resulting overall predictor is obtained for the current block by a sample-wise weighted superposition. Motion parameters of each additional prediction hypothesis can be signaled either explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference, or implicitly by specifying a merge index. These explicit or implicit signaling are considered sub-optimal because none of these signaling consider the mode nor motion vector already in use by the current block. It is desirable to propose solutions allowing to overcome the above issues. In particular, it is desirable to propose solutions improving the MHP mode by considering the base prediction of the current block. 4. BRIEF SUMMARY In a first aspect, one or more of the present embodiments provide a method for encoding a current block of a current picture in video data comprising: obtaining a first motion information of a first predictor for the current block; deriving a second motion information for at least one second predictor from the first motion information of the first predictor; and, generating a final predictor for the current block using at least the first and each second predictor. In a second aspect, one or more of the present embodiments provide a method for decoding a current block of a current picture from video data comprising: obtaining a first motion information of a first predictor for the current block; deriving a second motion information for at least one second predictor from the first motion information of the first predictor; and, generating a final predictor for the current block using at least the first and each second predictor.
2023PF00496 In an embodiment of the first or second aspect, the first predictor is a base predictor obtained first for the current block or an additional predictor obtained for the current block after the base predictor. In an embodiment of the first or second aspect, the second motion information comprises a second motion vector the coordinates of which depending on coordinates of a first motion vector comprised in the first motion information weighted by a weighting factor. In an embodiment of the first or second aspect, the second motion information comprises a second reference picture index on a second reference picture, the second reference picture index depending on a first reference picture index on a first reference picture comprised in the first motion information and on the weighting factor. In an embodiment of the first or second aspect, the second motion information results from an application of a refinement process to an intermediate motion information derived from the first motion information, the refinement process allowing obtaining a motion information refinement of the intermediate motion information, the second motion information being a sum of the intermediate motion information and the motion information refinement. In an embodiment of the first or second aspect, the motion information refinement is obtained by minimizing a sum of absolute difference between the first predictor and the second predictor. In an embodiment of the first or second aspect, the motion information refinement is obtained by minimizing a sum of absolute difference between the current block and the second predictor, the motion information refinement being signaled in the video data. In an embodiment of the first or second aspect, the motion information refinement is obtained by a template matching process. In an embodiment of the first or second aspect, the first motion information is a bi-prediction motion information comprising two first pairs, each first pair comprising a motion vector and a reference picture index, the second motion information comprising two second pairs, each second pair being derived from one of the first pairs, a second predictor being obtained from each second pair. In an embodiment of the first or second aspect, the final predictor is based on at least one of the second predictors.
2023PF00496 In an embodiment of the first or second aspect, the final predictor is based on one of the second predictors selected based on a criterion. In an embodiment of the first or second aspect, the criterion depends on a mode used for obtaining the first predictor. In an embodiment of the first or second aspect, a syntax element signals that the second predictor is obtained based on second motion information derived from the first motion information. In an embodiment of the first or second aspect, a use of the second predictor obtained based on second motion information derived from the first motion information is allowed by a high level syntax. In a third aspect, one or more of the present embodiments provide a device for encoding a current block of a current picture in video data comprising electronic circuitry configured for: obtaining a first motion information of a first predictor for the current block; deriving a second motion information for at least one second predictor from the first motion information of the first predictor; and, generating a final predictor for the current block using at least the first and each second predictor. In a fourth aspect, one or more of the present embodiments provide a device for decoding a current block of a current picture from video data comprising electronic circuitry configured for: obtaining a first motion information of a first predictor for the current block; deriving a second motion information for at least one second predictor from the first motion information of the first predictor; and, generating a final predictor for the current block using at least the first and each second predictor. In an embodiment of the third or fourth aspect, the first predictor is a base predictor obtained first for the current block or an additional predictor obtained for the current block after the base predictor.
2023PF00496 In an embodiment of the third or fourth aspect, the second motion information comprises a second motion vector the coordinates of which depending on coordinates of a first motion vector comprised in the first motion information weighted by a weighting factor. In an embodiment of the third or fourth aspect, the second motion information comprises a second reference picture index on a second reference picture, the second reference picture index depending on a first reference picture index on a first reference picture comprised in the first motion information and on the weighting factor. In an embodiment of the third or fourth aspect, the second motion information results from an application of a refinement process to an intermediate motion information derived from the first motion information, the refinement process allowing obtaining a motion information refinement of the intermediate motion information, the second motion information being a sum of the intermediate motion information and the motion information refinement. In an embodiment of the third or fourth aspect, the motion information refinement is obtained by minimizing a sum of absolute difference between the first predictor and the second predictor. In an embodiment of the third or fourth aspect, the motion information refinement is obtained by minimizing a sum of absolute difference between the current block and the second predictor, the motion information refinement being signaled in the video data. In an embodiment of the third or fourth aspect, the motion information refinement is obtained by a template matching process. In an embodiment of the third or fourth aspect, the first motion information is a bi-prediction motion information comprising two first pairs, each first pair comprising a motion vector and a reference picture index, the second motion information comprising two second pairs, each second pair being derived from one of the first pairs, a second predictor being obtained from each second pair. In an embodiment of the third or fourth aspect, the final predictor is based on at least one of the second predictors. In an embodiment of the third or fourth aspect, the final predictor is based on one of the second predictors selected based on a criterion.
2023PF00496 In an embodiment of the third or fourth aspect, the criterion depends on a mode used for obtaining the first predictor. In an embodiment of the third or fourth aspect, a syntax element signals that the second predictor is obtained based on second motion information derived from the first motion information. In an embodiment of the third or fourth aspect, a use of the second predictor obtained based on second motion information derived from the first motion information is allowed by a high level syntax. In a fifth aspect, one or more of the present embodiments provide a non- transitory information storage medium storing program code instructions for implementing the method according to the first or the second aspect. In a sixth aspect, one or more of the present embodiments provide a computer program comprising program code instructions for implementing the method according to the method according to the first or the second aspect. In a seventh aspect, one or more of the present embodiments provide a signal generated by the method of the first aspect or by the device of the third aspect. 5. BRIEF SUMMARY OF THE DRAWINGS Fig. 1 illustrates an example of context in which various embodiments may be implemented; Fig. 2 illustrates schematically an example of partitioning undergone by a picture of pixels of an original video; Fig.3 depicts schematically a method for encoding a video stream; Fig.4 depicts schematically a method for decoding an encoded video stream; Fig. 5A illustrates schematically an example of hardware architecture of a processing module able to implement an encoding module or a decoding module in which various aspects and embodiments are implemented; Fig. 5B illustrates a block diagram of an example of a first system in which various aspects and embodiments are implemented;
2023PF00496 Fig.5C illustrates a block diagram of an example of a second system in which various aspects and embodiments are implemented; Figs.6A and 6B and illustrates schematically spatial and temporal positions considered for constructing a list of merge candidates; Fig.7 represents schematically a process for decoding the syntax elements of the multi- hypothesis inter prediction mode; Fig. 8 represents schematically a process for generating a multi-hypothesis inter prediction mode predictor according to an embodiment; Fig. 9 describes schematically a process for decoding syntax elements of the MHP mode compliant with various embodiments; and, Fig. 10 describes schematically a process for encoding syntax elements of the MHP mode compliant with various embodiments. 6. DETAILED DESCRIPTION The following examples of embodiments are described in the context of a video format similar to VVC (ISO/IEC 23090-3 – MPEG-I : Versatile Video Coding (VVC) / ITU-T H.266). However, these embodiments are not limited to the video coding/decoding method corresponding to VVC. These embodiments are in particular adapted to various video formats comprising for example HEVC (ISO/IEC 23008-2 – MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265)), AVC ((ISO/CEI 14496-10), EVC (Essential Video Coding/MPEG-5), AV1, AV2 and VP9. Fig.1 describes an example of a context in which following embodiments can be implemented. In Fig. 1, a system 11, that could be a camera, a storage device, a computer, a server or any device capable of delivering a video stream, transmits a video stream to a system 13 using a communication channel 12. The video stream is either encoded and transmitted by the system 11 or received and/or stored by the system 11 and then transmitted. The communication channel 12 is a wired (for example Internet or Ethernet) or a wireless (for example WiFi, 3G, 4G or 5G) network link. The system 13, that could be for example a set top box, receives and decodes the video stream to generate a sequence of decoded pictures. The obtained sequence of decoded pictures is then transmitted to a display
2023PF00496 system 15 using a communication channel 14, that could be a wired or wireless network. The display system 15 then displays said pictures. In an embodiment, the system 13 is comprised in the display system 15. In that case, the system 13 and display system 15 are comprised in a TV, a computer, a tablet, a smartphone, a head-mounted display, etc. Figs.2, 3 and 4 introduce an example of video format. Fig.2 illustrates an example of partitioning undergone by a picture of pixels 21 of an original video sequence 20. It is considered here that a pixel is composed of three components: a luminance component and two chrominance components. Other types of pixels are however possible comprising less or more components such as only a luminance component or an additional depth component or transparency component. A picture is divided into a plurality of coding entities. First, as represented by reference 23 in Fig. 2, a picture is divided in a grid of blocks called coding tree units (CTU). A CTU consists of an ^^ ൈ ^^ block of luminance samples together with two corresponding blocks of chrominance samples. N is generally a power of two having a maximum value of “128” for example. Second, a picture is divided into one or more groups of CTU. For example, it can be divided into one or more tile rows and tile columns, a tile being a sequence of CTU covering a rectangular region of a picture. In some cases, a tile could be divided into one or more bricks, each of which consisting of at least one row of CTU within the tile. Above the concept of tiles and bricks, another encoding entity, called slice, exists, that can contain at least one tile of a picture or at least one brick of a tile. In the example of Fig.2, as represented by reference 22, the picture 21 is divided into three slices S1, S2 and S3 of the raster-scan slice mode, each comprising a plurality of tiles (not represented), each tile comprising only one brick. As represented by reference 24 in Fig. 2, a CTU may be partitioned into the form of a hierarchical tree of one or more sub-blocks called coding units (CU). The CTU is the root (i.e., the parent node) of the hierarchical tree and can be partitioned in a plurality of CU (i.e. child nodes). Each CU becomes a leaf of the hierarchical tree if it is not further partitioned in smaller CU or becomes a parent node of smaller CU (i.e., child nodes) if it is further partitioned. In the example of Fig.2, the CTU 24 is first partitioned in “4” square CU using a quadtree type partitioning. The upper left CU is a leaf of the hierarchical tree since it
2023PF00496 is not further partitioned, i.e., it is not a parent node of any other CU. The upper right CU is further partitioned in “4” smaller square CU using again a quadtree type partitioning. The bottom right CU is vertically partitioned in “2” rectangular CU using a binary tree type partitioning. The bottom left CU is vertically partitioned in “3” rectangular CU using a ternary tree type partitioning. During the coding of a picture, the partitioning is adaptive, each CTU being partitioned so as to optimize a compression efficiency of the CTU criterion. In HEVC appeared the concept of prediction unit (PU) and transform unit (TU). Indeed, in HEVC, the coding entity that is used for prediction (i.e., a PU) and transform (i.e., a TU) can be a subdivision of a CU. For example, as represented in Fig.2, a CU of size 2 ^^ ൈ 2 ^^, can be divided in PU 2411 of size ^^ ൈ 2 ^^ or of size 2 ^^ ൈ ^^. In addition, said CU can be divided in “4” TU 2412 of size ^^ ൈ ^^ or in “16” TU of size ^ே ே ଶ^ ൈ ^ ଶ^. can note that in VVC, except in some particular cases, frontiers of the TU
and PU are aligned on the frontiers of the CU. Consequently, a CU comprises generally one TU and one PU. In the present application, the term “block” or “picture block” can be used to refer to any one of a CTU, a CU, a PU and a TU. In addition, the term “block” or “picture block” can be used to refer to a macroblock, a partition and a sub-block as specified in H.264/AVC or in other video coding standards, and more generally to refer to an array of samples of numerous sizes. In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture”, “sub-picture”, “slice” and “frame” may be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side. Fig.3 depicts schematically a method for encoding a video stream executed by an encoding module. Variations of this method for encoding are contemplated, but the method for encoding of Fig. 3 is described below for purposes of clarity without describing all expected variations. Before being encoded, a current original picture of an original video sequence may go through a pre-processing. For example, in a step 301, a color transform is applied to the current original picture (e.g., conversion from RGB 4:4:4 to YCbCr
2023PF00496 4:2:0), or a remapping is applied to the current original picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one of the color components). Pictures obtained by pre-processing are called pre-processed pictures in the following. The encoding of a pre-processed picture begins with a partitioning of the pre- processed picture during a step 302, as described in relation to Fig.2. The pre-processed picture is thus partitioned into CTU, CU, PU, TU, etc. For each block, the encoding module determines a coding mode between an intra prediction and an inter prediction. The intra prediction consists of predicting, in accordance with an intra prediction method, during a step 303, the pixels of a current block from a prediction block derived from pixels of reconstructed blocks situated in a causal vicinity of the current block to be encoded. The result of the intra prediction is a prediction mode indicating which pixels of the blocks in the vicinity to use, and a residual block resulting from a calculation of a difference between the current block and the prediction block. The inter prediction consists of predicting the pixels of a current block from a block of pixels, referred to as the reference block, of a picture preceding or following the current picture, this picture being referred to as the reference picture. During the coding of a current block in accordance with the inter prediction method, a block of the reference picture closest, in accordance with a similarity criterion, to the current block is determined by a motion estimation step 304. During step 304, a motion vector indicating the position of the reference block in the reference picture is determined. Said motion vector is used during a motion compensation step 305, involving interpolation operations between samples of the reference block, to generate a prediction block. A residual block is then calculated in the form of a difference between the current block and the prediction block. In first video compression standards, the mono-prediction inter mode described above was the only inter mode available. As video compression standards evolve, the family of inter modes has grown significantly and comprises now many different inter modes, such as bi-prediction modes in which a current is predicted from two reference blocks designated by two different motion information. A mentioned in introduction of this document, a new inter prediction mode called multi-hypothesis inter prediction (MHP) mode had been proposed. In the MHP mode, in addition to a traditional mono-prediction or bi-prediction predictor, one or
2023PF00496 more additional inter predictors (also called additional prediction hypothesis) are signaled. A resulting overall predictor is obtained for the current block by a sample- wise weighted superposition. With a mono-prediction or bi-prediction predictor signal ^^^^^^_^^ and a first additional prediction hypothesis ℎଷ, a resulting MHP predictor ^^ଷ is obtained as follows: ^^ଷ ൌ ^1 െ ^^^ ^^^^^^_^^ ^ ^^ℎଷ ^^ is a weighting factor specified for example by a syntax element add_hyp_weight_idx, according to a mapping represented in table TAB1: add_hyp_weight_idx ^^ 0 1/4 1 -1/8 Table TAB1 More than one additional prediction hypothesis can be used. The resulting overall MHP predictor is accumulated iteratively with each additional prediction hypothesis. ^^^ା^ ൌ ^ 1 െ ^^^ା^ ^ ^^^ ^ ^^^ା^ℎ^ା^ (eq. 1) The resulting MHP predictor is obtained as the last ^^^ (i.e., the ^^^ having the largest index ^^). In some implementations, up to two additional prediction hypothesis can be used (i.e., ^^ is limited to maxNumAddHyp=2), with up to a maximum of four different reference pictures for a single block. The motion parameters of each additional prediction hypothesis can be signaled either explicitly by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly by specifying a merge index (see below for a description of the merge mode). A separate multi-hypothesis merge flag distinguishes between these two signalling modes. The MHP process can either explicitly or implicitly signal the motion vector used for the additional prediction hypothesis. This is supposed to provide the encoder with a choice of a costly but accurate motion information (when using the explicit signaling), or a cheaper motion information (when using the implicit signaling). This is asserted to be sub-optimal because none of the two signaling processes consider the
2023PF00496 mode nor motion information already in use by the current block (i.e., the base prediction). During a selection step 306, the prediction mode optimising the compression performances, in accordance with a rate/distortion optimization criterion (i.e., RDO criterion), among the prediction modes tested (Intra prediction modes, Inter prediction modes), is selected by the encoding module. When the prediction mode is selected, the residual block is transformed during a step 307. In some implementations, a plurality of type of transforms can be applied to a transformed residual block. Indeed, in addition to DCT-II, a Multiple Transform Selection (MTS) scheme is used for both inter and intra predicted blocks. It uses multiple selected transforms from the DCT-VIII/DST-VII. The transformed block is then quantized during a step 309. Note that the encoding module can skip the transform and apply quantization directly to the non-transformed residual signal. The quantized residual block determined for the current block during an inter or intra prediction is encoded by an entropic encoder during a step 310. Note that the encoding module can bypass both transform and quantization, i.e., the entropic encoding is applied on the residual without the application of the transform or quantization processes. The result of the entry coding is inserted in the video data 311. When the current block is coded according to an intra prediction mode, the intra prediction mode is encoded by the entropic encoder during the step 310 in the video data 311. When the current block is encoded according to an inter prediction, a process is applied to encode the motion information. The output of this process is then encoded by the entropic encoder during the step 310 in the video data 311. Two processes are employed to encode the motion information: AMVP (Adaptive Motion Vector Prediction) or Merge. In each process, the motion information are predicted. In an implementation of the AMVP mode, a motion vector predictor (MVP) is selected, and a motion vector difference noted MVd relative to the selected MVP is computed. The MVP is selected in a list of AMVP candidates made of “2” candidates.
2023PF00496 The index of the chosen MVP and the MVd are then encoded by the entropic encoder during step 310 along with the transformed and quantized residual block resulting from the inter prediction of the current block. The AMVP candidate list is constructed first by deriving a first spatial candidate from a left block neighbouring the current block, if this block is available and inter coded. Then a second spatial candidate is derived from a top block neighbouring the current block, if this block is available and inter coded. Then, a temporal candidate is derived from a so-called collocated picture at a position collocated with the current block, if an inter block exist at this collocated position. Each derived MVP candidate is scaled according to a temporal distance between the reference picture associated to this MVP candidate and the reference picture considered for the current block. A redundancy check is then conducted between derived spatial candidates and, if a duplicate candidate exists, this candidate is discarded. The final AMVP candidate list contains the two first derived MVP candidates. If less than “2” MVP candidates are obtained through the above process, then the AMVP candidate list is completed with zero motion vectors. The merge mode consists in deriving motion information of a current block from a selected motion information predictor candidate. The motion information considered here includes all the inter prediction parameters of a block, that is to say: the uni-prediction or bi-prediction type, the reference picture index within each reference picture list and the motion vector(s). The selected motion information predictor candidate (i.e., the merge candidate) is selected in a list of motion information predictor candidates (i.e., in a list of merge candidates). When a block is encoded in merge mode, the index of the selected merge candidate is encoded. If no residual block is encoded for the current block, the current block is considered as encoded according to a particular merge mode called skip mode. In some implementations, the list of merge candidates is systematically made of “5” merge candidates. Up to “5” spatial positions are considered to retrieve some potential candidates for the list of merge candidates. Fig. 6A illustrates schematically the five spatial positions considered for constructing a list of merge candidates. These positions are investigated according to the following order: 1. Left (A1)
2023PF00496 2. Above (B1) 3. Above right (B0) 4. Left bottom (A0) 5. Above left (B2) Each spatial candidate is introduced in the list of merge candidates provided that the motion information corresponding to this candidate is not already present in the list of merge candidates. Then a temporal predictor noted TMVP is determined. Fig. 6B illustrates schematically collocated positions considered for determining the TMVP. The determination of the TMVP consists first in investigating position H and, if no motion information is available at position H, the position C is investigated. A last pruning process is then applied to ensure that the set of spatial and temporal candidates does not contain redundant candidates. In case of B-slice (slice allowing bi-predicted blocks), candidates of another type, called combined candidates, are introduced in the list of merge candidates if this list is not full. Finally, if the merge list is still not full then zero motion vectors are introduced in at the end of the merge list until it is full. Metadata such as SEI (supplemental enhancement information) messages can be attached to the encoded video stream 311. A SEI message as defined for example in standards such as AVC, HEVC or VVC is a data container associated to a video stream and comprising metadata providing information relative to the video stream. After the quantization step 309, the current block is reconstructed so that the pixels corresponding to that block can be used for future predictions. This reconstruction phase is also referred to as a prediction loop. An inverse quantization is therefore applied to the transformed and quantized residual block during a step 312 and an inverse transformation is applied during a step 313. According to the prediction mode used for the block obtained during a step 314, the prediction block of the block is reconstructed. If the current block is encoded according to an inter prediction mode, the encoding module applies, when appropriate, during a step 316, a motion compensation using the motion information of the current block in order to identify each reference block of the current block. If the current block is encoded according to an intra prediction mode, during a step 315, the intra prediction mode is used for
2023PF00496 reconstructing the prediction block of the current block. The prediction block and the reconstructed residual block are added in order to obtain the reconstructed current block. Following the reconstruction, an in-loop filtering intended to reduce the encoding artefacts is applied, during a step 317, to the reconstructed block. This filtering is called in-loop filtering since this filtering occurs in the prediction loop to obtain at the decoder the same reference pictures as the encoder and thus avoid a drift between the encoding and the decoding processes. In-loop filtering tools comprises deblocking filtering, SAO (Sample adaptive Offset) and ALF (Adaptive Loop Filtering). When a block is reconstructed, it is inserted during a step 318 into a reconstructed picture stored in a memory 319 of reconstructed pictures generally called Decoded Picture Buffer (DPB). The reconstructed pictures thus stored can then serve as reference pictures for other pictures to be coded. Fig. 4 depicts schematically a method for decoding the encoded video stream 311 encoded according to method described in relation to Fig.3 executed by a decoding module. Variations of this method for decoding are contemplated, but the method for decoding of Fig. 4 is described below for purposes of clarity without describing all expected variations. The decoding is done block by block. For a current block, it starts with an entropic decoding of the current block during a step 410. Entropic decoding allows to obtain, at least, the prediction mode of the block. If the current block has been encoded according to an inter prediction mode, the entropic decoding allows to obtain, when appropriate, information representative of a motion of the current block and a residual block. During a step 408, the motion information is reconstructed for the current block using the decoded information representative of the motion information. In AMVP, the information representative of the motion comprises an index of the AMVP MVP in the AMVP list and the MVd. The MVd is then added to the AMVP MVP to reconstruct the motion vector of the block. In merge, an index of the merge MVP in the merge list is obtained. The merge MVP corresponding to the index provides the motion information of the current block. If the current has been encoded using the MHP mode, syntax elements specific to the MHP mode are decoded.
2023PF00496 Fig. 7 represents schematically a process for decoding the syntax elements of the MHP mode. This process is for example applied during step 408 by the decoding module. In a step 701, a processing module of the decoding module initializes a variable numAddHyp to zero. In a step 702, the processing module compares the value of the variable numAddHyp to the maximum number of additional prediction hypothesis maxNumAddHyp. If numAddHyp = maxNumAddHyp the processing module stops the process in a step 704. Otherwise, in a step 703, the processing module read a flag MHP_Flag indicating if an additional prediction hypothesis numbered numAddHyp is to be used. In a step 705, the processing module determines from the value of the flag MHP_Flag if an additional prediction hypothesis numbered numAddHyp is to be used. If no, the processing module stops the process in the step 704. Otherwise, the processing module executes a step 706. During step 706, the processing module reads a MHP merge flag MHP_merge_flag. The flag MHP_merge_flag indicates which of the implicit or the explicit signaling process is used for signaling the motion information. If in a step 707 the processing module determines that the implicit signaling is used, the processing module continues with the step 708. During step 708, the processing module reads the merge index merge_idx of the MVP of the additional prediction hypothesis numbered numAddHyp. If the explicit signaling is used, the processing module continues with the step 709. During step 709, the processing module reads reference index ref_idx, the MVd mvd and the MVP index MVP_idx of the additional prediction hypothesis numbered numAddHyp. Steps 708 and 709 are followed by a step 710. During step 710, the processing module reads the MHP weight index add_hyp_weight_idx of the additional prediction hypothesis numbered numAddHyp. In a step 711, the processing module generates the additional prediction hypothesis numbered numAddHyp.
2023PF00496 In step 712, the processing module increments the variable numAddHyp of one unit. Step 712 is followed by step 701. If at least one additional prediction hypothesis was generated, in step 704, the processing module generates a MHP predictor using equation eq.1. The MHP predictor is then used to predict the current block. If the block has been encoded according to an intra prediction mode, entropic decoding allows to obtain the intra prediction mode and a residual block. Steps 412, 413, 414, 415, 416 and 417 implemented by the decoding module are in all respects identical respectively to steps 412, 413, 414, 415, 416 and 417 implemented by the encoding module. Decoded blocks are saved in decoded pictures and the decoded pictures are stored in a DPB 419 in a step 418. When the decoding module decodes a given picture, the pictures stored in the DPB 419 are identical to the pictures stored in the DPB 319 by the encoding module during the encoding of said given image. The decoded picture can also be outputted by the decoding module for instance to be displayed. The post-processing step 421 can comprise an inverse color transform (e.g. conversion from YCbCr 4:2:0 to RGB 4:4:4), an inverse mapping performing the inverse of the remapping process performed in the pre-processing of step 301 and a post-filtering for improving the reconstructed pictures based for example on filter parameters provided in a SEI message. Fig. 5A illustrates schematically an example of hardware architecture of a processing module 500 able to implement an encoding module or a decoding module capable of implementing respectively the method for encoding of Fig.3 and the method for decoding of Fig. 4 modified according to different aspects and embodiments. The encoding module is for example comprised in the system 11 when this apparatus is in charge of encoding the video stream. The decoding module is for example comprised in the system 13. The processing module 500 comprises, connected by a communication bus 5005: a processor or CPU (central processing unit) 5000 encompassing one or more microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples; a random access memory (RAM) 5001; a read only memory (ROM) 5002; a
2023PF00496 storage unit 5003, which can include non-volatile memory and/or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and/or optical disk drive, or a storage medium reader, such as a SD (secure digital) card reader and/or a hard disc drive (HDD) and/or a network accessible storage device; at least one communication interface 5004 for exchanging data with other modules, devices or equipment. The communication interface 5004 can include, but is not limited to, a transceiver configured to transmit and to receive data over a communication channel. The communication interface 5004 can include, but is not limited to, a modem or network card. If the processing module 500 implements a decoding module, the communication interface 5004 enables for instance the processing module 500 to receive encoded video streams and to provide a sequence of decoded pictures. If the processing module 500 implements an encoding module, the communication interface 5004 enables for instance the processing module 500 to receive a sequence of original picture data to encode and to provide an encoded video stream. The processor 5000 is capable of executing instructions loaded into the RAM 5001 from the ROM 5002, from an external memory (not shown), from a storage medium, or from a communication network. When the processing module 500 is powered up, the processor 5000 is capable of reading instructions from the RAM 5001 and executing them. These instructions form a computer program causing, for example, the implementation by the processor 5000 of a decoding method as described in relation with Fig.4, an encoding method described in relation to Fig.3, and methods described in relation to Figs.7 to 10, these methods comprising various aspects and embodiments described below in this document. All or some of the algorithms and steps of the methods of Figs.3, 4 and 7 to 10 may be implemented in software form by the execution of a set of instructions by a programmable machine such as a DSP (digital signal processor) or a microcontroller, or be implemented in hardware form by a machine or a dedicated component such as a FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit).
2023PF00496 As can be seen, microprocessors, general purpose computers, special purpose computers, processors based or not on a multi-core architecture, DSP, microcontroller, FPGA and ASIC are electronic circuitry adapted to implement at least partially the methods of Figs.3, 4, and 7 to 10. Fig. 5C illustrates a block diagram of an example of the system 13 in which various aspects and embodiments are implemented. The system 13 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances and head mounted display. Elements of system 13, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and/or discrete components. For example, in at least one embodiment, the system 13 comprises one processing module 500 that implements a decoding module. In various embodiments, the system 13 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports. In various embodiments, the system 13 is configured to implement one or more of the aspects described in this document. The input to the processing module 500 can be provided through various input modules as indicated in block 531. Such input modules include, but are not limited to, (i) a radio frequency (RF) module that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a component (COMP) input module (or a set of COMP input modules), (iii) a Universal Serial Bus (USB) input module, and/or (iv) a High Definition Multimedia Interface (HDMI) input module. Other examples, not shown in FIG.5C, include composite video. In various embodiments, the input modules of block 531 have associated respective input processing elements as known in the art. For example, the RF module can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down-converted and band-
2023PF00496 limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF module of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, down-converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF module and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down- converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and/or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF module includes an antenna. Additionally, the USB and/or HDMI modules can include respective interface processors for connecting system 13 to other electronic devices across USB and/or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within the processing module 500 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within the processing module 500 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to the processing module 500. Various elements of system 13 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangements, for example, an internal bus as known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards. For example, in the system 13, the processing module 500 is interconnected to other elements of said system 13 by the bus 5005. The communication interface 5004 of the processing module 500 allows the system 13 to communicate on the communication channel 12. As already mentioned
2023PF00496 above, the communication channel 12 can be implemented, for example, within a wired and/or a wireless medium. Data is streamed, or otherwise provided, to the system 13, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi- Fi signal of these embodiments is received over the communications channel 12 and the communications interface 5004 which are adapted for Wi-Fi communications. The communications channel 12 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 13 using the RF connection of the input block 531. As indicated above, various embodiments provide data in a non- streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network. The system 13 can provide an output signal to various output devices, including the display system 15, speakers 56, and other peripheral devices 57. The display system 15 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and/or a foldable display. The display system 15 can be for a television, a tablet, a laptop, a cell phone (mobile phone), a head mounted display or other devices. The display system 15 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 57 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and/or a lighting system. Various embodiments use one or more peripheral devices 57 that provide a function based on the output of the system 13. For example, a disk player performs the function of playing an output of the system 13. In various embodiments, control signals are communicated between the system 13 and the display system 15, speakers 56, or other peripheral devices 57 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 13 via dedicated connections through respective interfaces 532, 533, and 534. Alternatively,
2023PF00496 the output devices can be connected to system 13 using the communications channel 12 via the communications interface 5004 or a dedicated communication channel via the communication interface 5004. The display system 15 and speakers 56 can be integrated in a single unit with the other components of system 13 in an electronic device such as, for example, a television. In various embodiments, the display interface 532 includes a display driver, such as, for example, a timing controller (T Con) chip. The display system 15 and speaker 56 can alternatively be separate from one or more of the other components. In various embodiments in which the display system 15 and speakers 56 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs. Fig. 5B illustrates a block diagram of system 11. System 11 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, a camera and a server. Elements of system 11, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and/or discrete components. For example, in at least one embodiment, the system 11 comprises one processing module 500 that implements an encoding module. In various embodiments, the system 11 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports. In various embodiments, the system 11 is configured to implement one or more of the aspects described in this document. The input to the processing module 500 can be provided through various input modules as indicated in block 531 already described in relation to Fig.5C. Various elements of system 11 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangements, for example, an internal bus as known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards. For example, in the system 11, the processing module 500 is interconnected to other elements of said system 11 by the bus 5005. The communication interface 5004 of the processing module 500 allows the
2023PF00496 system 11 to communicate on the communication channel 12. Data is streamed, or otherwise provided, to the system 11, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi- Fi signal of these embodiments is received over the communications channel 12 and the communications interface 5004 which are adapted for Wi-Fi communications. The communications channel 12 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 11 using the RF connection of the input block 531. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network. The data provided to the system 11 can be provided in different format. In various embodiments these data are encoded and compliant with a known video compression format such as AV1, VP9, VVC, HEVC, AVC, etc. In various embodiments, these data are raw data provided for example by a picture and/or audio acquisition module connected to the system 11 or comprised in the system 11. In that case, the processing module take in charge the encoding of these data. The system 11 can provide an output signal to various output devices capable of storing and/or decoding the output signal such as the system 13. Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded video stream in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and prediction. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application, for example, for reconstructing a block predicted using the MHP mode. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on
2023PF00496 the context of the specific descriptions and is believed to be well understood by those skilled in the art. Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded video stream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, prediction, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application, for example, for encoding a block according to the MHP mode. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art. Note that the syntax elements names as used herein, are descriptive terms. As such, they do not preclude the use of other syntax element names. When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method/process. Various embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between a rate and a distortion is usually considered. The rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of a reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on a prediction or a prediction residual signal, not the reconstructed one. Mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible
2023PF00496 encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion. The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented, for example, in a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users. Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment. Additionally, this application may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, retrieving the information from memory or obtaining the information for example from another device, module or from user. Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the
2023PF00496 information, determining the information, predicting the information, or estimating the information. Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information. It is to be appreciated that the use of any of the following “/”, “and/or”, and “at least one of”, “one or more of” for example, in the cases of
“A and/or B” and “at least one of A and B”, “one or more of A and B” is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and/or C” and “at least one of A, B, and C”, “one or more of A, B and C” such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed. Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a use of some coding tools. In this way, in an embodiment the same parameters can be used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is
2023PF00496 realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun. As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the encoded video stream and SEI messages of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding an encoded video stream and modulating a carrier with the encoded video stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium. In the following, several embodiments are proposed allowing deriving additional prediction hypothesis from a base prediction of a block. Fig. 8 represents schematically a process for generating a MHP predictor according to an embodiment. The process of Fig.8 is executed by the processing module 500 of the encoding module when the MHP mode is tested as one of the potential inter prediction modes considered for the current block (during steps 304 and 305) or by the processing module 500 of the decoding module when the current block was encoded according to the MHP mode. In a step 801, the processing module 500 obtains a first motion information ^^ ^^^^^^௧_^^^ௗ of a first predictor for the current block. In a step 802, the processing module 500 derives a second motion information ^^ ^^^^^_^^^ௗ for at least one second predictor from the first motion information ^^ ^^^^^^௧_^^^ௗ of the first predictor. In a step 803, the processing module 500 generates a MHP predictor using at
2023PF00496 least the first and each second predictor. During step 803, all predictors obtained for the current block are combined using a sample-wise weighted superposition based on equation eq. 1. Other predictors not derived from the motion information of the first predictor can also be obtained and used for the generation of the MHP predictor. For instance, another predictor compliant with the explicit or implicit signalization described above can be obtained. The MHP predictor is then used to predict the current block. In a first embodiment, the first predictor is a predictor corresponding to the base prediction of the current block. In a second embodiment, the first predictor is an additional prediction hypothesis previously obtained for the current block. In a first variant of the first and the second embodiments, when the first predictor was obtained from a mono-prediction with a motion information ^^ ^^^^^^௧_^^^ௗ comprising a motion vector with coordinates (x,y) using a reference picture Pref_first_pred, then the motion information for the second predictor ^^ ^^^^^_^^^ௗ comprises a motion vector with coordinates (σ x; σ y) with reference picture Pref_sec_pred. In a sub-variant of the first variant of the first and second embodiments, Pref_sec_pred is selected as the reference picture temporarily closest to PCurr that is different from Pref_first_pred. In a sub-variant of this sub-variant, if the value of ^^ ^^^^^_^^^ௗ found for this reference picture corresponds to an existing motion vector in the merge list, then the reference picture that is second closest to PCurr that is different as Pref_first_pred is selected for Pref_sec_pred . In a sub-variant of the first variant of the first and second embodiments, σ < 0, for example when Pref_sec_pred is in the opposite direction than Pref_first_pred with respect to the current picture PCurr comprising the current block, In a sub-variant of the first variant of the first and second embodiments, | σ | < 1, for example when Pref_sec_pred is closer to the current picture PCurr than the reference picture Pref_first_pred. In an embodiment, the distance (i.e. the Picture Order Count (POC) difference) ^^^^^_^^ୡ_^^^ௗ between the reference picture Pref_sec_pred and the current picture PCurr is equal to ^^ ൈ ^^^^^_^୧୰^^_^^^ௗ where ^^^^^_^୧୰^^_^^^ௗ is the distance between the reference picture Pref_first_pred and the current picture PCurr ( ^^^^^_^^ୡ _^^^ௗ ൌ ^^. ^^^^^_^୧୰^^ _^^^ௗ). In a sub-variant of the first variant of the first and second embodiment, the
2023PF00496 distances ^^^^^_^^ୡ_^^^ௗ and ^^^^^_^୧୰^^_^^^ௗ needs to be the same to use the second predictor to generate the MHP predictor. In this case, the value of σ is always set to “- 1”. In a sub-variant of the first variant of the first and second embodiments, σ is a power of 2, ½, -2 or -1/2 so that the multiplications can be done as shift operations. In a sub-variant of the first variant of the first and second embodiments, a refinement process is applied to the motion information ^^ ^^^^^_^^^ௗ of the second predictor to refine it. A motion information refinement ∆MV is added to the motion information ^^ ^^^^^_^^^ௗ , i.e., ^^ ^^^^^_^^^ௗ ൌ ^^ ^^^^^_^^^ௗ ^ ∆MV. In an embodiment, the motion information refinement ∆MV is obtained by minimizing a sum of absolute difference (SAD) between the first predictor and a second predictor pointed by the refined motion information ^^ ^^^^^_^^^ௗ ^ ∆MV. In an embodiment, the processing module 500 of the encoding module determines the motion information refinement ∆MV by minimizing a sum of absolute difference (SAD) between the current block and a second predictor pointed by the refined motion information ^^ ^^^^^^^ௗ_^^^ௗ ^ ∆MV. The encoding module then signals an information representing the determined motion information refinement ∆MV in the encoded video data 311. The processing module 500 of the decoding module then decodes the information representing the determined motion information refinement ∆MV to determine the second predictor. In an embodiment, the motion information refinement ∆MV is obtained by applying a template matching process. In that case a template is defined in the neighborhood of the current block and in the neighborhood of the predictor designated by the motion information ^^ ^^^^^_^^^ௗ ^ ∆MV. A SAD is then computed between the two templates and the motion information refinement ∆MV minimizing the SAD is selected. The templates are for example L-shaped sets of samples neighboring respectively the current block and the block designated by the motion information ^^ ^^^^^_^^^ௗ ^ ∆MV. In an embodiment, the motion information refinement ∆MV is obtained by applying a search process similar to a search process applied in a mode called DMVR (decoder side motion vector refinement) described in section 3.4.10 of document JVET- T2002-V2, Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11), Jianle Chen, Yan Ye, Seung Hwan Kim, joint Video Experts Team (JVET) of ITU-
2023PF00496 T SG 16 WP 3 and ISO/IEC JTC 1/SC 29 20th Meeting, by teleconference, 7 – 16 October 2020. In an embodiment, the motion information ^^ ^^^^^_^^^ௗ ൌ ^^ ^^^^^^௧_^^^ௗ ^ ∆MV. In that case, only the use of the motion information refinement needs to be signaled (and eventually an information representative of the motion information refinement if the motion information refinement is explicitly signaled in the video data 311). In a second variant of the first and the second embodiments, the first predictor was obtained from a bi-prediction with two motion information ^^ ^^^^^^௧_^^^^^ௗ^ =(x1; y1) and ^^ ^^^^^^௧_^^^^^ௗଶ =(x2; y2) using reference pictures Pref_first_bipred1 and Pref_first_bipred2. In that case, a second motion information ^^ ^^^^^_^^^ௗ is derived from each motion information ^^ ^^^^^^௧_^^^^^ௗ^ and ^^ ^^^^^^௧_^^^^^ௗଶ of the first predictor. Therefore, two second predictors are generated. All sub-variants of the first variant of the first and second embodiments are applicable to the second variant of the first and the second embodiment for deriving a second motion information ^^ ^^^^^_^^^ௗ from each motion information ^^ ^^^^^^௧_^^^^^ௗ^ and ^^ ^^^^^^௧_^^^^^ௗଶ of the first predictor. In an embodiment, the MHP predictor is generated using at least the first predictor and the two second predictors. In an embodiment, the second predictor of the two second predictors the closest to the first predictor in terms of SAD is selected. The MHP predictor is generated using at least the first predictor and the selected second predictor. In the sub-variant based on the refinement process, the refinement process is applied independently to each derived second motion information ^^ ^^^^^_^^^ௗ . The refined motion information ^^ ^^^^^_^^^ௗ ^ ∆MV minimizing a SAD is then selected for generating the second predictor. The SAD is for instance computed between the first predictor and each second predictor pointed by the refined motion information ^^ ^^^^^_^^^ௗ ^ ∆MV. In a third variant of the first and the second embodiment, the first predictor was obtained using the BCW (Bi-prediction with CU-level Weight) mode described in section 3.4.8 of document JVET-T2002-V2, Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11), Jianle Chen, Yan Ye, Seung Hwan Kim, joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 2920th
2023PF00496 Meeting, by teleconference, 7 – 16 October 2020. A block predicted using the BCW mode is associated to two motion vectors and an index indicating a weight to apply to the predictors designated by the two motion vectors. In that case, only one second motion information ^^ ^^^^^_^^^ௗ is derived to generate the second predictor. In an embodiment, the single second motion information ^^ ^^^^^_^^^ௗ is for example derived from the motion vector designating the predictor associated with the highest BCW weight. In a sub-variant of the third variant of the first and the second embodiments, when the first predictor was obtained using the BCW mode, the second motion information ^^ ^^^^^_^^^ௗ comprises two motion vectors derived from the two motion vectors of the first predictor. The second predictor is obtained by applying the same weighting process than the one used in the BCW mode. In an embodiment, the weights applied in the weighting process are the same than the weights used to obtain the first predictor. In an embodiment, the weights applied in the weighting process are both equal to 0.5. In a fourth variant of the first and the second embodiment, the first predictor was obtained using the GPM (Geometric partition mode) mode described in section 3.4.11 of document JVET-T2002-V2, Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11), Jianle Chen, Yan Ye, Seung Hwan Kim, joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 2920th Meeting, by teleconference, 7 – 16 October 2020. A block predicted using the GPM mode is associated to two motion vectors and one partitioning index. The partitioning index indicates how partitioning a combined predictor block in a first and a second partition. The combined predictor is a combination of two predictors designated by the two motion vectors. Samples of the first partition are provided essentially by a first predictor designated by one of the two motion vectors. Samples of the second partition are provided essentially by a second predictor designated by the second of the two motion vectors. In that case, only one second motion vector ^^ ^^^^^_^^^ௗ is derived to generate the second predictor. In an embodiment, the single second motion vector ^^ ^^^^^_^^^ௗ is for example derived from the motion vector associated with the largest partition of the GPM mode.
2023PF00496 Some blocks could be predicted using an affine prediction such as described in section 3.4.4 of document JVET-T2002-V2, Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11), Jianle Chen, Yan Ye, Seung Hwan Kim, joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 2920th Meeting, by teleconference, 7 – 16 October 2020. In an embodiment, the MHP mode is deactivated when the first predictor was obtained using an affine prediction. In a fifth variant of the first and the second embodiment, the first predictor was obtained using an affine prediction. In the fifth variant, a top left Control Point Motion Vector (CPMV) as described in section 3.4.4 of document JVET-T2002-V2, Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11), Jianle Chen, Yan Ye, Seung Hwan Kim, joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 2920th Meeting, by teleconference, 7 – 16 October 2020 is used as the first motion information ^^ ^^^^^^௧_^^^ௗ from which is derived the second motion information ^^ ^^^^^_^^^ௗ . Fig. 9 describes schematically a process for decoding syntax elements of the MHP mode compliant with the various embodiments described above. This process is for example applied during step 408 by the processing module 500 of the decoding module. The process of Fig. 9 shares the same steps as the process of Fig. 7. Two additional steps are introduced between steps 705 and 706. If in step 705, the processing module 500 determines from the value of the flag MHP_Flag that an additional prediction hypothesis numbered numAddHyp is to be used, step 705 is followed by a step 901. During step 901, the processing module 500 reads a flag MHP_derive_flag indicating if the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block (as described in relation to Fig.8). In a step 902, if the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block, step 902 is followed by step 710 already explained in relation to Fig.7.
2023PF00496 If the additional prediction hypothesis numbered numAddHyp is not a second predictor determined from a first predictor already determined for the current block, step 902 is followed by step 706 already explained in relation to Fig.7. In the process of Fig. 9, the flag MHP_derive_flag is decoded before the flag flag MHP_merge_flag. In another embodiment, the flag MHP_derive_flag is decoded after the flag MHP_merge_flag. In that case, steps 901 and 902 are executed between steps 707 and 708 or between steps 707 and 709. If steps 901 and 902 are between steps 707 and 708, the semantic of the flag MHP_merge_flag changes slightly. Indeed, in that case, the flag MHP_merge_flag indicates with the first value that the explicit signaling is used for the additional prediction hypothesis numbered numAddHyp and with a second value that the implicit signaling is used for the additional prediction hypothesis numbered numAddHyp or that the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block. If steps 901 and 902 are between steps 707 and 709, the semantic of the flag MHP_merge_flag changes also slightly. Indeed, in that case, the flag MHP_merge_flag indicates with the first value that the implicit signaling is used for the additional prediction hypothesis numbered numAddHyp and with a second value that the explicit signaling is used for the additional prediction hypothesis numbered numAddHyp or that the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block. In an embodiment, if, for the additional prediction hypothesis numbered numAddHyp, the second motion information ^^ ^^^^^_^^^ௗ is in the merge list of the current block, then it is removed from the merge list for next additional prediction hypotheses. In an embodiment, if, for the additional prediction hypothesis numbered numAddHyp, the second motion information ^^ ^^^^^_^^^ௗ is in the merge list of the current block, then it is replaced by another second motion information different from any motion information present in the merge list. For instance, the second motion information ^^ ^^^^^_^^^ௗ derived in step 802 is replaced by a motion information with the same reference index but with a motion vector equal to (0,0). In an embodiment, the flag MHP_derive_flag is decoded only for the first additional prediction hypothesis (numAddHyp = 0) and is assumed to indicate that next
2023PF00496 additional prediction hypothesis (numAddHyp > 0) are not a second predictors determined from a first predictor already determined for the current block. In an embodiment, a high level syntax element, for instance a SPS or PPS or picture header or slice header syntax element, indicates if in the MHP mode additional prediction hypotheses can be second predictors determined from a first predictor already determined for the current block or not. If the high level syntax element indicates that additional prediction hypotheses cannot be second predictors determined from a first predictor already determined for the current block, the flag MHP_derive_flag is assumed to be equal to zero. In an embodiment, the value σ and/or the motion information refinement ∆MV are explicitly signaled when the flag MHP_derive_flag indicates that the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block to provide more flexibility for the encoder. In an embodiment, to reduce a signaling cost, the flag MHP_derive_flag is only signaled when the base prediction corresponds to certain modes. For example, the flag MHP_derive_flag is signaled only when the base prediction uses the merge mode. The flag MHP_derive_flag is inferred to indicate that any additional prediction hypothesis cannot be a second predictor determined from a first predictor already determined for the current block when the merge mode is not used by for the base prediction. In an embodiment, the mode wherein the additional prediction hypothesis is a second predictor determined from a first predictor already determined for the current block replaces the implicit or the explicit signaling. In that case, the flag MHP_derive_flag is not needed. The semantic of the flag MHP_merge_flag changes slightly. Indeed, in that case, the flag MHP_merge_flag indicates with the first value that the explicit (or implicit) signaling is used for the additional prediction hypothesis numbered numAddHyp and with a second value that the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block. In a variant of this embodiment, a first high level syntax element, for instance a SPS or PPS or picture header or slice header syntax element, indicates if one of the implicit signaling or explicit signaling of the MHP mode is replaced by the mode wherein additional prediction hypotheses are second predictors determined from a first predictor already determined for the current
2023PF00496 block. By default, the implicit signaling is replaced. In an additional variant, if the first high level syntax element indicates that one of the implicit signaling or explicit signaling of the MHP mode is replaced by the mode wherein additional prediction hypotheses are second predictors determined from a first predictor already determined for the current block, a second high level syntax element indicates which of the implicit signaling or explicit signaling is replaced. Fig. 10 describes schematically a process for encoding syntax elements of the MHP mode compliant with the various embodiments described above. This process is for example applied during step 310 by the processing module 500 of the encoding module. It is supposed here that the processing module 500 has executed step 306 during which it has decided to encode the current block using the MHP mode. Furthermore, at least one additional prediction hypothesis is a second predictor determined from a first predictor already determined for the current block. In a step 1001, the processing module 500 initializes a variable numAddHyp to zero. In a step 1002, the processing module 500 compares the value of the variable numAddHyp to the maximum number of additional prediction hypothesis maxNumAddHyp. If numAddHyp = maxNumAddHyp the processing module stops the process in a step 1014. Otherwise, in a step 1003, the processing module 500 determines if an additional prediction hypothesis numbered numAddHyp is to be used. If no, step 1003 is followed by the step 1004 during which the processing module 500 signals the flag MHP_Flag in the video data 311 with a value indicating that no additional prediction hypothesis is used. Step 1004 is followed by step 1014. Otherwise, step 1003 is followed by a step 1005 during which the processing module 500 signals the flag MHP_Flag in the video data with a value indicating that an additional prediction hypothesis numbered numAddHyp is used. In a step 1006, the processing module 500 determines if the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block.
2023PF00496 If yes, step 1006 is followed by step 1007 during which the processing module 500 signals the flag MHP_derive_flag in the video data 311 with a value indicating that the additional prediction hypothesis numbered numAddHyp is a second predictor determined from a first predictor already determined for the current block. Step 1007 is followed by a step 1012 during which the processing module 500 the MHP weight index add_hyp_weight_idx of the additional prediction hypothesis numbered numAddHyp in the video data 311. In a step 1013, the processing module 500 increments the variable numAddHyp of one unit. Step 1013 is followed by step 1001. If the additional prediction hypothesis numbered numAddHyp is not a second predictor determined from a first predictor already determined for the current block, step 1006 is followed by a step 1008. During step 1008 the processing module 500 signals the flag MHP_derive_flag with a value indicating that the additional prediction hypothesis numbered numAddHyp is not a second predictor determined from a first predictor already determined for the current block. Step 1008 is followed by a step 1009 during which the processing module 500 determines if the implicit or the explicit signaling is used for the current block. If the implicit signaling is used, the processing module 500 signals the MHP merge flag MHP_merge_flag with a value indicating the use of the implicit signaling in a step 1010. During step 1010, the processing module 500 signals the merge index merge_idx of the MVP of the additional prediction hypothesis numbered numAddHyp. If the explicit signaling is used, the processing module 500 signals the MHP merge flag MHP_merge_flag with a value indicating the use of the explicit signaling mode in a step 1011. During step 1011, the processing module 500 signals the reference index ref_idx, the MVd mvd and the MVP index MVP_idx of the additional prediction hypothesis numbered numAddHyp Steps 1010 and 1011 are followed by step 1012 already described. Of course, all embodiment and variants of the process for decoding syntax elements of the MHP mode described in relation to Fig. 9 are applied reciprocally by the process for encoding syntax elements.
2023PF00496 We described above a number of embodiments. Features of these embodiments can be provided alone or in any combination. Further, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types: ^ A bitstream or signal that includes one or more of the described syntax elements, or variations thereof. ^ Creating and/or transmitting and/or receiving and/or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof. ^ A TV, set-top box, cell phone, tablet, or other electronic device that performs at least one of the embodiments described. ^ A TV, set-top box, cell phone, tablet, or other electronic device that performs at least one of the embodiments described, and that displays (e.g. using a monitor, screen, or other type of display) a resulting picture. ^ A TV, set-top box, cell phone, tablet, or other electronic device that tunes (e.g. using a tuner) a channel to receive a signal including an encoded video stream, and performs at least one of the embodiments described. ^ A TV, set-top box, cell phone, tablet, or other electronic device that receives (e.g. using an antenna) a signal over the air that includes an encoded video stream, and performs at least one of the embodiments described. ^ A server, camera, cell phone, tablet or other electronic device that transmits (e.g. using an antenna) a signal over the air that includes an encoded video stream, and performs at least one of the embodiments described. ^ A server, camera, cell phone, tablet or other electronic device that tunes (e.g. using a tuner) a channel to transmit a signal including an encoded video stream, and performs at least one of the embodiments described.