EP4512092A1 - Verfahren und vorrichtungen zur kandidatenableitung für affinen zusammenführungsmodus in der videocodierung - Google Patents
Verfahren und vorrichtungen zur kandidatenableitung für affinen zusammenführungsmodus in der videocodierungInfo
- Publication number
- EP4512092A1 EP4512092A1 EP23792452.7A EP23792452A EP4512092A1 EP 4512092 A1 EP4512092 A1 EP 4512092A1 EP 23792452 A EP23792452 A EP 23792452A EP 4512092 A1 EP4512092 A1 EP 4512092A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- area
- block
- affine
- restricted
- current
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/55—Motion estimation with spatial constraints, e.g. at image or region borders
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
- H04N19/137—Motion inside a coding unit, e.g. average field, frame or block difference
- H04N19/139—Analysis of motion vectors, e.g. their magnitude, direction, variance or reliability
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
- H04N19/517—Processing of motion vectors by encoding
- H04N19/52—Processing of motion vectors by encoding by predictive encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/119—Adaptive subdivision aspects, e.g. subdivision of a picture into rectangular or non-rectangular coding blocks
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/129—Scanning of coding units, e.g. zig-zag scan of transform coefficients or flexible macroblock ordering [FMO]
Definitions
- Audio Video Coding which refers to digital audio and digital video compression standard
- AVS Audio Video Coding
- Most of the existing video coding standards are built upon the famous hybrid video coding framework i.e., using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy present in video images or sequences and using transform coding to compact the energy of the prediction errors.
- An important goal of video coding techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradations to video quality.
- a decoder may obtain a restricted area that is not adjacent to a current coding unit (CU) according to a value associated with the restricted area. Additionally, the decoder may obtain one or more motion vector (MV) candidates from a plurality of non-adjacent CUs to the current CU based on the restricted area. Furthermore, the decoder may one or more control point motion vectors (CPMVs) for the current CU based on the one or more MV candidates.
- CPMVs control point motion vectors
- an encoder may obtain a restricted area that is not adjacent to a current CU according to a value associated with the restricted area. Additionally, the encoder may obtain one or more MV candidates from a plurality of non-adjacent CUs to the current CU based on the restricted area. Furthermore, the encoder may one or more CPMVs for the current CU based on the one or more MV candidates.
- an apparatus for video decoding includes one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors, upon execution of the instructions, are configured to perform the method according to the first aspect above.
- an apparatus for video encoding includes one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors, upon execution of the instructions, are configured to perform the method according to the second aspect above.
- a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to receive a bitstream, and perform the method according to the first aspect above.
- a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the second aspect above to encode a current CU into a bitstream, and transmit the bitstream.
- FIG. 1A is a block diagram illustrating a system for encoding and decoding video blocks in accordance with some examples of the present disclosure.
- FIG. IB is a block diagram of an encoder in accordance with some examples of the present disclosure.
- FIG. 3D is a diagram illustrating block partitions in a multi-type tree structure in accordance with some examples of the present disclosure.
- FIG. 4D illustrates candidate pairs that are considered for redundancy check of spatial merge candidates in accordance with some examples of the present disclosure.
- FIG. 5 illustrates 6-parameter affine model in accordance with some examples of the present disclosure.
- FIG. 6 illustrates adjacent neighboring blocks for inherited affine merge candidates in accordance with some examples of the present disclosure.
- FIG 7 illustrates adjacent neighboring blocks for constructed affine merge candidates in accordance with some examples of the present disclosure.
- FIG. 10 illustrates perpendicular scanning of non-adjacent neighboring blocks in accordance with some examples of the present disclosure.
- FIG. 12 illustrates combined perpendicular and parallel scanning of non-adjacent neighboring blocks in accordance with some examples of the present disclosure.
- FIG. 13A illustrates neighbor blocks with the same size as the current block in accordance with some examples of the present disclosure.
- FIG. 13B illustrates neighbor blocks with a different size than the current block in accordance with some examples of the present disclosure.
- FIG. 14A illustrates an example of the bottom-left or top-right block of the bottommost or rightmost block in a previous distance is used as the bottommost or rightmost block of a current distance in accordance with some examples of the present disclosure.
- FIG. 14B illustrates an example of the left or top block of the bottommost or rightmost block in the previous distance is used as the bottommost or rightmost block of the current distance in accordance with some examples of the present disclosure.
- FIG. 15A illustrates scanning positions at bottom-left and top-right positions used for above and left non-adjacent neighboring blocks in accordance with some examples of the present disclosure.
- FIG. 15B illustrates scanning positions at bottom-right positions used for both above and left non-adjacent neighboring blocks in accordance with some examples of the present disclosure.
- FIG. 15C illustrates scanning positions at bottom-left positions used for both above and left non-adjacent neighboring blocks in accordance with some examples of the present disclosure.
- FIG. 15D illustrates scanning positions at top-right positions used for both above and left non-adjacent neighboring blocks in accordance with some examples of the present disclosure.
- FIG 16 illustrates a simplified scanning process for deriving constructed merge candidates in accordance with some examples of the present disclosure.
- FIG. 17B illustrates spatial neighbors for deriving constructed affine merge candidates in accordance with some examples of the present disclosure.
- FIG. 20 illustrates template and reference samples of a template for block with sub-block motion using the motion information of the subblocks of a current block in accordance with some examples of the present disclosure.
- FIG. 21 illustrates an example where non-adjacent spatial area is restricted to be within half CTU size on the area above and left of the current CTU in accordance with some examples of the present disclosure.
- FIG. 22A illustrates one storage method of directly saving affine motion information about an affined coded block CTU in accordance with some examples of the present disclosure.
- FIG. 23 illustrates an example of using center point to derive regular/translational motion at each 4x4 regular block in accordance with some examples of the present disclosure.
- FIG. 24 is a diagram illustrating a computing environment coupled with a user interface in accordance with some examples of the present disclosure.
- FIG. 25 illustrates an example of storing motion information of an affine-coded block at a granularity greater than the minimum affine block size in accordance with some examples of the present disclosure.
- FIG. 26 illustrates merge mode with MVD (MMVD) search points respectively in L0 reference and LI reference in accordance with some examples of the present disclosure.
- FIG. 27B illustrates an example of motion storage for non-adjacent spatial neighbors including affine neighbor CUs and non-affine neighbor CUs motion storage in line buffer in accordance with some examples of the present disclosure.
- FIG. 28 illustrates an example of projected or clipped non-adjacent neighbor positions when a scanned non-adjacent neighbor position is beyond the allowable spatial area in accordance with some examples of the present disclosure.
- FIG. 29 is a flow chart illustrating a method for video decoding in accordance with some examples of the present disclosure.
- FIG. 30 is a flow chart illustrating a method for video encoding corresponding to the method for video decoding as shown in FIG. 29 in accordance with some examples of the present disclosure.
- first,” “second,” “third,” etc. are all used as nomenclature only for references to relevant elements, e.g., devices, components, compositions, steps, etc., without implying any spatial or chronological orders, unless expressly specified otherwise.
- a “first device” and a “second device” may refer to two separately formed devices, or two parts, components, or operational states of a same device, and may be named arbitrarily.
- module may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors.
- a module may include one or more circuits with or without stored code or instructions.
- the module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to, or located adjacent to, one another.
- a method may comprise steps of: i) when or if condition X is present, function or action X’ is performed, and ii) when or if condition Y is present, function or action Y’ is performed.
- the method may be implemented with both the capability of performing function or action X’, and the capability of performing function or action Y’.
- the functions X’ and Y’ may both be performed, at different times, on multiple executions of the method.
- a unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software.
- the unit or module may include functionally related code blocks or software components, that are directly or indirectly linked together, so as to perform a particular function.
- FIG. 1A is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel in accordance with some implementations of the present disclosure.
- the system 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14.
- the source device 12 and the destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming device, or the like.
- the source device 12 and the destination device 14 are equipped with wireless communication capabilities.
- the destination device 14 may receive the encoded video data to be decoded via a link 16.
- the link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14.
- the link 16 may include a communication medium to enable the source device 12 to transmit the encoded video data directly to the destination device 14 in real time.
- the encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14.
- the communication medium may include any wireless or wired communication medium, such as a Radio Frequency (RF) spectrum or one or more physical transmission lines.
- RF Radio Frequency
- the communication medium may form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet.
- the communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from the source device 12 to the destination device 14.
- the encoded video data may be transmitted from an output interface 22 to a storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the destination device 14 via an input interface 28.
- the storage device 32 may include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Versatile Disks (DVDs), Compact Disc Read-Only Memories (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.
- the storage device 32 may correspond to a fde server or another intermediate storage device that may hold the encoded video data generated by the source device 12.
- the destination device 14 may access the stored video data from the storage device 32 via streaming or downloading.
- the fde server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14.
- Exemplary fde servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, Network Attached Storage (NAS) devices, or a local disk drive.
- FTP File Transfer Protocol
- NAS Network Attached Storage
- the destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on a file server.
- a wireless channel e.g., a Wireless Fidelity (Wi-Fi) connection
- a wired connection e.g., Digital Subscriber Line (DSL), cable modem, etc.
- the transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
- the destination device 14 may include the display device 34, which can be an integrated display device and an external display device that is configured to communicate with the destination device 14.
- the display device 34 displays the decoded video data to a user, and may include any of a variety of display devices such as a Liquid Crystal Display (LCD), a plasma display, an Organic Light Emitting Diode (OLED) display, or another type of display device.
- LCD Liquid Crystal Display
- OLED Organic Light Emitting Diode
- the intra BC unit 85 may use some of the received syntax elements, e.g., a flag, to determine that the current video block was predicted using the intra BC mode, construction information of which video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information to decode the video blocks in the current video frame.
- a flag e.g., a flag
- the video encoder 20 (or more specifically a partition unit in a prediction processing unit of the video encoder 20) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs.
- a video frame may include an integer number of CTUs ordered consecutively in a raster scan order from left to right and from top to bottom.
- Each CTU is a largest logical coding unit and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set, such that all the CTUs in a video sequence have the same size being one of 128x 128, 64x64, 32x32, and 16x 16. But it should be noted that the present application is not necessarily limited to a particular size.
- each CTU may include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to code the samples of the coding tree blocks.
- the syntax elements describe properties of different types of units of a coded block of pixels and how the video sequence can be reconstructed at the video decoder 30, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters.
- a CTU may include a single coding tree block and syntax elements used to code the samples of the coding tree block.
- a coding tree block may be an NxN block of samples.
- each leaf node of the quadtree corresponding to one CU of a respective size ranging from 32x32 to 8x8.
- each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of a frame of the same size, and syntax elements used to code the samples of the coding blocks.
- a CU may include a single coding block and syntax structures used to code the samples of the coding block.
- the video encoder 20 may use intra prediction or inter prediction to generate the predictive blocks for a PU. If the video encoder 20 uses intra prediction to generate the predictive blocks of a PU, the video encoder 20 may generate the predictive blocks of the PU based on decoded samples of the frame associated with the PU. If the video encoder 20 uses inter prediction to generate the predictive blocks of a PU, the video encoder 20 may generate the predictive blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.
- the video encoder 20 may generate a luma residual block for the CU by subtracting the CU’s predictive luma blocks from its original luma coding block such that each sample in the CU’s luma residual block indicates a difference between a luma sample in one of the CU's predictive luma blocks and a corresponding sample in the CU's original luma coding block.
- the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, such that each sample in the CU's Cb residual block indicates a difference between a Cb sample in one of the CU's predictive Cb blocks and a corresponding sample in the CU's original Cb coding block and each sample in the CU's Cr residual block may indicate a difference between a Cr sample in one of the CU's predictive Cr blocks and a corresponding sample in the CU's original Cr coding block.
- the video decoder 30 also reconstructs the coding blocks of the current CU by adding the samples of the predictive blocks for PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU After reconstructing the coding blocks for each CU of a frame, video decoder 30 may reconstruct the frame.
- the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to produce a Motion Vector Difference (MVD) for the current CU.
- MVD Motion Vector Difference
- a set of rules need to be adopted by both the video encoder 20 and the video decoder 30 for constructing a motion vector candidate list (also known as a “merge list”) for a current CU using those potential candidate motion vectors associated with spatially neighboring CUs and/or temporally co-located CUs of the current CU and then selecting one member from the motion vector candidate list as a motion vector predictor for the current CU.
- a motion vector candidate list also known as a “merge list”
- the 6-parameter affine mode has the following parameters: two parameters for translation movement in horizontal and vertical directions respectively, two parameters for zoom motion and rotation motion respectively in horizontal direction, another two parameters for zoom motion and rotation motion respectively in vertical direction.
- the 6-parameter affine motion model is coded with three CPMVs. As shown in FIG. 5, the three control points of one 6-paramter affine block are located at the top-left, top-right and bottom left comer of the block.
- the motion at topleft control point is related to translation motion
- the motion at top-right control point is related to rotation and zoom motion in horizontal direction
- the motion at bottom-left control point is related to rotation and zoom motion in vertical direction.
- the rotation and zoom motion in horizontal direction of the 6-paramter may not be same as those motion in vertical direction.
- the motion vector of each sub-block (v x , Vy) is derived using the three MVs at control points as:
- affine merge mode the CPMVs for the current block are not explicitly signaled but derived from neighboring blocks. Specifically, in this mode, motion information of spatial neighbor blocks is used to generate CPMVs for the current block.
- the affine merge mode candidate list has a limited size. For example, in the current VVC design, there may be up to five candidates.
- the encoder may evaluate and choose the best candidate index based on rate-distortion optimization algorithms. The chosen candidate index is then signaled to the decoder side.
- the affine merge candidates can be decided in three ways. In the first way, the affine merge candidates may be inherited from neighboring affine coded blocks. Tn the second way, the affine merge candidates may be constructed from translational MVs from neighboring blocks. In the third way, zero MVs are used as the affine merge candidates.
- the candidates are obtained from the neighboring blocks located at the bottom-left of the current block (e.g., scanning order is from A0 to Al as shown in FIG. 6) and from the neighboring blocks located at the topright of the current block (e g., scanning order is from B0 to B2 as shown in FIG. 6), if available.
- the candidates are the combinations of neighbor’s translational MVs, which may be generated by two steps.
- Step 1 obtain four translational MVs including MV1, MV2, MV3 and MV4 from available neighbors.
- MV1 MV from the one of the three neighboring blocks close to the top-left comer of the current block. As shown in FIG. 7, the scanning order is B2, B3 and A2.
- MV2 MV from the one of the one from the two neighboring blocks close to the top-right comer of the current block. As shown in FIG. 7, the scanning order is Bland BO.
- MV3 MV from the one of the one from the two neighboring blocks close to the bottomleft comer of the current block. As shown in FIG. 7, the scanning order is Aland AO.
- MV4 MV from the temporally collocated block of the neighboring block close to the bottom-right corner of current block. As shown in the Fig, the neighboring block is T.
- Step 2 derive combinations based on the four translational MVs from Step 1.
- Combination 1 MV1, MV2, MV3;
- Combination 2 MV1, MV2, MV4;
- Combination 3 MV1, MV3, MV4;
- Combination 4 MV2, MV3, MV4;
- Combination 6 MV1, MV3.
- Affine advanced motion vector prediction (AMVP) mode may be applied for CUs with both width and height larger than or equal to 16.
- An affine flag in CU level is signaled in the bitstream to indicate whether affine AMVP mode is used and then another flag is signaled to indicate whether 4-parameter affine or 6-parameter affine.
- the difference of the CPMVs of current CU and their CPMV predictors (CPMVPs) is signaled in the bitstream.
- the affine AVMP candidate list size is 2 and the affine AMVP candidate list is generated by using the following four types of CPMV candidate in order below:
- the checking order of inherited affine AMVP candidates is the same to the checking order of inherited affine merge candidates. The only difference is that, for AMVP candidate, only the affine CU that has the same reference picture as in current block is considered. No pruning process is applied when inserting an inherited affine motion predictor into the candidate list.
- Constructed AMVP candidate is derived from the same spatial neighbors as affine merge mode.
- the same checking order is used as done in affine merge candidate construction.
- reference picture index of the neighboring block is also checked.
- the first block in the checking order that is inter coded and has the same reference picture as in current CUs is used.
- the current CU is coded with 4-parameter affine mode, and mv 0 and mv 1 are both available, mv 0 and mv 1 are added as one candidate in the affine AMVP candidate list.
- the current CU is coded with 6-parameter affine mode, and all three CPMVs are available, they are added as one candidate in the affine AMVP candidate list. Otherwise, constructed AMVP candidate is set as unavailable.
- mv 0 , mv r and mv 2 will be added, in order, as the translational MVs to predict all control point MVs of the current CU, when available. Finally, zero MVs are used to fill the affine AMVP list if it is still not full.
- the regular inter merge candidate list is constructed by including the following five types of candidates in order:
- the size of merge list is signaled in sequence parameter set header and the maximum allowed size of merge list is 6.
- an index of best merge candidate is encoded using truncated unary binarization (TU).
- the first bin of the merge index is coded with context and bypass coding is used for other bins.
- the derivation of spatial merge candidates in VVC is same to that in HEVC except the positions of first two merge candidates are swapped. A maximum of four merge candidates are selected among candidates located in the positions depicted in FIG. 4C.
- the order of derivation is B0, A0, Bl, Al and B2.
- Position B2 is considered only when one or more than one CUs of position B0, A0, Bl, Al are not available (e.g., because it belongs to another slice or tile) or is intra coded.
- candidate at position Al is added, the addition of the remaining candidates is subject to a redundancy check which ensures that candidates with same motion information are excluded from the list so that coding efficiency is improved.
- the candidate derivation methods proposed for affine merge mode may be extended to other coding modes, such as affine AMVP mode and regular merge mode.
- the candidate derivation process for affine merge mode is extended by using not only adjacent neighboring blocks but also non-adjacent neighboring blocks.
- Detailed methods may be summarized in following aspects including affine merge candidate pruning, non-adjacent neighbor based derivation process for affine inherited merge candidates, non-adjacent neighbor based derivation process for affine constructed merge candidates, inheritance based derivation method for affine constructed merge candidates, HMVP based derivation method for affine constructed merge candidates, candidate derivation method for affine AMVP mode and regular merge mode, and motion information storage.
- top-left comer CPMV and top-right comer CPMV termed as V0 and VI
- the six parameters of a, b, c, d, e and f can be calculated as [00188]
- top-left corner CPMV and bottom-left corner CPMV termed as V0 and V2
- the six parameters of a, b, c, d, e and f can be calculated as
- Step 2 based on one or more pre-defined threshold values, similarity check is performed between the two sets of affine model parameters.
- a positive threshold value such as the value of 1
- the two candidates are considered to be similar and one of them can be pruned/removed and not put in the merge candidate list.
- the divisions or right shift operations in Step 1 may be removed to simplify the calculations in the CPMV pruning process.
- the model parameters may be converted to take the impact of the width and height into account.
- the approximated model parameters of c', d', e' may be calculated based on equation (8) below.
- the approximated model parameters of c', d' , e' and f may be calculated based on equation (9) below.
- threshold values are needed to evaluate the similarity between two candidate sets of CPMV.
- the threshold values may be defined per comparable parameter.
- Table 1 is one example in this embodiment showing threshold values defined per comparable model parameter.
- the threshold values may be defined by considering the size of the current coding block.
- Table 2 is one example in this embodiment showing threshold values defined by the size of the current coding block.
- the threshold values may be defined by considering the weight or the height of the current block.
- Table 3 and Table 4 are examples in this embodiment. Table 3 shows threshold values defined by the width of the current coding block and Table 4 shows threshold values defined by the height of the current coding block.
- the threshold values may be defined as a group of fixed values. In another embodiment, the threshold values may be defined by any combinations of above embodiments. In one example, the threshold values may be defined by considering different parameters and the weight and the height of the current block. Table 5 is one example in this embodiment showing threshold values defined by the height of the current coding block. Note that in any above proposed embodiments, the comparable parameters, if needed, may represent any parameters defined in any equations from equation (4) to equation (9).
- the benefits of using the converted affine model parameters for candidate redundancy check include that: it creates a unified similarity check process for candidates with different affine model types, e g., one merge candidate may user 6-parameter affine model with three CPMVs while another candidate may use 4-parameter affine model with two CPMVs; it considers the different impacts of each CPMV in a merge candidate when deriving the target MV at each sub-block; and it provides the similarity significance of two affine merge candidates related to the width and height of the current block.
- Non-Adjacent Neighbor Based Derivation Process for Affine Inherited Merge Candidates may be performed in three steps. Step 1 is for candidate scanning. Step 2 is for CPMV projection. Step 3 is for candidate pruning.
- Step 1 non-adjacent neighboring blocks are scanned and selected by following methods.
- non-adjacent neighboring blocks may be scanned from left area and above area of the current coding block.
- the scanning distance may be defined as the number of coding blocks from the scanning position to the left side or top side of the current coding blocks.
- FIG. 8 on either the left or above of the current coding block, multiple lines of non-adjacent neighboring blocks may be scanned.
- the distance shown in FIG. 8 represents the number of coding blocks from each candidate position to the left side or top side of the current block. For example, the area with “distance 2 (D2)” on the left side of the current block indicates that the candidate neighboring blocks located in this area are 2 blocks away from the current block. Similar indications may be applied to other scanning areas with different distances.
- the non-adj cent neighboring blocks at each distance may have the same block size as the current coding block, as shown in the FIG. 13A. As shown in FIG. 13A, the non-adjacent neighbor blocks 1301 on the left side and the non-adjacent neighbor blocks 1302 on the above side have the same size as the current block 1303. In some embodiments, the non-adjacent neighboring blocks at each distance may have a different block size as the current coding block, as shown in the FIG. 13B.
- the neighbor block 1304 is an adjacent neighbor block to the current block 1303. As shown in FIG. 13B, the non-adjacent neighbor blocks 1305 on the left side and the non-adjacent neighbor blocks 1306 on the above side have the same size as the current block 1307.
- the neighbor block 1308 is an adjacent neighbor block to the current block 1307.
- the value of the block size is adaptively changed according to the partition granularity at each different area in an image.
- the value of the block size may be predefined as a constant value, such as 4x4, 8x8 or 16x16.
- the 4x4 non-adjacent motion fields shown in FIG. 10 and FIG. 12 are examples in this case, where the motion fields may be considered as, but not limited to, special cases of sub-blocks.
- the non-adjacent coding blocks shown in FIG. 11 may have different sizes as well.
- the non-adjacent coding blocks may have the size as the current coding block, which is adaptively changed.
- the non-adjacent coding blocks may have a predefined size with a fixed value, such as 4x4, 8x8 or 16x16.
- the total size of the scanning area on either the left or above of the current coding clock may be determined by a configurable distance value.
- the maximum scanning distance on the left side and above side may use a same value or different values.
- FIG. 13 shows an example where the maximum distance on both the left side and above side shares a same value of 2.
- the maximum scanning distance value(s) may be determined by the encoder side and signaled in a bitstream Alternatively, the maximum scanning distance value(s) may be predefined as fixed value(s), such as the value of 2 or 4. When the maximum scanning distance is predefined as the value of 4, it indicates that the scanning process is terminated when the candidate list is full or all the non-adjacent neighboring blocks with at most distance 4 have been scanned, whichever comes first.
- the starting and ending neighboring blocks may be position dependent.
- the starting neighboring blocks may be the adjacent bottom -left block of the starting neighboring block of the adjacent scanning area with smaller distance.
- the starting neighboring block of the “distance 2” scanning area on the left side of the current block is the adjacent bottomleft neighboring block of the starting neighboring block of the “distance 1 (DI)” scanning area.
- DI, D2, D3 respectively indicates distance 1, distance 2, and distance 3.
- the ending neighboring blocks may be the adjacent left block of the ending neighboring block of the above scanning area with smaller distance.
- the ending neighboring block of the “distance 2” scanning area on the left side of the current block is the adjacent left neighboring block of the ending neighboring block of the “distance 1” scanning area above the current block.
- the starting neighboring blocks may be the adjacent top-right block of the starting neighboring block of the adjacent scanning area with smaller distance.
- the ending neighboring blocks may be the adjacent top-left block of the ending neighboring block of the adjacent scanning area with smaller distance.
- the left area may be scanned first, and then followed by scanning the above areas.
- three lines of non-adjacent areas e g., from distance 1 (DI) to distance 3 (D3)
- DI distance 1
- D3 distance 3
- a scanning order may be defined.
- the scanning may be started from the bottom neighboring block to the top neighboring block.
- the scanning may be started from the right block to the left block.
- X is set to be 1, which means the scanning is terminated for each distance if the first qualified candidate is found and the scanning process is restarted from a different distance of the same area or the same or different distance of a different area.
- the value of X may be set as the same value or different values for different distances. If the maximum number of qualified candidates are found from all allowable distances (e.g., regulated by a maximum distance) of an area, the scanning process for one area is completely terminated.
- the X may be defined for an area.
- X is set to be 3, which means the scanning is terminated for the whole area (e.g., left or above area of the current block) if the first 3 qualified candidates are found and the scanning process is restarted from the same or different distance of another area.
- the value of X may be set as the same value or different values for different areas. If the maximum number of qualified candidates are found from all areas, the whole scanning process is completely terminated.
- the values of X may be defined for both distance and areas. For example, for each area (e.g., left or above area of the current block), X is set to 3, and for each distance, X is set to 1. The values of X may be set as the same value or different values for different areas and distances.
- the scanning process may be performed continuously. For example, the scanning performed in a specific area at a specific distance may be stopped at the instance when all covered neighboring blocks are scanned and no more qualified candidates are identified or the maximum allowable number of candidates is reached.
- each candidate non-adjacent neighboring block is determined and scanned by following the above proposed scanning methods.
- each candidate non-adjacent neighboring block may be indicated or located by a specific scanning position. Once a specific scanning area and distance are decided by following above proposed methods, the scanning positions may be determined accordingly based on following methods.
- bottom-left and top-right positions are used for above and left non- adjacent neighboring blocks respectively, as shown in FIG. 15 A.
- bottom-right positions are used for both above and left non- adjacent neighboring blocks, as shown in FIG. 15B.
- bottom-left positions are used for both above and left nonadj acent neighboring blocks, as shown in FIG. 15C.
- top-right positions are used for both above and left non-adj acent neighboring blocks, as shown in FIG. 15D.
- each non-adjacent neighboring block is assumed to have the same block size as the current block. Without loss of generality, this illustration may be easily extended to non-adjacent neighboring blocks with different block sizes.
- Step 2 the same process of CPMV projection as used in the current AVS and VVC standards may be utilized.
- the current block is assumed to share the same affine model with the selected neighboring block, then two or three comer pixel’ s coordinates (e.g., if the current block uses 4-prameter model, two coordinates (top-left pixel/sample location and top-right pixel/sample location) are used; if the current block uses 6- prameter model, three coordinates (top-left pixel/sample location, top-right pixel/sample location and bottom-left pixel/sample location) are used) are plugged into equation (1) or (2), which depends on whether the neighboring block is coded with a 4-parameter or 6-parameter affine model, to generate two or three CPMVs.
- any qualified candidate that is identified in Step 1 and converted in Step 2 may go through a similarity check against all existing candidates that are already in the merge candidate list. The details of similarity check are already described in the section of “Affine Merge Candidate Pruning” above. If the newly qualified candidate is found to be similar with any existing candidate in the candidate list, this newly qualified candidate is removed/pruned.
- one neighboring block is identified at one time, where this single neighboring block needs to be coded in affine mode and may contain two or three CPMVs.
- two or three neighboring blocks may be identified at one time, where each identified neighboring block does not need to be coded in affine mode and only one translational MV is retrieved from this block.
- FIG. 9 presents an example where constructed affine merge candidates may be derived by using non-adjacent neighboring block.
- A, B and C are the geographical positions of three non-adjacent neighboring blocks.
- a virtual coding block is formed by using the position of A as the top-left comer, the position of B as the top-right comer, and the position of C as the bottom -left comer.
- the MVs at the positions of A', B' and C’ may be derived by following the equation (3), where the model parameters (a, b, c, d, e, ) may be calculated by the translational MV at the positions of A, B and C.
- the MVs at positions of A’, B’ and C’ may be used as the three CPMVs for the current block, and the existing process (the one used in the AVS and VVC standards) of generating constructed affine merge candidates may be used.
- non-adjacent neighbor based derivation process may be performed in five steps.
- the non-adjacent neighbor based derivation process may be performed in the five steps in an apparatus such as an encoder or a decoder.
- Step 1 is for candidate scanning.
- Step 2 is for affine model determination.
- Step 3 is for CPMV projection.
- Step 4 is for candidate generation.
- Step 5 is for candidate pruning.
- non-adjacent neighboring blocks may be scanned and selected by following methods.
- the scanning process is only performed for two non-adjacent neighboring blocks.
- the third non-adjacent neighboring block may be dependent on the horizontal and vertical positions of the first and second non- adjacent neighboring blocks.
- the scanning process is only performed for the positions of B and C.
- the position of A may be uniquely determined by the horizontal position of C and the vertical position of B.
- the position of A may need to be at least valid.
- the validity of position A may be defined as whether the motion information at the position A is available or not.
- the coding block located at the position A may need to be coded in inter-modes such that the motion information is available to form a virtual coding block.
- the scanning area and distance may be defined according to a specific scanning direction.
- the qualified candidate does not need to be affine coded since only translational MV is needed.
- the scanning process may be terminated when the first X qualified candidates are identified, where X is a positive value.
- X is a positive value.
- the scanning process in Step 1 may be only performed for identifying the non-adjacent neighboring blocks located at comers B and C, while the coordinate of A may be precisely determined by taking the horizontal coordinate of C and the vertical coordinate of B. In this way, the formed virtual coding block is restricted to be rectangle.
- the horizontal coordinate or vertical coordinate of C may be defined as the horizontal coordinate or vertical coordinate of the top-left point of the current block respectively.
- the comer B and/or corner C when the comer B and/or corner C is firstly determined from the scanning process in Step 1, the non-adjacent neighboring blocks located at comer B and/or C may be identified accordingly. Secondly, the position(s) of the corner B and/or C may be reset to pivot point within the corresponding non-adjacent neighboring blocks, such as the mass center of each non-adjacent neighboring block. For example, the mass center may be defined as the geometric center of each neighboring block.
- the process may be performed jointly or independently.
- independent scanning the previously proposed scanning methods may be applied separately on the comers B and C.
- joint scanning there may be different methods as follows.
- pairwise scanning may be performed.
- the candidate positions for corners B and C are simultaneously advanced.
- FIG. 17B it is to take FIG. 17B as an example.
- the scanning of comer B is started from the first non-adjacent neighboring block located on the above side of the current block, in a bottom-to-up direction.
- the scanning of corner C is started from the first non-adjacent neighboring block located on the left side of the current block, in a right-to-left direction. Therefore, in the example shown in FIG.
- pairwise scanning may be defined as that the candidate positions of B and C are both advanced with one unit of step size, where one unit of step size is defined as the height of the current coding block for corner B and defined as the width of the current coding block for comer C.
- alternative scanning may be performed. Tn one example of alternative scanning, the candidate positions for comers B and C are alternatively advanced. At one step, only the position of B or C may be advanced, while the position of C or B is not changed. In one example, the position of comer B may be progressively increased from the first non-adjacent neighboring block to the distance at the maximum number of non-adjacent neighboring blocks, while the position of corner C remains at the first non-adjacent neighboring block. In the next round, the position of the corner C moves to the second non-adjacent neighboring block, and the position of the comer B is traversed from the first to the maximum value again. The rounds are continued until all combinations are traversed.
- the methods of defining scanning area and distance, scanning order, and scanning termination proposed for deriving inherited merge candidates may completely or partially reused for deriving constructed merge candidates
- the same methods defined for inherited merge candidate scanning which include but no limited to scanning area and distance, scanning order and scanning termination, may be completely reused for constructed merge candidate scanning.
- the same methods defined for inherited merge candidate scanning may be partly reused for constructed merge candidate scanning.
- FIG. 16 shows an example in this case.
- the block size of each non-adjacent neighboring blocks is same as the current block, which is similarly defined as inherited candidate scanning, but the whole process is a simplified version since the scanning at each distance is limited to be only one block.
- FIGS. 17A-17B represent another example in this case. In FIGS. 17A-17B, both non-adjacent inherited merge candidates and non-adjacent constructed merge candidates are defined with the same block size as the current coding block, while the scanning order, scanning area, and scanning termination conditions may be defined differently.
- the maximum distance for left side non-adj acent neighbors is 4 coding blocks, while the maximum distance for above side non-adjacent neighbors is 5 coding blocks. Also, at each distance, the scanning direction is bottom-up for left side and right-to-left for above side. In FIG. 17B, the maximum distance of non-adjacent neighbors is 4 for both left side and above side. In addition, the scanning at a specific distance is unavailable because there is only one block at each distance. In FIG. 17A, the scanning operations within each distance may be terminated if M qualified candidates are identified.
- the value of M may be a predefined fixed value such as the value of 1 or any other positive integer, or a signaled value decided by the encoder, or a configurable value at the encoder or the decoder. In one example, the value of M may be the same as the merge candidate list size.
- the scanning operations at different distances may be terminated if N qualified candidates are identified.
- the value of N may be a predefined fixed value such as the value of 1 or any other positive integer, or a signaled value decided by the encoder, or a configurable value at the encoder or the decoder.
- the value of N may be the same as the merge candidate list size.
- the value of N may be the same as the value ofM.
- the non-adjacent spatial neighbors with closer distance to the current block may be prioritized, which indicates that non-adjacent spatial neighbors with distance i is scanned or checked before the neighbors with distance i+1, where i may be a nonnegative integer representing a specific distance.
- the positions of one left and above non-adjacent spatial neighbors are firstly determined independently. After that, the location of the top-left neighbor can be determined accordingly which can enclose a rectangular virtual block together with the left and above non-adjacent neighbors. Then, as shown in the FIG. 9, the motion information of the three non-adjacent neighbors is used to form the CPMVs at the top-left (A), top-right (B) and bottom-left (C) of the virtual block, which is finally projected to the current CU to generate the corresponding constructed candidates.
- Step 2 the translational MVs at the positions of the selected candidates after step 1 are evaluated and an appropriate affine model may be determined.
- FIG. 9 is used as an example again.
- the scanning process may be terminated before enough number of candidates are identified. For example, the motion information of the motion field at one or more of the selected candidates after Step 1 may be unavailable.
- the corresponding virtual coding block represents a 6-parameter affine model. If the motion information of one of the three candidates is unbailable, the corresponding virtual coding block represents a 4-parameter affine model. If the motion information of more than one of the three candidates is unbailable, the corresponding virtual coding block may be unable to represent a valid affine model.
- the virtual block may be set to be invalid and unable to represent a valid model, then Step 3 and Step 4 may be skipped for the current iteration.
- the virtual block may represent a valid 4-parameter affine model.
- Step 3 if the virtual coding block is able to represent a valid affine model, the same projection process used for inherited merge candidate may be used.
- the same projection process used for inherited merge candidate may be used.
- a 4-parameter model represented by the virtual coding block from Step 2 is projected to a 4-parameter model for the current block
- a 6-parameter model represented by the virtual coding block from Step 2 is projected to a 6-parameter model for the current block.
- the affine model represented by the virtual coding block from Step 2 is always projected to a 4-parameter model or a 6-parameter model for the current block.
- the type of the projected 4-parameter affine model is the same type of the 4-parameter affine model represented by the virtual coding block.
- the affine model represented by the virtual coding block from Step 2 is type A or B 4-parameter affine model
- the projected affine model for the current block is also type A or B respectively.
- the 4-parameter affine model represented by the virtual coding block from Step 2 is always projected to the same type of 4-parameter model for the current block.
- the type A or B of 4-parameter affine model represented by the virtual coding block is always projected to the type A 4-parameter affine model.
- Step 4 based on the projected CPMVs after Step 3, in one example, the same candidate generation process used in the current VVC or AVS standards may be used.
- the temporal motion vectors used in the candidate generation process for the current VVC or AVS standards may be not used for the non-adjacent neighboring blocks based derivation method. When the temporal motion vectors are not used, it indicates that the generated combinations do not contain any temporal motion vectors.
- Step 5 any newly generated candidate after Step 4 may go through a similarity check against all existing candidates that are already in the merge candidate list. The details of similarity check are already described in the section of “Affine merge candidate pruning.” If the newly generated candidate is found to be similar with any existing candidate in the candidate list, this newly generated candidate is removed or pruned.
- a virtual coding block is formed by determining three corner points A, B and C, and then the translational MVs of the 4x4 blocks located at the three corners are used to represent an affine model for the virtual coding block.
- the affine model of the virtual coding block is projected to the current coding block. This whole process may be used to derive the first type of affine candidates constructed from non-adjacent spatial neighbors (e.g., the sub-blocks located by the three comer points A, B and C are non-adjacent spatial neighbors).
- this method may be applied to an affine mode, such as affine merge mode and affine AMVP mode, and this method may be also applied to regular mode, such as regular merge mode and regular AMVP mode, because the projected affine model can be used to derive a translational MV based on a specific position (e.g., the center position) inside of a prediction block or a coding block.
- an affine mode such as affine merge mode and affine AMVP mode
- regular mode such as regular merge mode and regular AMVP mode
- the combination of inheritance and construction may be realized by separating the affine model parameters into different groups, where one group of affine parameters are inherited from one neighboring block, while other groups of affine parameters are inherited from other neighboring blocks.
- the parameters of one affine model may be constructed from two groups.
- an affine model may contain 6 parameters, including a, b, c, d , e and f .
- the translational parameters ⁇ a, b ⁇ may represent one group, while the non- translational parameters ⁇ c, d, e, f ⁇ may represent another group.
- the two groups of parameters may be independently inherited from two different neighboring blocks in the first step and then concatenated/ constructed to be a complete affine model in the second step.
- the group with non-translational parameters has to be inherited from one affine coded neighboring block, while the group with translational parameters may be from any inter-coded neighboring block, which may or may not be coded in affine mode.
- the affine coded neighboring block may be selected from adjacent affine neighboring blocks or non-adjacent affine neighboring blocks based on previously proposed scanning methods for affine inherited candidates, such as the methods shown in FIG.
- the affine coded neighboring block may be not physically existed, but virtually constructed from regular inter-coded neighboring blocks, such as the methods shown in FIG. 17B, that is the scanning method/rule including the scanning area and distance, scanning order, and scanning termination used in the Section of “Non-Adjacent Neighbor Based Derivation Process for Affine Constructed Merge Candidates.”
- the neighboring blocks associated with each group may be determined in different ways.
- the neighboring blocks for different groups of parameters may be all from non-adjacent neighboring/neighbor areas, while the scanning methods may be similarly designed as the previously proposed methods for non-adjacent neighbor based derivation process.
- the neighboring blocks for different groups of parameters may be all from adjacent neighboring/neighbor areas, while the scanning methods may be the same as the current VVC or AVS video standards.
- the neighboring blocks for different groups of parameters may be partly from adjacent areas and partly from non- adjacent neighboring/neighbor areas.
- the scanning process may be differently performed from the non-adjacent neighborbased derivation process for affine inherited candidates.
- the scanning area, distance and order may be similarly defined, but the scanning termination rule may be differently specified.
- the non-adjacent neighboring blocks may be exhaustively scanned within a defined maximum distance at each area. In this case, all non- adjacent neighboring blocks within a distance may be scanned by following a scanning order. In some embodiments, the scanning area may be different.
- the right bottom adjacent and non-adjacent area of the current coding block may be scanned to determine neighbors for generating translational or/and non- translational parameters.
- the neighbors scanned at the right bottom area may be used to find collocated temporal neighbors, instead of spatial neighbors.
- One scanning criteria may be conditionally based on whether the right-bottom collocated temporal neighbor(s) is/are already used for generating affine constructed neighbors. If used already, the scanning is not performed, otherwise the scanning is performed. Alternatively, if used already, which means the right-bottom collocated temporal neighbor(s) is/are available, the scanning is performed, otherwise the scanning is not performed.
- the associated neighboring block or blocks for each group may be checked whether to use the same reference picture for at least one direction or both directions.
- the associated neighboring block or blocks for each group may be checked whether use the same precision/resolution for motion vectors.
- the first X associated neighboring block(s) for each group may be used.
- the value of X may be defined as the same or different values for different groups of parameters.
- the first 1 or 2 neighboring blocks containing non-translational affine parameters may be used, while the first 3 or 4 neighboring blocks containing translational affine parameters may be used.
- the second is construction formula.
- the CPMVs of the new candidates may be derived in equation below: where (x, y) is a comer position within the current coding block (e g., (0, 0) for top-left comer CPMV, (width, 0) for top-right corner CPMV), ⁇ c, d, e, f ⁇ is one group of parameters from one neighboring block, ⁇ a, b ⁇ is another group of parameters from another neighboring block.
- the CPMVs of the new candidates may be derived in below equation: where the (Aw, Ah) is the distance between the top-left corner of the current coding block and the top-left corner of one of the associated neighboring block(s) for one group of parameters, such as the associated neighboring block of the group of ⁇ a, b ⁇ .
- the definitions of the other parameters in this equation are the same as the example above.
- the parameters may be grouped in another way: (a, b, c, d, e,f) are formed as one group, while the (Aw, Ah) are formed as another group. And the two groups of parameters are from two different neighboring blocks.
- the value of (Aw, Ah) may be predefined as fixed values such as (0, 0) or at any constant values, which is not dependent on the distance between a neighboring block and the current block.
- FIG. 18 shows an example of inheritance based derivation method for deriving affine constructed candidates.
- the encoder or the decoder may perform scanning of adjacent and non-adjacent neighboring blocks for each group.
- the encoder or the decoder may perform scanning of adjacent and non-adjacent neighboring blocks for each group.
- two groups are defined, where neighbor 1 is coded in affine mode and provides non- translational affine parameters, while neighbor 2 provides translational affine parameters.
- Neighbor 1 may be obtained according to the process in the Section of “Non- Adjacent Neighbor Based Derivation Process for Affine Inherited Merge Candidates” as shown in FIGS.
- neighbor 1 may be an adjacent or non-adjacent neighbor block of the current block.
- neighbor 2 may be obtained according to the process as shown in FIGS. 16 and 17B.
- the neighbor 1, which is coded in the affine mode may be scanned from adjacent or/and non-adjacent areas, by following above proposed scanning methods.
- the neighbor 2, which is coded in the affine or a non-affine mode may be also scanned from adjacent or non-adjacent areas.
- the neighbor 2 may be from one of the scanned adjacent or non-adjacent areas if the motion information is not already used for deriving some affine merge or AMVP candidates, or from right-bottom positions of the current block if a collocated TMVP candidate at this position is available or/and already used for deriving some affine merge or AMVP candidates.
- a small coordinate offset e.g., +1 or +2 or -1 or -2 for vertical or/and horizontal coordinates
- Step 2 with the parameters and positions decided in Step 1, a specific affine model may be defined, which can derive different CPMVs according to the coordinate (x, y) of a CPMV.
- the non-translational parameters ⁇ c, d, e, f ⁇ may be obtained based on neighbor 1 obtained in Stepl
- the translational parameters ⁇ a, b ⁇ may be obtained based on neighbor 2 obtained in Step 1.
- the distance parameters d w, J h may thus obtained based on the position of the current block (x lt y t ) and the position of neighbor 2 (x 2 ,y 2 )
- the distance parameters Aw, Ah may respectively indicate a horizontal distance and a vertical distance between the current block and neighbor 1 or neighbor 2.
- the distance parameters Aw, Ah may respectively indicate the horizontal distance ⁇ x 1 — x 2 ) between the current block and neighbor 2 and the vertical distance (y t — y 2 ) between the current block and neighbor 2.
- Step 3 two or three CPMVs are derived for the current coding block, which can be constructed to form a new affine candidate
- other prediction information may be further constructed.
- the prediction direction (e.g., bi or uni -predicted) and indexes of reference pictures may be the same as the associated neighboring blocks if neighboring blocks are checked to have the same directions and/or reference pictures.
- the prediction information is determined by reusing the minimum overlapped information among the associated neighboring blocks from different groups. For example, if only the reference index of one direction from one neighboring block is the same as the reference index of the same direction of the other neighboring block, the prediction direction of the new candidate is determined as uni -prediction, and the same reference index and direction are reused.
- an affine model may be constructed by combining model parameters from different inheritances.
- the translational model parameters may be inherited from translational blocks (e g., from adjacent or/and non-adjacent spatial neighboring 4x4 blocks), while the non-translational model parameters may be inherited from affine coded blocks (e.g., from adjacent or/and non-adjacent spatial neighboring affine coded blocks).
- the non-translational model parameters may be inherited from historically coded affine blocks instead of explicitly scanned non-adjacent spatial neighboring affine coded blocks, while the historically coded affine blocks may be adjacent or nob-adjacent spatial neighbors.
- This whole process may be used to derive the second type of affine candidates constructed from non- adjacent spatial neighbors (e.g., the non-translational model parameters may be inherited from non-adjacent spatial neighbors).
- this method may be applied to an affine mode, such as affine merge mode and affine AMVP mode, and this method may be also applied to regular mode, such as regular merge mode and regular AMVP mode, because the generated affine model can be used to derive a translational MV based on a specific position (e.g., the center position) inside of a prediction block or coding block.
- the HMVP merge mode is already adopted in the current VVC and AVS, where the translational motion information from neighboring blocks are already stored in a history table, as described in the introduction section.
- the scanning process may be replaced by searching the HMVP table.
- the translational motion information may be obtained from HMVP table, instead of the scanning method as shown in the FIG. 17B and FIG. 18.
- the position information, width, height and reference information are also needed, which may be accessible if the current HMVP table can be modified. Therefore, it is proposed to extend the HMVP table to store additional information in addition to the motion information of each history neighbor.
- the additional information may include positions of an affine or non-affine neighboring blocks, or affine motion information such as CPMVs or equivalent regular motion derived from CPMVs (e.g.., this regular motion may be from the internal sub-blocks of an affine coded neighboring block) reference index, etc.
- affine motion information such as CPMVs or equivalent regular motion derived from CPMVs (e.g.., this regular motion may be from the internal sub-blocks of an affine coded neighboring block) reference index, etc.
- the above provided non-adjacent neighbor based derivation process and inheritance based derivation process for affine mode may be replaced by modifying the existing HMVP table.
- the translational motion of the derived affine model may be directly obtained from the existing HMVP table, where the translational motion is previously saved when neighboring blocks are previously coded at regular inter mode.
- the translational motion of the derived affine model may be still obtained from the existing HMVP, where the translational motion is previously saved when neighboring blocks are previously coded at affine mode.
- the non-translational motion of the derived affine model is obtained from the existing HMVP table, where the non-translational motion is previously saved when neighboring blocks are previously code at affine mode.
- the reused non- translational motion may not be directly saved while original CPMVs of the previously coded affine neighboring blocks are saved. Tn this case, the existing HMVP table is updated with not only original CPMVs but also the position and size information of the previously coded affine neighboring blocks.
- the above proposed non-adjacent neighbor based derivation process and inheritance based derivation process for affine mode may be replaced by creating one or more new HMVP tables.
- an affine candidate list is also needed for deriving CPMV predictors.
- all the above proposed derivation methods may be similarly applied to affine AMVP mode.
- the selected neighboring blocks must have the same reference picture index as the current coding block.
- a candidate list is also constructed, but with only translational candidate MVs, not CPMVs.
- all the above proposed derivation methods can still be applied by adding an additional derivation step.
- this additional derivation step it is to derive a translation MV for the current block, which may be realized by selecting a specific pivot position (x, y) within the current block and then follow the same equation (3).
- the three corner positions of the block are used as the pivot position (x, y) in equation (3)
- the center position of the block may be used as the pivot position (x, y) in equation (3).
- the newly derived candidates may be inserted into the affine AMVP candidate list by following the order as below:
- the newly derived candidates may be inserted into the affine AMVP candidate list by following the order as below:
- the newly derived candidates may be inserted into the affine AMVP candidate list by following the order as below:
- the newly derived candidates may be inserted into the affine AMVP candidate list by following the order as below:
- the newly derived candidates may be inserted into the affine AMVP candidate list by following the order as below:
- the newly derived candidates may be inserted into the affine AMVP candidate list by following the order as below:
- the candidates constructed from non-adjacent spatial neighbors may be referred to as the first type or/and the second type of candidates constructed from non-adjacent spatial neighbors.
- the newly derived candidates may be inserted into the regular merge candidate list by following the order as below:
- the non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order below: 1. Subblock-based Temporal Motion Vector Prediction (SbTMVP) candidate, if available; 2. Inherited from adjacent neighbors; 3. Inherited from non-adjacent neighbors; 4. Constructed from adjacent neighbors; 5. Constructed from non-adjacent neighbors; 6. Zero MVs.
- SBTMVP Temporal Motion Vector Prediction
- the non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order below: 1. SbTMVP candidate, if available; 2. Inherited from adjacent neighbors; 3. Constructed from adjacent neighbors; 4. Inherited from non-adjacent neighbors; 5. Constructed from non-adjacent neighbors; 6. Zero MVs.
- the non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order below: 1. SbTMVP candidate, if available; 2. Inherited from adjacent neighbors; 3. Constructed from adjacent neighbors; 4. One set of zero MVs; 5. Inherited from non-adjacent neighbors; 6. Constructed from non-adjacent neighbors; 7. Remaining zero MVs, if the list is still not full.
- the non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order below: 1. SbTMVP candidate, if available; 2. Inherited from adjacent neighbors; 3. Inherited from non-adjacent neighbors with distance smaller than X; 4. Constructed from adjacent neighbors; 5. Constructed from non-adjacent neighbors; 6. Constructed from inherited translational and non-translational neighbors; 7. Zero MVs, if the list is still not full.
- the non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order below: 1. SbTMVP candidate, if available; 2. Inherited from adjacent neighbors; 3. Inherited from non-adjacent neighbors; 4. The first candidate constructed from adjacent neighbors; 5. The first X candidates constructed from inherited translational and non-translational neighbors; 6. Constructed from non-adjacent neighbors; 7. Other Y candidates constructed from inherited translational and non-translational neighbors; 8. Zero MVs, if the list is still not full.
- the value of X may be the same as the value of Y.
- the value of X may be different from the value of Y.
- the non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order below: 1. SbTMVP candidate, if available; 2. Inherited from adjacent neighbors; 3. Inherited from non-adjacent neighbors with distance smaller than X; 4. Constructed from adjacent neighbors; 5. Constructed from non-adjacent neighbors with distance smaller than Y; 6. Inherited from non-adjacent neighbors with distance bigger than X; 7. Constructed from non-adjacent neighbors with distance bigger than Y; 8. Zero MVs.
- the value X and Y may be a predefined fixed value such as the value of 2, or a signaled value decided by the encoder, or a configurable value at the encoder or the decoder.
- the value of X may be the same as the value of Y.
- the value of N may be different from the value of M.
- a new candidate is derived by using the inheritance based derivation method which constructs CPMVs by combining affine motion and translational MV
- the placement of this new candidate may be dependent on the placement the other constructed candidates.
- the reordering of the affine merge candidate list may follow the order as below:
- the reordering of the affine merge candidate list may follow the order below:
- the reordering of the affine merge candidates may be partially or completely interleaved among different categories of candidates (e.g., interleaving may indicate that the candidates from the same category may not be adjacently placed in the candidates list).
- interleaving may indicate that the candidates from the same category may not be adjacently placed in the candidates list.
- there may be seven categories of affine merge candidates placed in the affine merge candidate list:
- the specific order discussed above may be applied in any candidate list including an affine AMVP candidate list, a regular merge candidate list, and an affine merge candidate list.
- the order of the candidates may remain the same as the above insertion order.
- An adaptive reordering method may be applied to reorder the candidates afterwards; the adaptive reordering may be template based methods (ARMC) or non-template based method such as bilateral matched based methods.
- the order of the candidates may be reordered in a specific pattern.
- the specific pattern may be applied in any candidate list including an affine AMVP candidate list, a regular merge candidate list, and an affine merge candidate list.
- the reordering pattern may depend on the number of available candidates for each category.
- the reordering pattern may be defined as below:
- the first X inherited candidates from non-adjacent neighbors (e.g., X may be a prefixed number such as 1 or a signaled number);
- the first Y constructed candidates from the second type of constructed candidates from non-adjacent neighbors (e.g., Y may be similarly defined as X);
- the first Z constructed candidates from the first type of constructed candidates from non-adjacent neighbors (e.g., Z may be similarly defined as X);
- the remaining inherited candidates from non-adjacent neighbors (e.g., X may be a prefixed number such as 1 or a signaled number);
- the remaining constructed candidates from the second type of constructed candidates from non-adjacent neighbors (e.g., Y may be similarly defined as X);
- the reordering pattern may be an interleaved method which may merge different candidates from different categories.
- the interleaved pattern may be defined as below:
- the value of (Xi, Yi, Zi, Ki) may be a prefixed number such as 1 or a signaled number. If the number of available candidates of one category is smaller than other categories, the position of the candidates for this category is skipped and the remaining available candidates of other categories would take over this position.
- the reordering pattern may be a combined version which considers both availability and interleaving method.
- the combined pattern may be defined as below:
- the reordering methods may be selected based on the types of the video frames/slices. For example, for low-delay pictures or slices, all the candidates of the first type of constructed candidates from non-adjacent neighbors may be placed after all the constructed candidates from adjacent neighbors. While for non-low-delay pictures or slices, the first KI candidates of the first type of constructed candidates from non-adjacent neighbors may be placed after the first K2 constructed candidates from adjacent neighbors, and the remaining candidates of the first type of constructed candidates from non-adjacent neighbors may be placed after the remaining constructed candidates from adjacent neighbors.
- one or more candidates may be derived for an existed affine merge candidate list, or an affine AMVP candidate list, or a regular merge candidate list, where the size of the corresponding list may be statically (e.g., configurable size) or adaptively (e.g., dynamically changed according to availability at encoder and then signaled to decoder) adjusted.
- the new candidates are firstly derived as affine candidates, and then converted to translational motion vectors by using a pivot position (e.g., center sample or pixel position) within a coding block and associated affine models before insert into the regular merge candidate list.
- an adaptive reordering method such as ARMC may be applied to one or more of the above candidate lists after the candidate lists are updated or constructed by adding some new candidates which are derived by above proposed candidate derivation methods.
- a temporal candidate list may be created first, where the temporal candidate list may have a larger size than the existed candidate list (e.g., affine merge candidate list, affine AMVP candidate list, regular merge candidate list).
- an adaptive reordering method such as ARMC may be applied to reorder the temporal candidate list.
- the first N candidates of the temporal candidate list are inserted to the existed candidate lists.
- the value of N may be a fixed or configurable value. Tn one example, the value of N may be the same as the size of the existed candidate list, where the selected N candidates from of the temporal candidate list are located.
- a cost function such as the sum of absolute differences (SAD) between samples of a template of the current block and their corresponding reference samples may be used.
- the reference samples of the template may be located by the same motion information of the current block.
- an interpolation fdtering process may be used to generate prediction samples of the template. Since the generated prediction samples are just used to comparing the motion accuracy between different candidates, not for final block reconstructions, the prediction accuracy of the template samples may be relaxed by using an interpolation filter with smaller tap.
- a 2-tap or 4-tap any other shorter length (e.g., 6-tap, 8-tap) interpolation filter may be used to generate prediction samples for the selected template of the current block. Or even the nearest integer samples (completely skip the interpolation filtering process) may be used as the prediction samples of the template.
- An interpolation filter with smaller tap may be similarly used when a template matching method is used to adaptively reorder the candidates in other candidate list such as regular merge candidate list or affine AMVP candidate list.
- a cost function such as the SAD between samples of a template of the current block and their corresponding reference samples may be used.
- the corresponding reference samples may be located at integer positions or fractional positions. When fractional positions are located, a certain level of prediction accuracy may be achieved by performing an interpolation filter process. Due to the limited prediction accuracy, the calculated matching costs for different candidates may contain noise level differences. To reduce the impact of the noise level cost difference, the calculated matching costs may be adjusted by removing a few bits of the least significance bits before candidate sorting process.
- a candidate list may be padded with zero MVs at the end of each list, if not enough candidates could be derived by using different derivation methods.
- the candidate cost may be only calculated for the first zero MV, while the remaining zero MVs may be statically assigned with an arbitrarily large cost value, such that these repeated zero MVs are placed at the end of the corresponding candidate list.
- all zero MVs may be statically assigned with an arbitrarily large cost value, such that all zero MVs are placed at the end of the corresponding candidate list.
- an early termination method may be applied for a reordering method to reduce complexity at the decoder side.
- a candidate list when a candidate list is constructed, different types of candidates may be derived and inserted into the list. If one candidate or one type of candidates is not participated in the reorder process, but selected and signaled to the decoder, the reordering process, which is applied to other candidates, may be early terminated.
- the SbTMVP candidate in the case of applying ARMC for the affine merge candidate list, the SbTMVP candidate may be excluded from the reordering process. In this case, if the signaled merge index value for an affine coded block indicates a SbTMVP candidate at the decoder side, the ARMC process may be skipped or early terminated for this affine block.
- both the derivation process and the reorder process for this specific candidate or this specific type of candidates may be skipped.
- the skipped derivation process and reordering process are only applied to the specific candidate or the specific type of candidates, while the remaining candidates or types of candidates are still performed, where the derivation process is skipped indicates that the related operations of deriving the specific candidate or this specific type of candidates are skipped, but the predefined list position (e.g., according to a predefined insertion order) of the specific candidate or this specific type of candidates may be still kept, just the candidate content such as the motion information may be invalid due to skipped derivation process.
- the cost calculation of this specific candidate or this specific type of candidates may be skipped and the list position of this specific candidate or this specific type of candidates may be not changed after reordering other candidates.
- the selected non-adjacent spatial neighbors may be affine coded blocks or non-affine coded blocks (e.g., regular inter AMVP or merge coded blocks).
- the motion information may include translational MVs and corresponding reference index at each direction.
- the motion information may include CPMVs and corresponding reference index at each direction, and also the positions and the sizes of the affine coded blocks.
- the motion information of these blocks may need to be saved in a memory once these blocks have been coded.
- the non-adjacent spatial neighbors may be restricted to a certain area.
- the allowed non-adjacent area for scanning non-adjacent spatial neighboring blocks may be restricted to a limited area size.
- the restricted area may be applied to affine or non- affine spatial neighboring blocks.
- the size of the allowed non-adjacent area may be defined according to the size of current CTU, e.g., integer (e.g., 1 or 2 or other integer) or fractional number (e.g., 0.5 or 0.25 or other fractional number) of current CTU size.
- the size of the allowed non-adjacent area may be defined according to a fixed number of pixels or samples, e.g., 128 samples on the above of the current CTU or/and on the left of the current CTU.
- the size (e.g., according to the CTU size or number of samples) may be a prefixed value or a signaled value determined at the encoder and carried in the bit-stream.
- the size of the restricted area may be separately defined for top and left non-adjacent neighboring blocks.
- the above non-adjacent neighboring blocks may be restricted to be within the current CTU, or outside of the current CTU but within at most fixed number samples/pixels away from the top of the current CTU such that no additional line buffer is needed for saving the motion information of above non-adjacent neighboring blocks.
- the fixed number may be defined as 8, if 8 sample rows of neighboring/neighbor area away from the current CTU top is already covered by the existing line buffer.
- the left non-adjacent neighboring blocks may be restricted to be within the current CTU, or outside of the current CTU but within a predefined or a signaled number of samples/pixels away from the left boundary of the current CTU.
- the allowed non-adjacent area (for either non-adj acent affine neighbors or non-affine neighbors) above the current CU may have large memory cost if the allowed non-adjacent area is beyond the current CTU.
- the actual memory cost is proportionally increased with the picture width and the maximum allowable scanning distance in the vertical direction.
- the height of the above non-adjacent area outside of the current CTU may be limited to a value of h, as shown in FIG. 27A.
- this value of h may be configurable or signaled to decoder.
- affine motion and non-affine motion are stored in a separate buffer, for the example shown in FIG. 27B, there may be different methods to save the motion in the line buffer as follows.
- the line buffer used to store affine motion may indicate that the buffer area where the CU B is located is set to be invalid since CU B is not affine CU.
- the line buffer used to store affine motion may indicate that the buffer area where the CU B is located is set to be valid and the affine motion is copied from CU A, since CU A is CU B’s adjacent affine neighbor.
- the height value h and the width value w may be set to be multiples of 4 for easier implementation.
- the value h and w may be set to be the minimum value (e.g., 4) of a non-affine CU.
- the scanned non-adjacent neighbor position may be out of the allowed non-adjacent area. In this case, different methods may be used to solve this issue as follows.
- the scanning process may indicate that this scanned position has no valid neighbor information.
- the scanning process may project or clip this out-of-range position to another position which is within the allowed non-adjacent area.
- FIG. 28 there are two positions 2801 (i.e., the two dotted spots 2801) out of the range of the allowable non- adjacent area. These two positions 2801 are projected/clipped to another two positions 2802 which are at the same vertical/horizontal coordinate but within the allowable non-adjacent area, respectively.
- the projected/clipped new position 2802 may be located on the boundary of the allowable non-adjacent area which is closest to the original position.
- the projected/clipped new position 2802 may be interchangeably set to be on the one boundary or the other boundary because the buffer is so small that the motion information from only one CU may be stored and, in this case, there is no difference to be clipped to one boundary side or the other side.
- the motion information When motion information of an affine-coded block is saved in memory, the motion information, including CPMVs, reference index, block size and positions may be saved at the granularity of minimum affine block size (e.g., an 8x8 block). In case the current affine-coded block is a coding unit with larger size than the minimum affine block, the motion information may be saved in different methods.
- the motion information saved at each minimum affine block (e.g., 8x8 block) within the current block is just a repeated copy of the motion information of the current block.
- the position and size of the current block may need to be repeatedly saved at each minimum affine block (termed as sub-block in the FIGS. 22A-22B) as well.
- An example for this case is shown in the FIG. 22A, where the current block (termed as the parent block) is at the size of 24x16, and the minimum affine block (termed as sub-block) is at a fixed size of 8x8.
- the motion information saved at each minimum affine block is the motion information already projected to this minimum affine block. Since the position of each minimum affine block is known (the topleft corner of each minimum affine block), and the size of each minimum affine block is also known (the minimum size, 8x8), the position and size information of the current block (termed as the parent block in FIGS. 22A-22B) do not need to be saved. An example for this case is shown in the FIG. 22B.
- the regular/translational motion at each inside non-affine block may be computed as follows. [00372] Taking FTG. 23 as an illustrative example. For this minimum affine block, it has three already projected CPMVs following the method shown in FIG. 22B. Based on the affine model shown in the equation (2), the regular/translational motion may be derived as below (e.g., ignoring some precision related shifts). The examples provided below are based on the center point of each 4x4 block as shown in FIG. 23, but the present disclosure is not limited to using center point to derive translation MV for each block.
- MVl x e + (a » 2) + (c » 2)
- MV1 _y f + (b » 2) + (d » 2)
- a CPMV2_x - CPMVl_x
- b CPMV2_y - CPMVl_y
- c CPMV3_x - CPMVl_x
- d CPMV3_y - CPMVl_y
- e CPMVl_x
- f CPMVl_y.
- MV3_x MVl x + (c »1)
- MV3_y MVl_y + (d» 1).
- MV4_x MVl_x + ((a+c) »1)
- MV4_y MVl_y + ((b+d)»l ).
- each 16x16 block may only save one set of affine motion information, which includes two or three CPMVs and represents one single affine model, even though the four 8x8 sub-blocks within this 16x16 block may be from more than one affine blocks, which is shown in the FIG. 25.
- the four 8x8 sub-blocks A, B, C and D form a 16x16 block/area, and only one affine model information is saved.
- the four 8x8 sub-blocks are from four different affine blocks, which represent four affine models and include four sets of affine motion information. In this case, there may be different ways to derive and save one single set of affine information.
- one of the multiple sets of available affine motion information may be selected and saved
- the affine motion information at one fixed or configurable position e.g., the top-left minimum affine block
- an averaged affine motion information of multiple models may be calculated for motion storage.
- the affine motion information at a selected neighboring affine block may be simplified/compressed before storage.
- the selected neighboring affine block is always 4-parameter model and only two CPMVs are saved.
- the selected neighboring affine block is always uni -predicted, and only one direction of affine motion is saved.
- each saved CPMV may be compressed before storage to further reduce the memory size.
- One example is to use general techniques for data compression. For example, it is provided to save a compounded value from one exponent and mantissa to approximately represent each saved CPMV.
- methods may be applied for motion information storage in any combinations.
- the defined restricted area for non-adjacent neighboring blocks may be combined with the usage of compressed affine motion information.
- MMVD mode the best MVD information is selected at the encoder side based on rate-distortion optimization (RDO) method, and then signaled to the decoder side.
- RDO rate-distortion optimization
- MVD information to refine the existing candidates in the affine AMVP or/and affine merge candidate list.
- the new candidates after refinements are then inserted into the existing affine AMVP or/and affine merge candidate list.
- the available number of combinations for MVD information such as the motion magnitude (e.g., offset value) and motion direction (e.g., sign value), may be the same or different as the existing MMVD mode in the VVC.
- a smaller number of offset values such as ⁇ 1 , 2, 4, 8, 16 ⁇ , may be used.
- a different or the same set of direction values as the existing MMVD mode may be used.
- the base MV is any one of the candidates from the existing affine AMVP and/or merge candidate list, and there may be multiple ways to determine the selection of a potential base MV.
- the base MV may be selected from a candidate list before or after an adaptive reordering method such as ARMC is applied to this candidate list.
- a single base MV or multiple base MVs may be selected from a candidate list. For example, when a single base MV is selected, one or multiple combinations (e.g., Y combinations) of MVD information may be selected to refine this single base MV, which indicates that Y new candidates (e.g., each combination of MVD information is applied to the base MV and generate one new candidate) may be generated and inserted into the candidate list.
- Y new candidates e.g., each combination of MVD information is applied to the base MV and generate one new candidate
- MVD information may be selected to refine each selected base MV, which indicates X multiplied by Y new candidates may be generated and inserted into the candidates list.
- the index of each selected base MV may be determined in different ways. In one or more examples, the index of the selected base MV may be determined by avoiding the base MVs which are already selected in MMVD mode, if affine MMVD mode is enabled for the current coding process.
- the index of the selected base MV may be determined by following a predefined order. For example, N candidates from the beginning of the list are sequentially selected as the base MVs. If any base MV is already selected by the current affine MMVD mode, this base MV may be skipped.
- the newly generated candidates after different combinations of MVD refinements may be directly inserted into the existing candidate list.
- another round of reordering process may be applied to all the new candidates and the top Z candidates with smaller matching cost (e.g., template matching cost or bilateral matching cost) may be selected to be inserted into the candidate list.
- FIG. 24 shows a computing environment (or a computing device) 2410 coupled with a user interface 2460.
- the computing environment 2410 can be part of a data processing server.
- the computing device 2410 can perform any of various methods or processes (such as encoding/decoding methods or processes) as described hereinbefore in accordance with various examples of the present disclosure.
- the computing environment 2410 may include a processor 2420, a memory 2440, and an I/O interface 2450.
- the processor 2420 typically controls overall operations of the computing environment 2410, such as the operations associated with the display, data acquisition, data communications, and image processing.
- the processor 2420 may include one or more processors to execute instructions to perform all or some of the steps in the above-described methods.
- the processor 2420 may include one or more modules that facilitate the interaction between the processor 2420 and other components.
- the processor may be a Central Processing Unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.
- the memory 2440 is configured to store various types of data to support the operation of the computing environment 2410.
- Memory 2440 may include predetermine software 2442. Examples of such data include instructions for any applications or methods operated on the computing environment 2410, video datasets, image data, etc.
- the memory 2440 may be implemented by using any type of volatile or non-volatile memory devices, or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic or optical disk.
- SRAM static random access memory
- EEPROM electrically erasable programmable read-only memory
- EPROM erasable programmable read-only memory
- PROM programmable read-only memory
- ROM read-only memory
- magnetic memory a magnetic memory
- flash memory a magnetic
- the EO interface 2450 provides an interface between the processor 2420 and peripheral interface modules, such as a keyboard, a click wheel, buttons, and the like.
- the buttons may include but are not limited to, a home button, a start scan button, and a stop scan button.
- the EO interface 2450 can be coupled with an encoder and decoder.
- a non-transitory computer-readable storage medium including a plurality of programs, such as included in the memory 2440, executable by the processor 2420 in the computing environment 2410, for performing the abovedescribed methods.
- the non-transitory computer-readable storage medium may be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disc, an optical data storage device or the like.
- the non-transitory computer-readable storage medium has stored therein a plurality of programs for execution by a computing device having one or more processors, where the plurality of programs when executed by the one or more processors, cause the computing device to perform the above-described method for motion prediction.
- the computing environment 2410 may be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field- programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components, for performing the above methods.
- ASICs application-specific integrated circuits
- DSPs digital signal processors
- DSPDs digital signal processing devices
- PLDs programmable logic devices
- FPGAs field- programmable gate arrays
- GPUs graphical processing units
- controllers microcontrollers, microprocessors, or other electronic components, for performing the above methods.
- FIG. 29 is a flowchart illustrating a method for video decoding according to an example of the present disclosure.
- the processor 2420 may obtain a restricted area that is not adjacent to a current coding unit (CU) according to a value associated with the restricted area.
- the restricted area is a predefined area associated with the current coding unit. Such association may be a spatial relationship between the restrict area and the CU, or a mapping relationship predefined between the restrict area and the CU.
- the restricted area may be one of following areas: a first restricted neighboring/neighbor area above the current CU or a second restricted neighboring/neighbor area on the left of the current CU.
- the first restricted neighboring/neighbor area 2706 is in the non-adjacent area of the current CU 2701 and above the CTU 2704 that the current CU 2701 located in.
- the second restricted neighboring/neighbor area 2705 is in the non-adjacent area of the current CU 2701 and to the left of the CTU 2704 that the current CU 2701 is located in.
- the processor 2420 may determine that the value is a height value associated with the first restricted neighbor area in response to determining that the restricted neighbor area is the first restricted neighbor area and may determine that the value is a width value associated with the second restricted neighbor area in response to determining that the restricted neighbor area is the second restricted neighbor area. For example, as shown in FIG. 27A, the first restricted neighbor area 2706 has a height value h and the second restricted neighbor area 2705 has a width value w. [00405] In some examples, the processor 2420 may obtain the value associated with the restricted area signaled in a bitstream sent by an encoder.
- the processor 2420 may pre-define the value associated with restricted area.
- the processor 2420 may determine that a buffer area for storing the CU is invalid in response to determining that a CU obtained by scanning the restricted area is not an affine CU.
- the processor 2420 may obtain a second CU that is an affine CU and located adj acent to the first CU in response to determining that a first CU obtained by scanning the restricted area is not an affine CU, obtain affine motion information of the second CU, determine that a buffer area for storing the first CU is valid and store the affine motion information obtained from the second CU in a buffer area for storing the first CU.
- the first CU may be the non-affme CU 2703
- the second CU may be the affine CU 2702
- the affine CU 2702 is adjacently above the non-affme CU 2703.
- the processor 2420 may pre-define the value associated with the restricted area as a multiple of a minimum size of a non-affme CU.
- the processor 2420 may obtain a CU at a scanning position by scanning a neighbor area of the current CU and determine that no valid neighbor information exists at the scanning position in response to determining that the scanning position is not within the restricted area.
- the processor 2420 may obtain a CU at a scanning position by scanning a neighbor area of the current CU, obtain a projected position by projecting the CU to the restricted area in response to determining that the scanning position is not within the restricted area, and obtain motion information associated with the CU at the scanning position in a buffer area for storing a projected CU that is located at the projected position.
- the scanning position may be the position 2801 that is beyond the allowed spatial area, i.e., outside of the restricted area.
- the projected position may be the position 2802 that is within the allowed spatial area, i.e., within the restricted area.
- the projected position may be located at a boundary of the restricted area.
- the restricted area may be one of following areas: a first restricted neighbor area above the current CU or a second restricted neighbor area on the left of the current CU, and the projected position may be located at a boundary of the first restricted neighbor area or the second restricted neighbor area.
- the processor 2420 may obtain one or more MV candidates from a plurality of non-adjacent CUs to the current CU based on the restricted area.
- the non-adjacent CUs are the non-adjacent neighbor CUs to the current CU.
- the plurality of non-adjacent CUs may be located within the restricted area, and the one or more MV candidates may be obtained by scanning the restricted area in which the plurality of non- adjacent CUs are located.
- Non-adjacent CUs may be located on the boundary of the restricted area in some examples.
- the processor 2420 may obtain one or more CPMVs for the current CU based on the one or more MV candidates.
- FIG. 30 is a flowchart illustrating a method for video encoding corresponding the method for video decoding as shown in FIG. 29.
- the processor 2420 at the encoder side, may obtain a restricted area that is not adjacent to a current CU according to a value associated with the restricted area.
- the restricted area may be one of following areas: a first restricted neighbor area above the current CU or a second restricted neighbor area on the left of the current CU.
- the first restricted neighbor area 2706 is in the non-adjacent area of the current CU 2701 and above the CTU 2704 that the current CU 2701 located in.
- the second restricted neighbor area 2705 is in the non-adjacent area of the current CU 2701 and to the left of the CTU 2704 that the current CU 2701 is located in.
- the processor 2420 may determine that the value is a height value associated with the first restricted neighbor area in response to determining that the restricted neighbor area is the first restricted neighbor area and may determine that the value is a width value associated with the second restricted neighbor area in response to determining that the restricted area is the second restricted neighbor area. For example, as shown in FIG. 27A, the first restricted neighbor area 2706 has a height value h and the second restricted neighbor area 2705 has a width value w. [00420] In some examples, the processor 2420 may signal the value associated with the restricted area in a bitstream that is to be sent to a decoder.
- the processor 2420 may pre-define the value associated with restricted area.
- the processor 2420 may determine that a buffer area for storing the CU is invalid in response to determining that a CU obtained by scanning the restricted area is not an affine CU.
- the processor 2420 may obtain a second CU that is an affine CU and located adj acent to the first CU in response to determining that a first CU obtained by scanning the restricted area is not an affine CU, obtain affine motion information of the second CU, determine that a buffer area for storing the first CU is valid and store the affine motion information obtained from the second CU in a buffer area for storing the first CU.
- the first CU may be the non-affme CU 2703
- the second CU may be the affine CU 2704
- the affine CU 2704 is adjacently above the non-affme CU 2703.
- the processor 2420 may pre-define the value associated with the restricted area as a multiple of a minimum size of a non-affme CU.
- the processor 2420 may obtain a CU at a scanning position by scanning a neighbor area of the current CU and determine that no valid neighbor information exists at the scanning position in response to determining that the scanning position is not within the restricted area.
- the processor 2420 may obtain a CU at a scanning position by scanning a neighbor area of the current CU, obtain a projected position by projecting the CU to the restricted area in response to determining that the scanning position is not within the restricted area, and obtain motion information associated with the CU at the scanning position in a buffer area for storing a projected CU that is located at the projected position.
- the scanning position may be the position 2801 that is beyond the allowed spatial area, i.e., outside of the restricted area.
- the projected position may be the position 2802 that is within the allowed spatial area, i.e., within the restricted area.
- the projected position may be located at a boundary of the restricted area.
- the restricted area may be one of following areas: a first restricted neighbor area above the current CU or a second restricted neighbor area on the left of the current CU, and the projected position may be located at a boundary of the first restricted neighbor area or the second restricted neighbor area.
- the processor 2420 may obtain one or more MV candidates from a plurality of non-adjacent neighbor CUs to the current CU based on the restricted area.
- the non-adjacent CUs are the non-adjacent neighbor CUs to the current CU.
- the plurality of non-adjacent CUs may be located within the restricted area, and the one or more MV candidates may be obtained by scanning the restricted area in which the plurality of non-adjacent CUs are located.
- Non-adjacent CUs may be located on the boundary of the restricted area in some examples.
- the processor 2420 may obtain one or more CPMVs for the current CU based on the one or more MV candidates.
- an apparatus for video decoding includes a processor 2420 and a memory 2440 configured to store instructions executable by the processor; where the processor, upon execution of the instructions, is configured to perform any method as illustrated in FIG. 29.
- an apparatus for video encoding includes a processor 2420 and a memory 2440 configured to store instructions executable by the processor; where the processor, upon execution of the instructions, is configured to perform any method as illustrated in FIG. 30.
- a non-transitory computer readable storage medium having instructions stored therein.
- the instructions When the instructions are executed by a processor 2420, the instructions cause the processor to perform any method as illustrated in FIGS. 29-30.
- the plurality of programs may be executed by the processor 2420 in the computing environment 2410 to receive (for example, from the video encoder 20 in FIG. 1G) a bitstream or data stream including encoded video information (for example, video blocks representing encoded video frames, and/or associated one or more syntax elements, etc.), and may also be executed by the processor 2420 in the computing environment 2410 to perform the decoding method described above according to the received bitstream or data stream.
- the plurality of programs may be executed by the processor 2420 in the computing environment 2410 to perform the encoding method described above to encode video information (for example, video blocks representing video frames, and/or associated one or more syntax elements, etc.) into a bitstream or data stream, and may also be executed by the processor 2420 in the computing environment 2410 to transmit the bitstream or data stream (for example, to the video decoder 30 in FIG. 2B).
- the non-transitory computer-readable storage medium may have stored therein a bitstream or a data stream including encoded video information (for example, video blocks representing encoded video frames, and/or associated one or more syntax elements etc.) generated by an encoder (for example, the video encoder 20 in FIG.
- the non-transitory computer-readable storage medium may be, for example, a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, an optical data storage device or the like.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263332244P | 2022-04-18 | 2022-04-18 | |
| PCT/US2023/019002 WO2023205185A1 (en) | 2022-04-18 | 2023-04-18 | Methods and devices for candidate derivation for affine merge mode in video coding |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| EP4512092A1 true EP4512092A1 (de) | 2025-02-26 |
| EP4512092A4 EP4512092A4 (de) | 2026-04-15 |
Family
ID=88420468
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23792452.7A Pending EP4512092A4 (de) | 2022-04-18 | 2023-04-18 | Verfahren und vorrichtungen zur kandidatenableitung für affinen zusammenführungsmodus in der videocodierung |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4512092A4 (de) |
| CN (1) | CN119054289A (de) |
| WO (1) | WO2023205185A1 (de) |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2015106747A (ja) * | 2013-11-28 | 2015-06-08 | 富士通株式会社 | 動画像符号化装置、動画像符号化方法及び動画像符号化用コンピュータプログラム |
| EP3414900B1 (de) * | 2016-03-15 | 2025-08-06 | HFI Innovation Inc. | Verfahren und vorrichtung zur videocodierung mit affiner bewegungskompensation |
| US10602180B2 (en) * | 2017-06-13 | 2020-03-24 | Qualcomm Incorporated | Motion vector prediction |
| US11082708B2 (en) * | 2018-01-08 | 2021-08-03 | Qualcomm Incorporated | Multiple-model local illumination compensation |
| CN118264797A (zh) * | 2018-02-28 | 2024-06-28 | 弗劳恩霍夫应用研究促进协会 | 合成式预测及限制性合并 |
| US10863193B2 (en) * | 2018-06-29 | 2020-12-08 | Qualcomm Incorporated | Buffer restriction during motion vector prediction for video coding |
| US11057636B2 (en) * | 2018-09-17 | 2021-07-06 | Qualcomm Incorporated | Affine motion prediction |
| US11212550B2 (en) * | 2018-09-21 | 2021-12-28 | Qualcomm Incorporated | History-based motion vector prediction for affine mode |
-
2023
- 2023-04-18 WO PCT/US2023/019002 patent/WO2023205185A1/en not_active Ceased
- 2023-04-18 EP EP23792452.7A patent/EP4512092A4/de active Pending
- 2023-04-18 CN CN202380034718.XA patent/CN119054289A/zh active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2023205185A1 (en) | 2023-10-26 |
| CN119054289A (zh) | 2024-11-29 |
| EP4512092A4 (de) | 2026-04-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4399876A1 (de) | Bewegungskompensation unter berücksichtigung von aussergrenzenbedingungen in der videocodierung | |
| US20250024060A1 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| US20250047897A1 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| WO2023141177A1 (en) | Motion compensation considering out-of-boundary conditions in video coding | |
| JP2026021367A (ja) | ビデオ・コーディングにおけるアフィン・マージ・モードに対する候補導出のための方法及びデバイス | |
| US20250220210A1 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| US20240283921A1 (en) | Candidate derivation for affine merge mode in video coding | |
| EP4552329A1 (de) | Verfahren und vorrichtungen zur kandidatenableitung für affinen zusammenführungsmodus in der videocodierung | |
| EP4523412A1 (de) | Verfahren und vorrichtungen zur kandidatenableitung für affinen zusammenführungsmodus in der videocodierung | |
| WO2023133160A1 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| US12615360B2 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| US20250039365A1 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| US12621462B2 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| US12532015B2 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| EP4512092A1 (de) | Verfahren und vorrichtungen zur kandidatenableitung für affinen zusammenführungsmodus in der videocodierung | |
| WO2024206533A2 (en) | Methods and devices for candidate derivation for affine merge mode in video coding | |
| EP4494343A1 (de) | Interprädiktion in der videocodierung |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20241115 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| A4 | Supplementary search report drawn up and despatched |
Effective date: 20260318 |
|
| RIC1 | Information provided on ipc code assigned before grant |
Ipc: H04N 19/527 20140101AFI20260312BHEP Ipc: H04N 19/52 20140101ALI20260312BHEP Ipc: H04N 19/139 20140101ALI20260312BHEP Ipc: H04N 19/129 20140101ALI20260312BHEP Ipc: H04N 19/423 20140101ALI20260312BHEP Ipc: H04N 19/176 20140101ALI20260312BHEP |