EP3932073A1 - Advanced independent region boundary handling - Google Patents
Advanced independent region boundary handlingInfo
- Publication number
- EP3932073A1 EP3932073A1 EP20706330.6A EP20706330A EP3932073A1 EP 3932073 A1 EP3932073 A1 EP 3932073A1 EP 20706330 A EP20706330 A EP 20706330A EP 3932073 A1 EP3932073 A1 EP 3932073A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- block
- sub
- predetermined
- referenced
- inter
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
- H04N19/137—Motion inside a coding unit, e.g. average field, frame or block difference
- H04N19/139—Analysis of motion vectors, e.g. their magnitude, direction, variance or reliability
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/174—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a slice, e.g. a line of blocks or a group of blocks
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/55—Motion estimation with spatial constraints, e.g. at image or region borders
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/59—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial sub-sampling or interpolation, e.g. alteration of picture size or resolution
Definitions
- the present application relates to video coding.
- Constraint encodings of tiles, as independent regions, is a feature that is available in video coding standards for different use cases such as Region of Interest (Rol) or 360°-degree videos.
- Rol Region of Interest
- HEVC specifies Motion Constraint Tile Sets (MCTS), where an encoder carefully produces a bitstream where Motion Vectors (MVs) are properly set so that tiles do not need neighboring tiles as reference.
- MCTS Motion Constraint Tile Sets
- a better solution is to have within a decoding process handling of MVs in such a way that references to another tiles (when no desired) is implicitly restricted without requiring MV differences to be added into the bitstream which requiring spending additional bits.
- a block-based video codec supporting motion-compensated prediction is provided in a manner so that same allows for coding a video using motion-constrained regions in a more effective manner, such as coding a video in a manner so that one or more sub-videos thereof are coded independently from a remainder of the video, external to the corresponding sub-video.
- decoder and encoder perform a conflict check for a referenced block partition with respect to a predetermined picture region, namely whether a position thereof relative to a border of a predetermined picture region fulfills a predetermined criterion such as exceeding beyond the border, and that the decision is not used to dismiss inter-prediction for the referencing (predetermined) inter-predicted block at all, if the conflict check reveals a conflict and the predetermined criterion is not met, but to proceed with the inter-prediction by, responsive to the conflict check having revealed a conflict, subdividing the referencing (predetermined) inter-predicted block and performing the inter-prediction for a set of one or more sub-blocks resulting from the subdivision.
- the encoder needs not to subdivide the currently coded picture into smaller inter-predicted blocks, which measure is cost-intensive in terms of side information overhead, but the sub-divisioning is automatically activated or agreed between encoder and decoder upon encountering any motion-constraint conflict.
- the decoder is configured to split, or subdivide, a current picture into a plurality of blocks. These blocks are e.g. inter-predicted blocks, and the data stream received by the decoder contains motion information for these inter-predicted blocks. These motion information, which can comprise motion vectors, reference block portions in reference pictures for motion compensated prediction. The motion information reference is relative to the inter-predicted block.
- the decoder checks whether the position of a block portion in a reference picture, i.e. a referenced block portion, fulfills a certain criterion, which is discussed below.
- the decoder By checking whether the position fulfills the criterion, the decoder also checks whether the motion information reference a referenced block portion with a position that fulfils the criterion. And in turn, whether the inter-predicted block fulfils the criterion.
- the position is checked against the criterion in relation to the border of a predetermined picture region.
- the decoder proceeds in dependence of whether the motion information reference a block portion, that fulfils the criterion, and in particular, if the position of the referenced block portion fulfills the first predetermined criterion, i.e. the predetermined inter-predicted block fulfils the criterion, the predetermined inter-predicted block is inter-predicted by the decoder without further dividing inter-predicted block by using the referenced block portion.
- the decoder first subdivides the predetermined inter-predicted block into sub-blocks, and then inter-predicts one or more sub-blocks using the motion information for the predetermined inter-predicted block.
- the picture region might be a MCTS, motion-constrained tile set, or a motion-constrained region. This means it is a region of the reference picture which is associated with a co located region of the current picture. In the co-located region, the predetermined inter- predicted block is located. Possibly, the predetermined picture region is further associated with a co-located region of further pictures, together forming a spatiotemporal region, i.e. a sub-video, which is encoded and decoded independent from any surrounding of the colocated regions.
- a spatiotemporal region i.e. a sub-video
- a block-based video codec supporting motion-compensated prediction and using interpolation filtering for sub-pel motion vectors is provided in a manner so that same allows for coding a video using motion- constrained regions in a more effective manner, such as coding a video in a manner so that one or more sub-videos thereof are coded independently from a remainder of the video, external to the corresponding sub-video.
- decoder and encoder perform a conflict check for a referenced block partition with respect to a predetermined picture region, namely whether a position thereof relative to a border of a predetermined picture region fulfills a predetermined criterion such as becoming too close to the border so that an interpolation filter associated with the sub-pel motion vector would extend beyond the border, and that the decision is not used to dismiss inter-prediction for the referencing (predetermined) inter- predicted block at all, if the conflict check reveals a conflict and the predetermined criterion is not met, but to proceed with the inter-prediction using the sub-pel vector at hand by, responsive to the conflict check having revealed a conflict, using another interpolation filter having a smaller kernel reach instead.
- a block-based video decoder supporting motion- compensated prediction which subdivides a current picture into a plurality of blocks.
- these blocks are inter-predicted blocks, and for these inter-predicted blocks the data stream received by the decoder contains a sub-pel motion vector.
- These motion vectors, reference block portions in reference pictures for motion compensated prediction. The vector reference is relative to the inter-predicted block.
- the decoder checks whether the position of a block portion in a reference picture, i.e. a referenced block portion, fulfills a certain criterion, which is discussed below.
- the decoder By checking whether the position fulfills the criterion, the decoder also checks whether the motion vector reference a referenced block portion with a position that fulfils the criterion. And in turn, whether the inter-predicted block fulfils the criterion.
- the position is checked against the criterion in relation to the border of a predetermined picture region.
- the decoder proceeds in dependence of whether the motion vector references a block portion, that fulfils the criterion, and in particular, if the position of the referenced block portion fulfills the first predetermined criterion, i.e. the predetermined inter-predicted block fulfils the criterion, the predetermined inter-predicted block is inter-predicted by the decoder by sampling the reference picture at the referenced block portion using the interpolation filter associated with the sub-pel granularity of the sub-pel motion vector.
- the decoder inter-predicts the predetermined inter- predicted block by sampling the reference picture at the referenced block portion using a substitute interpolation filter having filter kernel reach which is smaller than the reach of the filter kernel of the interpolation filter associated with the sub-pel granularity of the sub-pel motion vector.
- the concepts can be implemented by methods for block-based decoding and/or block-based encoding according to embodiments of the present invention. These methods are based on the same considerations as the above-described decoder and encoder. However, it should be noted that the methods can be supplemented by any of the features, functionalities and details described herein, also with respect to the decoder and/or encoder. Moreover, the methods can be supplemented by the features, functionalities, and details of the decoder and/or encoder, both individually and taken in combination.
- the concepts can be used to produce an encoded data stream according to embodiments of the present invention.
- the data stream can also be supplemented by the features, functionalities, and details of the decoder and/or encoder, both individually and taken in combination.
- FIG. 1 shows an encoder for predictively coding a video composed of a sequence of pictures into a data stream using, exemplarily, transform-based residual coding according to one embodiment of the present application
- Fig. 2 shows a decoder for predictively decoding a data stream into a video composed of a sequence of pictures using, exemplarily, transform-based residual decoding according to one embodiment of the present application
- Fig. 3 illustrates the relationship between the reconstructed signal and the prediction signal according to one embodiment of the present application
- Fig. 4 shows an illustration of the concept of region / tile-independent coding according to one embodiment of the present application
- Fig. 5 shows a schematic data stream and decoder according to one embodiment of the present application
- Fig. 6 shows a flow chart of a method according to one embodiment of the present application
- Fig. 7 shows prediction of sub-blocks using substitute prediction according to one embodiment of the present application
- Fig. 8 shows inter-prediction of sub-blocks when the criterion is not fulfilled according to one embodiment of the present application
- Fig. 9 shows deriving motion information from the data stream according to one embodiment of the present application.
- Fig. 10 shows alternative derivation of the motion information according to one embodiment of the present application
- Fig. 11 shows correction of a motion vector predictor according to one embodiment of the present application.
- Fig. 12 is an example of splitting a block.
- similar reference signs denote similar elements and features.
- Fig. 1 shows an encoder for predictively coding a video composed of a sequence of pictures into a data stream using, exemplarily, transform-based residual coding according to one embodiment of the present application, in other words, Fig. 1 shows an apparatus for predictively coding a video 11 composed of a sequence of pictures 12 into a data stream 14 using, exemplarily, transform-based residual coding.
- the apparatus, or encoder is indicated using reference sign 10.
- Fig. 2 shows a corresponding decoder 20, i.e.
- an apparatus 20 configured to predictively decode a video 1 T composed of a sequence of pictures 12’ from the data stream 14 also using transform-based residual decoding, wherein the apostrophe has been used to indicate that the video 1 T and pictures 12' as reconstructed by decoder 20 deviate from picture 12 originally encoded by apparatus 10 in terms of coding loss introduced by a quantization of the prediction residual signal.
- Fig. 1 and Fig. 2 exemplarily use transform based prediction residual coding, although embodiments of the present application are not restricted to this kind of prediction residual coding. This is true for other details described with respect to Fig. 1 and 2, too, as will be outlined hereinafter.
- the encoder 10 is configured to subject the prediction residual signal to spatial-to-spectral transformation and to encode the prediction residual signal, thus obtained, into the data stream 14.
- the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and subject the prediction residual signal thus obtained to spectral- to-spatial transformation.
- the encoder 10 may comprise a prediction residual signal former 22 which generates a prediction residual 24 so as to measure a deviation of a prediction signal 26 from the original signal, i.e. the current picture 12.
- the prediction residual signal former 22 may, for instance, be a subtractor which subtracts the prediction signal from the original signal, i.e. current picture 12.
- the encoder 10 then further comprises a transformer 28 which subjects the prediction residual signal 24 to a spatial-to-spectral transformation to obtain a spectral-domain prediction residual signal 24’ which is then subject to quantization by a quantizer 32, also comprised by encoder 10.
- the thus quantized prediction residual signal 24” is coded into bitstream 14.
- encoder 10 may optionally comprise an entropy coder 34 which entropy codes the prediction residual signal as transformed and quantized into data stream 14.
- the prediction residual 26 is generated by a prediction stage 36 of encoder 10 on the basis of the prediction residual signal 24” decoded into, and decodable from, data stream 14.
- the prediction stage 36 may internally, as is shown in Fig. 1 , comprise a dequantizer 38 which dequantizes prediction residual signal 24” so as to gain spectral-domain prediction residual signal 24’”, which corresponds to signal 24’ except for quantization loss, followed by an inverse transformer 40 which subjects the latter prediction residual signal 24”’ to an inverse transformation, i.e.
- prediction residual signal 24 which corresponds to the original prediction residual signal 24 except for quantization loss.
- a combiner 42 of the prediction stage 36 then recombines, such as by addition, the prediction signal 26 and the prediction residual signal 24”” so as to obtain a reconstructed signal 46, i.e. a reconstruction of the original signal 12.
- Reconstructed signal 46 may correspond to signal 12’.
- a prediction module 44 of prediction stage 36 then generates the prediction signal 26 on the basis of signal 46 by using, for instance, spatial prediction, i.e. intra prediction, and/or temporal prediction, i.e. inter prediction.
- Fig. 2 shows a decoder for predictively decoding a data stream into a video composed of a sequence of pictures using, exemplarily, transform-based residual decoding according to one embodiment of the present application.
- decoder 20 may be internally composed of components corresponding to, and inter-connected in a manner corresponding to, prediction stage 36.
- entropy decoder 50 of decoder 20 may entropy decode the quantized spectral-domain prediction residual signal 24” from the data stream, whereupon dequantizer 52, inverse transformer 54, combiner 56 and prediction module 58, interconnected and cooperating in the manner described above with respect to the modules of prediction stage 36, recover the reconstructed signal on the basis of prediction residual signal 24” so that, as shown in Fig. 2, the output of combiner 56 results in the reconstructed signal, namely picture 12’.
- the encoder 10 may set some coding parameters including, for instance, prediction modes, motion parameters and the like, according to some optimization scheme such as, for instance, in a manner optimizing some rate and distortion related criterion, i.e. coding cost.
- encoder 10 and decoder 20 and the corresponding modules 44, 58, respectively may support different prediction modes such as intra-coding modes and inter-coding modes.
- the granularity at which encoder and decoder switch between these prediction mode types may correspond to a subdivision of picture 12 and 12’, respectively, into coding segments or coding blocks. In units of these coding segments, for instance, the picture may be subdivided into blocks being intra-coded and blocks being inter-coded.
- Intra-coded blocks are predicted on the basis of a spatial, already coded/decoded neighborhood of the respective block.
- Several intra-coding modes may exist and be selected for a respective intra-coded segment including directional or angular intra-coding modes according to which the respective segment is filled by extrapolating the sample values of the neighborhood along a certain direction which is specific for the respective directional intra-coding mode, into the respective intra-coded segment.
- the intra-coding modes may, for instance, also comprise one or more further modes such as a DC coding mode, according to which the prediction for the respective intra-coded block assigns a DC value to all samples within the respective intra-coded segment, and/or a planar intra-coding mode according to which the prediction of the respective block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of the respective intra-coded block with deriving tilt and offset of the plane defined by the two-dimensional linear function on the basis of the neighboring samples.
- inter-coded blocks may be predicted, for instance, temporally.
- motion information may be signaled within the data stream.
- the motion information may comprise vectors indicating the spatial displacement of the portion of a previously coded picture of the video to which picture 12 belongs, at which the previously coded/decoded picture is sampled in order to obtain the prediction signal for the respective inter-coded block. More complex motion models may be used as well.
- data stream 14 may have encoded thereinto coding mode parameters for assigning the coding modes to the various blocks, prediction parameters for some of the blocks, such as motion parameters for inter-coded blocks, and optional further parameters such as parameters controlling and signaling the subdivision of picture 12 and 12’, respectively, into the blocks.
- the decoder 20 uses these parameters to subdivide the picture in the same manner as the encoder did, to assign the same prediction modes to the blocks, and to perform the same prediction to result in the same prediction signal.
- Fig. 3 illustrates the relationship between the reconstructed signal, i.e. the reconstructed picture 12’, on the one hand, and the combination of the prediction residual signal 24”” as signaled in the data stream, and the prediction signal 26, on the other hand, according to one embodiment of the present application.
- the combination may be an addition.
- the prediction signal 26 is illustrated in Fig. 3 as a subdivision of the picture area into intra-coded blocks which are illustratively indicated using hatching, and inter-coded blocks which are illustratively indicated not-hatched.
- the subdivision may be any subdivision, such as a regular subdivision of the picture area into rows and columns of blocks or blocks, or a multi-tree subdivision of picture 12 into leaf blocks of varying size, such as a quadtree subdivision or the like, into blocks, wherein a mixture thereof is illustrated in Fig. 3 where the picture area is first subdivided into rows and columns of tree- root blocks which are then further subdivided in accordance with a recursive multi-tree sub- divisioning.
- data stream 14 may have an intra-coding mode coded thereinto for intra- coded blocks 80, which assigns one of several supported intra-coding modes to the respective intra-coded block 80.
- inter-coded blocks 82 the data stream 14 may have motion information coded thereinto such as motion information including one or more motion vectors. Details are set out hereinbelow.
- inter-coded blocks 82 are not restricted to being temporally coded.
- inter-coded blocks 82 may be any block predicted from previously coded portions beyond the current picture 12 itself, such as previously coded pictures of a video to which picture 12 belongs, or picture of another view or a hierarchically lower layer in the case of encoder and decoder being scalable encoders and decoders, respectively.
- the prediction residual signal 24”” in Fig. 3 is also illustrated as a subdivision of the picture area into blocks 84.
- FIG. 3 illustrates that encoder 10 and decoder 20 may use two different subdivisions of picture 12 and picture 12’, respectively, into blocks, namely one sub-divisioning into coding blocks 80 and 82, respectively, and another subdivision into blocks 84. Both subdivisions might be the same, i.e. each coding block 80 and 82, may concurrently form a transform block 84, but Fig.
- FIG. 3 illustrates the case where, for instance, a subdivision into transform blocks 84 forms an extension of the subdivision into coding blocks 80/82 so that any border between two blocks of blocks 80 and 82 overlays a border between two blocks 84, or alternatively speaking each block 80/82 either coincides with one of the transform blocks 84 or coincides with a cluster of transform blocks 84.
- the subdivisions may also be determined or selected independent from each other so that transform blocks 84 could alternatively cross block borders between blocks 80/82.
- similar statements are thus true as those brought forward with respect to the subdivision into blocks 80/82, i.e.
- the blocks 84 may be the result of a regular subdivision of picture area into blocks/blocks, arranged in rows and columns, the result of a recursive multi-tree sub-divisioning of the picture area, or a combination thereof or any other sort of blockation.
- blocks 80, 82 and 84 are not restricted to being of quadratic, rectangular or any other shape.
- an inter-predicted block 104 is representatively used to describe the specific details of the respective embodiment.
- This block 104 is may be one of the inter-predicted blocks 82.
- the other blocks mentioned in the subsequent figures may be any of the blocks 80 and 82.
- Fig. 3 illustrates that the combination of the prediction signal 26 and the prediction residual signal 24”” directly results in the reconstructed signal 12’.
- more than one prediction signal 26 may be combined with the prediction residual signal 24”” to result into picture 12’ in accordance with alternative embodiments.
- the transform segments 84 shall have the following significance.
- Transformer 28 and inverse transformer 54 perform their transformations in units of these transform segments 84.
- many codecs use some sort of DST or DCT for all transform blocks 84.
- Some codecs allow for skipping the transformation so that, for some of the transform segments 84, the prediction residual signal is coded in in the spatial domain directly.
- encoder 10 and decoder 20 are configured in such a manner that they support several transforms.
- the transforms supported by encoder 10 and decoder 20 could comprise:
- DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform
- transformer 28 would support all of the forward transform versions of these transforms, the decoder 20 or inverse transformer 54 would support the corresponding backward or inverse versions thereof:
- Figs. 1 - 3 have been presented as examples for an encoder which performs block-based video encoding using one or more of the concepts outlined in more detail below.
- the details set forth below may be transferred to the encoder of Fig. 1 or to another block-based video encoder so that the latter would be different from the encoder of Fig. 1 such as, for instance, in that same does not support intra-prediction, or in that the sub-division into blocks 80 and/or 82 is performed in a manner different than exemplified in Fig. 3, or even in that this encoder does not use transform prediction residual coding with coding the prediction residual, for instance, in spatial domain directly instead.
- decoders may perform block-based video decoding from data stream 14 using the any inter-prediction coding concept further outlined below, but may differ, for instance, from the decoder 20 of Fig. 2 in that same does not support intra-prediction, or in that same sub-divides picture 12’ into blocks in a manner different than described with respect to Fig. 3 and/or in that same does not derive the prediction residual from the data stream 14 in transform domain, but in spatial domain, for instance.
- Fig. 4 shows an illustration of the concept of region / tile-independent coding according to one embodiment of the present application, thus, after having described a potential implementation of block-based video encoders and video decoders, Fig. 4 is used to illustrate the concept of region/tile-independent coding.
- Tile-based independent coding may represent just one of several coding settings for the encoder. That is, the encoder may be set to adhere to tile-independent coding or may be set to not adhere to tile-independent coding. Naturally, it could be that the video encoder inevitably applies tile-independent coding.
- the pictures of the video 1 1 are partitioned into tiles 110.
- Fig. 4 illustratively, merely three pictures 12a, 12b and 12c out of video 11 are shown and they are exemplarily shown to be partitioned into six tiles 110, respectively, although a partition into any other number of tiles would be feasible as well such as two or more tiles or even one tile.
- Fig. 4 illustrates that the tile partitioning into tiles 110 is done in a manner so that the tiles 110 within one picture are regularly arranged in rows and columns with each tile 110 being a rectangular portion of the respective picture and the boundaries between the tiles 110 forming straight lines 92 leading through pictures 12a to 12c, it should be noted that alternative solutions exist as well where the tiles have another shape.
- each picture 12 into blocks is aligned to tile boundaries 92 and that none of the coding blocks, prediction blocks and/or transform blocks, crosses and tile boundary 92, but rather are exclusively in one of the tiles 1 10 of the respective picture.
- the tile partitioning may be done in a manner aligned with a partitioning of pictures 12 into tree root blocks, i.e. blocks which are individually subdivided by the encoder into blocks 80/82 by way of recursive multi-tree partitioning, to which also the representatively shown block 104 belongs, and which individual subdivision is signaled to the decoder in the data stream.
- Tile-independent coding means the following: the encoding of a respective portion of a tile, such as exemplarily shown for block 104 in Fig. 4, is done in a manner independent from any other tile within block 104 is not located.
- This coding independency does not only pertain to other tiles of the same picture, i.e., other tiles in the same picture which block 104 is part of, which intra-picture independency is illustrated in Fig. 4 using an exemplary arrow 96.
- the coding independency also relates to coding dependencies with respect to other pictures: all pictures 12a to 12c of video 1 1 are partitioned into the tiles 1 10 in the same manner. Thus, each tile 1 10 of a certain picture 12a has a corresponding, co-located tile in all other pictures.
- A“tile” does, accordingly, not only describe a spatial segment of a certain picture such as the segment of picture 12a of Fig. 4 containing block 104, but also the spatiotemporal segment of video 11 composed of all co-located tiles of all pictures of the video which are co-located to this tile in the certain picture. In Fig. 4, this spatiotemporal segment is illustrated by dashed lines with respect to the tile containing block 104.
- the encoder When coding block 104, the encoder will, accordingly, also restrict coding dependencies to other pictures than picture 12a containing block 104 in a manner so as to not cause coding dependencies to tiles outside the spatiotemporal segment 98. In Fig. 4, this is illustrated by interrupted arrow 99 which points from the tile containing block 104 to a different tile in a reference picture 12b.
- Fig. 5 shows a schematic data stream 14 and decoder according to one embodiment of the present application.
- the decoder receives the data stream 14 comprising motion information, mi..
- Current picture 12a and reference picture 12b which are part of video 1 1 , are shown.
- the motion vector 142 references the referenced block portion 106 in the predetermined picture region 1 10 of the reference picture 12b relative to the inter-predicted block 104 of the current picture 12a.
- the referenced block portion 106 extends beyond the border 108 in the reference picture 12b.
- the encoder would have refrained from having chosen that motion vector. Maybe, the encoder would have sub-divided the block and sent corresponding subdivision side information.
- encoder and decoder act differently.
- the referenced block portion 106 extends beyond the border 108 of the predetermined picture region 1 10.
- criteria further outlined below, which would be qualified as revealing a motion constraint conflict, partially including ones which take interpolation filter kernels into account, but such details are set out in more detail below.
- Fig. 6 shows a flow chart of a method according to one embodiment of the present application. The method would be performed by both decoder and encoder. Therein it can be seen that in step 150 motion information 109 is derived from the data stream. Examples for details of the derivation are discussed with reference to Fig. 9 below.
- the motion information may include a motion vector 142.
- step 111 On the basis of the derived motion information 109 for the predetermined inter-predicted block, whether it references, relative to the predetermined inter-predicted block, a referenced block portion 106 in a reference picture, a position of which relative to a border 108 of a predetermined picture region fulfills a first predetermined criterion.
- This criterion is exemplary considered to be fulfilled if, if the motion information requires the usage of a filter kernel 112, the referenced block portion 106 is distanced from the border of the predetermined picture region by more than a filter kernel reach 114 of the filter kernel 112.
- the criterion might alternatively considered as being fulfilled inevitably if the referenced block portion 106 is distanced from the border of the predetermined picture region by more than the filter kernel reach 1 14 of the filter kernel 112 although, for instance, the motion information for the current block was a full-pel motion vector and would, thus, not require the usage of the interpolation filter, i.e. might be considered to be fulfilled if the referenced block portion 106 is distanced from the border of the predetermined picture region by more than the filter kernel reach 114 of the filter kernel 1 12, independent from whether the motion information for the current block requires the usage of the interpolation filter or not.
- the conflict check may be relaxed a little bit by encoder and decoder allowing the extending beyond the border 108 and not considering it as raising a motion constraint conflict if the extension beyond the border is less a predetermined distance, the latter fraction would then be, in inter-predicting block 104 un-subdivided, derived at encoder and decoder in a manner not requiring any information from beyond the border 108 such as by border extension which might be implemented as an intra-prediction mode for portions of block 106 beyond border 108 using, for instance, an angle perpendicular to border 108.
- the predetermined inter-predicted block is inter-predicted in a manner unsubdivided, i.e. without dividing the predetermined inter- predicted block first, using the referenced block portion 106.
- the predetermined inter-predicted block 104 is first subdivided 120 into sub-blocks 122, and then one or more sub-blocks are inter-predicted 124 using the motion information 109.
- the motion vector 142 may be used for the one or more sub-blocks inter-predicted.
- the one or more sub-block 122 which are inter-predicted will be farer away from border 108 (or its co-located footprint in the current picture).
- the one or more other sub-blocks will be treated differently, such as intra-predicted sub-blocks using a predetermined intra-prediction mode or the like.
- the conflict check 111 might be re-used to decide which of the sub-blocks are to be inter-predicted, i.e. for which the referenced portion when using the motion information of block 104, does not raise any conflict, and which are to be handled differently, i.e. for which the referenced portion when using the motion information of block 104, would raise a conflict.
- motion clipping may be used for such blocks as described with respect to Fig. 8.
- Fig. 8 shows the step of inter-prediction 124 of sub-blocks when the first criterion is not fulfilled according to one embodiment of the present application. Then the decoder may classify 128 the sub-blocks 122 into sub-blocks 130 and 132 according to a second predetermined criterion.
- the second predetermined criterion inspects for each sub-block 130 and 132 as to which referenced sub-block portions 134 it references via the motion information, e.g. motion vector 142, signaled in the data stream for the predetermined inter-predicted block 106, relative to the respective sub-block 130, 132.
- the motion information e.g. motion vector 142
- the second predetermined criterion can for example be, whether the position of the first referenced sub-block portion 134 does extend beyond the border 108 of the predetermined picture region 1 10 or not, or by how much. In principles, similar examples as outlined with respect to block 106 hold true.
- the sub-blocks 120 are classified into first sub-blocks 130 that fulfil the second predetermined criterion and second sub-blocks 132 which do not fulfil the second predetermined criterion.
- the decoder then can inter-predict 136 the one or more first sub-blocks 130 using the one or more first referenced sub-block portions 106, and predict 138 the one or more second sub-blocks using a substitute prediction.
- Fig. 7 shows prediction 138 of sub-blocks 132 using substitute prediction according to one embodiment of the present application, i.e. the prediction 138 for the second sub-blocks 132 using the substitute prediction.
- Fig. 7 thus shows an example for treating conflicts of referenced portions such as referenced block 106 or referenced sub-blocks 134 with respect to border 108.
- Such treatment may also be used at encoder and decoder to deal with inter-predicted block 104 or sub-blocks 122 in order to account for the afore-mentioned predetermined amount (at which an extension beyond the border is allowed), or to deal with the remaining sub-blocks for which a conflict would apply if the motion information was applied as it is.
- the decoder clips 140a, 140b for each of the second sub-blocks a motion vector mv 142 comprised by the motion information 109.
- a clipped motion vector mv’ - and corresponding motion information 109’ - is obtained, which then references, relative to the respective second sub-block 132, a third referenced sub-block portion 134’ in the reference picture.
- the third referenced sub-block portion 134’ then does extend beyond the border 108 of the predetermined picture region 1 10 not at all or at least not by more than a predetermined distance.
- Fig. 9 shows deriving 150 motion information 109 from the data stream according to one embodiment of the present application, wherein the derivation is performed by decoding merge indicator 160 from the data stream 14.
- the merge indicator 160 indicates, as can be seen in item 152, whether the motion information 109, ml, for the predetermined inter-predicted block 104 is to be determined by merging, i.e. based on a motion information candidate list 158 (yes at item 152) or not (no in item 152).
- a predetermined motion information candidate 156 is determined out of the motion information candidate list 158 on the basis of further motion information signaled in the data stream for previous inter-predicted blocks, and indicated 154.
- the predetermined motion information candidate 156 is used 161 a as the motion information 109 or the motion information 109 is determined 161 b using the predetermined motion information candidate 156.
- the merge indicator 160 indicates that the motion information for the predetermined inter-predicted block is not to be determined based on the motion information candidate list, i.e. no in item 152
- further motion parameters 64 are decoded from the data stream 14 for the predetermined inter-predicted block, and the motion information is determined 162 based on the further motion parameters 164.
- An example for such further motion parameters 164 is a motion vector difference 172, mvd, and/or a reference picture index.
- Fig. 10 shows alternative derivation of the motion information according to one embodiment of the present application.
- the motion information 109 e.g. a motion vector 142
- the motion information 109 is derived 150 from the data stream 14 for the predetermined inter-predicted block 104 by determining a motion vector predictor 170’ and 170 by temporal or spatial prediction.
- a motion vector difference or offset 172, 180 is then decoded for the predetermined inter-predicted block 104 from the data stream 14, and the motion vector predictor 170’ and 170 is corrected 182 and 176, e.g. by clipping, using the motion vector difference or offset to obtain a motion vector 142 comprised by the motion information 109.
- the motion vector predictor 170’, 170 is clipped 182.
- Fig. 10 shows the clipping 182 in details whereas Fig. 11 shows the correction 176 in detail.
- Fig. 1 1 shows correction of a motion vector predictor according to one embodiment of the present application.
- a motion vector predictor 170 is corrected 176 using a motion vector difference 172, which e.g. is part of the further motion parameters 164, to obtain a motion vector 142 comprised by the motion information 109.
- the criterion can be held to be fulfilled if, for example, the following two conditions are met:
- the first condition is that the inter-predicting of the predetermined inter-predicted block in the manner unsubdivided, i.e. without dividing the inter-predicted block, does not involve a filter kernel 1 12.
- the second condition is that the referenced block portion 106 does not extend beyond the border 108 of the predetermined picture region 110 at all or at least not by more than a predetermined distance.
- the decoder can also check 111 whether the position of the referenced block portion fulfills the first predetermined criterion by checking whether the motion information requires a usage of a filter having a filter kernel 112. And if the result of this checking is that using a filter is not required, then the decoder checks whether the referenced block portion 106 does not extend beyond the border of the predetermined picture region at all or at least not by more than a predetermined distance.
- the decoder checks whether the referenced block portion 106 is distanced from the border of the predetermined picture region by more than a filter kernel reach 1 14 of the filter kernel or at least by the reach 114 less the predetermined distance.
- the decoder determines that the usage of a filter having a filter kernel is required if the motion information comprises a sub-pel motion vector, or if the motion information indicates a bidirectional optical flow mode or a decoder side motion vector refinement mode.
- the above-mentioned criterion can also be held to be fulfilled, for example, if the referenced block portion is distanced from the border of the predetermined picture region by more than a maximum filter kernel reach or by more than the maximum filter kernel reach less a predetermined distance.
- the decoder can also derive any sample of the reference picture lying outside the predetermined picture region and involved in the inter-predicting the predetermined inter- predicted block in a manner unsubdivided using the referenced block portion by padding.
- the decoder can - if the criterion is not fulfilled - classify 128 the sub-blocks 122 into first 130 and second 132 sub-blocks according to whether or not the motion information which is signaled in the data stream, references in relation to the sub-blocks 130 and 132 one or more references sub-block portions which have a position relative to the border 108 of the predetermined picture region that fulfills a second predetermined criterion.
- the decoder inter-predicts 136 the one or more first sub-blocks using the one or more first referenced sub-block portions, and predicts 138 the one or more second sub-blocks using a substitute prediction.
- the second predetermined criterion can be held to be fulfilled if the inter-predicting 136 of the one or more first sub-blocks does not involve a filter kernel reaching out beyond the border of the predetermined picture region and if the one or more first referenced sub block portions 134 do not extend beyond the border of the predetermined picture region at all or not by more than a predetermined distance.
- the decoder can check whether the motion information 109 requires a usage of a filter having a filter kernel 112.
- the decoder also can check whether the first referenced sub-block portion 134 extends beyond the border 108 at all or by more than a predetermined distance. And the decoder can check whether the respective first referenced sub-block portion 134 is distanced from the border 108 of the predetermined picture region by more than the filter kernel reach 1 14 of the filter kernel 1 12 or more than the filter kernel reach 1 14 of the filter kernel 1 12 less a predetermined distance.
- the decoder then appoints the respective sub-block 122 a first sub-block, if the usage of the filter kernel is required and the referenced sub-block portion 134 is distanced from the border 108 by more than the filter kernel reach 114 or by more than the filter kernel reach 1 14 less the predetermined distance.
- the decoder also appoints the respective sub-block 122 a first sub-block, if the usage of the filter kernel is not required and the referenced sub-block portion 134 does not extend beyond the border 108 at all or by more than a predetermined distance.
- the decoder appoints the respective sub-block 122 a second sub-block, if the usage of the filter kernel is required and the referenced sub-block portion 134 is not distanced from the border 108 by more than the filter kernel reach 1 14 or by more than the filter kernel reach 1 14 less the predetermined distance.
- the decoder also appoints the respective sub-block 122 a second sub-block, if the usage of the filter kernel is not required and the referenced sub-block portion 134 does extend beyond the border 108 at all or by more than a predetermined distance.
- the decoder checks for each of the sub blocks 122, whether the respective first referenced sub-block portion 134 is distanced from the border of the predetermined picture region relative to the respective sub-block by more than a filter kernel reach of the filter kernel or by more than the filter kernel reach less a predetermined distance.
- the second predetermined criterion can also be held to be fulfilled if the first referenced sub-block portions 134 are distanced from the border 108 of the predetermined picture region 1 10 by more than a maximum filter kernel reach or by more than the maximum filter kernel reach less a predetermined distance.
- sample of the reference picture lying outside the predetermined picture region 1 10 which is involved in the inter-predicting of the first sub-blocks can be derived using the first referenced sub-block portions by padding.
- the prediction 138 of second sub-blocks 132 using the substitute prediction, as discussed above, can be performed by clipping 140a, 140b, for each of the second sub-blocks, a motion vector 142 which is comprised by the motion information 109. Thereby a clipped motion vector of a further motion information 109’ can be obtained which references, relative to the respective second sub-block 132, a third referenced sub-block portion 134’ in the reference picture.
- the third referenced sub-block portion 134’ does not extend beyond the border 108 of the predetermined picture region 110 at all or at least not by more than a predetermined distance. Then the second sub-block is inter-predicted using the third referenced sub-block portion 134'.
- the clipping 140a,b can be implemented as one of the following.
- the motion vector 142 comprised by the motion information can be clipped to an integer motion vector.
- the motion vector 142 comprised by the motion information 109 can be clipped so that the third referenced sub-block portion is distanced from the border 108 of the predetermined picture region 110 at least as far as a filter kernel reach 1 14 of the filter kernel 112 or at least as far as a filter kernel reach 114 of the filter kernel 1 12 less the predetermined distance.
- the motion vector 142 comprised by the motion information 109 can be clipped so that the position of the third referenced sub-block portion 134’ fulfills the second predetermined criterion.
- the subdividing 120 of the predetermined inter-predicted block 104 into sub-blocks 122 can be performed by sub-dividing the predetermined inter-predicted block 104 into a predetermined number of equally sized portions or by using one or more split lines determined based on a spatial relative location between the predetermined inter-predicted block 104 and the border 108 of the predetermined picture region 110.
- the motion information 109 can be derived 150 from the data stream 14 for the predetermined inter-predicted block 104 by decoding a merge indicator 160 from the data stream 14, wherein the merge indicator 160 indicates whether 152 the motion information 109 for the predetermined inter-predicted block 104 is to be determined based on a motion information candidate list 158 or not. If the merge indicator indicates that the motion information for the predetermined inter-predicted block is to be determined based on the motion information candidate list 158, a predetermined motion information candidate 156 out of the motion information candidate list 158 is determined on the basis of further motion information signaled in the data stream for previous inter-predicted blocks and indicated 154. Then the predetermined motion information candidate 156 is uses 161a as the motion information 109 or the motion information 109 is determined 161b using the predetermined motion information candidate 156.
- the decoder can decode one or more further motion parameters 164 from the data stream for the predetermined inter-predicted block, and determine 162 the motion information based on the one or more further motion parameters 164.
- the decoder can determine a motion vector predictor 170 by temporal or spatial prediction, and correct 176 the motion vector predictor 170 using a motion vector difference 172 comprised by the one or more further motion parameters 164 to obtain a motion vector 142 comprised by the motion information.
- the decoder can also derive the motion information 109 from the data stream for the predetermined inter-predicted block by determining a motion vector predictor 170’, 170 by temporal or spatial prediction, decoding a motion vector difference or offset 172, 180 for the predetermined inter-predicted block from the data stream, and correcting 182, 176 the motion vector predictor using the motion vector difference to obtain a motion vector 142 comprised by the motion information 109. And if the motion vector difference 172, 180 was set to zero, and if the position of the referenced block portion 106 does not fulfill the first predetermined criterion, the decoder can clip 182 the motion vector predictor 170’, 170.
- the decoder can optionally clip the motion vector predictor 170’, 170 to an integer motion vector predictor.
- the decoder can clip 182 the motion vector predictor 170’; 170 so that, if the motion vector difference 172, 180 was set to zero, the further referenced block portion is distanced from the border of the predetermined picture region at least as far as a filter kernel reach of the filter kernel, or at least as far as a filter kernel reach of the filter kernel less a predetermined distance.
- the decoder can clip the motion vector predictor 170’; 170 so that, if the motion vector difference 172, 180 was set to zero, the position of the referenced block portion 106 fulfills the first predetermined criterion.
- the decoder can check for each sub-block, if the motion vector difference 172 was set to zero and if the motion information signaled in the data stream would reference, relative to the respective sub-block 122, a third referenced sub-block portion 134 in the reference picture, a position of which relative to the border 108 of the predetermined picture region fulfills a second predetermined criterion.
- the decoder clips the motion vector predictor 170’; 170 to a sub-block individual version of the motion vector predictor so that a sub-block individual motion information which comprises a sub-block individual motion vector resulting from correcting the sub-block individual version of the motion vector predictor using the motion vector difference referenced a fourth referenced sub-block portion which fulfills the second predetermined criterion.
- the decoder inter-predicts the respective sub-block using the fifth referenced sub-block portion. If the position of the fifth referenced sub-block portion does not fulfill the second predetermined criterion, the decoder clips the sub-block individual motion vector so that the fifth referenced sub-block portion fulfills the second predetermined criterion, and inter-predicts the respective subblock using the fifth referenced sub-block portion.
- the decoder clips when the position of the referenced block portion 106 does not fulfill the first predetermined criterion, a motion vector of the motion information
- the decoder inter-predicts the predetermined inter-predicted block 104 in a manner unsubdivided using the further referenced block portion 106.
- the first predetermined criterion is held to be fulfilled if the referenced block portion 106 does not extend beyond the border 108 of a predetermined picture region 1 10 at all or by more than a predetermined distance and a distance of the referenced block portion 106 relative to the border is smaller than a reach of the filter kernel of an interpolation filter associated with a sub-pel granularity of the sub-pel motion vector, or smaller than the reach less the predetermined distance.
- the clipping of the motion vector that references outside of the region can then be performed on a sub-block basis.
- the first option is that the motion vectors that are applied to the block point at an integer sample position.
- a second option is that the motion vectors applied to the block point at a sub-pel sample position.
- a third option is that no distinction is made as whether the motion vectors point at a sub-pel sample position or integer sample position.
- two regions can be defined, namely an inside region and an outside region, which is separated by the border of the region. For all blocks within the inside region full-block processing can be carried out. For blocks which are at least partly in the outside region, sub-block processing is carried out.
- the first region is the inside region, but with a security margin.
- Other spans of the security margin are possible.
- the outside region is defined with a security margin.
- the security margin can for example span 4 samples at the right and bottom borders and 3 samples at the left and top border.
- the security margin can be the same for the inside and outside region, but also be defined differently.
- full-block processing can be carried out.
- sub-block processing is carried out for block at least party within the outside region and/or its security margin.
- the first mode is the merge mode.
- the merge mode no motion vector difference is added to the motion vector prediction.
- an offset can be indicated.
- the motion vector is a result of the motion vector predictor, to which the optional offset is added if an offset is used. The resulting motion vector is then used for each of the sub blocks independently.
- the respective motion vector for each of the sub-blocks is then clipped as required. It should be noted that this clipping is not a simple cut-off operation, but can involve rounding operations, for example if the motion vector resulting from the clipping alone would point to a sub-pel position that is not allowed.
- the second mode is called advanced motion vector prediction mode, sometimes also called regular mode.
- a prediction is derived either from neighboring blocks or blocks at a previously encoded frame.
- a motion vector difference is added to the predictor. Similar to the merge mode, the individual sub-blocks can be clipped separately by clipping the final motion vector, which is the result of the sum of the predictor and the motion vector difference.
- the motion vector predictor can first be clipped. The clipped predictor can then be added to the motion vector difference, and the result therefrom, i.e. the final vector, can then be clipped again.
- the motion vector predictor could be clipped block-wise.
- Block wise clipping means that contrary to the alternatives detailed above, the same motion vector predictor can be used for all sub-blocks, after the predictor has been clipped for the sub-block with highest required restriction, i.e. worst-case clipping. Then the motion vector difference would be applied to the clipped predictor and the final motion vector is then clipped sub-block wise.
- the third mode is called skip mode, which is also similar to the merge mode, but does not no result in a residual.
- the sub-block processing in the skip mode can be carried out similar to the processing described for the merge mode.
- CUs i.e. the blocks to be processed
- the skip mode discontinuities which can for example be a result of overlapping sub-blocks, cannot be corrected. For this the additional residual would be required.
- Another option would be to process the block similar to the processing described above for the advanced motion vector prediction mode or regular mode. That means no clipping on a sub-block basis. However, since no residual and no motion vector difference are added, in skip mode no sub-block processing is carried out.
- affine prediction which uses also sub-block, is to average the motion vectors of the corresponding 2x2 chroma values and apply the averaged motion vector to a 4x4 chroma block. This is done since processing of blocks smaller than 4x4 are undesirable due to complexity issues.
- Another solution is instead of having fixed sub-blocks of 4x4 luma and chroma, the subblocks are 8x8, luma and therefore 4x4 chroma. As a result, there is no necessity to use the averaged motion vectors for chroma.
- each block could be divided into a variable number of subblocks so that no fixed size is implied.
- the sub-block partitioning is dependent on the motion vectors and regions boundaries, and can be specified at a parameter set indicating the subblock size or the number of sub-blocks per block. This number can also be different for different block sizes.
- VVC Versatile Video Coding
- the current Versatile Video Coding (VVC) specification includes a mode that applies decoder side motion derivation. This means that for the merge mode, when no offset is signaled, the decoder searches at a small region close to the motion vector predictor for the best“match”.
- the sub-pels that are used for this decoder search use a sub-pel interpolation filter which is smaller than the 8-tap interpolation filter used for other modes. In particular, this filter is a bilinear filter.
- an 8-tap filter is used to compute the proper sub-pel values. Since the bilinear filter allows a less restrictive mode, where less restrictions would be required, in the present embodiment, a bilinear filter could be used instead of the 8-tap filter, which means that the sub-block clipping uses a different security margin and in addition the search that is done at the decoder side is limited by the region boundaries (or tile boundaries).
- BDOF is not allowed when the block is at the region boundaries or in the security margin if sub-pel interpolation is used as described in the prior art.
- the additional samples required for BDOF are always chosen as integer sample positions, the BDOF operation is only not allowed for cases where the block is at the region borders, i.e. tiles boundaries.
- BDOF is disabled for sub-block processing, i.e. BDOF is not used when affine motion vectors are in use.
- affine motion vectors can lead to overlapping subblocks or discontinuous sub-blocks, for which BDOF computation may not make much sense and instead just adds some complexity.
- the idea therefore is to have sub-block motion vector similarity based BDOF enabling/disabling, in other words, depending on whether sub-block motion vectors are similar to each other, BDOF is enabled or disabled.
- a further embodiment related to the discussed technique relates to switching sub-pel interpolation filters as required for all inter prediction modes. I.e., clipping (including a rounding operation - for sub-pel reference prohibition - happens for a given interpolation filter.
- clipping including a rounding operation - for sub-pel reference prohibition - happens for a given interpolation filter.
- such a filter is an 8-tap filter, which means that a margin of 3 or 4 samples from the boundary (depending on boundary) are to take into account for the sub-pel interpolation process at the motion vector redirection step.
- the sub-pel part of the vector is rounded.
- DMVR sub-pel interpolation filter
- the interpolation filter could be changed from the 8 tap filter to the bilinear case, or to a smaller filter that reduces the region where no sub-pel positions are allowed. Under some circumstances limited padding might be desirable for some tools. For instance, DMVR uses two sample padding for the referenced block.
- an initial motion vector is used to determine the initial reference position to start a search at.
- a number of samples equal to CUsize + 7 (in each picture dimension) is moved to the search buffer for inter-prediction of the current block.
- the motion vector refinement is determined, which is limited to an area of -2, +2 integer sample distance of the initial position, the motion refinement is added to the initial motion vector and a proper 8-tap filter is used.
- the modified position could require samples outside the CUsize + 7 samples taken for the search buffer and therefore, in order not to get more samples from the reference, the required samples (at most 2) are padded. This is a compromise on how many samples can be padded without incurring a big processing overhead. Therefore, all clipping and corresponding rounding operation described above and corresponding security margins (i.e. predetermined distance) can be relaxed to allow a limited amount of padding (of e.g. 2 samples).
- the clipping is done as per sub-block.
- Option 1 MVs applied to the block point at an integer sample position.
- Regions are defined inside the region and outside the region. For the first region, full-block processing is carried out. For the second region sub-block processing is carried out.
- Option 2 MVs applied to the block point at a sub-pel sample position.
- Regions are defined inside the region with a security margin (typically 4 samples from the right and bottom region borders and 3 from left and top borders) and outside the region with a security margin.
- a security margin typically 4 samples from the right and bottom region borders and 3 from left and top borders
- For the first region full- block processing is carried out.
- Option 3 No distinction is done as whether the MVs applied to the block points at a sub-pel sample position or integer sample position.
- Regions are defined inside the region with a security margin (typically 4 samples from the right and bottom region borders and 3 from left and top borders) and outside the region with a security margin.
- a security margin typically 4 samples from the right and bottom region borders and 3 from left and top borders
- For the first region full- block processing is carried out.
- Merge mode In the merge mode no MV difference is added to the MV prediction. At most an offset among a predefined number of possibilities can be indicated.
- the MV result of the predictor and potentially additional offset is used for each of the sub-block independently and clipping of the MV as required for each sub-block is performed.
- clipping is not simple a clip operation but could entail as well as rounding operation if the clipped MV points to a sub-pel position that is not allowed.
- Advanced MV prediction mode (or regular mode): In this mode a prediction is derived either from neighboring blocks or blocks at a previously encoded frame and a MV difference is added.
- the sub-blocks could be clipped separately either:
- the MV predictor could be clipped Block-wise, which means that contrary to what is described above in 2, the same MV predictor would be used for all sub-block, with the predictor being clipped for the worst-case (i.e. sub-block with highest required restriction). Then the MV difference would be applied to the clipped version of the predictor and then the final MV would be clipped sub-block wise.
- Skip mode Similar to merge but no residual
- the sub-block processing here could be carried out as described for merge. However, since no residual is added to Coding Units (CUs) encoded using the skip mode. Discontinuities (e.g. result of overlapping sub-blocks) cannot be corrected with the additional residual.
- CUs Coding Units
- One of the issues of using sub-blocks is that when a 4:2:0 format is used the sub-block of 4x4 lead to a 2x2 chroma values.
- One option that is currently used in affine prediction, which uses also sub-block, is to average the MVs of the corresponding 2x2 chroma values and applied the averaged MV to a 4x4 chroma block, since processing of blocks smaller than 4x4 are undesirable due to complexity issues.
- the sub-block processing depending on MVs and region/tile boundaries is only allowed for blocks bigger or equal than 8x8.
- Another embodiment would be to instead of having fixed sub-blocks of 4x4 (luma and chroma), the sub-blocks were of 8x8 (luma and therefore 4x4 chroma). This would result to not requiring using the average MVs for chroma.
- each block would be divided into a variable number of subblocks so that no fixed size is implied.
- the sub-block partitioning dependent on MVs and regions boundaries is specified at a parameter set indicated the sub-block size or the number of sub-blocks per block, which could be different for different block sizes.
- VVC Versatile Video Coding
- the sub-pels that are used for this decoder search use a sub-pel interpolation filter which is smaller than the 8-tap interpolation filter used for other modes. More concretely, it is a bilinear filter.
- an 8-tap filter is used to compute the proper sub-pel values. Since the bi-linear filter allows a less restrictive mode, where less restrictions would be required, in the present embodiment, a bilinear filter could be used instead of the 8-tap filter, which means that the sub-block clipping uses a different“security range” and in addition the search that is done at the decoder side is limited by the tile/region boundaries.
- sub-block prediction BDOF is not allowed when the Block is at the tiles boundaries or in the security margin if sub-pel interpolation is used as described in the above-mentioned previous EP application.
- additional samples required for BDOF are always chosen as integer sample positions, the BDOF operation is only not allowed for cases where the Block is at the tiles boundaries.
- BDOF is disabled in general for sub-block processing, i.e. BDOF is not used when affine MVs are in use.
- affine MVs can lead to overlapping sub-blocks, discontinuous sub-blocks, for which BDOF computation may not make much sense and just add some complexity.
- the idea therefore is to have Sub-block MV similarity based BDOF enabling/disabling.
- a further embodiment related to the discussed technique relates to switching sub-pel interpolation filters as required for all inter prediction modes. I.e., in the current described ideas and as described in the above-mentioned previous EP application, clipping (including a rounding operation - sub-pel prohibition) happens for a given interpolation filter. In HEVC or WC such a filter is an 8-tap filter, which means that a margin of 3 or 4 samples from the boundary (depending on boundary) are to take into account for the sub-pel interpolation process at the MV redirection step.
- the sub-pel part of the vector is rounded. Since, as discussed above there is a mode (i.e. DMVR) that uses a different sub-pel interpolation filter, instead of rounding the MVs to integer if they point to the area within the margin of 3 or 4 samples from the boundary, the interpolation filter could be changed from the 8 tap filter to the bilinear case, or to a smaller filter that reduces the are within no sub-pel positions are allowed.
- DMVR sub-pel interpolation filter
- DMVR uses 2 sample padding for the referenced block. More concretely, an initial motion vector is used to determine the initial reference position to start a search at. A number of samples equal to CUsize + 7 (in each picture dimension) is moved to the search buffer for inter-prediction of the current block. Once the Motion vector refinement is determined, which is limited to an area of -2, +2 integer sample distance of the initial position, the motion refinement is added to the initial motion vector and proper 8-tap filter is used. The modified position could require samples outside the CUsize + 7 samples taken for the search buffer and therefore, in order not to get more samples from the reference, the required samples (at most 2) are padded.
- the concept can be implemented by methods for block-based decoding and/or block-based encoding according to embodiments of the present invention. These methods are based on the same considerations as the above-described decoder and encoder. However, it should be noted that the methods can be supplemented by any of the features, functionalities and details described herein, also with respect to the decoder and/or encoder. Moreover, the methods can be supplemented by the features, functionalities, and details of the decoder and/or encoder, both individually and taken in combination.
- the concept can be used to produce an encoded data stream according to embodiments of the present invention.
- the data stream can also be supplemented by the features, functionalities, and details of the decoder and/or encoder, both individually and taken in combination.
- aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a device or a part thereof corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding apparatus or part of an apparatus or item or feature of a corresponding apparatus.
- Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
- embodiments of the invention can be implemented in hardware or in software.
- the implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
- Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
- embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.
- the program code may for example be stored on a machine-readable carrier.
- inventions comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
- an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
- a further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein.
- the data carrier, the digital storage medium or the recorded medium are typically tangible and/or nontransitionary.
- a further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.
- the data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
- a further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
- a processing means for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
- a further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
- a further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver.
- the receiver may, for example, be a computer, a mobile device, a memory device or the like.
- the apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
- a programmable logic device for example a field programmable gate array
- a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein.
- the methods are preferably performed by any hardware apparatus.
- the apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
- the apparatus described herein, or any components of the apparatus described herein may be implemented at least partially in hardware and/or in software.
- the methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP19160422 | 2019-03-01 | ||
| PCT/EP2020/055273 WO2020178170A1 (en) | 2019-03-01 | 2020-02-28 | Advanced independent region boundary handling |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP3932073A1 true EP3932073A1 (en) | 2022-01-05 |
Family
ID=65729087
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP20706330.6A Pending EP3932073A1 (en) | 2019-03-01 | 2020-02-28 | Advanced independent region boundary handling |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP3932073A1 (en) |
| WO (1) | WO2020178170A1 (en) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7345573B2 (en) * | 2019-06-24 | 2023-09-15 | 鴻穎創新有限公司 | Apparatus and method for coding moving image data |
| US12206861B2 (en) | 2022-01-12 | 2025-01-21 | Tencent America LLC | Motion vector restriction for out-of-frame boundary conditions |
| CN115941971A (en) * | 2022-11-21 | 2023-04-07 | 腾讯科技(深圳)有限公司 | Video processing method, device and computer equipment, storage medium, program product |
-
2020
- 2020-02-28 EP EP20706330.6A patent/EP3932073A1/en active Pending
- 2020-02-28 WO PCT/EP2020/055273 patent/WO2020178170A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2020178170A1 (en) | 2020-09-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12160568B2 (en) | Apparatus for selecting an intra-prediction mode for padding | |
| CN114009028B (en) | Encoder, decoder, method and computer program with improved transform-based scaling | |
| US12501053B2 (en) | Inter-prediction concept using tile-independency constraints | |
| US20160073107A1 (en) | Method and apparatus for video encoding/decoding using intra prediction | |
| US12095989B2 (en) | Region based intra block copy | |
| EP3939272A1 (en) | Encoders, decoders, methods, and video bit streams, and computer programs for hybrid video coding | |
| EP3994879A1 (en) | Encoder, decoder, methods and computer programs for an improved lossless compression | |
| WO2020178170A1 (en) | Advanced independent region boundary handling | |
| US11553207B2 (en) | Contour mode prediction | |
| US20250071292A1 (en) | Multi-hypothesis prediction | |
| WO2020234376A1 (en) | Encoder and decoder, encoding method and decoding method for drift-free padding and hashing of independent coding regions | |
| US20250317564A1 (en) | Usage of coded subblock flags along with transform switching including a transform skip mode | |
| HK40114578A (en) | Encoder, decoder, methods and computer programs with an improved transform based scaling | |
| HK40067202B (en) | Encoder, decoder, methods and computer programs with an improved transform based scaling | |
| HK40067202A (en) | Encoder, decoder, methods and computer programs with an improved transform based scaling |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20210831 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| RAP3 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V. |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20240502 |