WO2025200081A1 - Method and apparatus for template-based intra mode derivation and encoder/decoder including the same - Google Patents
Method and apparatus for template-based intra mode derivation and encoder/decoder including the sameInfo
- Publication number
- WO2025200081A1 WO2025200081A1 PCT/CN2024/091020 CN2024091020W WO2025200081A1 WO 2025200081 A1 WO2025200081 A1 WO 2025200081A1 CN 2024091020 W CN2024091020 W CN 2024091020W WO 2025200081 A1 WO2025200081 A1 WO 2025200081A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- block
- coding
- intra
- parametrizations
- candidate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/11—Selection of coding mode or of prediction mode among a plurality of spatial predictive coding modes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/146—Data rate or code amount at the encoder output
- H04N19/147—Data rate or code amount at the encoder output according to rate distortion criteria
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/157—Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
- H04N19/159—Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/46—Embedding additional information in the video signal during the compression process
- H04N19/463—Embedding additional information in the video signal during the compression process by compressing encoding parameters before transmission
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/593—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
Definitions
- the present disclosure generally relates to the field of encoding/decoding pictures, images or videos, and embodiments of the present disclosure concern improvements regarding the Intra prediction, more specifically, improvements regarding a template-based intra mode derivation, TIMD, process. More specific embodiments of the present disclosure relate to a merge mode for template-based intra mode derivation (TIMD-Merge) .
- TIMD-Merge merge mode for template-based intra mode derivation
- the encoding and decoding of a picture, an image or a video is performed in accordance with a certain standard, for example, in accordance with the advanced video coding, AVC standard (see reference [1] ) , the high efficiency video coding, HEVC, standard (see reference [2] ) or the versatile video coding, VVC, standard (see reference [3] ) .
- AVC standard see reference [1]
- HEVC high efficiency video coding
- VVC versatile video coding
- the encoder 100 comprises a prediction stage 124 for determining the prediction signal 112, which includes a de-quantizer or inverse quantizer 126, an inverse transformer 128, a combiner 130, an in-loop filter 134, a motion estimator 138, an intra/inter mode selector 140, an inter predictor 142 and an intra predictor 144.
- motion compensation and motion estimation are performed by the motion estimator 138 which searches, in one or more reference pictures provided by the picture buffer 136 and used to predictively code the current picture, a CU that is a good predictor of a current CU.
- a good predictor of a current CU is a predictor which is similar to the current CU, i.e., the distortion between the two CUs is low or below a certain threshold.
- the motion estimation may also account for the rate cost of signaling the predictor to optimize a rate-distortion tradeoff.
- the output of the motion estimation step is one or more motion vectors and reference indices associated with the current CU.
- the encoder 100 may skip the transform stage 116 and apply the quantization directly to the non-transformed residual signal 110 in a so-called transform-skip coding mode.
- the encoder decodes the CU and reconstructs it so as to obtain the reconstructed signal 132 that may serve as a reference data for predicting future CUs or blocks to encode.
- the quantized transform coefficients 110’ are de-quantized and inverse transformed leading to a decoded prediction CU or block 110”, and the decoded prediction residuals and the predicted block are then combined at 130, typically summed, so as to provide the reconstructed block or CU 132.
- the in-loop filters 134 are applied to the reconstructed picture to reduce compensation artifacts.
- a deblocking filter For example, a deblocking filter, a sample adaptive offset, SAO, filter, and an adaptive loop filter, ALF, may be applied to reduce encoding artifacts.
- the filtered picture is stored in the buffer 136, also referred to as the decoded picture buffer, DPB, so that it may be used as a reference picture for coding subsequent pictures.
- Fig. 2 is a block diagram of a video decoder 150 for predictively decoding from a data stream 152 a picture or video which is provided at an output 154 of the decoder 150.
- the decoder 150 includes an entropy decoder 156, a partitioning block 158, an inverse quantizer 160, an inverse transformer 162, a combiner 164, an in-loop filter 166, optionally a post-decoding processor 168, and a prediction module 170.
- the prediction module 170 includes a decoded picture buffer 180 a motion compensator 182 an intra predictor 184.
- An encoded picture of a video sequence is decompressed and decoded by the decoder 150 as follows.
- the input bit stream 152 is entropy decoded by the decoder 156 which provides, for example, the block partitioning information, the coding mode for each coding unit, the transform coefficients contained in each transform block, prediction information, like intra prediction mode, motion vectors, reference picture indices, and other coding information.
- the block partitioning information indicates how the picture is partitioned and the decoder 150 may divide the input picture into coding tree units, CTUs, typically of a size of 64x64 or 128x128 pixels and divide each CTU into rectangular or square coding units, CUs, according to the decoded partitioning information.
- CABAC Context-Adaptive Binary Arithmetic Coding
- the present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a first aspect by testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set and selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors even, which steps belong to a first tool, even if the corresponding first tool is not used, and taking into account the first set in establishing a merge list of candidate coding parametrizations for a second tool, namely by comparing, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, and modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the
- the first aspect avoids that the prediction parameters that are derived from the merge list in the second tool, are similar to those that the first tool such as regular TIMD would provide.
- the present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a second aspect by construing the merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between a respective second block on the one hand and the first block on the other hand.
- a size or aspect-ratio comparison between a respective second block on the one hand and the first block on the other hand By this measure, it is possible to suppress, or lower the degree of, the contribution of second blocks’ coding parametrizations to the merge list, whose size or aspect-ratio does not equal, or is not similar enough, to the first block.
- the second aspect avoids blindly including all neighbors (second blocks) in a merge list such as the TIMD-Merge list and, thus, avoids, or renders it less likely that the current block to code (first block) would derive its prediction parameters from an irrelevant candidate block (second block) .
- the present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a third aspect by populating the test set of at least one or more of the intra prediction modes used in a first tool using the merge list of candidate coding parametrizations used for the second tool so that the effective first tool tests intra prediction modes which are more likely to be efficient.
- the second aspect avoids blindly including all neighbors (second blocks) in a merge list such as the TIMD-Merge list and, thus, avoids, or renders it less likely that the current block to code (first block) would derive its prediction parameters from an irrelevant candidate block (second block) .
- the present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a fourth aspect by reconstructing the first block based on the selected coding parametrization, selected from the merge list, by selecting a predetermined residual transform domain of a prediction residual for the first block out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems.
- transform domain signaling is not necessary with the transform domain implicitly chosen being, nevertheless, efficient at pretty high likelihood.
- the present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a fifth aspect by forming a first list of horizontal residual transforms based on the second blocks in the neighborhood of a first block, and forming a second list of vertical residual transforms based on the second blocks, with performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream.
- Fig. 1 is a block diagram of a conventional video encoder
- Fig. 2 is a block diagram of a conventional video decoder
- Fig. 3 shows neighboring reconstructed samples used for DIMD chroma mode
- Fig. 4 shows a proposed scheme using non-adjacent spatial neighboring candidates for DIMD merge
- Fig. 5 shows above and left templates used for TIMD mode derivation
- Fig. 6 shows an example of TIMD merge list and example information of each candidate within the list
- Fig. 7 shows an example of adjacent neighbors in TIMD merge mode
- Fig. 8 shows a division method for angular modes
- Fig. 9 shows a table of modified weights used for angular modes
- Fig. 10 shows a syntax design for a fusion of chroma intra prediction modes
- Fig. 11 shows a diversification of modes in the TIMD merge list, by using the modes derived from TIMD legacy
- Fig. 12 shows a two-step eligibility based on coding mode and block size
- Fig. 13 shows an extending of the MPM list of TIMD legacy by including modes from TIMD merge list
- Fig. 14 shows an inheriting horizontal and vertical transform types from the same TIMD merge candidate (CU 4) that provides the intra prediction modes;
- Fig. 15 shows an implicit transform type inheritance allowing combinations of transform types that are not explicitly provided by ECM.
- Fig. 16 shows a block diagram illustrating an electronic device 900 according to embodiments of the present disclosure.
- the term "and/or" is intended to cover all possible combinations and sub-combinations of the listed elements, including any one of the listed elements alone, any sub-combination, or all of the elements, and without necessarily excluding additional elements.
- the phrase "at least one of... or..." is intended to cover any one or more of the listed elements, including any one of the listed elements alone, any sub-combination, or all of the elements, without necessarily excluding any additional elements, and without necessarily requiring all of the elements.
- coding refers to "encoding” or to “decoding” as becomes apparent from the context of the described embodiments.
- coder refers to " an encoder” or to “adecoder” .
- DIMD decoder side intra mode derivation
- the weight for each of the five derived modes is modified if the one the above or left histogram magnitudes is twice larger than the other one.
- the weights are location dependent and computed as follows:
- wDimd i is the unmodified uniform weight of the DIMD selected as in JVET-O0449
- ⁇ i is pre-defined and set to 10.
- Derived intra modes are included into the primary list of intra most probable modes (MPM) , so the DIMD process is performed before the MPM list is constructed.
- the primary derived intra mode of a DIMD block is stored with a block and is used for MPM list construction of the neighboring blocks.
- region of neighboring reconstructed samples used for computing the histogram of gradients is modified compared to JVET-O0449 method, depending on reconstructed samples availability.
- the region of decoded reference samples of current WxH luma CB is extended towards the above-right side if available, up to W additional columns. It is extended towards the bottom-left side if available, up to H additional rows.
- the DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the neighboring reconstructed Y, Cb and Cr samples in the second neighboring row and column as shown in Figure 3. Specifically, a horizontal gradient and a vertical gradient are calculated for each collocated reconstructed luma sample of the current chroma block, as well as the reconstructed Cb and Cr samples, to build a HoG. Then the intra prediction mode with the largest histogram amplitude values is used for performing chroma intra prediction of the current chroma block.
- Figure 3 shows neighboring reconstructed samples used for DIMD chroma mode.
- Fig. 3a shows the neighboring collocated reconstructed Y samples
- Fig. 3b shows the neighboring reconstructed Cb samples
- Fig. 3c shows the neighboring reconstructed Cr samples.
- the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode.
- a CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied.
- DIMD merge mode is proposed in JVET-AF0120.
- DIMD Merge the DIMD information extracted from neighbouring blocks is used to compute the intra prediction for the current block.
- a new Merged Histogram of Gradients (MHoG) is computed for the current block based on the HoGs of neighbouring blocks. Only neighbouring blocks encoded with DIMD or with DIMD Merge are considered.
- the MHoG is used to compute intra-prediction modes and weights, as in conventional DIMD.
- the directional modes and their weights corresponding to the five highest amplitudes in the MHoG are selected, and the corresponding predictors are blended as in conventional DIMD.
- JVET-AF0106 A method of using non-adjacent spatial candidates for DIMD merge is proposed in JVET-AF0106. It is proposed to use non-adjacent spatial candidates for DIMD merge candidates list construction. As shown in Figure 4, the distances between non-adjacent candidates and current block are defined based on the width and height of current coding block. When using DIMD Merge, the DIMD information extracted from neighbouring blocks and non-adjacent spatial blocks is used to compute the intra prediction for the current block.
- Figure 4 shows the proposed scheme using non-adjacent spatial neighboring candidates for DIMD merge.
- costMode2 ⁇ 2*costMode1.
- the division operations are conducted using the same lookup table (LUT) based integerization scheme used by the CCLM.
- LUT lookup table
- Figure 5 illustrates the above and left templates used for TIMD mode derivation.
- TIMD merge list with associated information of each TIMD candidate is shown in below.
- Figure 6 shows an example of TIMD merge list and example information of each candidate within the list.
- Figure 7 illustrates an example of adjacent neighbors in TIMD merge mode and Figure 4 illustrates an example of non-adjacent neighbors in TIMD merge mode.
- the encoder may test one or more of the TIMD merge candidates in the list and the best performing candidate (and associated information) according to the rate-distortion optimization (RDO) mechanism may be selected and used for prediction of the current block.
- RDO rate-distortion optimization
- Intra prediction fusion derives predicted samples as a weighted combination of multiple predictors generated from different reference lines. In this process multiple intra predictors are generated and then fused by weighted averaging. The process of deriving the predictors to be used in the fusion process is described as follows:
- the number of predictors selected for a weighted average is increased from 3 to 6.
- Intra prediction fusion method is applied to luma blocks when angular intra mode has non-integer slope (required reference samples interpolation) and the block size is greater than 16, it is used with MRL and not applied for ISP coded blocks.
- PDPC is applied for the intra prediction mode using the closest to the current block reference line.
- CIIP Intra prediction signal predicted using CIIP-TM merge candidate and an intra prediction signal predicted using TIMD derived intra prediction mode.
- the method is only applied to coding blocks with an area less than or equal to 1024.
- CIIP-TM a CIIP-TM merge candidate list is built for the CIIP-TM mode.
- the merge candidates are refined by template matching.
- the CIIP-TM merge candidates are also reordered by the ARMC method as regular merge candidates.
- the maximum number of CIIP-TM merge candidates is equal to two.
- chroma intra prediction modes there is also fusion of chroma intra prediction modes.
- two chroma intra prediction signals can be fused together.
- One of the two chroma intra prediction signals is predicted using one of the DM mode, DIMD chroma mode and the four default modes (non-LM mode) .
- the other chroma intra prediction signal is predicted using cross-component linear prediction modes (LM mode) . Two different methods are supported.
- pred0 (i, j) is the predictor obtained by applying the non-LM mode
- pred1 (i, j) is the predictor obtained by applying the LM mode
- pred C (i, j) is the final predictor of the current chroma block.
- Two template costs are calculated by fusing the angular chroma prediction with MM-CCLM or MM-CCCM, respectively, and the one of the two CCPs which provides a smaller template cost is utilized to derive pred1.
- the LM mode can be either MMLM or CCLM mode
- pred0 (i, j) is the predictor obtained by applying the non-LM mode
- rec′ L (i, j) is the set of downsampled reconstructed luma samples at co-located positions
- pred C (i, j) is the final predictor of the current chroma block.
- ⁇ is a fixed value and is set equal to 512 for 10-bit content.
- the three weights, ⁇ 0 , ⁇ 1 and ⁇ 2 are derived from the adjacent luma and chroma samples using the same LDL derivation method as in CCCM.
- the non-LM mode can be DM mode, DIMD chroma mode and the four default modes.
- DIMD chroma mode is allowed to be fused with LM modes.
- MTS Multiple Transform Sets
- MTS Multiple Transform Sets
- MTS is a tool that allows different transform types to be applied within a CU. This means that instead of using a single transform type for the entire CU, different CUs can be encoded using different transform types based on their characteristics. Explicit MTS enables this by allowing the encoder to signal which transform type to use for each part of the CU.
- the MTS index is a parameter that determines which transform type is applied to different horizontal and vertical partitions within the CU. This index is signaled along with other coding parameters and is used by the decoder to reconstruct the original signal.
- the encoder When encoding a CU, the encoder tries different MTS indexes for the CU, and for each tested MTS index, it translates the index into a pair of horizontal and vertical transform types. Finally, the index that represents the best transform pair is signaled in the bitstream.
- a first Current TIMD-Merge can be sub-optimal in the following cases:
- Constraints are added when including neighboring candidates that disallow some blocks. These constraints are based on characteristics of current block to be coded (e.g. block size)
- the embodiments relate to the first aspect of diversification of modes in TIMD legacy and TIMD merge. Given an arbitrary metric of similarity, one can determine whether intra modes of TIMD merge candidates are similar to those of TIMD legacy. If so, an additional process is applied to diversify the superset of modes provided by TIMD legacy and TIMD merge.
- the similarity is measure between two blocks: 1) current block if it were coded as TIMD legacy at one hand, and 2) a generic TIMD block that is contributing to TIMD merge list construction of current block, on the other hand.
- the similarity measurement step involves defining a similarity metric which can be computed on different variables.
- the intra prediction mode index is the principle variable to be used for this purpose. Furthermore, one can take into account also the flag determining the use wide-angle mode. Since intra modes of each TIMD block are ordered based on their weights, one can also take into account the order modes. For example, if the intra prediction modes indexes between the TIMD legacy algorithm and a generic TIMD block in the merge list are identical, however, their fusion weights have put them in different orders in the two blocks (e.g. ⁇ 18, 34> and ⁇ 34, 18>) , one might (or might not) consider that the two block similar. Finally, the fusion weight itself can be used in the similarity measurement. For example, if both index and order of the modes are identical, but their fusion weights are different in the two blocks, one might (or might not) consider them similar.
- a threshold when comparing a given variable between two blocks. Let’s say when comparing intra prediction mode index (i.e. variable) of two TIMD blocks, one encounters the modes ⁇ 18, 34> for the first block and ⁇ 19, 37> for the second blocks. At this point with a threshold of 2 or greater, these two blocks are considered similar, since the maximum difference between each element of the modes pair is smaller than or equal to 2.
- the mode distinction step diversifies them be modifying the way the generic TIMD block is added to the merge list. There are several ways to modify the above process. Here are some examples.
- Figure 11 shows a diversification of modes in the TIMD merge list, by using the modes derived from TIMD legacy.
- only the intra prediction indexes are compared to determine whether or not the merge list should be diversified.
- prediction information such as blending weights, wide-angle, order of prediction modes etc. are also used to determine whether or not the merge list should be diversified.
- the intra prediction mode indexes are similar but their order is different, then they are not considered as similar.
- a threshold larger than 0 is used to apply a certain level of flexibility. This means that even if the indexes are not identical, but they have a difference less than the defined threshold, then they are considered similar.
- the embodiments described next relate to the second aspect of block size based eligibility of TIMD-Merge candidates.
- the TIMD merge algorithm of the state-of-the-art does take into account some block characteristics (e.g. block size) of neighboring generic TIMD blocks when adding them to the TIMD merge list of current block.
- block characteristics e.g. block size
- the second aspect involves, according to the present embodiment, identifying some generic TIMD blocks in the neighborhood non-eligible, due to the difference between their block size and size of the current block.
- the size-eligibility algorithm compares current block to a generic TIMD block, by taking into account some sort of size-based metric.
- Figure 12 shows an example of using a metric based on block size in order to exclude a neighboring CU candidate.
- neighbor CU 1 has been excluded from the list, since its block size (i.e. 64 ⁇ 32) is different from the block size of current CU (i.e. 32 ⁇ 16) .
- Figure 12 shows two-step eligibility based on coding mode and block size.
- the size metric is a pair composed of the width and the height of the CU.
- Figure 13 shows an example of using modes from the TIMD-merge list to extend the MPM list of TIMD legacy algorithm.
- a current block is to be coded by TIMD legacy algorithm, which includes an MPM list construction stage.
- a TIMD-merge list is also computed on the current block to provide a new source of prediction modes.
- the two lists i.e. MPM and TIMD-merge
- the extended MPM list has a size (i.e. M) larger than its original size in TIMD legacy (e.g. N) .
- Figure 13 illustrates an extending of the MPM list of TIMD legacy by including modes from TIMD merge list.
- a condition is used to determine whether the TIMD merge modes should be included in the MPM list.
- the condition is based on the presence of other TIMD merge blocks in the neighborhood around the current CU. For instance, if there is at least on TIMD merge block in the neighborhood of the current block, then the TIMD merge list modes of current block are included in the MPM list of current block for TIMD legacy algorithm.
- the fourth aspect pertaining to implicit transform type inheritance.
- Current TIMD merge algorithm in the ECM uses regular Multiple Transform Sets (MTS) to determine the separable transform type for horizontal and vertical residual transformation. This includes an encoder-side search for the best index and then signaling it at the block level.
- MTS Multiple Transform Sets
- a method is proposed to derive the transform type along with prediction information (intra modes, wide-angle etc. ) from one of the candidates in the TIMD merge list.
- the encoder no longer needs to exhaustively search for the optimal transform index, which makes the encoding faster.
- the transform index no longer needs to be signaled, which can save bitrate and potentially increase the compression efficiency.
- the transform derivation can possibly result in a combination of horizontal and vertical transform types that are not presented by either of available transform indexes. This can potentially result in better compression efficiency.
- a merge list of transform candidates that is associated to the TIMD merge list and is filled accordingly.
- This list can be based on the same neighboring CUs that contribute to the construction of the TIMD merge list.
- a pair of its transform types i.e. horizontal and vertical
- the transform pair of the best candidate is also inherited to be used for the current TIMD merge CU.
- a TIMD merge candidate is identified in the neighborhood, its horizontal and vertical transforms are considered separately to construct two merge lists of transform candidates, namely horizontal transform merge list and vertical transform merge list.
- the particular step needed for this approach is to associate a cost to each transform type when adding them to the transform merge lists. To this end, one can use the same template-based cost that is computed for the TIMD merge candidate.
- the final transform types of current blocks can be inherited separately from each list. For example, one can take the first element of each transform merge list for horizontal and vertical transform type of current CU.
- Figure 14 shows an example of inheriting horizontal and vertical transform types from a neighbor CU.
- first neighbor CU 4 is selected by the TIMD merge algorithm to provide intra prediction modes.
- the second stage also inherits both transform types from CU 4 to be used for current TIMD merge block.
- Figure 14 illustrates an inheriting of horizontal and vertical transform types from the same TIMD merge candidate (CU 4) that provides the intra prediction modes.
- Figure 15 shows an example of inheriting a pair of transform type that is not explicitly provided for current CU.
- the neighboring candidate CU that has been selected from the TIMD merge of current block has a smaller block size. Therefore, the set of possible transform type pairs that are allowed by ECM is different from current CU that is larger.
- current CU inherits its transform type information along with prediction information from the selected candidate, then it will end up with a pair of transform types (i.e. DCT5 and DCT7 in this simplified example) that are not typically allowed by ECM for current block, due to the imitations of explicit transform type signaling.
- Figure 15 illustrates an implicit transform type inheritance allowing combinations of transform types that are not explicitly provided by ECM.
- two separate merge lists are considered for horizontal and vertical transform types, which are both detached from the construction of the TIMD merge list.
- every time a new TIMD merge CU is identified its horizontal and vertical transform types separately contribute to their corresponding transform merge lists.
- the transform inheritance process for current TIMD merge CU will consist of separately selecting one transform type from each transform merge list. For example, one can take the first element in each merge list.
- the testing the test set of intra prediction modes comprises determining for each intra prediction mode of the test set a prediction error by applying the respective intra prediction mode to a first previously decoded neighboring portion of the picture, neighboring the first block so as to predict a second previously decoded portion of the picture, neighboring the first block and the first previously decoded neighboring portion, and measuring the prediction error with respect to the second portion.
- the method further comprises, if the selected intra-coding tool is the second tool, decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and reconstructing the first block based on the selected coding parametrization.
- the comparing is performed so that the similarity depends on one or more of mode indices of the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively, fusion weights and/or ranks associated with the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively; wide-angle flags associated with the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively.
- the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities comprises thresholding the similarities and not performing the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes in case of the similarities not exceeding a predetermined threshold.
- the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities comprises one or more of excluding candidate coding parametrizations from the merge list of candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, reducing a weight of candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, at which same contribute to the merge list of candidate coding parametrizations, and modifying candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, with respect to the one or more candidate intra prediction modes.
- An embodiment relates to a method of block-based encoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors in the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and encoding one or more tool-selection syntax elements for the first block into a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encoding the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication
- An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and a decoding module configured to decode one or more tool-selection syntax elements for the first block from a data stream, wherein the decoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstruct the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first
- An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a testing module configured to a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors in the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and an encoding module configured to encode one or more tool-selection syntax elements for the first block into a data stream, wherein the encoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encode the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block
- An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
- An embodiment relates to a method of block-based decoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and reconstructing the first block using the merge list of candidate coding parametrizations, wherein the construing a merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- the construing the merge list of candidate coding parametrizations comprises populating the merge list of candidate coding parametrizations with coding parametrizations of one or more blocks of the second blocks in the neighborhood, whose size or aspect ratio fulfills a predetermined criterion with respect to a size or aspect ratio of the first block, and excluding coding parametrizations of one or more blocks of the second blocks in the neighborhood, whose size or aspect ratio does not fulfill the predetermined criterion with respect to the size or aspect ratio of the first block.
- whether the predetermined criterion is fulfilled depends on one or more of equality in number of pixels equality in width and equality in height equality in aspect ratio and equality in number of pixels, an absolute difference, or one minus a ratio in number of pixels, falling below a first threshold, an absolute difference, or one minus a ratio in width, falling below a second threshold; an absolute difference, or one minus a ratio in height, falling below a third threshold.
- An embodiment relates to a method of block-based encoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and encoding the first block using the merge list of candidate coding parametrizations, wherein the construing a merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and a reconstructing module configured to reconstruct the first block using the merge list of candidate coding parametrizations, wherein the construing module is configured to construe the merge list of candidate coding parametrizations depending on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
- An embodiment relates to a method of block-based decoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, decoding one or more tool-selection syntax elements for the first block from a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstructing the first block using the first set of one or more intra prediction modes,
- the testing the test set of intra prediction modes comprises determining for each intra prediction mode of the test set a prediction error by applying the respective intra prediction mode to a first previously decoded neighboring portion of the picture, neighboring a first block so as to predict a second previously decoded portion of the picture, neighboring the first block and the first previously decoded neighboring portion, and measuring the prediction error with respect to the second portion.
- the method further comprises, if the selected intra-coding tool is the second tool, decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and reconstructing the first block based on the selected coding parametrization.
- the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations by checking a similarity between the candidate coding parametrizations in the merge list of candidate coding parametrizations one the one hand and the list of most probable intra prediction modes, on the other hand, and adding one or more of the candidate coding parametrizations in the merge list of candidate coding parametrizations to the test set of at least one or more of the intra prediction modes if the similarity falls below a certain threshold.
- the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations by checking a number of second blocks in the neighborhood of the first block, or of candidate coding parametrizations in the merge list of candidate coding parametrizations, and adding one or more of the candidate coding parametrizations in the merge list of candidate coding parametrizations to the test set of at least one or more of the intra prediction modes if the number exceeds a certain threshold.
- An embodiment relates to a method of block-based encoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, encoding one or more tool-selection syntax elements for the first block into a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encoding the first block using the first set of one or more intra
- An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and a construing module configured to construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, a decoding module configured to decode one or more tool-selection syntax elements for the first block from a data stream, where the decoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the
- An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and a construing module configured to construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, an encoding module configured to encode one or more tool-selection syntax elements for the first block into a data stream, wherein encoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for
- An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
- An embodiment relates to a method of block-based decoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and reconstructing the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected
- the method further comprises decoding one or more tool-selection syntax elements for the first block from a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstructing the first block by intra-predicting the first block to obtain a predictor for the first block, decoding a transform domain indictor for the first block, from the data stream, selecting, using the transform domain pointer, a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, and decoding a prediction residual for the first block from the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual, and if the selected intra-coding tool is a second tool, reconstructing the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting
- the selecting, using the transform domain pointer, the predetermined residual transform domain out of the plurality of residual transforms domains supported by the decoder takes place by forming a list of residual transforms domains so as to comprise a proper subset of the residual transforms domains of the plurality of residual transforms domains, and applying the transform domain pointer onto the list of residual transforms domains so as to select the predetermined residual transform domain, and wherein the list of residual transforms domains does not comprise the residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems.
- An embodiment relates to a method of block-based encoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and encoding a merge candidate indicator into the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and encoding the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the encoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for encoding a second block from
- An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and a decoding module configured to decode a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and a reconstructing module configured to reconstruct the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the
- An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and an encoding module configured to encode a merge candidate indicator into the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and an encoding module configured to encode the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by
- An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
- An embodiment relates to a method of block-based decoding of a picture, the method comprising: forming a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, forming a second list of vertical residual transforms based on the second blocks, performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, decoding a prediction residual for the first block from the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, and correcting a predictor for the first block using the prediction residual.
- the method further comprises decoding from the data stream a first pointer and a second pointer for the first block, and performing the first selection based on the first pointer and the second selection based on the second pointer.
- the method further comprises testing different combinations of a horizontal and a vertical residual transform with respect to rate-distortion in an already decoded neighbourhood of the first block to obtain rate-distortion values for the different combination; and performing the first selection and the second selection based on the rate-distortion values.
- An embodiment relates to a method of block-based encoding of a picture, the method comprising: forming a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, forming a second list of vertical residual transforms based on the second blocks, performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, encoding a prediction residual for the first block into the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, and correcting a predictor for the first block using the prediction residual.
- An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a forming module configured to form a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, a forming module configured to form a second list of vertical residual transforms based on the second blocks, a performing module configured to perform a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, a decoding module configured to decode a prediction residual for the first block from the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected
- An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a forming module configured to form a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, a forming module configured to form a second list of vertical residual transforms based on the second blocks, a performing module configured to perform a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, an encoding module configured to encode a prediction residual for the first block into the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected
- An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
- aspects of the disclosed concept have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or a device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
- the device 900 includes a computing unit 901 to perform various appropriate actions and processes according to computer program instructions stored in a read only memory (ROM) 902, or loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data for the operation of the storage device 900 can also be stored.
- the computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904.
- An input/output (I/O) interface 905 is also connected to the bus 904.
- Components in the device 900 are connected to the I/O interface 905, including: an input unit 906, such as a keyboard, a mouse; an output unit 907, such as various types of displays, speakers; a storage unit 908, such as a disk, an optical disk; and a communication unit 909, such as network cards, modems, wireless communication transceivers, and the like.
- the communication unit 909 allows the device 900 to exchange information/data with other devices through a computer network such as the Internet and/or various telecommunication networks.
- the computing unit 901 may be formed of various general-purpose and/or special-purpose processing components with processing and computing capabilities.
- the computing unit 901 include, but are not limited to, a central processing unit (CPU) , graphics processing unit (GPU) , various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processor (DSP) , and any suitable processor, controller, microcontroller, etc.
- the computing unit 901 performs various methods and processes described above, such as an image processing method.
- the image processing method may be implemented as computer software programs that are tangibly embodied on a machine-readable medium, such as the storage unit 908.
- part or all of the computer program may be loaded and/or installed on the device 900 via the ROM 902 and/or the communication unit 909.
- the computing unit 901 When a computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image processing method described above may be performed.
- the computing unit 901 may be configured to perform the image processing method in any other suitable manner (e.g., by means of firmware) .
- Various implementations of the systems and techniques described herein above may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA) , application specific integrated circuits (ASIC) , application specific standard products (ASSP) , system-on-chip (SOC) , complex programmable logic device (CPLD) , computer hardware, firmware, software, and/or combinations thereof.
- FPGA field programmable gate arrays
- ASIC application specific integrated circuits
- ASSP application specific standard products
- SOC system-on-chip
- CPLD complex programmable logic device
- programmable processor may be a special-purpose or general-purpose programmable processor, and may receive data and instructions from a storage system, at least one input device and at least one output device, and may transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
- Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general computer, a dedicated computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions and/or operations specified in the flow diagrams and/or block diagrams is performed.
- the program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on a machine and partly on a remote machine or entirely on a remote machine or server.
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
- the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
- machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM) , read-only memories (ROM) , erasable programmable read-only memories (EPROM or flash memory) , fiber optics, compact disc read-only memories (CD-ROM) , optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
- RAM random access memories
- ROM read-only memories
- EPROM or flash memory erasable programmable read-only memories
- CD-ROM compact disc read-only memories
- magnetic storage devices or any suitable combination of the foregoing.
- the systems and techniques described herein may be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) ) for displaying information for the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide an input to the computer.
- a display device e.g., a cathode ray tube (CRT) or liquid crystal display (LCD)
- LCD liquid crystal display
- keyboard and pointing device e.g., a mouse or trackball
- Other types of devices can also be used to provide interaction with the user, for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback) ; and may be in any form (including acoustic input, voice input, or tactile input) to receive the input from the user.
- the computer system may include a client and a server.
- the Client and server are generally remote from each other and usually interact through a communication network.
- the relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other.
- the server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business expansion in traditional physical hosts and virtual private servers ( "VPS" for short) .
- the server may also be a server of a distributed system, or a server combined with a blockchain.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Different concepts for improving merge related intra prediction coding tools are described.
Description
Cross-Reference to Related Applications
This application is based on and claims priority to the European Patent Application No. “24166515.7” , filed on March 26, 2024, the entire content of which is incorporated herein by reference.
The present disclosure generally relates to the field of encoding/decoding pictures, images or videos, and embodiments of the present disclosure concern improvements regarding the Intra prediction, more specifically, improvements regarding a template-based intra mode derivation, TIMD, process. More specific embodiments of the present disclosure relate to a merge mode for template-based intra mode derivation (TIMD-Merge) .
The encoding and decoding of a picture, an image or a video is performed in accordance with a certain standard, for example, in accordance with the advanced video coding, AVC standard (see reference [1] ) , the high efficiency video coding, HEVC, standard (see reference [2] ) or the versatile video coding, VVC, standard (see reference [3] ) .
A block diagram of a standard video compression system 100 operating in accordance with the VVC standard is illustrated in Fig. 1. The standard video coder 100 compresses and encodes a picture 102 of a video sequence. The picture 102 to be encoded is partitioned into blocks 104 also referred to as coding units, CUs. The encoder 100 comprises a pre-encode filter 106 and a prediction residual signal former 108 which generates a prediction residual signal 110 so as to measure a deviation of a prediction signal 112 from the signal 114 output by the filter 106. The encoder further comprises a transformer 116, a quantizer 118 and provides an output bitstream or data stream 120 using an entropy coder 122. Further, the encoder 100 comprises a prediction stage 124 for determining the prediction signal 112, which includes a de-quantizer or inverse quantizer 126, an inverse transformer 128, a combiner 130, an in-loop filter 134, a motion estimator 138, an intra/inter mode selector 140, an inter predictor 142 and an intra predictor 144.
The video coder 100 as described with reference to Fig. 1 compresses and encodes a picture 102 of a video sequence as follows. The picture 102 to be encoded is partitioned into the blocks or CUs 104. Each coding unit 104 is encoded using either an intra or inter coding mode. When a CU is encoded in the intra mode, intra prediction is performed by the intra predictor 144. The intra prediction comprises predicting the current CU 114 being encoded by means of already coded, decoded and reconstructed picture samples located around the current CU, e.g., on the top and on the left of the current CU. The intra prediction is performed in the spatial domain. In an inter mode, motion compensation and motion
estimation are performed by the motion estimator 138 which searches, in one or more reference pictures provided by the picture buffer 136 and used to predictively code the current picture, a CU that is a good predictor of a current CU. For example, a good predictor of a current CU is a predictor which is similar to the current CU, i.e., the distortion between the two CUs is low or below a certain threshold. The motion estimation may also account for the rate cost of signaling the predictor to optimize a rate-distortion tradeoff. The output of the motion estimation step is one or more motion vectors and reference indices associated with the current CU. The motion compensation then predicts the current CU by means of the one or more motion vectors and reference pictures indices as determined by the motion estimator 138. Basically, the block or CU contained in the selected reference picture and pointed to by the determined motion vector is used as the prediction block for the current CU. The encoder 100, by means of the selector 140, selects one of the intra coding mode or the inter coding mode to use for encoding the CU and indicates the intra/inter decision, for example, by means of a prediction mode flag. Prediction residuals 110 are then transformed and quantized by blocks 116 and 118, and the quantized transform coefficients as well as the motion vectors and other syntax elements are entropy encoded and written into the output bitstream 120. The encoder 100 may skip the transform stage 116 and apply the quantization directly to the non-transformed residual signal 110 in a so-called transform-skip coding mode. After a block or CU has been encoded, the encoder decodes the CU and reconstructs it so as to obtain the reconstructed signal 132 that may serve as a reference data for predicting future CUs or blocks to encode. The quantized transform coefficients 110’ are de-quantized and inverse transformed leading to a decoded prediction CU or block 110”, and the decoded prediction residuals and the predicted block are then combined at 130, typically summed, so as to provide the reconstructed block or CU 132. The in-loop filters 134 are applied to the reconstructed picture to reduce compensation artifacts. For example, a deblocking filter, a sample adaptive offset, SAO, filter, and an adaptive loop filter, ALF, may be applied to reduce encoding artifacts. The filtered picture is stored in the buffer 136, also referred to as the decoded picture buffer, DPB, so that it may be used as a reference picture for coding subsequent pictures.
Fig. 2 is a block diagram of a video decoder 150 for predictively decoding from a data stream 152 a picture or video which is provided at an output 154 of the decoder 150. The decoder 150 includes an entropy decoder 156, a partitioning block 158, an inverse quantizer 160, an inverse transformer 162, a combiner 164, an in-loop filter 166, optionally a post-decoding processor 168, and a prediction module 170. The prediction module 170 includes a decoded picture buffer 180 a motion compensator 182 an intra predictor 184.
An encoded picture of a video sequence is decompressed and decoded by the decoder 150 as follows. The input bit stream 152 is entropy decoded by the decoder 156 which provides, for example, the block partitioning information, the coding mode for each coding unit, the transform coefficients contained in each transform block, prediction information, like intra prediction mode, motion vectors, reference picture indices, and other coding information. The block partitioning information indicates how the picture is partitioned and the decoder 150 may divide the input picture into coding tree units, CTUs, typically of a size of 64x64 or 128x128 pixels and divide each CTU into rectangular or square coding
units, CUs, according to the decoded partitioning information. The entropy decoded quantized coefficients 172 are de-quantized 160 and inverse transformed 162 so as to obtain the decoded residual picture or CU 174. The decoded prediction parameters are used to predict the current block or CU, i.e., whether the predicted block is to be obtained through its intra prediction or through its motion-compensated temporal prediction. The prediction process performed at the decoder side is the same as the one performed at the encoder side. The decoded residual blocks 174 are added to the predicted block 176, thereby yielding the reconstructed current image block 164. The in-loop filters 166 are applied to the reconstructed picture or image which is also stored in the decoded picture buffer 180 to serve with the reference picture for future pictures to decode. As mentioned above, the decoded picture may further go through a post-decoding processing, for example for performing an inverse color transformation, for example a conversion from YCbCr 4: 2: 0 to RGB 4: 4: 4.
In all above steps, the entropy encoding or entropy decoding of syntax elements representing decisions at the encoder side such as block partitioning information, prediction modes/parameters, quantized transform coefficients, etc. is typically carried out by using a context-adaptive entropy coding, such as Context-Adaptive Binary Arithmetic Coding (CABAC) . To use CABAC, each syntax element is binarized to be represented with a series of bins and each bin is associated with a CABAC context model, that keeps track of binary values of that particular bin in the past, in order to more efficiently model its probability distribution.
Currently, regular TMID and TMID merge tools are under discussion, but despite the efficiency of these tools, it is still desirable to increase the efficiency.
The present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a first aspect by testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set and selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors even, which steps belong to a first tool, even if the corresponding first tool is not used, and taking into account the first set in establishing a merge list of candidate coding parametrizations for a second tool, namely by comparing, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, and modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes
on the other hand more diverse. By rendering the tools’ selectable parametrizations more diverse, the coding is made more effective. In other words, the first aspect avoids that the prediction parameters that are derived from the merge list in the second tool, are similar to those that the first tool such as regular TIMD would provide.
The present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a second aspect by construing the merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between a respective second block on the one hand and the first block on the other hand. By this measure, it is possible to suppress, or lower the degree of, the contribution of second blocks’ coding parametrizations to the merge list, whose size or aspect-ratio does not equal, or is not similar enough, to the first block. In other words, the second aspect avoids blindly including all neighbors (second blocks) in a merge list such as the TIMD-Merge list and, thus, avoids, or renders it less likely that the current block to code (first block) would derive its prediction parameters from an irrelevant candidate block (second block) .
The present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a third aspect by populating the test set of at least one or more of the intra prediction modes used in a first tool using the merge list of candidate coding parametrizations used for the second tool so that the effective first tool tests intra prediction modes which are more likely to be efficient. In other words, the second aspect avoids blindly including all neighbors (second blocks) in a merge list such as the TIMD-Merge list and, thus, avoids, or renders it less likely that the current block to code (first block) would derive its prediction parameters from an irrelevant candidate block (second block) .
The present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a fourth aspect by reconstructing the first block based on the selected coding parametrization, selected from the merge list, by selecting a predetermined residual transform domain of a prediction residual for the first block out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems. Thus, transform domain signaling is not necessary with the transform domain implicitly chosen being, nevertheless, efficient at pretty high likelihood.
The present disclosure addresses the above drawbacks and/or provides a better coding efficiency according to a fifth aspect by forming a first list of horizontal residual transforms based on the second blocks in the neighborhood of a first block, and forming a second list of vertical residual transforms based on the second blocks, with performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which
differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream. By this measure likely well fitting transforms are recruited from neighboring blocks in both directions separately, and a new combination, not having been used for the neighboring blocks, which might be optimal for the first block, might be chosen for coding the first block, thereby increasing efficiency.
It is to be understood that the content described in this section is not intended to identify key or critical features of the embodiment of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure become readily apparent from the following description.
The drawings are explanatory and serve to explain the present disclosure, and are not to be construed to limit the present disclosure to the illustrated embodiments.
Fig. 1 is a block diagram of a conventional video encoder;
Fig. 2 is a block diagram of a conventional video decoder;
Fig. 3 shows neighboring reconstructed samples used for DIMD chroma mode;
Fig. 4 shows a proposed scheme using non-adjacent spatial neighboring candidates for DIMD merge;
Fig. 5 shows above and left templates used for TIMD mode derivation;
Fig. 6 shows an example of TIMD merge list and example information of each candidate within the list;
Fig. 7 shows an example of adjacent neighbors in TIMD merge mode;
Fig. 8 shows a division method for angular modes;
Fig. 9 shows a table of modified weights used for angular modes;
Fig. 10 shows a syntax design for a fusion of chroma intra prediction modes;
Fig. 11 shows a diversification of modes in the TIMD merge list, by using the modes derived from TIMD legacy;
Fig. 12 shows a two-step eligibility based on coding mode and block size;
Fig. 13 shows an extending of the MPM list of TIMD legacy by including modes from TIMD merge list;
Fig. 14 shows an inheriting horizontal and vertical transform types from the same TIMD merge candidate (CU 4) that provides the intra prediction modes;
Fig. 15 shows an implicit transform type inheritance allowing combinations of transform types that are not explicitly provided by ECM; and
Fig. 16 shows a block diagram illustrating an electronic device 900 according to embodiments of the present disclosure.
Illustrative embodiments of the present disclosure are described below with reference to the drawings, where various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered as illustrative only. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope of the present disclosure. Also, descriptions of well-known functions and constructions are omitted from the following description for clarity and conciseness.
In the present disclosure, the term "and/or" is intended to cover all possible combinations and sub-combinations of the listed elements, including any one of the listed elements alone, any sub-combination, or all of the elements, and without necessarily excluding additional elements.
In the present disclosure, the phrase "at least one of... or... " is intended to cover any one or more of the listed elements, including any one of the listed elements alone, any sub-combination, or all of the elements, without necessarily excluding any additional elements, and without necessarily requiring all of the elements.
In the present disclosure, the term “coding” refers to "encoding” or to “decoding” as becomes apparent from the context of the described embodiments. Likewise, the term “coder” refers to " an encoder” or to “adecoder” .
Before describing some embodiments of the present application, the discussion of currently discussed coding tools is resumed. It is noted that the details set out in this discussion might be used to provide details which are combinable with the improved embodiments described thereinafter. In other words, such details may form more detailed embodiments by being combined with the subsequently described embodiments. In this regard, it shall also be noted that even the details set out in the introductory portions may be used as a reservoir to provide such further combinable features for the subsequently described embodiments.
As to intra prediction tools, there is a tool called decoder side intra mode derivation (DIMD) . When JVET DIMD is applied, up to five intra modes are derived from the reconstructed neighbor samples, and those five predictors are combined with the planar mode predictor with the weights derived from the histogram of gradients as described in JVET-O0449 . The division operations in weight derivation are performed utilizing the same lookup table (LUT) based integerization scheme used by the CCLM. For example, the division operation in the orientation calculation:
Orient=Gy/Gx
Orient=Gy/Gx
is computed by the following LUT-based scheme:
x = Floor (Log2 (Gx) )
normDiff = ( (Gx<< 4) >> x) & 15
x += (3 + (normDiff ! = 0) ? 1 : 0)
Orient = (Gy* (DivSigTable [normDiff] | 8) + (1<< (x-1) ) ) >> x
x = Floor (Log2 (Gx) )
normDiff = ( (Gx<< 4) >> x) & 15
x += (3 + (normDiff ! = 0) ? 1 : 0)
Orient = (Gy* (DivSigTable [normDiff] | 8) + (1<< (x-1) ) ) >> x
where
DivSigTable [16] = {0, 7, 6, 5 , 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0 } .
DivSigTable [16] = {0, 7, 6, 5 , 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0 } .
For a block of size W×H, the weight for each of the five derived modes is modified if the one the above or left histogram magnitudes is twice larger than the other one. In this case, the weights are location dependent and computed as follows:
If the above histogram is twice the left, then:
If the left histogram is twice the above, then:
where wDimdi is the unmodified uniform weight of the DIMD selected as in JVET-O0449, Δi is pre-defined and set to 10.
Derived intra modes are included into the primary list of intra most probable modes (MPM) , so the DIMD process is performed before the MPM list is constructed. The primary derived intra mode of a DIMD block is stored with a block and is used for MPM list construction of the neighboring blocks.
Finally, note the region of neighboring reconstructed samples used for computing the histogram of gradients is modified compared to JVET-O0449 method, depending on reconstructed samples availability. The region of decoded reference samples of current WxH luma CB is extended towards the above-right side if available, up to W additional columns. It is extended towards the bottom-left side if available, up to H additional rows.
The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the neighboring reconstructed Y, Cb and Cr samples in the second neighboring row and column as shown in Figure 3. Specifically, a horizontal gradient and a vertical gradient are calculated for each collocated reconstructed luma sample of the current chroma block, as well as the reconstructed Cb and Cr samples, to build a HoG. Then the intra prediction mode with the largest histogram amplitude values is used for performing chroma intra prediction of the current chroma block.
Figure 3 shows neighboring reconstructed samples used for DIMD chroma mode. Fig. 3a shows the neighboring collocated reconstructed Y samples, Fig. 3b shows the neighboring reconstructed Cb samples and Fig. 3c shows the neighboring reconstructed Cr samples.
When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is signaled to indicate whether the
proposed DIMD chroma mode is applied.
Finally, the luma region of reconstructed samples used for computing the histogram of gradients for chroma DIMD mode is modified compared to JVET-O0449. For a WxH pair of chroma CBs to predict, to build the histogram of gradients associated to the collocated luma CB, the pairs of a vertical gradient and a horizontal gradient are extracted from the second and third lines in this luma CB instead of being extracted from the regular set of DIMD decoded reference samples around this luma CB.
DIMD merge mode is proposed in JVET-AF0120. When using DIMD Merge, the DIMD information extracted from neighbouring blocks is used to compute the intra prediction for the current block. A new Merged Histogram of Gradients (MHoG) is computed for the current block based on the HoGs of neighbouring blocks. Only neighbouring blocks encoded with DIMD or with DIMD Merge are considered.
When a single DIMD or DIMD Merge neighbouring block is available, then its histogram of gradients is used to form the MHoG for the current block. If more than one DIMD or DIMD Merge neighbouring blocks are available, the corresponding histograms are combined by means of amplitude averaging to derive the MHoG. Up to maximum 13 CUs in the surrounding of the current block are considered to extract DIMD information.
Finally, the MHoG is used to compute intra-prediction modes and weights, as in conventional DIMD. The directional modes and their weights corresponding to the five highest amplitudes in the MHoG are selected, and the corresponding predictors are blended as in conventional DIMD.
A method of using non-adjacent spatial candidates for DIMD merge is proposed in JVET-AF0106. It is proposed to use non-adjacent spatial candidates for DIMD merge candidates list construction. As shown in Figure 4, the distances between non-adjacent candidates and current block are defined based on the width and height of current coding block. When using DIMD Merge, the DIMD information extracted from neighbouring blocks and non-adjacent spatial blocks is used to compute the intra prediction for the current block. Figure 4 shows the proposed scheme using non-adjacent spatial neighboring candidates for DIMD merge.
A fusion is used for template-based intra mode derivation (TIMD) . For each intra prediction mode in MPMs, as well as the wide-angle modes if the above-right and/or bottom-left reference samples are available, SATD between the prediction and reconstruction samples of the template is calculated. First two intra prediction modes with the minimum SATD are selected as the TIMD modes. These two TIMD modes are fused with the weights after applying PDPC process, and such weighted intra prediction is used to code the current CU. Position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.
The costs of the two selected modes are compared with a threshold, in the test the cost factor of 2 is applied as follows:
costMode2 < 2*costMode1.
costMode2 < 2*costMode1.
If this condition is true, the fusion is applied, otherwise the only mode1 is used.
Weights of the modes are computed from their SATD costs as follows:
weight1 = costMode2 / (costMode1+ costMode2)
weight2 = 1 -weight1.
weight1 = costMode2 / (costMode1+ costMode2)
weight2 = 1 -weight1.
The division operations are conducted using the same lookup table (LUT) based integerization scheme used by the CCLM.
Figure 5 illustrates the above and left templates used for TIMD mode derivation.
TIMD-Merge is a an intra coding tool based on regular TIMD (described above) , with an exception that TIMD-Merge keeps a list of merge candidates that are collected as already coded blocks from causal neighbourhood. Block candidates in this list can be of type of regular TIMD blocks, TIMD-merge or other coding modes if they meet certain requirements. The candidate list contains different information for each, notably their prediction modes, whether or not the candidate uses any sort of prediction blending. If blending, their blending etc. Then, when coding a current block with TIMD-Merge mode, the algorithm chooses at least one of the merge candidates and inherits its stored prediction parameters for current block. Finally, the use of TIMD-Merge mode is signaled at the block level to inform the decoder about the selected candidate.
A list of N candidates which are collected from blocks (de) coded in TIMD mode may be generated. The TIMD candidates in the list contain TIMD information that are used for (de) coding the previous blocks. Such information could be one or more of the below:
-Primary TIMD mode
-TIMD fusion mode (s) which are used for blending with the primary TIMD mode,
-Fusion weights for primary and fusion modes,
-Blending flag indicator,
-Template cost of the primary and/or fusion modes,
-Transform type (s) ,
-Wide angle conditions for primary and fusion modes.
An example of TIMD merge list with associated information of each TIMD candidate is shown in below.
Figure 6 shows an example of TIMD merge list and example information of each candidate within the list.
The list may contain only the TIMD candidates from its immediate spatially neighboring blocks (aka adjacent candidates) , or the list may contain only candidates from non-adjacent blocks. Alternatively, the list may contain both adjacent and non-adjacent candidates. Figures 7 and 4 show examples of adjacent and non-adjacent blocks that are checked for generating the TIMD merge list candidates.
Figure 7 illustrates an example of adjacent neighbors in TIMD merge mode and Figure 4 illustrates an example of non-adjacent neighbors in TIMD merge mode.
The encoder may test one or more of the TIMD merge candidates in the list and the best performing candidate (and associated information) according to the rate-distortion optimization (RDO) mechanism may be selected and used for prediction of the current block.
An intra prediction method called Intra prediction fusion derives predicted samples as a weighted combination of multiple predictors generated from different reference lines. In this process multiple intra predictors are generated and then fused by weighted averaging. The process of deriving the predictors to be used in the fusion process is described as follows:
1) For angular intra prediction modes including the single mode case of TIMD and DIMD, the proposed method derives intra prediction by weighting intra predictions obtained from multiple reference lines represented as pfusion=w0pline+w1pline+1, where pline is the intra prediction from the default reference line and pline+1 is the prediction from the line above the default reference line. The weights are set as w0=3/4 and w1=1/4.
2) For TIMD mode with blending, pline is used for the first mode (w0=1, w1=0) and pline+1 is used for the second mode (w0=0, w1=1) .
3) For DIMD mode with blending, the number of predictors selected for a weighted average is increased from 3 to 6.
Intra prediction fusion method is applied to luma blocks when angular intra mode has non-integer slope (required reference samples interpolation) and the block size is greater than 16, it is used with MRL and not applied for ISP coded blocks. In the method studied in the sub-test a, PDPC is applied for the intra prediction mode using the closest to the current block reference line.
There is also a combination of CIIP with TIMD and TM merge. In CIIP mode, the prediction samples are generated by weighting an inter prediction signal predicted using CIIP-TM merge candidate and an intra prediction signal predicted using TIMD derived intra prediction mode. The method is only applied to coding blocks with an area less than or equal to 1024.
The TIMD derivation method is used to derive the intra prediction mode in CIIP. Specifically, the intra prediction mode with the smallest SATD values in the TIMD mode list is selected and mapped to one of the 67 regular intra prediction modes.
In addition, it is also proposed to modify the weights (wIntra, wInter) for the two tests if the derived intra prediction mode is an angular mode. For near-horizontal modes (2 <= angular mode index < 34) , the current block is vertically divided as shown in Figure 8a; for near-vertical modes (34 <= angular mode index <= 66) , the current block is horizontally divided as shown in Figure 8b. Fig. 8 shows the division method for angular modes. The (wIntra, wInter) for different sub-blocks are shown in Fig. 9. Fig. 9 shows a table of modified weights used for angular modes.
With CIIP-TM, a CIIP-TM merge candidate list is built for the CIIP-TM mode. The merge candidates are refined by template matching. The CIIP-TM merge candidates are also reordered by the ARMC method as regular merge candidates. The maximum number of CIIP-TM merge candidates is equal to two.
In ECM, there is also fusion of chroma intra prediction modes. Here, two chroma intra prediction signals can be fused together. One of the two chroma intra prediction signals is predicted using one of the DM mode, DIMD chroma mode and the four default modes (non-LM mode) . The other chroma intra prediction signal is predicted using cross-component linear prediction modes (LM mode) . Two different methods are supported.
In the first method, the LM mode can be either MM-CCLM or MM-CCCM, and the final predictor is derived as follows:
predC (i, j) = (w0×pred0 (i, j) +w1×pred1 (i, j) + (1<< (shift-1) ) ) >>shift
predC (i, j) = (w0×pred0 (i, j) +w1×pred1 (i, j) + (1<< (shift-1) ) ) >>shift
where pred0 (i, j) is the predictor obtained by applying the non-LM mode, pred1 (i, j) is the predictor obtained by applying the LM mode and predC (i, j) is the final predictor of the current chroma block. The two weights, w0 and w1 are determined by the intra prediction mode of adjacent chroma blocks and shift is set equal to 2. Specifically, when the above and left adjacent blocks are both coded with LM modes, {w0, w1 } = {1, 3} ; when the above and left adjacent blocks are both coded with non-LM modes, {w0, w1} = {3, 1} ; otherwise, {w0, w1} = {2, 2} . Two template costs are calculated by fusing the angular chroma prediction with MM-CCLM or MM-CCCM, respectively, and the one of the two CCPs which provides a smaller template cost is utilized to derive pred1.
In the second method, the LM mode can be either MMLM or CCLM mode, and the final predictor is derived as follows:
predC (i, j) = α0×pred0 (i, j) + α1×rec′L (i, j) +α2×β
predC (i, j) = α0×pred0 (i, j) + α1×rec′L (i, j) +α2×β
where pred0 (i, j) is the predictor obtained by applying the non-LM mode, rec′L (i, j) is the set of downsampled reconstructed luma samples at co-located positions and predC (i, j) is the final predictor of the current chroma block. β is a fixed value and is set equal to 512 for 10-bit content. The three weights, α0, α1 and α2 are derived from the adjacent luma and chroma samples using the same LDL derivation method as in CCCM.
For the syntax design, one index is signaled to indicate whether fusion is applied, and which method is used, e.g., see Fig. 10. It is noted that for I slices, the non-LM mode can be DM mode, DIMD chroma mode and the four default modes. For non-I slices, only DIMD chroma mode is allowed to be fused with LM modes.
In VVC as well as ECM, there is a mechanism for an explicit transform type selection using Multiple Transform Sets (MTS) . Multiple Transform Sets (MTS) are used. MTS is a tool that allows different transform types to be applied within a CU. This means that instead of using a single transform type for the entire CU, different CUs can be encoded using different transform types based on their characteristics. Explicit MTS enables this by allowing the encoder to signal which transform type to use
for each part of the CU.
The MTS index is a parameter that determines which transform type is applied to different horizontal and vertical partitions within the CU. This index is signaled along with other coding parameters and is used by the decoder to reconstruct the original signal. When encoding a CU, the encoder tries different MTS indexes for the CU, and for each tested MTS index, it translates the index into a pair of horizontal and vertical transform types. Finally, the index that represents the best transform pair is signaled in the bitstream.
Despite all this tools an options, there is a need to provide further improvements.
In particular, a first Current TIMD-Merge can be sub-optimal in the following cases:
Case A) When the prediction parameters that are derived from the merge list are similar to those that regular TIMD would provide, a design redundancy might occur. This means that for the same setting of prediction, there will be two different ways of signaling; one with the regular TIMD algorithm and the other with the TIMD-Merge. This would result in waste of bitrate.
Case B) Blindly including all neighbors in the TIMD-Merge list might cause that current block to code would derive its prediction parameters from an irrelevant candidate block. For instance, deriving prediction parameters for a 32x32 from a small 4x4 candidate can be considered irrelevant.
Case C) Current TIMD-Merge can potentially collect valuable information about the neighbourhood texture. However, current design does not consider using those information to help coding of other types of intra coding modes.
Case D) The choice of residual transform is signaled explicitly in the bitstream of TIMD-Merge blocks. However, this might overlook possible correlations between the transform type of candidate blocks and current block.
According to embodiments, the above shortcomings of the current TIMD-Merge design are addressed as follows. Specifically, the following items are the novelties brought:
Solution for case A) Comparisons are made between prediction parameters of candidates and those provided by the regular TIMD-Merge mode. If they are similar given an arbitrary metric, then they might be treated differently.
Solution for case B) Constraints are added when including neighboring candidates that disallow some blocks. These constraints are based on characteristics of current block to be coded (e.g. block size)
Solution for case C) Some information collected by the TIMD-Merge method is reused for coding of other intra coding tools. Specifically, some intra prediction modes in the merge list are included in the MPM list of other blocks if they are in the same causal neighborhood.
Solution for case D) transform types of TIMD-Merge blocks are derived from candidates in the merge list, instead of explicitly being signaled in the bitstream.
In the rest of this document, the TIMD algorithm is referred to as “TIMD legacy” to be distinguished from “TIMD merge” algorithm. Furthermore, the term “generic TIMD” is used to refer both algorithms. In particular, a generic TIMD block is a block that is coded either with TIMD legacy, or TIMD
merge mode. For example, blocks that contribute to construction of the merge list of a TIMD merge block are all generic TIMD blocks, meaning that they can either be TIMD merge or TIMD legacy blocks.
In order to improve the performance of the intra prediction using Fusion for Merge-based template-based intra mode derivation (TIMD-Merge) the following four improvements have been added.
The embodiments are described relate to the first aspect of diversification of modes in TIMD legacy and TIMD merge. Given an arbitrary metric of similarity, one can determine whether intra modes of TIMD merge candidates are similar to those of TIMD legacy. If so, an additional process is applied to diversify the superset of modes provided by TIMD legacy and TIMD merge.
It is important to note that the diversification process compares A) intra modes of a neighboring generic TIMD blocks that is contributing to merge list construction of current TIMD merge block and B) intra modes of TIMD legacy algorithm on current block, if it was coded as TIMD legacy. The philosophy behind this comparison is make TIMD legacy and TIMD merge algorithms distinct enough for current block, so that they cover a larger different set of possible ways to code current block. If no diversification is applied, technically it would be possible that, for a given current block, the two algorithms of TIMD legacy and TIMD merge would end up with exactly the same intra modes, resulting in two different ways of coding the block. And this would waste the signaling rate.
The diversification process therefore consists of two steps of 1) similarity measurement and 2) mode distinction.
The similarity is measure between two blocks: 1) current block if it were coded as TIMD legacy at one hand, and 2) a generic TIMD block that is contributing to TIMD merge list construction of current block, on the other hand. Given two such blocks, the similarity measurement step involves defining a similarity metric which can be computed on different variables.
The intra prediction mode index is the principle variable to be used for this purpose. Furthermore, one can take into account also the flag determining the use wide-angle mode. Since intra modes of each TIMD block are ordered based on their weights, one can also take into account the order modes. For example, if the intra prediction modes indexes between the TIMD legacy algorithm and a generic TIMD block in the merge list are identical, however, their fusion weights have put them in different orders in the two blocks (e.g. <18, 34> and <34, 18>) , one might (or might not) consider that the two block similar. Finally, the fusion weight itself can be used in the similarity measurement. For example, if both index and order of the modes are identical, but their fusion weights are different in the two blocks, one might (or might not) consider them similar.
Regardless of the variables involved, it is also possible to apply a threshold when comparing a given variable between two blocks. Let’s say when comparing intra prediction mode index (i.e. variable) of two TIMD blocks, one encounters the modes <18, 34> for the first block and <19, 37> for the second blocks. At this point with a threshold of 2 or greater, these two blocks are considered similar, since the maximum difference between each element of the modes pair is smaller than or equal to 2.
Once a generic TIMD block in the list is considered “similar” to current TIMD legacy block, the
mode distinction step diversifies them be modifying the way the generic TIMD block is added to the merge list. There are several ways to modify the above process. Here are some examples.
·The generic TIMD block is ignored and not added to the merge list.
·The generic TIMD block is added, but with a reduced weight.
·The intra modes of the generic TIMD block are modified before adding to the merge list.
Figure 11 shows one example of applying mode diversification. In this example, first the eligible neighbor CUs (i.e. Generic TIMD blocks) are identified and added to the list. Then, TIMD legacy modes are derived from the template of current CU. The diversification step consists of a similarity measure step that identifies TIMD merge candidates in which their difference of intra prediction mode indexes (in any order) from the TIMD legacy pair is higher than or equal to a threshold (e.g. 3) . Then the mode distinction step applies an exclusion strategy and removes those pairs from the final TIMD merge list.
Figure 11 shows a diversification of modes in the TIMD merge list, by using the modes derived from TIMD legacy.
In an embodiment of the mode diversification algorithm, only the intra prediction indexes are compared to determine whether or not the merge list should be diversified.
In another embodiment, other prediction information such as blending weights, wide-angle, order of prediction modes etc. are also used to determine whether or not the merge list should be diversified.
In an embodiment of the mode diversification, if the intra prediction mode indexes are similar but their order is different, then they are not considered as similar.
In an embodiment of the mode diversification, when comparing intra prediction mode index of two blocks, a threshold larger than 0 is used to apply a certain level of flexibility. This means that even if the indexes are not identical, but they have a difference less than the defined threshold, then they are considered similar.
In an embodiment of the mode diversification, when intra prediction modes of a candidate CU is considered similar to those of TIMD legacy, then that candidate CU is excluded from the TIMD merge list.
The embodiments described next relate to the second aspect of block size based eligibility of TIMD-Merge candidates. The TIMD merge algorithm of the state-of-the-art does take into account some block characteristics (e.g. block size) of neighboring generic TIMD blocks when adding them to the TIMD merge list of current block. However, one might argue that it is more likely that a block derives its intra prediction modes from blocks of similar size than blocks with sizes that are significantly different from that of current block.
The second aspect involves, according to the present embodiment, identifying some generic TIMD blocks in the neighborhood non-eligible, due to the difference between their block size and size of the current block. In particular, the size-eligibility algorithm compares current block to a generic TIMD block, by taking into account some sort of size-based metric.
Different approaches can be used to identify a neighboring block as non-eligible based on its size.
First (and the most strict) approach is to put a constraint on widths and heights of the two blocks so that if they do not perfectly match (i.e. width1 = width2 and height1 = height2) , then the neighbor is non-eligible. Alternatively, one can make the constraint more generic by not considering the order of width and height. That is to say, even the case where width1 = height2 and width2 = height1, then the neighbor is considered as eligible. In a different approach, one might use the number of pixels in the block (i.e. width × height) as a metric of size. Given this metric, one can also apply different levels of constraints when determining whether two blocks are of similar size. For example, one can apply a threshold on the number pixels when comparing their size.
Figure 12 shows an example of using a metric based on block size in order to exclude a neighboring CU candidate. In particular, neighbor CU 1 has been excluded from the list, since its block size (i.e. 64×32) is different from the block size of current CU (i.e. 32×16) . Figure 12 shows two-step eligibility based on coding mode and block size.
In an embodiment of the proposed block-size based TIMD merge candidate eligibility, a size metric is designed based on the size of the CU and it is computed on current block and one neighbor candidate CU. Then a size difference is designed and computed on the size metric. Finally, if the size difference metric is higher than a threshold, then the neighbor CU is excluded from the TIMD merge list of current block.
In an embodiment, the size metric is on the number of pixels in the CU.
In an embodiment, the size metric is a pair composed of the width and the height of the CU.
In an embodiment, the size metric includes the aspect ratio of the CU.
In an embodiment, if the neighbor CU has a large size difference value, then instead of excluding it from the TIMD merge list, its corresponding candidate is moved lower in the merge list. This can be done, for example, by increasing its cost before the candidate sorting.
Next, embodiments are described which relate to a third aspect of using TIMD-Merge modes in MPM of TIMD legacy. State-of-the-art TIMD legacy algorithm uses the Most Probable Modes (MPM) list to search for the best TIMD intra modes. It is proposed that some information from TIMD merge list would be also added this search list. More precisely, we propose that when a current block is coded (or decoded) with TIMD legacy mode, the merge list construction of the TIMD merge algorithm is also called to compute and provide additional modes to TIMD legacy mode. This is carried out by including some modes from TIMD merge list to the MPM of the TIMD legacy.
One can consider a conditional design of the above TIMD merge mode inclusion, so that TIMD merge modes are included in the MPM of the TIMD legacy mode if certain conditions are met. For example, one condition could be the presence of at least N blocks of TIMD merge in the neighborhood of current block. The logic behind this condition is that a significant presence of TIMD merge blocks in a neighborhood might indicate that the template pixels of that neighborhood were not strongly correlated to the texture of their corresponding blocks. Therefore, the encoder opted to code that block with TIMD merge mode which actually uses intra modes that are derived from templates that are relative further to
current block. Another possible condition can be based on the similarity of modes that already exist in the MPM list and the modes provided by the TIMD merge list. For example, if one observes that there are several common modes in the two lists, however, the TIMD merge list contains modes that are not present in the MPM list, then it might be logical to also include those absent mode to the MPM list.
Regardless of the above aspects, the TIMD merge mode inclusion in the MPM list of TIMD legacy can take place in different MPM lists, depending on the underlying codec. For instance, in ECM, the high-level MPM is split into “primary” , “secondary” and “remaining” lists. Some of above aspects are applicable on all these three lists.
Figure 13 shows an example of using modes from the TIMD-merge list to extend the MPM list of TIMD legacy algorithm. In this figure, a current block is to be coded by TIMD legacy algorithm, which includes an MPM list construction stage. In parallel, a TIMD-merge list is also computed on the current block to provide a new source of prediction modes. Then the two lists (i.e. MPM and TIMD-merge) are conditionally combined to provide an extended MPM list for the TIMD legacy algorithm. As a result, the extended MPM list has a size (i.e. M) larger than its original size in TIMD legacy (e.g. N) . Figure 13 illustrates an extending of the MPM list of TIMD legacy by including modes from TIMD merge list.
In an embodiment, the prediction modes from TIMD merge list are added to the MPM list if they do not already exist.
In an embodiment, the TIMD merge list modes are used to extend the primary MPM list of ECM.
In another embodiment, the TIMD merge list modes are used to extend the secondary MPM list of ECM.
In an embodiment, a condition is used to determine whether the TIMD merge modes should be included in the MPM list.
In an embodiment, the condition is based on the presence of other TIMD merge blocks in the neighborhood around the current CU. For instance, if there is at least on TIMD merge block in the neighborhood of the current block, then the TIMD merge list modes of current block are included in the MPM list of current block for TIMD legacy algorithm.
Next, embodiments relating to the fourth and fifth aspects are described, the fourth aspect pertaining to implicit transform type inheritance. Current TIMD merge algorithm in the ECM uses regular Multiple Transform Sets (MTS) to determine the separable transform type for horizontal and vertical residual transformation. This includes an encoder-side search for the best index and then signaling it at the block level. However, a method is proposed to derive the transform type along with prediction information (intra modes, wide-angle etc. ) from one of the candidates in the TIMD merge list.
There are mainly three advantages in implicitly deriving the transform type. First, the encoder no longer needs to exhaustively search for the optimal transform index, which makes the encoding faster. Second, the transform index no longer needs to be signaled, which can save bitrate and potentially increase the compression efficiency. Third, the transform derivation can possibly result in a combination of horizontal and vertical transform types that are not presented by either of available transform indexes.
This can potentially result in better compression efficiency.
In order to inherit transform type from a neighbor CU, one can construct a merge list of transform candidates that is associated to the TIMD merge list and is filled accordingly. This list can be based on the same neighboring CUs that contribute to the construction of the TIMD merge list. In particular, when a CU is added to the TIMD merge list of current CU, a pair of its transform types (i.e. horizontal and vertical) is also associated to the candidate. Then, once the best candidate is selected based on the algorithm of the TIMD merge, the transform pair of the best candidate is also inherited to be used for the current TIMD merge CU.
Alternatively, one can construct an independent list of transform candidates that are detached from the TIMD merge list. In this approach, once a TIMD merge candidate is identified in the neighborhood, its horizontal and vertical transforms are considered separately to construct two merge lists of transform candidates, namely horizontal transform merge list and vertical transform merge list. The particular step needed for this approach is to associate a cost to each transform type when adding them to the transform merge lists. To this end, one can use the same template-based cost that is computed for the TIMD merge candidate. Finally, once all TIMD merge candidates are processed and their transform types have contributed to the two transform merge lists, the final transform types of current blocks can be inherited separately from each list. For example, one can take the first element of each transform merge list for horizontal and vertical transform type of current CU.
Currently in ECM, an MTS index is associated with each CU which can take different values up to 6. Then, depending on CU characteristics (e.g. coding mode, block size etc. ) , this index is translated into a pair of horizontal and vertical transform types. This means that current design is not allowing all possible combinations of the horizontal and vertical transform types for a given block. However, combinations of the horizontal and vertical transform types are allowed that are currently not provided by ECM. This is due to the fact that the merge list of transform types is constructed based on neighboring CUs with different CU characteristics, for which the set of possible transform types combinations might be different than the current CU. As a result, current algorithm can implicitly inherit new combinations of transform types.
Figure 14 shows an example of inheriting horizontal and vertical transform types from a neighbor CU. In this figure, first neighbor CU 4 is selected by the TIMD merge algorithm to provide intra prediction modes. Then the second stage also inherits both transform types from CU 4 to be used for current TIMD merge block. Figure 14 illustrates an inheriting of horizontal and vertical transform types from the same TIMD merge candidate (CU 4) that provides the intra prediction modes.
Furthermore, Figure 15 shows an example of inheriting a pair of transform type that is not explicitly provided for current CU. In this figure, the neighboring candidate CU that has been selected from the TIMD merge of current block has a smaller block size. Therefore, the set of possible transform type pairs that are allowed by ECM is different from current CU that is larger. As a result, when current CU inherits its transform type information along with prediction information from the selected candidate,
then it will end up with a pair of transform types (i.e. DCT5 and DCT7 in this simplified example) that are not typically allowed by ECM for current block, due to the imitations of explicit transform type signaling. Figure 15 illustrates an implicit transform type inheritance allowing combinations of transform types that are not explicitly provided by ECM.
In an embodiment of the transform inherit method, each TIMD merge candidate CU also stores in the merge list, its choice of horizontal and vertical transform. When current CU selects a candidate from this list, it also inherits the pair of transform type that is associated with that candidate in the merge list.
In an embodiment, when current CU inherits a pair of horizontal and vertical transform types from a selected TIMD merge candidate, the inherited pair of transform is not included in the list of transform pairs that are allowed by MTS indexes in ECM.
In an embodiment, two separate merge lists are considered for horizontal and vertical transform types, which are both detached from the construction of the TIMD merge list. In this embodiment, every time a new TIMD merge CU is identified, its horizontal and vertical transform types separately contribute to their corresponding transform merge lists. Finally, the transform inheritance process for current TIMD merge CU will consist of separately selecting one transform type from each transform merge list. For example, one can take the first element in each merge list.
With reference to Fig. 11 to Fig. 15 embodiments of the present disclosure are described in other words in the following wherein the details set out above may, individually or in combination, be combined with any of the embodiments set out below to result into further embodiments. Further, all embodiments may be combined with each other even those relating to different ones of the aspects of the present application.
An embodiment relates to a method of block-based decoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and decoding one or more tool-selection syntax elements for the first block from a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstructing the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, reconstructing the first block using the merge list of candidate coding parametrizations, comparing, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, and modifying
the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
According to the embodiment, the testing the test set of intra prediction modes comprises determining for each intra prediction mode of the test set a prediction error by applying the respective intra prediction mode to a first previously decoded neighboring portion of the picture, neighboring the first block so as to predict a second previously decoded portion of the picture, neighboring the first block and the first previously decoded neighboring portion, and measuring the prediction error with respect to the second portion.
According to the embodiment, the method further comprises, if the selected intra-coding tool is the second tool, decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and reconstructing the first block based on the selected coding parametrization.
According to the embodiment, the comparing is performed so that the similarity depends on one or more of mode indices of the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively, fusion weights and/or ranks associated with the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively; wide-angle flags associated with the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively.
According to the embodiment, the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities comprises thresholding the similarities and not performing the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes in case of the similarities not exceeding a predetermined threshold.
According to the embodiment, the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities comprises one or more of excluding candidate coding parametrizations from the merge list of candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, reducing a weight of candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, at which same contribute to the merge list of candidate coding parametrizations, and modifying candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, with respect to the one or
more candidate intra prediction modes.
An embodiment relates to a method of block-based encoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors in the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and encoding one or more tool-selection syntax elements for the first block into a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encoding the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, encoding the first block using the merge list of candidate coding parametrizations, comparing, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, and modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and a decoding module configured to decode one or more tool-selection syntax elements for the first block from a data stream, wherein the decoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstruct the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, reconstruct the first block using the merge list of candidate coding parametrizations, a comparing module configured to compare, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, and a modifying module configured to modify the
merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a testing module configured to a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors in the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and an encoding module configured to encode one or more tool-selection syntax elements for the first block into a data stream, wherein the encoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encode the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, encode the first block using the merge list of candidate coding parametrizations, a comparing module configured to compare, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, and a modifying module configured to modify the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
An embodiment relates to a method of block-based decoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and reconstructing the first block using the merge list of candidate coding parametrizations, wherein the construing a merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and
the first block on the other hand.
According to the embodiment, the construing the merge list of candidate coding parametrizations comprises populating the merge list of candidate coding parametrizations with coding parametrizations of one or more blocks of the second blocks in the neighborhood, whose size or aspect ratio fulfills a predetermined criterion with respect to a size or aspect ratio of the first block, and excluding coding parametrizations of one or more blocks of the second blocks in the neighborhood, whose size or aspect ratio does not fulfill the predetermined criterion with respect to the size or aspect ratio of the first block.
According to the embodiment, whether the predetermined criterion is fulfilled depends on one or more of equality in number of pixels equality in width and equality in height equality in aspect ratio and equality in number of pixels, an absolute difference, or one minus a ratio in number of pixels, falling below a first threshold, an absolute difference, or one minus a ratio in width, falling below a second threshold; an absolute difference, or one minus a ratio in height, falling below a third threshold.
An embodiment relates to a method of block-based encoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and encoding the first block using the merge list of candidate coding parametrizations, wherein the construing a merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and a reconstructing module configured to reconstruct the first block using the merge list of candidate coding parametrizations, wherein the construing module is configured to construe the merge list of candidate coding parametrizations depending on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and an encoding module configured to encode the first block using the merge list of candidate coding
parametrizations, wherein the construing module is configured to construe the merge list of candidate coding parametrizations depending on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
An embodiment relates to a method of block-based decoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, decoding one or more tool-selection syntax elements for the first block from a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstructing the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, reconstructing the first block based on the merge list of candidate coding parametrizations, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
According to the embodiment, the testing the test set of intra prediction modes comprises determining for each intra prediction mode of the test set a prediction error by applying the respective intra prediction mode to a first previously decoded neighboring portion of the picture, neighboring a first block so as to predict a second previously decoded portion of the picture, neighboring the first block and the first previously decoded neighboring portion, and measuring the prediction error with respect to the second portion.,
According to the embodiment, the method further comprises, if the selected intra-coding tool is the second tool, decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and reconstructing the first block based on the selected coding parametrization.
According to the embodiment, the method further comprises construing a list of most probable intra prediction modes for the first block based on neighboring second blocks, if the selected intra-coding tool is a third tool, decoding from the data stream a pointer and reconstructing the first block using an intra-prediction mode selected out of the list of most probable intra prediction modes using the pointer, wherein the test set of at least one or more of the intra prediction modes is further populated using the list of most probable intra prediction modes.
According to the embodiment, the test set of at least one or more of the intra prediction modes is
populated using the merge list of candidate coding parametrizations by checking a similarity between the candidate coding parametrizations in the merge list of candidate coding parametrizations one the one hand and the list of most probable intra prediction modes, on the other hand, and adding one or more of the candidate coding parametrizations in the merge list of candidate coding parametrizations to the test set of at least one or more of the intra prediction modes if the similarity falls below a certain threshold.
According to the embodiment, the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations by checking a number of second blocks in the neighborhood of the first block, or of candidate coding parametrizations in the merge list of candidate coding parametrizations, and adding one or more of the candidate coding parametrizations in the merge list of candidate coding parametrizations to the test set of at least one or more of the intra prediction modes if the number exceeds a certain threshold.
An embodiment relates to a method of block-based encoding of a picture, the method comprising: testing a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors the test set, selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, and construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, encoding one or more tool-selection syntax elements for the first block into a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encoding the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, encoding the first block based on the merge list of candidate coding parametrizations, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and a construing module configured to construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, a decoding module configured to decode one or more tool-selection syntax elements for the first block from a data stream, where the decoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, reconstruct the first block using the first set of one or
more intra prediction modes, and if the selected intra-coding tool is a second tool, reconstruct the first block based on the merge list of candidate coding parametrizations, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors the test set, a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, and a construing module configured to construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, an encoding module configured to encode one or more tool-selection syntax elements for the first block into a data stream, wherein encoding module is configured to depending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block, if the selected intra-coding tool is a first tool, encode the first block using the first set of one or more intra prediction modes, and if the selected intra-coding tool is a second tool, encode the first block based on the merge list of candidate coding parametrizations, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
An embodiment relates to a method of block-based decoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and reconstructing the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, and decoding a prediction residual for the first block from the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual.
According to the embodiment, the method further comprises decoding one or more tool-selection syntax elements for the first block from a data stream, depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block, if the selected intra-coding
tool is a first tool, reconstructing the first block by intra-predicting the first block to obtain a predictor for the first block, decoding a transform domain indictor for the first block, from the data stream, selecting, using the transform domain pointer, a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, and decoding a prediction residual for the first block from the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual, and if the selected intra-coding tool is a second tool, reconstructing the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, and decoding a prediction residual for the first block from the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual.
According to the embodiment, the selecting, using the transform domain pointer, the predetermined residual transform domain out of the plurality of residual transforms domains supported by the decoder takes place by forming a list of residual transforms domains so as to comprise a proper subset of the residual transforms domains of the plurality of residual transforms domains, and applying the transform domain pointer onto the list of residual transforms domains so as to select the predetermined residual transform domain, and wherein the list of residual transforms domains does not comprise the residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems.
An embodiment relates to a method of block-based encoding of a picture, the method comprising: construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and encoding a merge candidate indicator into the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and encoding the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the encoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for encoding a second block from which the selected coding parametrization stems, and encoding a prediction residual for the first block into the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual.
An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus
comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, and a decoding module configured to decode a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and a reconstructing module configured to reconstruct the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, and decoding a prediction residual for the first block from the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual.
An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, and an encoding module configured to encode a merge candidate indicator into the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, and an encoding module configured to encode the first block based on the selected coding parametrization by applying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block, selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, and encoding a prediction residual for the first block into the data stream in the predetermined residual transform domain and correcting the predictor for the first block using the prediction residual.
An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
An embodiment relates to a method of block-based decoding of a picture, the method comprising: forming a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, forming a second list of vertical residual transforms based on the second blocks, performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to
result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, decoding a prediction residual for the first block from the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, and correcting a predictor for the first block using the prediction residual.
According to the embodiment, the method further comprises decoding from the data stream a first pointer and a second pointer for the first block, and performing the first selection based on the first pointer and the second selection based on the second pointer.
According to the embodiment, the method further comprises testing different combinations of a horizontal and a vertical residual transform with respect to rate-distortion in an already decoded neighbourhood of the first block to obtain rate-distortion values for the different combination; and performing the first selection and the second selection based on the rate-distortion values.
An embodiment relates to a method of block-based encoding of a picture, the method comprising: forming a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, forming a second list of vertical residual transforms based on the second blocks, performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, encoding a prediction residual for the first block into the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, and correcting a predictor for the first block using the prediction residual.
An embodiment relates to an apparatus for block-based decoding of a picture, the apparatus comprising: a forming module configured to form a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, a forming module configured to form a second list of vertical residual transforms based on the second blocks, a performing module configured to perform a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data
stream, a decoding module configured to decode a prediction residual for the first block from the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, and a correcting module configured to correct a predictor for the first block using the prediction residual.
An embodiment relates to an apparatus for block-based encoding of a picture, the apparatus comprising: a forming module configured to form a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block, a forming module configured to form a second list of vertical residual transforms based on the second blocks, a performing module configured to perform a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream, an encoding module configured to encode a prediction residual for the first block into the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, and a correcting module configured to correct a predictor for the first block using the prediction residual.
An embodiment relates to a data stream having a picture encoded thereinto using the method of block-based encoding of a picture.
Although some aspects of the disclosed concept have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or a device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
Fig. 16 is a block diagram illustrating an electronic device 900 according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as a laptop, a desktop, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are described as examples only, and are not intended to limit implementations of the present disclosure described and/or claimed herein. The device 900 includes a computing unit 901 to perform various appropriate actions and processes according to computer program instructions stored in a read only memory (ROM) 902, or loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data for the operation of the storage device 900 can also be stored.
The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input/output (I/O) interface 905 is also connected to the bus 904.
Components in the device 900 are connected to the I/O interface 905, including: an input unit 906, such as a keyboard, a mouse; an output unit 907, such as various types of displays, speakers; a storage unit 908, such as a disk, an optical disk; and a communication unit 909, such as network cards, modems, wireless communication transceivers, and the like. The communication unit 909 allows the device 900 to exchange information/data with other devices through a computer network such as the Internet and/or various telecommunication networks. The computing unit 901 may be formed of various general-purpose and/or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU) , graphics processing unit (GPU) , various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processor (DSP) , and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as an image processing method. For example, in some embodiments, the image processing method may be implemented as computer software programs that are tangibly embodied on a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program may be loaded and/or installed on the device 900 via the ROM 902 and/or the communication unit 909. When a computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image processing method described above may be performed. In some embodiments, the computing unit 901 may be configured to perform the image processing method in any other suitable manner (e.g., by means of firmware) .
Various implementations of the systems and techniques described herein above may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA) , application specific integrated circuits (ASIC) , application specific standard products (ASSP) , system-on-chip (SOC) , complex programmable logic device (CPLD) , computer hardware, firmware, software, and/or combinations thereof. These various embodiments may include being implemented in one or more computer programs executable and/or interpretable on a programmable system including at least one programmable processor, and the programmable processor may be a special-purpose or general-purpose programmable processor, and may receive data and instructions from a storage system, at least one input device and at least one output device, and may transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general computer, a dedicated computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions and/or operations specified in the flow diagrams and/or block diagrams is performed. The
program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on a machine and partly on a remote machine or entirely on a remote machine or server.
In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM) , read-only memories (ROM) , erasable programmable read-only memories (EPROM or flash memory) , fiber optics, compact disc read-only memories (CD-ROM) , optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
To provide interaction with a user, the systems and techniques described herein may be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) ) for displaying information for the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide an input to the computer. Other types of devices can also be used to provide interaction with the user, for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback) ; and may be in any form (including acoustic input, voice input, or tactile input) to receive the input from the user.
The systems and techniques described herein may be implemented on a computing system that includes back-end components (e.g., as a data server) , or a computing system that includes middleware components (e.g., an application server) , or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein) , or a computer system including such a backend components, middleware components, front-end components or any combination thereof. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network) . Examples of the communication network includes: Local Area Networks (LAN) , Wide Area Networks (WAN) , the Internet and blockchain networks.
The computer system may include a client and a server. The Client and server are generally remote from each other and usually interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business expansion in traditional physical hosts and virtual private servers ( "VPS" for short) . The server may also be a server of a distributed system, or a server combined with a blockchain.
It should be understood that the steps may be reordered, added or deleted by using the various forms of flows shown above. For example, the steps described in the present disclosure may be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions in the present disclosure can be achieved, and no limitation is imposed herein.
The above-mentioned specific embodiments do not limit the scope of protection of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and replacements may be made depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present disclosure should be included within the protection scope of the present disclosure.
References
H. 264: Advanced video coding for generic audiovisual services, https: //www. itu. int/rec/T-REC-
H. 264-202108-P/en
H. 265: High efficiency video coding, https: //www. itu. int/rec/T-REC-H. 265-202108-P/en
H. 266: Versatile video coding, https: //www. itu. int/rec/T-REC-H. 266-202008-I/en AV1 Bitstream & Decoding Process Specification, http://aomedia.org/av1/specification/
Claims (41)
- A method of block-based decoding of a picture, the method comprising:testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set,selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, anddecoding one or more tool-selection syntax elements for the first block from a data stream,depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,reconstructing the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block,reconstructing the first block using the merge list of candidate coding parametrizations,comparing, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, andmodifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
- The method of claim 1, wherein the testing the test set of intra prediction modes comprisesdetermining for each intra prediction mode of the test set a prediction error by applying the respective intra prediction mode to a first previously decoded neighboring portion of the picture, neighboring the first block so as to predict a second previously decoded portion of the picture, neighboring the first block and the first previously decoded neighboring portion, and measuring the prediction error with respect to the second portion.
- The method of claim 1 or 2, further comprisingif the selected intra-coding tool is the second tool,decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, andreconstructing the first block based on the selected coding parametrization.
- The method of any one of claims 1 to 3, wherein the comparing is performed so that the similarity depends on one or more ofmode indices of the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively,fusion weights and/or ranks associated with the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively;wide-angle flags associated with the one or more candidate intra-prediction modes and the first set of one or more intra prediction modes, respectively.
- The method of any one of claim 1 to 4, wherein the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities comprisesthresholding the similarities and not performing the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes in case of the similarities not exceeding a predetermined threshold.
- The method of any one of claim 1 to 4, wherein the modifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities comprises one or more ofexcluding candidate coding parametrizations from the merge list of candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes,reducing a weight of candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, at which same contribute to the merge list of candidate coding parametrizations, andmodifying candidate coding parametrizations whose one or more candidate intra prediction modes are, according to the similarity determined for same, closer than a predetermined threshold to the first set of one or more intra prediction modes, with respect to the one or more candidate intra prediction modes.
- A method of block-based encoding of a picture, the method comprising:testing a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors in the test set,selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, andencoding one or more tool-selection syntax elements for the first block into a data stream,depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,encoding the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,construing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block,encoding the first block using the merge list of candidate coding parametrizations,comparing, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, andmodifying the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
- An apparatus for block-based decoding of a picture, the apparatus comprising:a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors in the test set,a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, anda decoding module configured to decode one or more tool-selection syntax elements for the first block from a data stream, wherein the decoding module is configured todepending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,reconstruct the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block,reconstruct the first block using the merge list of candidate coding parametrizations,a comparing module configured to compare, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, anda modifying module configured to modify the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
- An apparatus for block-based encoding of a picture, the apparatus comprising:a testing module configured to a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors in the test set,a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, andan encoding module configured to encode one or more tool-selection syntax elements for the first block into a data stream, wherein the encoding module is configured todepending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,encode the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block,encode the first block using the merge list of candidate coding parametrizations,a comparing module configured to compare, for each candidate coding parametrization, one or more candidate intra-prediction modes of the respective candidate coding parametrization with the first set of one or more intra prediction modes so as to obtain a similarity for each candidate coding parametrization, anda modifying module configured to modify the merge list of candidate coding parametrizations and/or the first set of one or more intra prediction modes depending on the similarities obtained for the candidate coding parametrizations so as to render the one or more candidate intra-prediction modes of the candidate coding parametrizations in the merge list of candidate coding parametrizations on the one hand and the first set of one or more intra prediction modes on the other hand more diverse.
- Data stream having a picture encoded thereinto using the method of block-based encoding of a picture according to claim 7.
- A method of block-based decoding of a picture, the method comprising:construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, andreconstructing the first block using the merge list of candidate coding parametrizations,wherein the construing a merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- The method of claim 11, the construing the merge list of candidate coding parametrizations comprisingpopulating the merge list of candidate coding parametrizations with coding parametrizations of one or more blocks of the second blocks in the neighborhood, whose size or aspect ratio fulfills a predetermined criterion with respect to a size or aspect ratio of the first block, and excluding coding parametrizations of one or more blocks of the second blocks in the neighborhood, whose size or aspect ratio does not fulfill the predetermined criterion with respect to the size or aspect ratio of the first block.
- The method of claim 12, wherein whether the predetermined criterion is fulfilled depends on one or more of1) equality in number of pixels2) equality in width and equality in height3) equality in aspect ratio and equality in number of pixels,4) an absolute difference, or one minus a ratio in number of pixels, falling below a first threshold,5) an absolute difference, or one minus a ratio in width, falling below a second threshold;6) an absolute difference, or one minus a ratio in height, falling below a third threshold.
- A method of block-based encoding of a picture, the method comprising:construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, andencoding the first block using the merge list of candidate coding parametrizations,wherein the construing a merge list of candidate coding parametrizations depends on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- An apparatus for block-based decoding of a picture, the apparatus comprising:a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, anda reconstructing module configured to reconstruct the first block using the merge list of candidate coding parametrizations,wherein the construing module is configured to construe the merge list of candidate coding parametrizations depending on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- An apparatus for block-based encoding of a picture, the apparatus comprising:a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, andan encoding module configured to encode the first block using the merge list of candidate coding parametrizations,wherein the construing module is configured to construe the merge list of candidate coding parametrizations depending on, for each of the one or more second blocks in the neighborhood of the first block, a size or aspect-ratio comparison between the respective second block on the one hand and the first block on the other hand.
- Data stream having a picture encoded thereinto using the method of block-based encoding of a picture according to claim 14.
- A method of block-based decoding of a picture, the method comprising:testing a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors the test set,selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, andconstruing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block,decoding one or more tool-selection syntax elements for the first block from a data stream,depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,reconstructing the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,reconstructing the first block based on the merge list of candidate coding parametrizations,wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
- The method of claim 18, wherein the testing the test set of intra prediction modes comprisesdetermining for each intra prediction mode of the test set a prediction error by applying the respective intra prediction mode to a first previously decoded neighboring portion of the picture, neighboring a first block so as to predict a second previously decoded portion of the picture, neighboring the first block and the first previously decoded neighboring portion, and measuring the prediction error with respect to the second portion.,
- The method of claim 18 or 19, further comprisingif the selected intra-coding tool is the second tool,decoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, andreconstructing the first block based on the selected coding parametrization.
- The method of any of claims 18 to 20, further comprising:construing a list of most probable intra prediction modes for the first block based on neighboring second blocks,if the selected intra-coding tool is a third tool,decoding from the data stream a pointer andreconstructing the first block using an intra-prediction mode selected out of the list of most probable intra prediction modes using the pointer,wherein the test set of at least one or more of the intra prediction modes is further populated using the list of most probable intra prediction modes.
- The method of claim 21, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations bychecking a similarity between the candidate coding parametrizations in the merge list of candidate coding parametrizations one the one hand and the list of most probable intra prediction modes, on the other hand, and adding one or more of the candidate coding parametrizations in the merge list of candidate coding parametrizations to the test set of at least one or more of the intra prediction modes if the similarity falls below a certain threshold.
- The method of any one of claims 18 to 22, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations bychecking a number of second blocks in the neighborhood of the first block, or of candidate coding parametrizations in the merge list of candidate coding parametrizations, and adding one or more of the candidate coding parametrizations in the merge list of candidate coding parametrizations to the test set of at least one or more of the intra prediction modes if the number exceeds a certain threshold.
- A method of block-based encoding of a picture, the method comprising:testing a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors the test set,selecting a first set of one or more intra prediction modes out of the test set based on the prediction errors, andconstruing a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block,encoding one or more tool-selection syntax elements for the first block into a data stream,depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,encoding the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,encoding the first block based on the merge list of candidate coding parametrizations,wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
- An apparatus for block-based decoding of a picture, the apparatus comprising:a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already decoded neighborhood of a first block so as to obtain prediction errors the test set,a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, anda construing module configured to construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block,a decoding module configured to decode one or more tool-selection syntax elements for the first block from a data stream, where the decoding module is configured todepending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,reconstruct the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,reconstruct the first block based on the merge list of candidate coding parametrizations,wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
- An apparatus for block-based encoding of a picture, the apparatus comprising:a testing module configured to test a test set of intra prediction modes with respect to a prediction error in an already encoded neighborhood of a first block so as to obtain prediction errors the test set,a selecting module configured to select a first set of one or more intra prediction modes out of the test set based on the prediction errors, anda construing module configured to construe a merge list of candidate coding parametrizations for the first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block,an encoding module configured to encode one or more tool-selection syntax elements for the first block into a data stream, wherein encoding module is configured todepending on the one or more tool-selection syntax elements, select one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool,encode the first block using the first set of one or more intra prediction modes, andif the selected intra-coding tool is a second tool,encode the first block based on the merge list of candidate coding parametrizations, wherein the test set of at least one or more of the intra prediction modes is populated using the merge list of candidate coding parametrizations.
- Data stream having a picture encoded thereinto using the method of block-based encoding of a picture according to claim 24.
- A method of block-based decoding of a picture, the method comprising:construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, anddecoding a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, andreconstructing the first block based on the selected coding parametrization byapplying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block,selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, anddecoding a prediction residual for the first block from the data stream in the predetermined residual transform domain andcorrecting the predictor for the first block using the prediction residual.
- The method of claim 28, further comprisingdecoding one or more tool-selection syntax elements for the first block from a data stream,depending on the one or more tool-selection syntax elements, selecting one of a plurality of intra-coding tools for the first block,if the selected intra-coding tool is a first tool, reconstructing the first block byintra-predicting the first block to obtain a predictor for the first block,decoding a transform domain indictor for the first block, from the data stream,selecting, using the transform domain pointer, a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, anddecoding a prediction residual for the first block from the data stream in the predetermined residual transform domain andcorrecting the predictor for the first block using the prediction residual, andif the selected intra-coding tool is a second tool, reconstructing the first block based on the selected coding parametrization byapplying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block,selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, anddecoding a prediction residual for the first block from the data stream in the predetermined residual transform domain andcorrecting the predictor for the first block using the prediction residual.
- The method of claim 29,wherein the selecting, using the transform domain pointer, the predetermined residual transform domain out of the plurality of residual transforms domains supported by the decoder takes place byforming a list of residual transforms domains so as to comprise a proper subset of the residual transforms domains of the plurality of residual transforms domains, andapplying the transform domain pointer onto the list of residual transforms domains so as to select the predetermined residual transform domain, andwherein the list of residual transforms domains does not comprise the residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems.
- A method of block-based encoding of a picture, the method comprising:construing a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, andencoding a merge candidate indicator into the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, andencoding the first block based on the selected coding parametrization byapplying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block,selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the encoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for encoding a second block from which the selected coding parametrization stems, andencoding a prediction residual for the first block into the data stream in the predetermined residual transform domain andcorrecting the predictor for the first block using the prediction residual.
- An apparatus for block-based decoding of a picture, the apparatus comprising:a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for decoding the respective second block, anda decoding module configured to decode a merge candidate indicator from the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, anda reconstructing module configured to reconstruct the first block based on the selected coding parametrization byapplying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block,selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, anddecoding a prediction residual for the first block from the data stream in the predetermined residual transform domain andcorrecting the predictor for the first block using the prediction residual.
- An apparatus for block-based encoding of a picture, the apparatus comprising:a construing module configured to construe a merge list of candidate coding parametrizations for a first block based on a coding parametrization of one or more second blocks in a neighborhood of the first block, wherein, for each of the one or more second blocks, the coding parametrization includes an indication of one or more intra-prediction modes used for encoding the respective second block, andan encoding module configured to encode a merge candidate indicator into the data stream, indicating a selected coding parametrization out of the merge list of candidate coding parametrizations, andan encoding module configured to encode the first block based on the selected coding parametrization byapplying one or more intra-prediction modes of the selected coding parametrization to the first block to obtain a predictor for the first block,selecting a predetermined residual transform domain out of a plurality of residual transforms domains supported by the decoder, so as to coincide with a residual transform domain being indicated by the selected coding parametrization as used for decoding a second block from which the selected coding parametrization stems, andencoding a prediction residual for the first block into the data stream in the predetermined residual transform domain andcorrecting the predictor for the first block using the prediction residual.
- Data stream having a picture encoded thereinto using the method of block-based encoding of a picture according to claim 31.
- A method of block-based decoding of a picture, the method comprising:forming a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block,forming a second list of vertical residual transforms based on the second blocks,performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream,decoding a prediction residual for the first block from the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, andcorrecting a predictor for the first block using the prediction residual.
- The method of claim 35, comprisingdecoding from the data stream a first pointer and a second pointer for the first block, andperforming the first selection based on the first pointer and the second selection based on the second pointer.
- The method of claim 35, comprisingtesting different combinations of a horizontal and a vertical residual transform with respect to rate-distortion in an already decoded neighbourhood of the first block to obtain rate-distortion values for the different combination; andperforming the first selection and the second selection based on the rate-distortion values.
- A method of block-based encoding of a picture, the method comprising:forming a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block,forming a second list of vertical residual transforms based on the second blocks,performing a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream,encoding a prediction residual for the first block into the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, andcorrecting a predictor for the first block using the prediction residual.
- An apparatus for block-based decoding of a picture, the apparatus comprising:a forming module configured to form a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block,a forming module configured to form a second list of vertical residual transforms based on the second blocks,a performing module configured to perform a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream,a decoding module configured to decode a prediction residual for the first block from the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, anda correcting module configured to correct a predictor for the first block using the prediction residual.
- An apparatus for block-based encoding of a picture, the apparatus comprising:a forming module configured to form a first list of horizontal residual transforms based on second blocks in the neighborhood of a first block,a forming module configured to form a second list of vertical residual transforms based on the second blocks,a performing module configured to perform a first selection out of the first list and a second selection out of the second list so as to obtain a selected horizontal residual transform and a selected vertical residual transform, wherein the first and second selection allow the selected horizontal residual transform and the selected vertical residual transform to result, by a concatenation of the selected horizontal residual transform and the selected vertical residual transform, into a separable residual transform domain which differs, for each second block, from a residual transform domain in which a prediction residual of the respective second block is coded in a data stream,an encoding module configured to encode a prediction residual for the first block into the data stream in the separable residual transform domain resulting from the concatenation of the selected horizontal residual transform and the selected vertical residual transform, anda correcting module configured to correct a predictor for the first block using the prediction residual.
- Data stream having a picture encoded thereinto using the method of block-based encoding of a picture according to claim 38.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24166515 | 2024-03-26 | ||
| EP24166515.7 | 2024-03-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025200081A1 true WO2025200081A1 (en) | 2025-10-02 |
Family
ID=90482062
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/091020 Pending WO2025200081A1 (en) | 2024-03-26 | 2024-04-30 | Method and apparatus for template-based intra mode derivation and encoder/decoder including the same |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025200081A1 (en) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210195227A1 (en) * | 2018-03-30 | 2021-06-24 | Electronics And Telecommunications Research Institute | Image encoding/decoding method and device, and recording medium in which bitstream is stored |
| WO2023044900A1 (en) * | 2021-09-27 | 2023-03-30 | Oppo广东移动通信有限公司 | Coding/decoding method, code stream, coder, decoder, and storage medium |
| WO2023194105A1 (en) * | 2022-04-07 | 2023-10-12 | Interdigital Ce Patent Holdings, Sas | Intra mode derivation for inter-predicted coding units |
| WO2024044404A1 (en) * | 2022-08-26 | 2024-02-29 | Beijing Dajia Internet Information Technology Co., Ltd. | Methods and devices using intra block copy for video coding |
| CN117730531A (en) * | 2021-08-30 | 2024-03-19 | 北京达佳互联信息技术有限公司 | Method and apparatus for decoder side intra mode derivation |
-
2024
- 2024-04-30 WO PCT/CN2024/091020 patent/WO2025200081A1/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210195227A1 (en) * | 2018-03-30 | 2021-06-24 | Electronics And Telecommunications Research Institute | Image encoding/decoding method and device, and recording medium in which bitstream is stored |
| CN117730531A (en) * | 2021-08-30 | 2024-03-19 | 北京达佳互联信息技术有限公司 | Method and apparatus for decoder side intra mode derivation |
| WO2023044900A1 (en) * | 2021-09-27 | 2023-03-30 | Oppo广东移动通信有限公司 | Coding/decoding method, code stream, coder, decoder, and storage medium |
| WO2023194105A1 (en) * | 2022-04-07 | 2023-10-12 | Interdigital Ce Patent Holdings, Sas | Intra mode derivation for inter-predicted coding units |
| WO2024044404A1 (en) * | 2022-08-26 | 2024-02-29 | Beijing Dajia Internet Information Technology Co., Ltd. | Methods and devices using intra block copy for video coding |
Non-Patent Citations (2)
| Title |
|---|
| P. ANDRIVON (OFINNO), M. BLESTEL (OFINNO): "EE2-1.20: TIMD fusion with non-angular predictor", 33. JVET MEETING; 20240117 - 20240126; TELECONFERENCE; (THE JOINT VIDEO EXPLORATION TEAM OF ISO/IEC JTC1/SC29/WG11 AND ITU-T SG.16 ), 16 January 2024 (2024-01-16), XP030313954 * |
| R. G. YOUVALARI (XIAOMI), M. ABDOLI (XIAOMI): "AHG 12: TIMD merge mode", 33. JVET MEETING; 20240117 - 20240126; TELECONFERENCE; (THE JOINT VIDEO EXPLORATION TEAM OF ISO/IEC JTC1/SC29/WG11 AND ITU-T SG.16 ), 17 January 2024 (2024-01-17), XP030313981 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12621463B2 (en) | Method and apparatus for processing a video signal | |
| JP7474365B2 (en) | Improved predictor candidates for motion compensation. | |
| TWI616089B (en) | Decoder, encoder, and associated methods and computer programs | |
| CN115037946B (en) | Method and apparatus for encoding or decoding video data | |
| WO2023193551A9 (en) | Method and apparatus for dimd edge detection adjustment, and encoder/decoder including the same | |
| US20230388484A1 (en) | Method and apparatus for asymmetric blending of predictions of partitioned pictures | |
| WO2023193550A9 (en) | Method and apparatus for dimd region-wise adaptive blending, and encoder/decoder including the same | |
| CN120826907A (en) | Video encoding and decoding method, device, equipment, system, and storage medium | |
| WO2026044857A1 (en) | Method and apparatus using intra prediction fusion | |
| US20220295046A1 (en) | Method and device for processing video signal | |
| EP4642026A1 (en) | Method and apparatus for intra block copy fusion, and encoder/decoder including the same | |
| EP4730772A1 (en) | Method and apparatus for predicting sub-partitions of a block of a picture | |
| WO2026081281A1 (en) | Method and apparatus for processing sub-partitions of a block of a picture | |
| WO2025200069A1 (en) | Method and apparatus for occurrence-based intra coding, and encoder/decoder including the same | |
| EP4664878A1 (en) | Method and apparatus for obtaining one or more virtual intra prediction modes (vipms) | |
| EP4676046A1 (en) | Method and apparatus mutually excluding use of transform-based distortion metric and transform-less residual coding | |
| WO2026065632A1 (en) | Method and apparatus for partitioning of one or more blocks of a picture | |
| EP4664879A1 (en) | Method and apparatus for obtaining one or more intra prediction modes (ipms) | |
| WO2026081284A1 (en) | Method and apparatus using angular intra prediction | |
| KR20210036822A (en) | Method and apparatus for encoding/decoding image, recording medium for stroing bitstream |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24932678 Country of ref document: EP Kind code of ref document: A1 |