WO2025199802A1 - 编解码方法、码流、编码器、解码器以及存储介质 - Google Patents

编解码方法、码流、编码器、解码器以及存储介质

Info

Publication number
WO2025199802A1
WO2025199802A1 PCT/CN2024/084081 CN2024084081W WO2025199802A1 WO 2025199802 A1 WO2025199802 A1 WO 2025199802A1 CN 2024084081 W CN2024084081 W CN 2024084081W WO 2025199802 A1 WO2025199802 A1 WO 2025199802A1
Authority
WO
WIPO (PCT)
Prior art keywords
current block
transform
candidate
determining
mode
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/084081
Other languages
English (en)
French (fr)
Inventor
王凡
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Guangdong Oppo Mobile Telecommunications Corp Ltd
Original Assignee
Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Guangdong Oppo Mobile Telecommunications Corp Ltd filed Critical Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority to PCT/CN2024/084081 priority Critical patent/WO2025199802A1/zh
Publication of WO2025199802A1 publication Critical patent/WO2025199802A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/12Selection from among a plurality of transforms or standards, e.g. selection between discrete cosine transform [DCT] and sub-band transform or selection between H.263 and H.264
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding

Definitions

  • an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:
  • a transform coefficient of the current block is determined, and the transform coefficient of the current block is transformed according to the transform kernel to determine a residual block of the current block.
  • an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:
  • the transform coefficients of the current block are coded and the resulting coded bits are written into the bitstream.
  • an embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following: a quantization coefficient of a current block, a transform kernel index of the current block, a transform kernel group index of the current block, a minimum sample threshold, a minimum size threshold, a value of a first syntax element, a value of a second syntax element, a value of a third syntax element, a value of a fourth syntax element, a value of a fifth syntax element, a value of a sixth syntax element, and a value of a seventh syntax element;
  • the first syntax element is used to indicate the transform kernel group index of the current block
  • the second syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform kernel index
  • the third syntax element is used to indicate whether the current sequence allows the use of multiple transform kernel group selection technology
  • the fourth syntax element is used to indicate whether the current image allows the use of multiple transform kernel group selection technology
  • the fifth syntax element is used to indicate whether the current slice allows the use of multiple transform kernel group selection technology
  • the sixth syntax element is used to indicate the minimum sample threshold
  • the seventh syntax element is used to indicate the minimum size threshold.
  • an encoder comprising a first memory and a first processor; wherein,
  • a first memory for storing a computer program capable of running on the first processor
  • a second determining unit configured to determine a transform core group index of the current block, and determine a transform core group of the current block according to the transform core group index of the current block; and further configured to determine a transform core of the current block according to the transform core group;
  • an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first aspect or the method as described in the second aspect.
  • FIG1 is a flow chart diagram of a hybrid coding framework
  • FIG10 is a schematic diagram of a prediction process of a MIP mode
  • FIG12 is a schematic diagram of a histogram of gradient and intra prediction mode
  • FIG15 is a schematic diagram of weights of various modes under a GPM mode
  • FIG16A is a typical filter schematic diagram 1
  • FIG16B is a second schematic diagram of a typical filter
  • FIG17C is a third schematic diagram of a reconstructed sample region for training a filter
  • FIG21 is a schematic diagram of a process flow with LFNST transformation
  • FIG31 is a fourth flow chart of a decoding method provided in an embodiment of the present application.
  • FIG32 is a fifth flow chart of a decoding method provided in an embodiment of the present application.
  • FIG35 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application.
  • FIG37 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application.
  • FIG38 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application.
  • a coding block In video images, a coding block (CB) is generally represented by a first color component, a second color component, and a third color component. These three color components are a luminance component, a blue chrominance component, and a red chrominance component. Specifically, the luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbols Cb or U, and the red chrominance component is usually represented by the symbols Cr or V. Thus, video images can be represented in either the YCbCr or YUV format.
  • VTM VVC Test Model
  • JVET Joint Video Experts Team
  • LCU Largest Coding Unit
  • TMVP Temporal Motion Vector Prediction
  • DST Discrete Sine Transform
  • the basic process of a video codec is as follows: On the encoder side, an image is divided into blocks. Intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The prediction block is subtracted from the initial block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain a quantization coefficient matrix. This quantization coefficient matrix is entropy coded and output to the bitstream. On the decoder side, intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The bitstream is then parsed to obtain a quantization coefficient matrix. This quantization coefficient matrix is inversely quantized and inversely transformed to obtain a residual block.
  • the Template Matching (TM) method was first used in inter-frame prediction. It exploits the correlation between adjacent samples and uses areas surrounding the current block as templates.
  • the current block is encoded or decoded, its left and top sides have already been encoded and decoded according to the encoding order.
  • existing hardware decoder implementations cannot guarantee that the left and top sides have already been decoded when the current block begins decoding.
  • This refers to inter-frame blocks.
  • intra-frame coded blocks must use reconstructed samples from the left and top sides as reference samples. Theoretically, the left and top sides are available, which means that hardware design adjustments can be made to achieve this. Relatively speaking, the right and bottom sides are not available under current standards such as VVC's encoding order.
  • the degree of match can be measured using distortion metrics such as the sum of absolute differences (SAD), the sum of absolute transformed differences (SATD), and the mean-square error (MSE).
  • SATD generally uses the Hadamard transform. Smaller values of SAD, SATD, and MSE indicate a higher degree of match.
  • the cost is calculated using the prediction block of the template corresponding to the position and the reconstructed block of the template around the current block. In addition to searching the entire sample position, you can also search the sub-sample position, and determine the motion information of the current block based on the position with the highest degree of matching found. By utilizing the correlation between adjacent samples, the motion information that is appropriate for the template may also be the appropriate motion information for the current block.
  • the template matching method may not necessarily be applicable to all blocks, so some methods can be used to determine whether the current block uses the above template matching method, such as using a control switch in the current block to indicate whether the template matching method is used.
  • This template matching method is called decoding-end motion matching. Decoder-side Motion Vector Derivation (DMVD). Both the encoder and decoder can use templates to search and derive motion information or find better motion information based on the existing motion information. This method does not require the transmission of specific motion vectors or motion vector differences. Instead, both the encoder and decoder perform the same search rules to ensure consistent encoding and decoding. Template matching can improve compression performance, but it also requires a "search" on the decoder side, which introduces a certain degree of decoding complexity.
  • the reference sample on the lower left is also unavailable.
  • available reference samples or certain values or methods can be used for filling, or no filling can be performed.
  • MRL intra prediction methods can use more reference samples to improve coding efficiency.
  • Figure 4 shows an example of using four reference rows/columns.
  • IBC significantly improves the compression efficiency of screen content coding (SCC), and is therefore used in screen content coding from HEVC to VVC.
  • Screen content unlike camera-captured content, is computer-generated. Screen content is noise-free and contains text, computer graphics, and has sharp boundaries. Screen content often contains a large amount of repetitive content, as shown in Figure 9.
  • IBC can be considered an intra-frame prediction method, or a different type of prediction method independent of intra-frame and inter-frame prediction. IBC is highly efficient for encoding screen content and can also improve compression efficiency in natural sequences captured by cameras.
  • the MIP To predict a block of width W and height H, the MIP requires H reconstructed samples in the column to the left of the current block and W reconstructed samples in the row above the current block as input.
  • the MIP generates the prediction block in three steps: (a) reference sample averaging, (b) matrix multiplication, and (c) interpolation.
  • the core of the MIP is considered to be matrix multiplication. It can be thought of as a process that uses input samples (reference samples) to generate a prediction block using a matrix multiplication.
  • the MIP provides a variety of matrices, and different prediction methods are reflected in different matrices. Using different matrices for the same input sample will produce different results.
  • reference sample averaging and interpolation processes are a compromise between performance and complexity. For larger blocks, reference sample averaging can achieve a similar downsampling effect, allowing the input to fit into a smaller matrix, while interpolation achieves an upsampling effect. This eliminates the need to provide MIP matrices for every block size; instead, matrices of one or a few specific sizes are sufficient. As the demand for compression performance increases and hardware capabilities improve, more complex MIPs may appear in the next generation of standards.
  • MIP is somewhat similar to PLANAR, but it is obviously more complex and more flexible than PLANAR.
  • TIMD uses the diagonally filled area shown in Figure 11 as the template, and the template reference area (the grid-filled area) is the template's reference sample.
  • the decoder can use a specific intra-prediction mode to predict on the template and compare the predicted value with the reconstructed value to obtain the cost of the intra-prediction mode on the template. Examples include SAD, SATD, and SSE.
  • TIMD predicts several candidate intra-prediction modes on the template, obtains their costs on the template, and selects the one or two intra-prediction modes with the lowest costs as the intra-prediction mode for the current block.
  • TIMD uses the prediction performance of intra-frame prediction modes on a template to select the appropriate intra-frame prediction mode, and can weight the two intra-frame prediction modes based on their cost on the template.
  • the advantage of TIMD is that if the current block selects TIMD mode, the decoder does not need to indicate the specific intra-frame prediction mode to be used. Instead, the decoder can derive the selected intra-frame prediction mode through the above process, which saves some overhead.
  • DIMD analyzes the gradient of black points, such as horizontal gradient and vertical gradient, and adapts an intra-frame prediction mode according to its gradient. Analyzing all the points that need to be checked can obtain a result similar to the bar graph below. That is, the statistics of the number of points matched by each intra-frame prediction mode. Of course, the so-called bar graph is just to help understanding, and it can be implemented in a variety of simple forms in practice.
  • the current DIMD selects the two highest intra-frame prediction modes in the histogram, plus the PLANAR mode, and the prediction values of a total of three intra-frame prediction modes are weighted. The weights are related to the analysis results.
  • the rectangular areas to the left and above the current block are used as templates.
  • the height of the left template portion is generally the same as the height of the current block, and the width of the upper template portion is generally the same as the width of the current block, but can also be different.
  • the best matching position of the template is found in the reference image to determine the motion information, or motion vector, of the current block. This process can be roughly described as starting from a starting position in a reference image (Ref0) and searching within a certain range around it. Search rules, such as the search range and search step size, can be predefined. At each position, the degree of match between the template corresponding to that position and the templates surrounding the current block is calculated. The degree of match can be measured using distortion costs, such as SAD or SATD.
  • SATD generally uses transforms such as the Hadamard transform and MSE. Lower values of SAD, SATD, and MSE indicate a higher degree of match.
  • the cost is calculated using the predicted block of the template corresponding to that position and the reconstructed block of the template surrounding the current block. In addition to searching at whole-sample positions, searches can also be performed at sub-sample positions.
  • the motion information of the current block is determined based on the position with the highest degree of match. By using the correlation between adjacent samples, the motion information that is suitable for the template may also be the motion information that is suitable for the current block.
  • the template matching method may not necessarily be applicable to all blocks, so Some methods are used to determine whether the current block uses the above-mentioned template matching method, such as using a control switch in the current block to indicate whether the template matching method is used.
  • a classic template matching technology is called DMVD (Decoder side Motion Vector Derivation).
  • DMVD Decoder side Motion Vector Derivation
  • Both the encoder and the decoder can use templates to search to derive motion information or find better motion information based on the original motion information. It does not require the transmission of specific motion vectors or motion vector differences. Instead, both the encoder and the decoder perform searches using the same rules to ensure consistency in encoding and decoding.
  • the template matching method can improve compression performance, but it requires “searching" in the decoder as well, which brings a certain degree of decoder complexity.
  • ITMP can be considered a combination of IBC and TM.
  • applying TM to inter-frames can reduce the overhead of encoding MVs.
  • applying TM to IBC can reduce the overhead of encoding BVs.
  • An example is to use the matching block found by TM as the ITMP prediction block for the current block, without encoding BVs.
  • Figure 14 shows an example of ITMP.
  • the inverted L-shaped area in the upper left corner of the current block is used as a template.
  • the search is performed within a point-filled search range, which covers the reconstructed area.
  • the point-filled area shown in Figure 14 includes the current CTU (R1), the CTU to the upper left of R2, the CTU above R3, and the CTU to the left of R4. This is just an example; the search range may vary in actual applications. In this example, the best matching block is found in R2.
  • the VVC video codec standard includes an inter-frame prediction mode called Geometric Partitioning Mode (GPM).
  • GPSM Geometric Partitioning Mode
  • AVS3 video codec standard includes an inter-frame prediction mode called Angular Weighted Prediction (AWP). While these two modes have different names and implementations, they share common principles.
  • Traditional unidirectional prediction uses only one reference block of the same size as the current block.
  • Traditional bidirectional prediction uses two reference blocks of the same size as the current block, and the sample value of each point in the prediction block is the average of the corresponding positions in the two reference blocks, that is, all points in each reference block account for 50%.
  • Bidirectional weighted prediction allows the proportions of the two reference blocks to be different, such as all points in the first reference block account for 75% and all points in the second reference block account for 25%. However, all points in the same reference block have the same proportion.
  • Other optimization methods such as decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BIO) will cause some changes to the reference samples or prediction samples, but they are not related to the principles described above.
  • DMVR decoder-side motion vector refinement
  • BIO bidirectional optical flow
  • BIO can also be abbreviated as BDOF.
  • GPM or AWP also uses two reference blocks of the same size as the current block, but some sample positions use 100% of the sample values from the first reference block, while others use 100% of the sample values from the second reference block. In the boundary area, or transition zone, sample values from both reference blocks are used in a certain proportion. The weights in the boundary area also transition gradually. The specific distribution of these weights is determined by the GPM or AWP mode. The weight of each sample position is determined based on the GPM or AWP mode. Of course, in some cases, such as very small block sizes, some GPM or AWP modes may not guarantee that some sample positions use 100% of the sample values from the first reference block, while others use 100% of the sample values from the second reference block. Alternatively, GPM or AWP uses two reference blocks of different sizes, each taking a required portion as the reference block. Specifically, the portion with a non-zero weight is used as the reference block, while the portion with a zero weight is discarded.
  • GPM and AWP use different weighting methods.
  • GPM determines the angle and offset for each mode and then calculates a weight matrix for each mode.
  • AWP first creates a one-dimensional weight line and then uses a method similar to intra-frame angle prediction to fill the entire matrix with this one-dimensional weight line.
  • GPM is an inter-frame technology in VVC, but it can also use intra-frame prediction.
  • the two prediction modes of GPM can both be inter-frame prediction modes, one can be inter-frame prediction mode and the other can be intra-frame prediction mode, or both can be intra-frame prediction modes.
  • the general logic is to write them separately in the code stream as shown in Table 1.
  • the syntax elements indicating these three modes are input, namely partition_mode_idx, intra_pred_mode0_idx, and intra_pred_mode1_idx as shown in Table 1.
  • SGPM uses the template around the current block to sort the combinations of the three modes, obtaining a candidate list of combined modes. Only the candidate index of SGPM needs to be written into the bitstream.
  • SGPM can construct the candidate list of combined modes and derive one "partition" mode and two intra-frame prediction modes partition_mode_idx, intra_pred_mode0_idx, and intra_pred_mode1_idx based on the candidate index of SGPM, such as sgpm_cand_idx in Table 2.
  • ECM includes an EIP mode.
  • Figures 16A, 16B, and 16C illustrate several typical filters in this EIP mode.
  • the sample positions in the grid-filled areas represent the input samples (input of EIP), and the sample positions in the white-filled areas represent the output samples (output of EIP).
  • the predicted value for the sample in the lower right corner is usually required.
  • the values of the 15 samples to its left, upper left, and upper sides are known.
  • Each grid-filled position in the filter has a preset coefficient.
  • the sample value at each input position is multiplied by the coefficient at the corresponding position. All the multiplication results are summed up and normalized to obtain the predicted value for the position to be predicted.
  • the EIP mode may provide a pre-trained filter or a filter trained based on the reconstructed sample area surrounding the current block.
  • Figures 17A, 17B, and 17C illustrate the three types of reconstructed sample areas that can be selected for filter training in the EIP mode. Assuming that the current block and the reconstructed sample area used to train the filter coefficients follow similar patterns, the filter trained using the reconstructed sample area can be applied to the current block.
  • hybrid coding frameworks During encoding, commonly used hybrid coding frameworks first perform a prediction. This prediction leverages spatial or temporal correlation to produce an image identical or similar to the current block. While it's possible for the predicted block to be identical to the current block for a given block, it's difficult to guarantee this for all blocks in a video, especially in natural video or video captured by a camera. Irregular motion, distortion, occlusion, and brightness changes in video are difficult to fully predict. Therefore, hybrid coding frameworks subtract the predicted image from the original image of the current block to produce a residual image, or, in other words, subtract the predicted block from the current block to produce a residual block. This residual block is typically much simpler than the original image, so prediction can significantly improve compression efficiency.
  • the residual block isn't encoded directly; instead, it's usually first transformed. This transform converts the residual image from the spatial domain to the frequency domain to remove correlation. After the residual image is transformed to the frequency domain, since most of the energy is concentrated in the low-frequency region, the non-zero coefficients are concentrated in the upper left corner. Quantization is then used to further compress the image. Furthermore, since the human eye is less sensitive to high frequencies, a larger quantization step size can be used in high-frequency regions.
  • Figure 18 is a schematic diagram of a DCT transform. As shown in Figure 18, after the DCT transform of the original image, only the upper left corner has non-zero coefficients. Of course, this example performs a DCT transform on the entire image. In video codecs, images are processed by dividing them into blocks, so the transform is also performed on a block-by-block basis.
  • Transforms are very useful in typical video compression, but not all blocks require transforms. In some cases, transforms can even yield worse compression results than non-transforms. Therefore, in some standards, such as VVC, the encoder can choose whether to use transforms for the current block.
  • DCT-II is the most commonly used transform in video compression standards, and its base image is shown in Figure 19.
  • VVC can also use DCT8 type (DCT-VIII) and DST7 type (DST-VII).
  • DCT8 type DCT-VIII
  • DST7 type DST-VII
  • Table 3 shows the basic transform formulas of DCT2, DCT8 and DST7 for N-point input.
  • the DCT2, DCT8, and DST7 transforms used in the standards were split into two steps: horizontal and vertical one-dimensional transforms. For example, the horizontal transform was performed first, followed by the vertical transform, or the vertical transform was performed first, followed by the horizontal transform.
  • VVC supports transform kernels such as DCT2, DCT8, and DST7.
  • the DCT2, DCT8, and DST7 kernels used in VVC are horizontally and numerically separable, allowing independent horizontal and vertical transforms.
  • the encoder selects the appropriate transform kernel and transmits its index to the bitstream.
  • the decoder uses the index to determine the inverse transform kernel.
  • Different transform kernels can be selected for the horizontal and vertical directions, such as using DCT8 horizontally and DST7 vertically. This technique is generally referred to as MTS.
  • LFNST transform is used in VVC.
  • the above transforms such as DCT2, DCT8, and DST7 are called primary transforms.
  • LFNST is used after DCT2 transform and before quantization.
  • LFNST is used after inverse quantization and before inverse DCT2 transform. Because it is a transform based on DCT2 (primary transform), LFNST is a secondary transform.
  • Figure 20 is a schematic diagram of the encoding and decoding process without LFNST (secondary transform)
  • Figure 21 is a schematic diagram of the encoding and decoding process with LFNST (secondary transform).
  • the encoding side can directly inverse quantize the saved quantization coefficients instead of entropy decoding, because entropy coding is lossless.
  • Figure 22 shows a detailed encoding and decoding process involving LFNST (secondary transform).
  • LFNST performs a secondary transform on the low-frequency coefficients in the upper left corner after the base transform.
  • the base transform decorrelates the image, concentrating energy in the upper left corner.
  • the secondary transform further decorrelates the low-frequency coefficients of the base transform.
  • the result is intuitively shown in Figure 22.
  • 16 coefficients are input to the 4 ⁇ 4 LFNST, and the output is 8 coefficients.
  • 48 coefficients are input to the 8 ⁇ 8 LFNST, and the output is 8 coefficients for the 8 ⁇ 8 block and 16 coefficients for other blocks.
  • 8 coefficients are input to the 4 ⁇ 4 inverse LFNST, and the output is 16 coefficients.
  • 8 coefficients are input to the 8 ⁇ 8 block and 16 coefficients for other blocks.
  • the coefficients are input to the 8 ⁇ 8 inverse LFNST, and the output is 48 coefficients.
  • LFNST is only applied to intra-coded blocks.
  • Angular prediction tiles the values of reference samples at a specified angle onto the current block as the prediction value. This means that the predicted block will have a distinct directional texture, and the residual of the current block after angular prediction will also statistically exhibit significant angular characteristics. Therefore, the transform kernel selected by LFNST can be tied to the intra-prediction mode. That is, once the intra-prediction mode is determined, LFNST can only use the set of transform kernels corresponding to that intra-prediction mode.
  • the LFNST in VVC has a total of 4 groups of transform kernels, and each group can select 2 transform kernels.
  • Table 5 gives the intra prediction mode and Correspondence between transform kernel groups.
  • the cross-component prediction modes used for chroma intra prediction are 81 to 83, while these modes are not present in luma intra prediction.
  • LFNST's transform kernel can be transposed to handle more angles with a single transform kernel group. For example, modes 13 to 23 and 45 to 55 both correspond to transform kernel group 2, but 13 to 23 is clearly closer to the horizontal mode, while 45 to 55 is clearly closer to the vertical mode.
  • VVC's LFNST uses four sets of transform kernels, with the intra-prediction mode specifying which set to use. This leverages the correlation between the intra-prediction mode and the LFNST transform kernel, reducing the transmission of the selected LFNST transform kernel in the bitstream. Whether the current block uses LFNST, and if so, whether to use the first or second set within a set, is determined by the bitstream and certain conditions.
  • LFNST has more transform kernel groups, 35 in ECM.
  • the correspondence between the transform kernel group index (LFNST set index) and the intra prediction mode (Intra pred.mode) is shown in Table 6.
  • Each transform kernel group is more efficient for textures at the corresponding angle.
  • each transform kernel group can select three transform kernels.
  • LFNST is a horizontally and vertically inseparable transform. Because it involves a secondary transform, DCT2 can be called the base transform. This approach of performing DCT2 before LFNST is a compromise between performance and complexity. While directly performing the inseparable base transform is more efficient, it also incurs higher complexity, such as increased computational effort and storage space required for the transform kernel.
  • both NSPT and LFNST process transforms for textures at various angles. They may have multiple transform kernels, and a transform kernel may be specifically optimized for a specific angular texture.
  • NSPT and LFNST also include transform kernels for gradient textures.
  • these transform kernels can also be said to be trained KLTs (Karhunen-Loeve Transforms).
  • KLTs Kerhunen-Loeve Transforms
  • NSPT and LFNST have multiple transform kernels, each designed for a specific texture, including angular textures, gradient textures, and so on.
  • gradient textures can be further expanded to include horizontal gradient textures, vertical gradient textures, and diagonal gradient textures.
  • They have multiple transform kernel groups, and each intra-frame prediction mode can correspond to a transform kernel group. Each intra-frame prediction mode actually represents a texture feature. Therefore, the intra-frame prediction mode index is also a texture feature index.
  • a block In VVC and the current ECM, a block (CU or TU) can only derive a single texture feature index for deriving the LFNST/NSPT transform kernel set, thereby deriving a unique LFNST/NSPT transform kernel set.
  • This texture feature can be considered the link for using LFNST/NSPT.
  • this texture feature index is the intra prediction mode used by the current block.
  • DIMD and TIMD both weight the prediction values of two or more intra prediction modes.
  • SGPM weights the prediction values of two intra prediction modes using a weight matrix.
  • FIG25 is a schematic diagram of a network architecture for video encoding and decoding provided in an embodiment of the present application.
  • the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein electronic devices 13 to 1N can perform video interaction via the communication network 01.
  • the electronic devices can be various types of devices with video encoding and decoding capabilities.
  • the electronic devices can include mobile phones, tablet computers, personal computers, personal digital assistants, navigators, digital phones, video phones, televisions, sensor devices, servers, etc., and the embodiments of the present application do not limit this.
  • a network architecture of a video encoding and decoding system including a decoding method and an encoding method is provided.
  • the decoder or encoder in the embodiment of the present application can be the aforementioned electronic device.
  • the electronic device in the embodiment of the present application has video encoding and decoding capabilities, and generally includes a video encoder (i.e., encoder) and a video decoder (i.e., decoder).
  • FIG26 is a block diagram of the system composition of an encoder provided in an embodiment of the present application.
  • the encoder 100 may include: a segmentation unit 101, a prediction unit 102, a first adder 107, a transform unit 108, a quantization unit 109, an inverse quantization unit 110, an inverse transform unit 111, a second adder 112, a filtering unit 113, a decoded picture buffer (DPB) unit 114, and an entropy coding unit 115.
  • DPB decoded picture buffer
  • segmentation unit 101 transmits the CTU to prediction unit 102.
  • prediction unit 102 may be composed of block segmentation unit 103, motion estimation (ME) unit 104, motion compensation (MC) unit 105, and intra prediction unit 106.
  • block segmentation unit 103 iteratively uses quadtree segmentation, binary tree segmentation, and ternary tree segmentation to further divide the input CTU into smaller coding units (CUs).
  • Prediction unit 102 may use ME unit 104 and MC unit 105 to obtain inter-frame prediction blocks for the CU.
  • Intra-frame prediction unit 106 may use various intra-frame prediction modes, including MIP mode, to obtain intra-frame prediction blocks for the CU.
  • a rate-distortion optimized motion estimation method may be used by ME unit 104 and MC unit 105 to obtain inter-frame prediction blocks
  • a rate-distortion optimized mode determination method may be used by intra-frame prediction unit 106 to obtain intra-frame prediction blocks.
  • the prediction unit 102 outputs the prediction block of the CU
  • the first adder 107 calculates the difference between the CU in the output of the segmentation unit 101 and the prediction block of the CU, i.e., the residual CU.
  • the transform unit 108 reads the residual CU and performs one or more transform operations on the residual CU to obtain coefficients.
  • the quantization unit 109 quantizes the coefficients and outputs the quantized coefficients (i.e., levels).
  • the filtering unit 113 includes one or more filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a luminance mapping and chroma scaling (LMCS) filter, and a neural network-based filter.
  • filters such as a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a luminance mapping and chroma scaling (LMCS) filter, and a neural network-based filter.
  • the filtering unit 113 determines that the CU is not used as a reference for encoding other CUs, the filtering unit 113 performs loop filtering on one or more target samples in the CU.
  • the output of the filtering unit 113 is a decoded picture or sub-picture, which is cached to the DPB unit 114.
  • the DPB unit 114 outputs the decoded picture or sub-picture according to the timing and control information.
  • the picture stored in the DPB unit 114 can also be used as a reference for the prediction unit 102 to perform inter-frame prediction or intra-frame prediction.
  • the entropy coding unit 115 converts the parameters necessary for decoding the picture from the encoder 100 (such as control parameters and supplementary information) into a decoded picture. etc.) into a binary form, and writes such a binary form into a code stream according to the syntax structure of each data unit, that is, the encoder 100 finally outputs the code stream.
  • the encoder 100 can be a device having a first processor and a first memory for recording a computer program. When the first processor reads and executes the computer program, the encoder 100 reads the input video and generates a corresponding bitstream.
  • the encoder 100 can be a computing device having one or more chips. These units implemented as integrated circuits on the chip have similar connection and data exchange functions as the corresponding units in FIG26 .
  • Figure 27 is a block diagram of the system components of a decoder provided in an embodiment of the present application.
  • the decoder 200 may include: a parsing unit 201, a prediction unit 202, an inverse quantization unit 205, an inverse transform unit 206, an adder 207, a filtering unit 208, and a decoded image buffer unit 209.
  • the input of the decoder 200 is a bitstream representing a compressed version of a video or a still image
  • the output of the decoder 200 may be a decoded video consisting of a series of images or a decoded still image.
  • the input codestream to decoder 200 may be the codestream generated by encoder 100.
  • Parsing unit 201 parses the input codestream and obtains syntax element values from the input codestream. Parsing unit 201 converts the binary representation of the syntax elements into digital values and sends the digital values to units within decoder 200 to obtain one or more decoded pictures. Parsing unit 201 may also parse one or more syntax elements from the input codestream to display decoded pictures.
  • prediction unit 202 passes relevant parameters from parsing unit 201 to intra prediction unit 204 to obtain an intra prediction block.
  • Dequantization unit 205 has the same functionality as dequantization unit 110 in encoder 100. Dequantization unit 205 performs a scaling operation on the quantization coefficients (i.e., levels) from parsing unit 201 to obtain reconstructed coefficients.
  • the inverse transform unit 206 has the same function as the inverse transform unit 111 in the encoder 100.
  • the inverse transform unit 206 performs one or more transform operations (i.e., the inverse of the one or more transform operations performed by the inverse transform unit 111 in the encoder 100) to obtain a reconstructed residual.
  • the adder 207 performs an addition operation on its input (the prediction block from the prediction unit 202 and the reconstructed residual from the inverse transform unit 206) to obtain a reconstructed block of the current decoded block.
  • the reconstructed block is also sent to the prediction unit 202 to be used as a reference for other blocks encoded in the intra prediction mode.
  • the decoder 200 can be a second memory having a second processor and a computer program. When the first processor reads and runs the computer program, the decoder 200 reads the input code stream and generates a corresponding decoded video.
  • the decoder 200 can also be a computing device having one or more chips. These units implemented as integrated circuits on the chip have similar connection and data exchange functions as the corresponding units in Figure 27.
  • the "current block” specifically refers to the current block to be encoded in the video image (which can also be simply referred to as the "encoding block”); when the embodiment of the present application is applied to the decoder 200, the “current block” specifically refers to the current block to be decoded in the video image (which can also be simply referred to as the "decoding block”).
  • FIG28 is a flowchart diagram of a decoding method provided by the embodiment of the present application. As shown in FIG28 , the method may include:
  • S2801 Determine a transform core group index of a current block, and determine a transform core group of the current block according to the transform core group index of the current block.
  • the method is applied to a decoder.
  • the decoding method of the embodiments of the present application is primarily applied to intra-predicted blocks.
  • the optimization scheme proposed here is mainly for the NSPT and LFNST transforms in intra-prediction mode to improve compression efficiency.
  • NSPT and LFNST are both transformations that process textures of various angles. They may have multiple transformation kernels, and a transformation kernel may be specifically optimized for a certain specific angle texture.
  • NSPT and LFNST also include transformation kernels for processing gradient textures.
  • these transformation kernels can also be said to be trained KL transforms (Karhunen-Loeve Transform, KLT).
  • KLT Kerhunen-Loeve Transform
  • gradient textures can be further expanded to include horizontal gradient textures, vertical gradient textures, and diagonal gradient textures.
  • the MTSS approach is not limited to non-separable transforms like NSPT and LFNST; it can also be applied to separable transforms optimized for specific textures.
  • the intra-frame prediction block can have multiple transform kernel groups, and each intra-frame prediction mode can correspond to a transform kernel group.
  • each intra-frame prediction mode actually represents a texture feature.
  • the intra-frame prediction mode index is also a texture feature index.
  • DC and PLANAR correspond to gradient texture features
  • a certain angle prediction mode corresponds to the texture feature of this angle.
  • the texture feature index can avoid the occurrence of intra-frame prediction mode in "inter-frame", and on the other hand, it is also more conducive to possible expansion.
  • an intra-frame prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradient texture, vertical gradient texture, oblique gradient texture, etc.
  • some intra-frame prediction modes are not based on simple texture features, but may contain two or more texture features. Therefore, the Multiple Transform Set Selection (MTSS) technique can be used.
  • MTSS Multiple Transform Set Selection
  • the intra-frame prediction mode of the current block is a special intra-frame prediction mode, more than one transform set can be selected.
  • some additional conditions may be added to determine whether the current block uses the MTSS technology.
  • these conditions may include: the number of samples in the current block, the size of the current block, and the prediction mode of the current block.
  • the number of samples in the current block may be selected to determine whether the current block uses the MTSS technology, and/or the size of the current block may be selected to determine whether the current block uses the MTSS technology, and/or the prediction mode of the current block may be selected to determine whether the current block uses the MTSS technology.
  • the prediction mode of the current block may be selected to determine whether the current block uses the MTSS technology.
  • the transform core group of the current block may be determined according to the transform core group index of the current block.
  • a block may choose one from two or more LFNST/NSPT transform core groups.
  • a block in VVC has only one available transform core group. Therefore, one possible implementation method is to limit the size of the block to which the MTSS technology is applicable, so that the MTSS technology can be disabled at the block size where the MTSS technology requires a lot of additional calculations but the compression efficiency is not significantly improved.
  • the embodiment of the present application can limit the block size to which the MTSS technology is applicable.
  • the method may further include: determining the number of samples of the current block; and when the number of samples of the current block is greater than or equal to a minimum sample threshold, performing the step of determining a transform core group index of the current block.
  • the method may further include: determining the size of the current block, wherein the size of the current block includes height and width; when the height and width of the current block are both greater than or equal to a minimum size threshold, executing the step of determining the transform core group index of the current block.
  • a minimum size threshold for applying MTSS technology can be set, represented by MIN_SIZE. If the width or height of a block is less than MIN_SIZE, MTSS technology cannot be used for this block. Otherwise, if the width and height of this block are both greater than or equal to MIN_SIZE, MTSS technology can be used for this block.
  • the value of MIN_SIZE may be 8, 16, etc.
  • the value of MIN_PIX may be equal to 8.
  • the method may further include: determining a prediction mode of the current block; and when the prediction mode of the current block is one of the first prediction mode set, performing the step of determining a transform core group index of the current block.
  • the first prediction mode set may include at least one of the following prediction modes: DC mode, PLANAR mode, and other prediction modes other than angular prediction mode.
  • the mode for combining the prediction values of at least two intra-frame prediction modes may specifically be a DIMD mode, a TIMD mode, and an SGPM mode; wherein the intra-frame prediction mode includes but is not limited to a DC mode, a PLANAR mode, and an angle prediction mode, etc.
  • the transform core group index is used to indicate the number of the transform core group of the current block in at least two candidate transform core groups.
  • the transform core group index can be represented by lfnst_feature_idx, or it can also be represented by lfnst_set_idx.
  • the transform core group index of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, etc.
  • the value of the first syntax element is equal to 0, it indicates that the first candidate transform core group is selected as the transform core group of the current block; if the value of the first syntax element is equal to 1, it indicates that the second candidate transform core group is selected as the transform core group of the current block; if the value of the first syntax element is equal to 2, it indicates that the third candidate transform core group is selected as the transform core group of the current block.
  • determining the transform core group of the current block according to the transform core group index of the current block may include: if the transform core group index is the i-th value, determining the transform core group of the current block based on the i-th intra-frame prediction mode derived from the prediction mode of the current block, where i is a positive integer.
  • the second prediction mode set may include at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, wherein the intra-frame prediction mode includes at least one of the following: a DC mode, a PLANAR mode, and an angular prediction mode.
  • the second prediction mode set may include at least one of the following prediction modes: DIMD mode, TIMD mode, and SGPM mode.
  • MTSS can be applied to DIMD mode, TIMD mode, and SGPM mode. If the transform core group index is a first value, the transform core group of the current block is determined based on the first intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is a second value, the transform core group of the current block is determined based on the second intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is a third value, the transform core group of the current block is determined based on the third intra-frame prediction mode derived from the prediction mode of the current block, and so on.
  • the first value may be 0, the second value may be 1, and the third value may be 2. That is, for the prediction mode in the second prediction mode set, after determining the transform core group index of the current block, the corresponding transform core group may be directly determined according to the transform core group index.
  • the DIMD mode since the DIMD mode itself uses gradient-derived texture features and it can derive multiple intra-frame prediction modes, there is no need for the MTSS technology to additionally use gradient-derived texture features.
  • the corresponding transform core group can be determined according to the transform core group index of the current block.
  • the TIMD mode since the TIMD mode itself can derive multiple intra-frame prediction modes, there is no need for the MTSS technology to additionally use gradients to derive texture features, and the corresponding transform core group can be determined according to the transform core group index of the current block.
  • the corresponding transform core group can be determined according to the transform core group index of the current block.
  • the transform core group index corresponding to the first intra prediction mode derived by the DIMD mode is 0, which is recorded as candFeature0; the transform core group index corresponding to the second intra prediction mode derived by the DIMD mode is 1, which is recorded as candFeature1.
  • the transform core group index is 0, the transform core group of the current block is determined based on the first intra prediction mode derived by the DIMD mode; if If the transform core group index is 1, the transform core group of the current block is determined based on the second intra prediction mode derived from the DIMD mode.
  • the transform core group index corresponding to the first intra-frame prediction mode derived by the TIMD mode is 0, denoted as candFeature0; and the transform core group index corresponding to the second intra-frame prediction mode derived by the TIMD mode is 1, denoted as candFeature1.
  • the transform core group index is 0, the transform core group for the current block is determined based on the first intra-frame prediction mode derived by the TIMD mode; if the transform core group index is 1, the transform core group for the current block is determined based on the second intra-frame prediction mode derived by the TIMD mode.
  • the transform core group index corresponding to the partition mode of the SGPM mode is 0, denoted as candFeature0; and the transform core group index corresponding to the first intra-frame prediction mode derived from the SGPM mode is 1, denoted as candFeature1.
  • the transform core group index is 0, the transform core group for the current block is determined based on the partition mode of the SGPM mode; if the transform core group index is 1, the transform core group for the current block is determined based on the first intra-frame prediction mode derived from the SGPM mode.
  • the first value may be 0 and the second value may be 1. That is, for the prediction mode in the third prediction mode set, after determining the transform core group index of the current block, the corresponding transform core group may be directly determined according to the transform core group index.
  • determining the first candidate list for the current block may specifically include: when the prediction mode of the current block is one of the items in the third prediction mode set, determining an intra-frame prediction mode derived based on the prediction mode of the current block; determining two candidate transform core groups according to the intra-frame prediction mode and the PLANAR mode, and adding the two candidate transform core groups to the first candidate list.
  • determining a first candidate list for the current block may further include: when the first candidate list is not full, determining candidate samples for deriving texture feature indexes; determining one or more candidate texture feature indexes of the current block based on the candidate samples; determining one or more candidate transform core groups based on the one or more candidate texture feature indexes, and adding the one or more candidate transform core groups to the first candidate list.
  • a prediction block of the current block may be determined; and at least part of the samples in the prediction block may be used as candidate samples.
  • the candidate samples used for inter-frame derivation of candidate texture feature indexes can be all samples in the prediction block or part of the samples in the prediction block.
  • the number of candidate samples used to derive the candidate texture feature index may be at least one, for example, 1, 2, 3 or more.
  • the method may further include: determining the number of candidate samples according to a size parameter of the current block.
  • the number of candidate samples used can be determined by the size parameter of the current block. For example, if the size of the current block is small, then all available samples can be counted; if the size of the current block is large, then the current block can be downsampled and counted, such as counting one sample out of every 2, or 4, or 8 samples in the horizontal and/or vertical directions.
  • the size of one of the horizontal or vertical directions of the current block is less than or equal to 8, then all available samples in that direction are counted; otherwise, if the size of one of the horizontal or vertical directions of the current block is less than or equal to 16, then one sample out of every 2 samples in that direction is counted; otherwise, one sample out of every 4 samples in that direction is counted, and there is no specific limitation here.
  • determining one or more candidate texture feature indexes of the current block based on the candidate samples may include: determining the horizontal gradient value and the vertical gradient value of the candidate sample; determining the texture feature index and the gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and the vertical gradient value of the candidate sample; determining a texture feature statistics table based on the texture feature index and the gradient intensity value corresponding to the candidate sample; and determining one or more candidate texture feature indexes of the current block based on the texture feature statistics table.
  • the texture feature index and gradient intensity value corresponding to the candidate sample when determining the texture feature index and gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and vertical gradient value of the candidate sample, it can include: performing angle mapping based on the horizontal gradient value and vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample; and performing gradient intensity calculation based on the horizontal gradient value and vertical gradient value of the candidate sample to determine the gradient intensity value corresponding to the candidate sample.
  • performing angle mapping based on the horizontal gradient value and the vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample may include: determining the texture feature index corresponding to the candidate sample using a preset lookup table based on the horizontal gradient value and the vertical gradient value of the candidate sample.
  • abs(grad x ) is equal to 2 times abs(grad y ) and grad x and grad y have the same sign, it corresponds to intra-frame prediction mode 40 in VVC.
  • other situations can be determined by looking up the table according to the same principle.
  • performing gradient strength calculation based on the horizontal gradient value and the vertical gradient value of the candidate sample to determine the gradient strength value corresponding to the candidate sample may include: performing an addition operation on the absolute value of the horizontal gradient value and the absolute value of the vertical gradient value to determine the gradient strength value corresponding to the candidate sample.
  • the gradient strength value corresponding to the candidate sample can be recorded as amp.
  • amp abs(grad x )+abs(grad y ).
  • DIMD itself will derive one or several intra-frame prediction modes for weighting
  • TIMD itself will also derive one or several intra-frame prediction modes for weighting
  • SGPM not only has two intra-frame prediction modes, but also has a "partitioning" mode that can also find the corresponding intra-frame prediction mode, and the residual often appears in the boundary area of the "partition”.
  • the prediction mode of the current block is the TIMD mode
  • the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.
  • the first candidate transformation core group is determined based on the reference texture feature index in the first position in the candidate texture feature list. Specifically, it can be: the first candidate transformation core group is determined based on the reference texture feature index with the largest gradient intensity accumulated value in the candidate texture feature list.
  • the method may include: determining a first candidate transform core group based on a first intra-frame prediction mode derived from a prediction mode of the current block; determining a second candidate transform core group based on the first reference texture feature index when a second condition is satisfied between the first reference texture feature index in the candidate texture feature list and the first candidate transform core group; and determining the first candidate list based on the first candidate transform core group and the second candidate transform core group.
  • the first candidate list not only considers the candidate texture feature index derived from the horizontal gradient value and the vertical gradient value of the candidate sample, but also considers the intra-frame prediction mode derived from the prediction mode of the current block itself. The following describes this in detail with several examples.
  • a prediction block or a prediction block plus a template can be used to derive a candidate texture feature index (or intra-frame prediction mode).
  • the candidate samples used by the MTSS technology to derive N reference texture feature indexes are the same as the candidate samples used by the DIMD mode, then the first intra-frame prediction mode derived by DIMD and the first reference texture feature index derived by the MTSS technology are the same.
  • determining the transform core of the current block can be further determined.
  • determining the transform core of the current block based on the transform core group may include: determining a transform core index of the current block; and determining the transform core of the current block based on the transform core group and the transform core index.
  • the method may further include: determining a transform core index of the current block; and determining the transform core of the current block according to the first candidate list and the transform core index.
  • the first candidate list indicates the transform cores included in two candidate transform core groups
  • the first candidate list may indicate 5 candidate transform cores
  • the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 3 candidate transform cores
  • the first candidate list may indicate 6 candidate transform cores.
  • the transform core of the current block can be determined in the first candidate list based on the transform core index.
  • the transform core index of the current block may be a positive integer, such as 1, 2, 3, 4, 5, 6, etc.
  • the transform core index of the current block may be determined by directly decoding the bitstream, or may be determined by decoding the value of the second syntax element.
  • one possible implementation is to decode the bitstream and determine the transform kernel index of the current block.
  • another possible implementation is to decode the bitstream and determine the value of the second syntax element; and when the second syntax element indicates that the current block uses the first transform mode, determine the transform kernel index of the current block based on the value of the second syntax element.
  • the value of the second syntax element is the third value, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is the fourth value, it is determined that the current block uses the first transform mode and the corresponding transform kernel index.
  • the third value can be set to 0, and the fourth value can be set to a non-zero value, such as 1, 2, 3, 4, 5, 6, etc.
  • the transform core index of the current block can also be represented by lfnst_idx.
  • lfnst_idx being 0 indicates that the current block does not use LFNST/NSPT.
  • Each transform core group of LFNST/NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST/NSPT.
  • the prediction mode of the current block is a special intra-frame prediction mode
  • it has more than one selectable transform core group.
  • the possible values of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6.
  • 1, 2, and 3 correspond to the three transform cores of the first transform core group
  • 4, 5, and 6 correspond to the three transform cores of the second transform core group.
  • the third binary symbol that is, the binary symbol with BinIdx being 2, can also be understood as selecting the first candidate transformation core group or the second candidate transformation core group.
  • the transform core group here can be one of the at least two candidate transform core groups indicated by the first candidate list.
  • the transform core group index of the current block can be represented by lfnst_feature_idx, or by lfnst_set_idx. Among them, the transform core group index of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, etc.
  • lfnst_feature_idx if the value of lfnst_feature_idx is equal to 0, it indicates that the first candidate transform core group is selected as the transform core group of the current block; if the value of lfnst_feature_idx is equal to 1, it indicates that the second candidate transform core group is selected as the transform core group of the current block.
  • the transform core of the current block can be determined based on the transform core group. Specifically, this can be done by first determining the transform core index of the current block, and then determining the transform core of the current block based on the transform core group and the transform core index.
  • one possible implementation is to decode the bitstream and determine the transform kernel index of the current block.
  • another possible implementation is to decode the bitstream and determine the value of the second syntax element; when the second syntax element indicates that the current block uses the first transform mode, determine the transform kernel index of the current block based on the value of the first syntax element.
  • Each transform core group of LFNST in VVC has two transform cores, so a lfnst_idx value of 1 or 2 indicates that the current block uses the first or second transform core of the selected transform core group of LFNST.
  • lfnst_idx may have four values, namely 0, 1, 2, and 3. Where lfnst_idx is 0, which means that the current block does not use LFNST/NSPT.
  • Each transform core group of LFNST/NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST/NSPT.
  • the binary symbol correspondence table of lfnst_idx is shown in Table 8.
  • determining the transformation coefficient of the current block may include: decoding the code stream to determine the quantization coefficient of the current block; and dequantizing the quantization coefficient of the current block to determine the transformation coefficient of the current block.
  • when transforming the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block it can include: performing an inseparable basic transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block; or performing a low-frequency inseparable transform on the transform coefficients of the current block according to the transform kernel to determine the transform block of the current block, and performing a discrete cosine transform on the transform block of the current block to determine the residual block of the current block.
  • the size parameter of the current block satisfies the first condition, including: the size parameter of the current block is relatively small, for example, the size parameter of the current block is less than a certain threshold.
  • an NSPT transform kernel is used, i.e., an inverse NSPT transform is performed on the transform coefficients of the current block according to the transform kernel to determine a residual block for the current block.
  • the size parameter of the current block satisfies the second condition, including: the size parameter of the current block is relatively large, for example, the size parameter of the current block is greater than a certain threshold.
  • an LFNST transform kernel is used. Specifically, an inverse LFNST transform is performed on the transform coefficients of the current block based on the transform kernel to determine the transform block of the current block; and an inverse DCT2 transform is performed on the transform block of the current block to determine the residual block of the current block.
  • the "inverse transformation” of the transform coefficients at the decoding end may also be referred to as “transformation” in the standard text.
  • the “transformation” and “inverse transformation” in this article correspond to two opposite processes. For example, if the “transformation” converts the numerical values in the spatial domain to the coefficients in the frequency domain, then the “inverse transformation” converts the coefficients in the frequency domain to the numerical values in the spatial domain. "Inverse” is relative to “positive”, and they are essentially both transformations. It should be noted that if the standard only stipulates decoding, then the "transformation” in the standard text is the decoding part, specifically referring to the “inverse transformation” in this article. The “inverse transformation” of the transform coefficients at the decoding end may also be referred to as "transformation” in the standard text.
  • the method may further include:
  • S3102 Determine a reconstructed block of the current block according to the prediction block of the current block and the residual block of the current block.
  • an addition operation may be performed on the prediction block of the current block and the residual block of the current block to determine the reconstructed block of the current block.
  • the decoder derives candidate texture feature indices based on the prediction block. The decoder then uses these indices to determine the transform core group for NSPT/LFNST. If a transform core group has multiple selectable transform cores, the decoder determines the transform core by decoding the lfnst_nspt_idx field in the bitstream. The decoding of lfnst_nspt_idx in the bitstream is independent of the process of determining the transform core group. Quantized coefficients are obtained from the bitstream through entropy decoding.
  • the quantized coefficients are inversely quantized to obtain decoded transform coefficients
  • the decoded transform coefficients are inversely NSPT transformed to obtain decoded residual blocks
  • a reconstructed block is obtained based on the decoded residual block and the prediction block.
  • the quantized coefficients are inversely quantized to obtain decoded transform coefficients
  • the decoded transform coefficients are inversely LFNST transformed
  • inverse DCT2 transformed to obtain a decoded residual block
  • a reconstructed block is obtained based on the decoded residual block and the prediction block.
  • the method may further include:
  • S3202 Determine a transformation kernel of the current block according to the texture feature index.
  • determining the transform core of the current block according to the texture feature index may include: determining the transform core group of the current block according to the texture feature index; decoding the code stream to determine the transform core index of the current block; and determining the transform core of the current block according to the transform core group and the transform core index.
  • the method further includes: when the number of samples of the current block is less than a minimum sample threshold, executing the step of determining a texture feature index of the current block.
  • the size of the current block includes height and width; the method further comprises: when the height or width of the current block is less than a minimum size threshold, performing a step of determining a texture feature index of the current block.
  • the prediction mode of the current block when the prediction mode of the current block is a prediction mode outside the first prediction mode set, the prediction mode of the current block may be one of the prediction modes in the fourth prediction mode set.
  • the fourth prediction mode set includes at least: a DC mode, a PLANAR mode, and an angular prediction mode.
  • the modes in the fourth prediction mode set are relatively simple prediction modes. For example, a classification can be made, where the DC mode, PLANAR mode, and various angle prediction modes are classified as the fourth prediction mode set, and the DIMD mode, TIMD mode, MIP mode, EIP mode, SGPM mode, ITMP mode, and IBC mode are classified as the first prediction mode set.
  • the first prediction mode set may include only one or more of the DIMD mode, TIMD mode, MIP mode, EIP mode, SGPM mode, ITMP mode, and IBC mode.
  • the first prediction mode set includes the DIMD mode, TIMD mode, and SGPM mode.
  • the other modes belong to the fourth prediction mode set.
  • the encoding method of lfnst_idx of the fourth prediction mode set is the same as that of the related art, while the encoding method of lfnst_idx of the first prediction mode set is different from that of the related art.
  • each transform core group of LFNST/NSPT has 3 transform cores.
  • Blocks using the fourth prediction mode set can only select one transform core group.
  • Blocks using the first prediction mode set can select two transform core groups, each with 3 transform cores. If the prediction mode of the current block belongs to the fourth prediction mode set, it can only select one texture feature index, so the possible value of lfnst_idx is 0. 1, 2, 3. If the prediction mode of the current block belongs to the first prediction mode set, then the possible values of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6.
  • 1, 2, and 3 correspond to the three transform cores of the transform core group corresponding to the first texture feature index
  • 4, 5, and 6 correspond to the three transform cores of the transform core group corresponding to the second texture feature index.
  • the binary symbol correspondence table of lfnst_idx for blocks using the fourth prediction mode set is shown in Table 8
  • the binary symbol correspondence table of lfnst_idx for blocks using the first prediction mode set is shown in Table 7.
  • the third syntax element can be represented by sps_mtss_enabled_flag, and the third syntax element is a syntax element in a sequence parameter set (SPS).
  • SPS sequence parameter set
  • the value of the third syntax element is the fifth value, it is determined that the third syntax element indicates that the current sequence allows the use of MTSS technology; if the value of the third syntax element is the sixth value, it is determined that the third syntax element indicates that the current sequence does not allow the use of MTSS technology.
  • the fifth value is different from the sixth value, and the fifth value and the sixth value can be in parameter form or in digital form.
  • the third syntax element can be a parameter written in the profile or the value of a flag, which is not specifically limited here.
  • the fifth value can be 1 and the sixth value can be 0; or, the fifth value can be 0 and the sixth value can be 1; or, the fifth value can be true and the sixth value can be false; or, the fifth value can be false and the sixth value can be true.
  • the fifth value is 1 and the sixth value is 0.
  • the method further includes: decoding the code stream to determine the value of a third syntax element; when the third syntax element indicates that the current sequence allows the use of a multi-transform core group selection mode, decoding the code stream to determine the value of a fourth syntax element; when the fourth syntax element indicates that the current image allows the use of a multi-transform core group selection technology, executing the step of determining a transform core group index of the current block.
  • the current sequence may include a current picture, and the current picture includes a current block.
  • the fourth syntax element may be represented by ph_inter_lfnst_nspt_enabled_flag, and the fourth syntax element is a picture-level syntax element.
  • the value of the fourth syntax element is the fifth value, it is determined that the fourth syntax element indicates that the current image allows the use of MTSS technology; if the value of the fourth syntax element is the sixth value, it is determined that the fourth syntax element indicates that the current image does not allow the use of MTSS technology.
  • the method further includes: decoding the bitstream and determining the value of a third syntax element; when the third syntax element indicates that the current sequence allows the use of a multi-transform core group selection mode, decoding the bitstream and determining the value of a fifth syntax element; when the fifth syntax element indicates that the current slice allows the use of a multi-transform core group selection technology, performing the step of determining a transform core group index of the current block; wherein the current sequence includes the current slice, and the current slice includes the current block.
  • the current sequence may include a current slice, and the current slice includes a current block.
  • the fifth syntax element may be represented by sh_inter_lfnst_nspt_enabled_flag, and the fifth syntax element is a slice-level syntax element.
  • the value of the fifth syntax element is the fifth value, it is determined that the fifth syntax element indicates that the current slice allows the use of MTSS technology; if the value of the fifth syntax element is the sixth value, it is determined that the fifth syntax element indicates that the current slice does not allow the use of MTSS technology.
  • the fifth value is different from the sixth value, and the fifth value and the sixth value can be in parameter form or in digital form.
  • the fourth syntax element or the fifth syntax element can be a parameter written in the profile, or it can be the value of a flag, which is not specifically limited here.
  • the fifth value can be 1 and the sixth value can be 0; or, the fifth value can be 0 and the sixth value can be 1; or, the fifth value can be true and the sixth value can be false; or, the fifth value can be false and the sixth value can be true.
  • the fifth value is 1 and the sixth value is 0.
  • sps_mtss_enabled_flag in addition to the sequence-level syntax element sps_mtss_enabled_flag, other levels of syntax can also be used to achieve more flexible control, such as the flag of the Picture Parameter Set (PPS), or the flag of the picture header or slice header.
  • PPS Picture Parameter Set
  • the minimum sample threshold in a possible implementation, it may include: decoding the code stream, determining the minimum sample threshold, Alternatively, in another possible implementation, the method may include: decoding a bitstream and determining a value of a third syntax element; when the third syntax element indicates that the current sequence allows the use of a multi-transform kernel group selection mode, decoding the bitstream and determining a value of a sixth syntax element; and determining a minimum sample threshold based on the value of the sixth syntax element.
  • the value of the seventh syntax element can be equal to the minimum size threshold, or decoding can be performed according to a certain mapping rule. For example, if the minimum size threshold is equal to 4, the value of the seventh syntax element is determined to be equal to 0; if the minimum sample threshold is equal to 16, the value of the seventh syntax element is determined to be equal to 1. In this way, when the value of the seventh syntax element obtained during decoding is equal to 1, it can be determined that the minimum size threshold is equal to 16.
  • a high-level syntax can be used to set a minimum size threshold for applying MTSS technology, such as sps_mtss_min_size.
  • a minimum size threshold for applying MTSS technology such as sps_mtss_min_size.
  • the decoder parses sps_mtss_min_size to determine the minimum size threshold for applying MTSS technology.
  • a trade-off between coding complexity and compression efficiency can be made based on demand. That is, under this method, the decoder needs to support all possible cases of sps_mtss_min_size, but the encoder can configure the sps_mtss_min_size required to encode the current bitstream.
  • a relatively small value such as 4 can be set for sps_mtss_min_size.
  • a relatively large value such as 16, can be set for sps_mtss_min_size.
  • An embodiment of the present application provides a decoding method, specifically an intra-frame LFNST/NSPT multi-angle selection scheme.
  • the transform core group index of the current block is determined, and the transform core group of the current block is determined based on the transform core group index of the current block; then, based on the transform core group, the transform kernel of the current block is determined; then, the transform coefficient of the current block is determined, and the transform coefficient of the current block is transformed based on the transform kernel to determine the residual block of the current block.
  • the transform core group of the current block is determined according to the multi-transform core group selection technology, and then the transform kernel of the current block is determined therefrom.
  • NSPT and LFNST are both transformations for processing textures of various angles. They may have multiple transformation kernels, and one transformation kernel may be specially optimized for a certain specific angle texture.
  • NSPT and LFNST also include transformation kernels for processing gradient textures.
  • these transformation kernels can also be said to be trained KL transforms (Karhunen-Loeve Transform, KLT). That is to say, NSPT and LFNST both have multiple transformation kernels, each of which is designed for a specific texture, and the specific texture includes angular textures.
  • the intra-frame prediction block can have multiple transform kernel groups, and each intra-frame prediction mode can correspond to a transform kernel group.
  • each intra-frame prediction mode actually represents a texture feature.
  • the intra-frame prediction mode index is also a texture feature index.
  • DC and PLANAR correspond to gradient texture features
  • a certain angle prediction mode corresponds to the texture feature of this angle.
  • the texture feature index can avoid the occurrence of intra-frame prediction mode in "inter-frame", and on the other hand, it is also more conducive to possible expansion.
  • an intra-frame prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradient texture, vertical gradient texture, oblique gradient texture, etc.
  • some intra-frame prediction modes are not based on simple texture features, but may contain two or more texture features. Therefore, the Multiple Transform Set Selection (MTSS) technique can be used.
  • MTSS Multiple Transform Set Selection
  • the intra-frame prediction mode of the current block is a special intra-frame prediction mode, more than one transform set can be selected.
  • the MTSS technology increases the complexity of the encoder due to the increase in candidates for the transform core group.
  • a block may select one from two or more LFNST/NSPT transform core groups.
  • a block in VVC has only one available transform core group. Therefore, one possible implementation method is to limit the size of the block to which the MTSS technology is applicable, so that the MTSS technology can be disabled at the block size where the MTSS technology requires a lot of additional calculations but does not significantly improve the compression efficiency.
  • the embodiment of the present application can limit the block size to which the MTSS technology is applicable.
  • the method may further include: determining the number of samples of the current block; and when the number of samples of the current block is greater than or equal to a minimum sample threshold, executing the step of determining a transform core group of the current block.
  • a threshold for the minimum number of samples for applying the MTSS technology (i.e., the "minimum sample threshold”) can be set, represented by MIN_PIX. If the number of samples (width ⁇ height) of a block is less than MIN_PIX, the MTSS technology cannot be used for this block. Otherwise, if the number of samples of this block is greater than or equal to MIN_PIX, the MTSS technology can be used for this block.
  • the value of MIN_PIX may be 32, 64, 256, etc. For example, the value of MIN_PIX may be equal to 256.
  • the method may further include: determining the size of the current block, wherein the size of the current block includes height and width; when the height and width of the current block are both greater than or equal to a minimum size threshold, executing the step of determining a transform core group for the current block.
  • the method may further include: determining a prediction mode of the current block; and when the prediction mode of the current block is one of the first prediction mode set, performing the step of determining a transform core group of the current block.
  • the first prediction mode set may include at least one of the following prediction modes: DC mode, PLANAR mode, and other prediction modes other than angular prediction mode.
  • the first prediction mode set may include at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, the intra-frame prediction mode including at least one of the following: DC mode, PLANAR mode and angle prediction mode; a mode for prediction by copying intra-frame blocks; a mode for prediction using an interpolation filter; a mode for prediction using matrix operations.
  • the mode for combining the prediction values of at least two intra-frame prediction modes may specifically be a DIMD mode, a TIMD mode, and an SGPM mode; wherein the intra-frame prediction mode includes but is not limited to a DC mode, a PLANAR mode, and an angle prediction mode, etc.
  • the mode for predicting the copied intra-frame block may specifically be the ITMP mode and the IBC mode.
  • the mode for using the extrapolation filter for prediction may specifically be the EIP mode.
  • the mode for prediction using matrix operations may specifically be a MIP mode.
  • the modes in the first prediction mode set are relatively complex prediction modes, and the corresponding textures may contain two or more texture features.
  • both the DIMD mode and the TIMD mode can weight the prediction values of two or more intra-frame prediction modes
  • the SGPM mode weights the prediction values of the two intra-frame prediction modes using a weight matrix
  • the MIP mode is based on a matrix operation.
  • the EIP mode uses an extrapolation filter, while the ITMP and IBC modes rely on replicating a reconstructed reference block.
  • Their texture features are not as simple as those in the DC and PLANAR modes. That is, when the prediction mode of the current block is any of the first prediction mode set, MTSS can be used for the current block.
  • determining the transform core group for the current block may include: calculating the encoding cost for the current block based on at least two candidate transform core groups, determining cost results corresponding to each of the at least two candidate transform core groups; determining the minimum cost result among the cost results corresponding to the at least two candidate transform core groups, and determining the candidate transform core group corresponding to the minimum cost result as the transform core group for the current block.
  • the method further includes: determining a transform core group index of the current block; encoding the transform core group index of the current block, and writing the obtained encoding bits into a bitstream.
  • lfnst_feature_idx if the value of lfnst_feature_idx is equal to 0, it indicates that the first candidate transform core group is selected as the transform core group of the current block; if the value of lfnst_feature_idx is equal to 1, it indicates that the second candidate transform core group is selected as the transform core group of the current block.
  • the first syntax element can be represented by lfnst_feature_idx or lfnst_set_idx.
  • the first syntax element can be used to indicate the transform core group index of the current block, specifically the number of the transform core group of the current block in at least two candidate transform core groups.
  • the value of the first syntax element can be an integer greater than or equal to zero, such as 0, 1, 2, etc.
  • the value of the first syntax element is equal to 0, it indicates that the first candidate transform core group is selected as the transform core group of the current block; if the value of the first syntax element is equal to 1, it indicates that the second candidate transform core group is selected as the transform core group of the current block; if the value of the first syntax element is equal to 2, it indicates that the third candidate transform core group is selected as the transform core group of the current block.
  • determining the transform core group of the current block according to the transform core group index of the current block may include: if the transform core group index is the i-th value, determining the transform core group of the current block based on the i-th intra-frame prediction mode derived from the prediction mode of the current block, where i is a positive integer.
  • the second prediction mode set may include at least one of the following prediction modes: DIMD mode, TIMD mode, and SGPM mode.
  • MTSS can be applied to DIMD mode, TIMD mode, and SGPM mode. If the transform core group index is a first value, the transform core group of the current block is determined based on the first intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is a second value, the transform core group of the current block is determined based on the second intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is a third value, the transform core group of the current block is determined based on the third intra-frame prediction mode derived from the prediction mode of the current block, and so on.
  • the first value may be 0, the second value may be 1, and the third value may be 2. That is, for the prediction mode in the second prediction mode set, after determining the transform core group index of the current block, the corresponding transform core group may be directly determined according to the transform core group index.
  • the transform core group index corresponding to the first intra-frame prediction mode derived by the DIMD mode is 0, denoted as candFeature0; and the transform core group index corresponding to the second intra-frame prediction mode derived by the DIMD mode is 1, denoted as candFeature1.
  • the transform core group index is 0, the transform core group for the current block is determined based on the first intra-frame prediction mode derived by the DIMD mode; if the transform core group index is 1, the transform core group for the current block is determined based on the second intra-frame prediction mode derived by the DIMD mode.
  • the transform core group index corresponding to the first intra-frame prediction mode derived by the TIMD mode is 0, denoted as candFeature0; and the transform core group index corresponding to the second intra-frame prediction mode derived by the TIMD mode is 1, denoted as candFeature1.
  • the transform core group index is 0, the transform core group for the current block is determined based on the first intra-frame prediction mode derived by the TIMD mode; if the transform core group index is 1, the transform core group for the current block is determined based on the second intra-frame prediction mode derived by the TIMD mode.
  • the transform core group index corresponding to the partition mode of the SGPM mode is 0, denoted as candFeature0; and the transform core group index corresponding to the first intra-frame prediction mode derived from the SGPM mode is 1, denoted as candFeature1.
  • the transform core group index is 0, the transform core group for the current block is determined based on the partition mode of the SGPM mode; if the transform core group index is 1, the transform core group for the current block is determined based on the first intra-frame prediction mode derived from the SGPM mode.
  • determining the transform core group of the current block according to the transform core group index of the current block may include: if the transform core group index is a first value, determining the transform core group of the current block based on the first intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is a second value, determining the transform core group of the current block based on the PLANAR mode.
  • determining the transform core group of the current block according to the transform core group index of the current block may include: if the transform core group index is a first value, determining the transform core group of the current block based on the first intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is a second value, determining the transform core group of the current block based on the second intra-frame prediction mode derived from the prediction mode of the current block.
  • the third prediction mode set may include at least one of the following prediction modes: a mode for performing prediction using an extrapolation filter and a mode for performing prediction using a matrix operation.
  • the third prediction mode set may include at least one of the following prediction modes: MIP mode and EIP mode.
  • MTSS can be applied to both the MIP and EIP modes. Since the MIP and EIP modes are widely applicable to blocks with texture gradients, their features are similar to those of the PLANAR mode. Therefore, the PLANAR mode can also be used as a candidate texture feature for both the MIP and EIP modes.
  • the first value may be 0 and the second value may be 1. That is, for the prediction mode in the third prediction mode set, after determining the transform core group index of the current block, the corresponding transform core group may be directly determined according to the transform core group index.
  • the transform core group index corresponding to the first intra-frame prediction mode derived from the MIP mode or EIP mode is 0, denoted as candFeature0; and the transform core group index corresponding to the PLANAR mode is 1, denoted as candFeature1.
  • the transform core group index is 0, the transform core group for the current block is determined based on the first intra-frame prediction mode derived from the prediction mode of the current block; if the transform core group index is 1, the transform core group for the current block is determined based on the PLANAR mode.
  • step S3301 may further include: determining a first candidate list for the current block, the first candidate list indicating at least two candidate transform core groups; and determining a transform core group for the current block based on the first candidate list and the transform core group index.
  • the first candidate list may include at least two candidate texture feature indexes, or the first candidate list may include at least two transform core groups.
  • each candidate texture feature index corresponds to a transform core group. Therefore, it can be said that the first candidate list indicates at least two candidate transform core groups.
  • determining a first candidate list for the current block may include: determining one or more intra-frame prediction modes derived based on a prediction mode of the current block; determining one or more candidate transform core groups based on the one or more intra-frame prediction modes, and adding the one or more candidate transform core groups to the first candidate list.
  • determining a first candidate list for the current block may specifically include: when the prediction mode of the current block is one of the items in the second prediction mode set, determining multiple intra-frame prediction modes derived based on the prediction mode of the current block; determining multiple candidate transform core groups based on the multiple intra-frame prediction modes, and adding the multiple candidate transform core groups to the first candidate list.
  • the second prediction mode set may include at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, wherein the intra-frame prediction mode includes at least one of the following: a DC mode, a PLANAR mode, and an angular prediction mode; illustratively, the second prediction mode set may include at least one of the following prediction modes: a DIMD mode, a TIMD mode, and a SGPM mode.
  • the MTSS technology can be applied to the DIMD mode, the TIMD mode, and the SGPM mode.
  • the gradient-derived texture features may not be additionally used, and the first candidate list may be constructed based solely on one or more intra-frame prediction modes derived from the prediction mode itself.
  • a first candidate list is constructed by selecting multiple intra prediction modes.
  • the TIMD mode since the TIMD mode itself can derive multiple intra-frame prediction modes, there is no need for the MTSS technology to additionally use gradient-derived texture features to construct the first candidate list based on the multiple intra-frame prediction modes derived from the TIMD mode itself.
  • the first candidate list is constructed based on a "partitioning" mode and 2 intra-frame prediction modes derived from the SGPM mode itself.
  • determining a first candidate list for the current block may further include: when the first candidate list is not full, determining a preset texture feature index for the current block; determining one or more candidate transform core groups based on the preset texture feature index, and adding the one or more candidate transform core groups to the first candidate list.
  • the preset texture feature index may be a PLANAR mode.
  • the method may further include: determining multiple candidate transform core groups according to multiple intra prediction modes and the PLANAR mode, and adding the multiple candidate transform core groups to the first candidate list.
  • the texture features corresponding to the PLANAR mode can also be used as candidate texture features.
  • the intra-frame prediction mode may be ITMP or IBC, or EIP or MIP mode.
  • the texture features corresponding to the PLANAR mode can also be used as candidate texture features.
  • the first candidate list can be constructed based on multiple intra-frame prediction modes derived from the DIMD mode, TIMD mode, and SGPM mode and the PLANAR mode.
  • determining the first candidate list for the current block may specifically include: when the prediction mode of the current block is one of the items in the third prediction mode set, determining an intra-frame prediction mode derived based on the prediction mode of the current block; determining two candidate transform core groups according to the intra-frame prediction mode and the PLANAR mode, and adding the two candidate transform core groups to the first candidate list.
  • the third prediction mode set may include at least one of the following prediction modes: a mode using an extrapolation filter for prediction and a mode using a matrix operation for prediction.
  • the third prediction mode set may include at least one of the following prediction modes: a MIP mode and an EIP mode.
  • the MTSS technology can be applied to the MIP mode and the EIP mode.
  • the gradient-derived texture features may not be used additionally, and the first candidate list can be constructed based solely on the first intra-frame prediction mode and the PLANAR mode derived from the prediction mode itself.
  • the candidate samples used for inter-frame derivation of candidate texture feature indexes can be all samples in the prediction block or part of the samples in the prediction block.
  • the method may further include: determining the number of candidate samples according to a size parameter of the current block.
  • the number of candidate samples used can be determined by the size parameter of the current block. For example, if the size of the current block is small, then all available samples can be counted; if the size of the current block is large, then the current block can be downsampled and counted, such as counting one sample out of every 2, or 4, or 8 samples in the horizontal and/or vertical directions.
  • performing angle mapping based on the horizontal gradient value and the vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample may include: determining the texture feature index corresponding to the candidate sample using a preset lookup table based on the horizontal gradient value and the vertical gradient value of the candidate sample.
  • abs(grad x ) is equal to 2 times abs(grad y ) and grad x and grad y have the same sign, it corresponds to intra-frame prediction mode 40 in VVC.
  • other situations can be determined by looking up the table according to the same principle.
  • the horizontal gradient value and the vertical gradient value of the candidate sample can be calculated using the Sobel operator.
  • the Sobel operator is as follows:
  • the candidate samples may be set to exclude the samples in the outermost row, column, and top, bottom, left, and right of the current block.
  • the embodiment of the present application may be set to exclude the gradients of the samples in the outermost row, column, and top, bottom, left, and right of the current block.
  • the gradients of all or part of the samples in the prediction block are calculated.
  • the horizontal gradient value and the vertical gradient value can be calculated.
  • the Sobel operator can be used to calculate the gradient value.
  • the texture direction of the sample can be inferred based on its horizontal gradient value and vertical gradient value. For example, if the horizontal gradient value is non-zero and the vertical gradient value is zero, then the texture of the sample is in the vertical direction. Conversely, if the horizontal gradient value is zero and the vertical gradient value is non-zero, then the texture of the sample is in the horizontal direction.
  • the texture of the sample is 45 degrees.
  • the texture direction of the sample can be determined based on their ratio.
  • the gradient intensity value of each sample can be mapped to the corresponding texture feature index.
  • a texture feature statistics table is constructed, and the gradient intensity value of each calculated sample is added to the corresponding texture feature index item in the statistics table to obtain the final texture feature statistics table. Then, one or more candidate texture feature indexes of the current block can be determined based on the texture feature statistics table.
  • determining one or more candidate texture feature indexes for a current block based on a texture feature statistics table may include: sorting the texture feature statistics table from high to low according to gradient intensity accumulation values, determining reference texture feature indexes corresponding to top N gradient intensity accumulation values, and determining the N reference texture feature indexes as the one or more candidate texture feature indexes for the current block, where N is a positive integer.
  • the texture feature statistics table can also be sorted in descending order according to the accumulated gradient intensity accumulation values, and then only the top N reference texture feature indexes are selected.
  • N can be 2, 3, 4, 5, ..., 10, etc., and there is no specific limitation on the value of N here.
  • the method may also include: pruning the N reference texture feature indexes in the candidate texture feature list to determine one or more candidate texture feature indexes for the current block.
  • pruning N reference texture feature indexes to determine one or more candidate texture feature indexes for the current block may include: determining a first candidate texture feature index based on a reference texture feature index at a first position among the N reference texture feature indexes; when a second condition is satisfied between other reference texture feature indexes other than the first position among the N reference texture feature indexes and the i-th candidate texture feature index, determining an i+1-th candidate texture feature index based on the other reference texture feature indexes to determine one or more candidate texture feature indexes for the current block; wherein i is an integer greater than zero and less than N.
  • the second condition may include: the difference between the i-th candidate texture feature index and the i+1-th candidate texture feature index satisfies a preset threshold, i.e., there is a certain difference between two adjacent candidate texture feature indexes.
  • a preset threshold i.e., there is a certain difference between two adjacent candidate texture feature indexes.
  • determining the first candidate texture feature index based on the reference texture feature index in the first position among the N reference texture feature indexes may include: determining the first candidate texture feature index based on the reference texture feature index with the largest accumulated gradient strength value among the N reference texture feature indexes.
  • the preset threshold can be represented by THR
  • the i-th candidate texture feature index can be represented by candFeature(i)
  • the i+1-th candidate texture feature index can be represented by candFeature(i+1).
  • the difference between the i-th candidate texture feature index and the i+1-th candidate texture feature index meets the preset threshold, which can include: candFeature(i+1)+THR ⁇ candFeature(i)
  • the THR value may be 3, 4, 5, 6, etc.
  • this pruning method will preferentially select candidate texture feature indices with a certain degree of discrimination.
  • THR is equal to 0 or is not set, that is, only the candidate texture feature indices need to be different.
  • pruning can also be omitted, that is, pruning is not a necessary operation step.
  • the above method does not take into account some intra-frame prediction modes derived from the DIMD, TIMD, SGPM, and other modes themselves. Therefore, in some embodiments, the method may further include: determining one or more intra-frame prediction modes derived from the prediction mode of the current block; determining one or more candidate transform core groups based on the one or more intra-frame prediction modes, and adding the one or more candidate transform core groups to the first candidate list.
  • the method further includes: when the first candidate list is not full, determining a preset texture feature index of the current block; determining one or more candidate transform core groups according to the preset texture feature index, and adding the one or more candidate transform core groups to the first candidate list.
  • the method further includes: when the first candidate list is not filled, determining candidate samples for deriving texture feature indexes; determining one or more candidate texture feature indexes of the current block based on the candidate samples; determining one or more candidate transform core groups based on the one or more candidate texture feature indexes, and adding the one or more candidate transform core groups to the first candidate list.
  • DIMD itself will derive one or several intra-frame prediction modes for weighting
  • TIMD itself will also derive one or several intra-frame prediction modes for weighting
  • SGPM not only has two intra-frame prediction modes, but also has a "partitioning" mode that can also find the corresponding intra-frame prediction mode, and the residual often appears in the boundary area of the "partition”.
  • one possible implementation method is to determine the first candidate list based on the intra-frame prediction mode derived from the DIMD, TIMD, SGPM and other modes themselves and the candidate texture feature index derived by the above method; or another possible implementation method is to determine the first candidate list based on the intra-frame prediction mode derived from the DIMD, TIMD, SGPM and other modes themselves and the default texture feature index.
  • the prediction mode of the current block is the DIMD mode
  • the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.
  • the prediction mode of the current block is the TIMD mode
  • the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.
  • the intra-frame prediction mode corresponding to the SGPM "partition" mode is used as candFeature0.
  • the method may further include: determining the at least two candidate transformation cores indicated by the first candidate list; performing encoding cost calculation on the current block based on the at least two candidate transformation cores, and determining the cost results corresponding to each of the at least two candidate transformation cores; determining the minimum cost result among the cost results corresponding to each of the at least two candidate transformation cores, and determining the candidate transformation core corresponding to the minimum cost result as the transformation core of the current block.
  • the cost calculation here can be determined based on the cost result of Rate Distortion Optimization (RDO), or based on the cost result of Sum of Absolute Difference (SAD), or even based on the cost result of Sum of Absolute Transformed Difference (SATD), but no limitation is made here.
  • the transform core index of the current block can be used to indicate the number of the transform core of the current block in the transform core group or the first candidate list of the current block.
  • the transform core index of the current block can be a positive integer, such as 1, 2, 3, 4, 5, 6, etc.
  • the transform core index can be written directly into the bitstream or written into the bitstream via the value of the second syntax element.
  • a transform core index of a current block is determined; the transform core index of the current block is coded, and the obtained coded bits are written into a bitstream.
  • a value of a second syntax element is determined; wherein the second syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform kernel index used; the value of the second syntax element is encoded, and the obtained encoded bits are written into the bitstream.
  • the value of the second syntax element is a third value, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is a fourth value, it is determined that the current block uses the first transform mode and the corresponding transform kernel index.
  • the third value can be set to 0, and the fourth value can be set to a non-zero value, such as 1, 2, 3, 4, 5, 6, etc.
  • the binary symbol correspondence table of lfnst_idx is shown in the aforementioned Table 7.
  • the third binary symbol, ie, the binary symbol with BinIdx being 2 can also be understood as selecting the first candidate transform core group or the second candidate transform core group.
  • the transform core group here can be one of the at least two candidate transform core groups indicated by the first candidate list.
  • the transform core group index of the current block can be represented by lfnst_feature_idx, or by lfnst_set_idx. Among them, the transform core group index of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, etc.
  • the transform core index of the current block can be used to indicate the number of the transform core of the current block in the transform core group of the current block.
  • the transform core index of the current block can be a positive integer, such as 1, 2, 3, etc.
  • the transform core index can be written directly into the bitstream or written into the bitstream via the value of the second syntax element.
  • a value of a second syntax element is determined; wherein the second syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform kernel index used; the value of the second syntax element is encoded, and the obtained encoded bits are written into the bitstream.
  • S3303 Determine a residual block of the current block, and transform the residual block of the current block according to the transformation kernel to determine a transformation coefficient of the current block.
  • S3304 Encode the transform coefficients of the current block and write the obtained coded bits into the bitstream.
  • the method when encoding the transform coefficients of the current block, may include: quantizing the transform coefficients of the current block to determine the quantization coefficients of the current block; encoding the quantization coefficients of the current block and writing the obtained coded bits into the bitstream.
  • the "transformation" of the residual block by the encoder can also be called a “forward transform,” specifically referring to the transformation from the spatial domain to the frequency domain to remove residual correlation. It should be noted that if the standard only specifies decoding, then the “transformation” in the standard text refers to the decoding part, specifically referring to the "inverse transform” in this article.
  • the method may include:
  • S3401 Perform intra-frame prediction on the current block to determine a prediction block for the current block.
  • S3402 Determine a residual block of the current block according to the initial block of the current block and the predicted block of the current block.
  • steps S3401 to S3402 can be operated in parallel with steps S3301 to S3302, or can be executed before steps S3301 to S3302, or can be executed after steps S3301 to S3303.
  • the order of the steps is not specifically limited here.
  • the size parameter of the current block satisfies the first condition, including: the size parameter of the current block is relatively small, for example, the size parameter of the current block is less than a certain threshold.
  • an NSPT transform kernel is used for relatively small blocks. That is, an NSPT transform is performed on the residual block of the current block according to the transform kernel to determine the transform coefficients of the current block.
  • the size parameter of the current block satisfies the second condition, including: the size parameter of the current block is relatively large, for example, the size parameter of the current block is greater than a certain threshold.
  • the LFNST transform kernel is used. Specifically, a basic DCT2 transform is first performed on the residual block of the current block, and then an LFNST transform is performed on the transform block of the current block based on the transform kernel to determine the transform coefficients of the current block.
  • the residual block can be forward transformed using NSPT to obtain transform coefficients, the transform coefficients are quantized to obtain quantized coefficients, and then the quantized coefficients are entropy encoded. Entropy encoding can be used to determine the overhead cost in the bitstream for this transform kernel.
  • the quantized coefficients are inversely quantized to obtain decoded transform coefficients, and the decoded transform coefficients are inversely NSPT transformed to obtain the decoded residual block.
  • the decoded transform coefficients may be different from the original transform coefficients because quantization is lossy.
  • the decoded residual block may also be different from the original residual block.
  • a reconstructed block is obtained based on the decoded residual block and the prediction block.
  • the distortion cost can be determined based on the reconstructed block and the original image of the current block.
  • the cost of encoding using the current NSPT transform kernel is the sum of the overhead cost and the distortion cost.
  • the costs of several transform kernels are compared, and the smallest one is selected as the optimal NSPT for the current block.
  • the residual block can be forward transformed using DCT2, then forward transformed using LFNST to obtain transform coefficients. These transform coefficients are quantized to obtain quantized coefficients, and then entropy encoded. Entropy encoding can be used to determine the overhead cost in the bitstream for this transform kernel. The quantized coefficients are then inversely quantized to obtain decoded transform coefficients. These decoded transform coefficients are then subjected to an inverse LFNST transform and then an inverse DCT2 transform to obtain the decoded residual block. The decoded transform coefficients may differ from the original transform coefficients because quantization is lossy. Similarly, the decoded residual block may also differ from the original residual block.
  • the decoded residual block and the prediction block are used to obtain a reconstructed block.
  • the distortion cost can be determined based on the reconstructed block and the original image of the current block.
  • the cost of encoding using the current NSPT transform kernel is the sum of the overhead cost and the distortion cost.
  • the costs of several transform kernels are compared, and the one with the smallest cost is selected as the optimal NSPT for the current block.
  • the method may also include: determining the texture feature index of the current block; determining the transformation kernel of the current block based on the texture feature index; determining the transformation coefficient of the current block, and transforming the transformation coefficient of the current block based on the transformation kernel to determine the residual block of the current block.
  • determining the transformation kernel of the current block according to the texture feature index may include: determining the transformation kernel group of the current block according to the texture feature index; and determining the transformation kernel of the current block according to the transformation kernel group.
  • the method further includes: when the number of samples of the current block is less than a minimum sample threshold, executing the step of determining a texture feature index of the current block.
  • the size of the current block includes height and width; the method further comprises: when the height or width of the current block is less than a minimum size threshold, performing a step of determining a texture feature index of the current block.
  • the method further comprises: when the prediction mode of the current block is a prediction mode outside the first prediction mode set, performing a step of determining a texture feature index of the current block.
  • the prediction mode of the current block when the prediction mode of the current block is a prediction mode outside the first prediction mode set, the prediction mode of the current block can be one of the fourth prediction mode set.
  • the fourth prediction mode set includes at least: DC mode mode, PLANAR mode and angle prediction mode.
  • the modes in the fourth prediction mode set are relatively simple prediction modes. For example, a classification can be made, where the DC mode, PLANAR mode, and various angle prediction modes are classified as the fourth prediction mode set, and the DIMD mode, TIMD mode, MIP mode, EIP mode, SGPM mode, ITMP mode, and IBC mode are classified as the first prediction mode set.
  • the first prediction mode set may include only one or more of the DIMD mode, TIMD mode, MIP mode, EIP mode, SGPM mode, ITMP mode, and IBC mode.
  • the first prediction mode set includes the DIMD mode, TIMD mode, and SGPM mode.
  • the other modes belong to the fourth prediction mode set.
  • the encoding method of lfnst_idx of the fourth prediction mode set is the same as that of the related art, while the encoding method of lfnst_idx of the first prediction mode set is different from that of the related art.
  • each transform core group in LFNST/NSPT has three transform cores.
  • Blocks using the fourth prediction mode set can only select one transform core group, while blocks using the first prediction mode set can select two transform core groups, each with three transform cores. If the prediction mode of the current block belongs to the fourth prediction mode set, it can only select one texture feature index, and the possible values of lfnst_idx are 0, 1, 2, or 3. If the prediction mode of the current block belongs to the first prediction mode set, the possible values of lfnst_idx are 0, 1, 2, 3, 4, 5, or 6.
  • the method further includes: determining a value of a third syntax element; encoding the value of the third syntax element, and writing the resulting encoded bits into the bitstream.
  • the third syntax element is used to indicate whether the current sequence allows the use of the multi-transform core group selection technology.
  • the third syntax element can be represented by sps_mtss_enabled_flag, and the third syntax element is a syntax element in a sequence parameter set (SPS).
  • the value of the third syntax element is determined to be the fifth value; if the current sequence does not allow the use of MTSS technology, the value of the third syntax element is determined to be the sixth value.
  • the method further includes: when the current sequence allows the use of the MTSS technology, executing the step of determining a transformation core group of the current block, wherein the current sequence includes the current block.
  • the fifth value is different from the sixth value, and the fifth value and the sixth value can be in parameter form or in digital form.
  • the third syntax element can be a parameter written in the profile or the value of a flag, which is not specifically limited here.
  • the fifth value can be 1 and the sixth value can be 0; or, the fifth value can be 0 and the sixth value can be 1; or, the fifth value can be true and the sixth value can be false; or, the fifth value can be false and the sixth value can be true.
  • the fifth value is 1 and the sixth value is 0.
  • the embodiment of the present application can use a high-level syntax to control the switch of the present technical solution.
  • a sequence-level flag is used, such as adding a syntax element sps_mtss_enabled_flag in the sequence parameter set. If the value of sps_mtss_enabled_flag is 1, the current sequence allows the use of MTSS technology; if the value of sps_mtss_enabled_flag is 0, the current sequence does not allow the use of MTSS technology.
  • the method further includes: determining a value of a third syntax element and a value of a fourth syntax element; wherein the third syntax element is used to indicate whether the current sequence allows the use of a multi-transform kernel group selection technique, and the fourth syntax element is used to indicate whether the current image allows the use of a multi-transform kernel group selection technique; encoding the values of the third syntax element and the fourth syntax element, and writing the obtained encoded bits into the bitstream.
  • the method further includes: when the current sequence allows the use of MTSS technology, determining whether the current image allows the use of MTSS technology; when the current image allows the use of MTSS technology, executing the step of determining the transformation core group of the current block.
  • the current sequence may include a current picture, and the current picture includes a current block.
  • the fourth syntax element may be represented by ph_inter_lfnst_nspt_enabled_flag, and the fourth syntax element is a picture-level syntax element.
  • the value of the fourth syntax element is determined to be the fifth value; if the current image does not allow the use of MTSS technology, the value of the fourth syntax element is determined to be the sixth value.
  • the method further includes: determining a value of a third syntax element and a value of a fifth syntax element; wherein the third syntax element is used to indicate whether the current sequence allows the use of the multi-transform core group selection technology, and the fifth syntax element is used to indicate whether the current slice allows the use of the multi-transform core group selection technology; encoding the values of the third syntax element and the values of the fifth syntax element, and writing the obtained coded bits into the bitstream.
  • the method further comprises: when the current sequence allows the use of the MTSS technology, determining whether the current slice is Whether the MTSS technology is allowed to be used; when the current slice allows the MTSS technology to be used, the step of determining the transformation core group of the current block is executed.
  • the current sequence may include a current slice, and the current slice includes a current block.
  • the fifth syntax element may be represented by sh_inter_lfnst_nspt_enabled_flag, and the fifth syntax element is a slice-level syntax element.
  • the value of the fifth syntax element is determined to be the fifth value; if the current slice does not allow the use of MTSS technology, the value of the fifth syntax element is determined to be the sixth value.
  • the fifth value is different from the sixth value, and the fifth value and the sixth value can be in parameter form or in digital form.
  • the fourth syntax element or the fifth syntax element can be a parameter written in the profile, or it can be the value of a flag, which is not specifically limited here.
  • the fifth value can be 1 and the sixth value can be 0; or, the fifth value can be 0 and the sixth value can be 1; or, the fifth value can be true and the sixth value can be false; or, the fifth value can be false and the sixth value can be true.
  • the fifth value is 1 and the sixth value is 0.
  • sps_mtss_enabled_flag in addition to the sequence-level syntax element sps_mtss_enabled_flag, other levels of syntax can also be used to achieve more flexible control, such as the flag of the Picture Parameter Set (PPS), or the flag of the picture header or slice header.
  • PPS Picture Parameter Set
  • the method may include: determining the minimum sample threshold; encoding the minimum sample threshold, and writing the resulting coded bits into the bitstream.
  • the method may include: determining the value of a sixth syntax element when the current sequence allows the use of a multi-transform kernel group selection technique; wherein the sixth syntax element is used to indicate the minimum sample threshold; encoding the value of the sixth syntax element, and writing the resulting coded bits into the bitstream.
  • the minimum sample threshold may be written directly into the bitstream, or written into the bitstream in the form of a sixth syntax element.
  • the sixth syntax element may be represented by sps_mtss_min_pix, which may be a sequence-level syntax element.
  • the value of the sixth syntax element can be equal to the minimum sample threshold, or it can be encoded according to a certain mapping rule. For example, if the minimum sample threshold is 32, the value of the sixth syntax element is determined to be 0; if the minimum sample threshold is 64, the value of the sixth syntax element is determined to be 1; if the minimum sample threshold is 256, the value of the sixth syntax element is determined to be 2. In this way, after the value of the sixth syntax element is written into the bitstream, if the subsequent decoding end obtains the value of the sixth syntax element equal to 2 during decoding, it can be determined that the minimum sample threshold is 256.
  • a high-level syntax can be used to set a minimum sample threshold for applying MTSS technology, such as sps_mtss_min_pix.
  • a minimum sample threshold for applying MTSS technology such as sps_mtss_min_pix.
  • the encoder encodes sps_mtss_min_pix to determine the minimum sample threshold for applying MTSS technology.
  • the encoder needs to support all possible cases of sps_mtss_min_pix, but the encoder can configure the sps_mtss_min_pix required to encode the current bitstream. For example, when encoding a bitstream, if better compression efficiency is required but encoding time is not particularly important, a relatively small value, such as 16, can be set for sps_mtss_min_pix. If encoding time is particularly important and a certain degree of compression efficiency can be sacrificed, a relatively large value, such as 256, can be set for sps_mtss_min_pix.
  • the method may include: determining the minimum size threshold; encoding the minimum size threshold, and writing the resulting coded bits into the bitstream.
  • the method may include: determining the value of the seventh syntax element when the current sequence allows the use of the multi-transform kernel group selection technique; wherein the seventh syntax element is used to indicate the minimum size threshold; encoding the value of the seventh syntax element, and writing the resulting coded bits into the bitstream.
  • the minimum size threshold may be written directly into the bitstream, or written into the bitstream in the form of the seventh syntax element.
  • the seventh syntax element may be represented by sps_mtss_min_size, which may be a sequence-level syntax element.
  • the value of the seventh syntax element can be equal to the minimum size threshold, or it can be encoded according to a certain mapping rule. For example, if the minimum size threshold is equal to 4, the value of the seventh syntax element is determined to be equal to 0; if the minimum sample threshold is equal to 16, the value of the seventh syntax element is determined to be equal to 1. In this way, after the value of the seventh syntax element is written into the bitstream, if the subsequent decoding end obtains the value of the seventh syntax element equal to 1 during decoding, it can be determined that the minimum size threshold is equal to 16.
  • a high-level syntax can be used to set a minimum size threshold for applying MTSS technology, such as sps_mtss_min_size.
  • a minimum size threshold for applying MTSS technology such as sps_mtss_min_size.
  • the encoder encodes sps_mtss_min_size to determine the minimum size threshold for applying MTSS technology.
  • a trade-off between coding complexity and compression efficiency can be made based on demand. That is, under this method, the encoder needs to support all possible cases of sps_mtss_min_size, but the encoder can configure the sps_mtss_min_size required to encode the current bitstream.
  • a relatively small value such as 4 can be set for sps_mtss_min_size.
  • a relatively large value such as 16, can be set for sps_mtss_min_size.
  • the embodiment of the present application provides a code stream, wherein the code stream is generated by bit encoding according to the information to be encoded; wherein the information to be encoded includes at least one of the following: a quantization coefficient of the current block, a transform kernel index of the current block, The transform kernel group index, the minimum sample threshold, the minimum size threshold, the value of the first syntax element, the value of the second syntax element, the value of the third syntax element, the value of the fourth syntax element, the value of the fifth syntax element, the value of the sixth syntax element, and the value of the seventh syntax element of the current block are described.
  • the first syntax element is used to indicate the transform kernel group index of the current block
  • the second syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform kernel index
  • the third syntax element is used to indicate whether the current sequence allows the use of multiple transform kernel group selection techniques
  • the fourth syntax element is used to indicate whether the current image allows the use of multiple transform kernel group selection techniques
  • the fifth syntax element is used to indicate whether the current slice allows the use of multiple transform kernel group selection techniques
  • the sixth syntax element is used to indicate the minimum sample threshold
  • the seventh syntax element is used to indicate the minimum size threshold.
  • the second syntax element if the value of the second syntax element is 0, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is non-zero, it is determined that the current block uses the first transform mode and the corresponding transform kernel index. For example, if the value of the second syntax element is 1, it is determined that the current block uses the first transform kernel; if the value of the second syntax element is 2, it is determined that the current block uses the second transform kernel; if the value of the second syntax element is 4, it is determined that the current block uses the fourth transform kernel, and so on, without any limitation here.
  • the embodiment of the present application provides a coding method, specifically an intra-frame LFNST/NSPT multi-angle selection scheme.
  • the transform kernel group of the current block is determined; then, based on the transform kernel group, the transform kernel of the current block is determined; then, the residual block of the current block is determined, and the residual block of the current block is transformed according to the transform kernel to determine the transform coefficient of the current block; finally, the transform coefficient of the current block is encoded, and the obtained coded bits are written into the bitstream.
  • the transform kernel group of the current block is determined according to the multi-transform kernel group selection technology, and then the transform kernel of the current block is determined therefrom.
  • multiple candidate texture features derived by the multi-transform kernel group selection technology can be used to guide the transform, thereby improving the accuracy of the transform prediction, thereby improving the compression efficiency, and further improving the encoding and decoding performance.
  • the LFNST/NSPT transform core index can be represented by the syntax element lfnst_idx.
  • lfnst_idx can have three values, namely 0, 1, and 2. Among them, lfnst_idx is 0, which means that the current block does not use LFNST.
  • Each transform core group of LFNST in VVC has 2 transform cores, so the value of lfnst_idx is 1 or 2, which means that the current block uses the first transform core or the second transform core of the selected transform core group of LFNST.
  • lfnst_idx can have four values, namely 0, 1, 2, and 3. Where lfnst_idx is 0, which means that the current block does not use LFNST/NSPT.
  • Each transform core group of LFNST/NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST/NSPT.
  • the intra prediction mode of the current block is a special intra prediction mode, it can select more than one transform kernel group. For example, it can select two transform kernel groups.
  • These special intra prediction modes include but are not limited to DIMD, TIMD, MIP, EIP, SGPM, ITMP, and IBC.
  • a classification is made here: DC, PLANAR, and various angle prediction modes are classified into the first intra-frame prediction mode set (equivalent to the "fourth prediction mode set" of the aforementioned embodiment), and DIMD, TIMD, MIP, EIP, SGPM, ITMP, IBC, etc. are classified into the second intra-frame prediction mode set (equivalent to the "first prediction mode set" of the aforementioned embodiment).
  • Each transform kernel group in LFNST/NSPT has three transform kernels. Blocks using the first intra-frame prediction mode set can only select one transform kernel group, while blocks using the second intra-frame prediction mode set can select two transform kernel groups, each with three transform kernels. If the prediction mode of the current block belongs to the first intra-frame prediction mode set, it can only select one texture feature, so the possible values of lfnst_idx are 0, 1, 2, and 3. If the prediction mode of the current block belongs to the second intra-frame prediction mode set, the possible values of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6. Among them, 1, 2, and 3 correspond to the three transform kernels of the transform kernel group corresponding to the first texture feature index, and 4, 5, and 6 correspond to the three transform kernels of the transform kernel group corresponding to the second texture feature index.
  • the binary symbol correspondence table of lfnst_idx of the block using the first intra prediction mode set is shown in the aforementioned Table 8
  • the binary symbol correspondence table of lfnst_idx of the block using the second intra prediction mode set is shown in the aforementioned Table 7.
  • the third binary symbol that is, the binary symbol with BinIdx being 2
  • the second intra-frame prediction mode set may include only one or more of DIMD, TIMD, MIP, EIP, SGPM, ITMP, and IBC.
  • the second intra-frame prediction mode set includes DIMD, TIMD, and SGPM.
  • the other modes belong to the first intra-frame prediction mode set.
  • the lfnst_idx encoding method of the first intra-frame prediction mode set is the same as that of the related art, while the lfnst_idx encoding method of the second intra-frame prediction mode set is different from that of the related art.
  • a dedicated syntax element such as lfnst_feature_idx
  • the possible values of lfnst_feature_idx are 0 or 1, indicating which candidate texture feature is selected, or which candidate transform core group is selected. If the intra prediction mode of the current block belongs to the second intra prediction mode set and lfnst_idx>0, lfnst_feature_idx is parsed, and the result is equivalent to the above embodiment.
  • this syntax element can also be called lfnst_set_idx, that is, the candidate transform core group index.
  • the MTSS technology can use the reconstructed
  • the candidate texture feature index is derived from the predicted area, which is similar to the existing DIMD approach. Because the reconstructed areas on the left and above are not the current block but are adjacent to the current block, for example, when the textures are connected, they can be used to estimate the texture of the current block to a certain extent. Another possibility is to use the predicted block of the current block to derive the candidate texture feature index. Another possibility is to use the predicted block of the current block and the reconstructed areas on the left and above the current block at the same time, so that more samples can be used to derive the candidate texture feature index.
  • candidate texture feature indexes (or candidate texture features) here can also be directly referred to as candidate transformation kernel groups.
  • One derivation method is to calculate the gradients of all or part of the samples in the selected area.
  • horizontal and vertical gradients can be calculated, and the Sobel operator can be used to calculate the gradients.
  • the texture direction is inferred based on its horizontal and vertical gradients. For example, if the horizontal gradient is non-zero and the vertical gradient is zero, the texture at that point is vertical. Conversely, if the horizontal gradient is zero and the vertical gradient is non-zero, the texture at that point is horizontal. For example, if the horizontal and vertical gradients are equal and non-zero, the texture at that point is 45 degrees.
  • abs(grad x ) and abs(grad y ) are not equal to 0: if abs(grad x ) is equal to abs(grad y ), and grad x and grad y have the same sign, it corresponds to intra-frame prediction mode 34 in VVC; if abs(grad x ) is equal to 2 times abs(grad y ), and grad x and grad y have the same sign, it corresponds to intra-frame prediction mode 40 in VVC.
  • DIMD also needs to perform similar statistical analysis to derive the intra-frame prediction mode. In terms of implementation, some logic can be reused with DIMD.
  • a sorted list of intra-frame prediction modes of length N is obtained.
  • a list of candidate texture features is then generated.
  • An optional operation is to prune the N intra-frame prediction modes, or to exclude intra-frame prediction modes that are too close to each other. Taking the intra-frame prediction modes in VVC as an example, some adjacent intra-frame prediction modes have very similar angles. If two candidate texture features are selected at very similar angles, the difference between them is not obvious. Therefore, a threshold can be set to represent the difference between the two intra-frame prediction modes.
  • the above method does not take into account some of the modes derived from the DIMD, TIMD, SGPM and other modes themselves.
  • DIMD itself will derive one or several intra-frame prediction modes for weighting
  • TIMD itself will also derive one or several intra-frame prediction modes for weighting
  • SGPM not only has two intra-frame prediction modes, but also a "partition" mode that can find the corresponding intra-frame prediction mode, and the residual often appears in the boundary area of the "partition”.
  • One possible method is to determine the candidate texture feature index based on the modes derived from the DIMD, TIMD, SGPM and other modes themselves and the modes derived by the above method.
  • candFeature1 is determined according to the above method. That is, the intra-frame prediction modes in the sorted intra-frame prediction mode list of length N are tried, starting from the first intra-frame prediction mode. If an intra-frame prediction mode meets the THR restriction, it is used as candFeature1.
  • the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.
  • the intra-frame prediction mode corresponding to the "partition" mode derived therefrom and the two intra-frame prediction modes used for prediction may be sequentially attempted to be determined as candidate texture feature indexes.
  • the MIP mode and EIP mode can be widely used in blocks with texture gradients, and their features are similar to those of the PLANAR mode. Therefore, the PLANAR mode can also be used as a candidate texture feature for the MIP mode and EIP mode.
  • the first derived texture feature is used as its candFeature0, and for PLANAR mode, the first derived texture feature is used as its candFeature1.
  • gradient statistics are performed on all available samples in that direction. Otherwise, if the horizontal or vertical size of the current block is less than or equal to 16, gradient statistics are performed on one of every two samples in that direction. Otherwise, gradient statistics are performed on one of every four samples in that direction.
  • the size of the blocks to which the MTSS technology is applicable can also be restricted. Specifically, the increase in the complexity of the decoder caused by the MTSS technology is reflected in the difference in parsing syntax elements, which is generally believed to have little effect on the complexity of the decoder. However, for the encoder, this will increase the complexity of the encoder due to the increase in candidates for the transform core group. Specifically, a block may choose one from two or more LFNST/NSPT transform core groups. However, a block of VVC has only one available transform core group. Therefore, one solution is to limit the size of the blocks to which the MTSS technology is applicable, so that the MTSS technology can be disabled at block sizes where the MTSS technology requires a lot of additional calculations but the compression efficiency improvement is not obvious.
  • a minimum sample threshold for applying MTSS can be set, called MIN_PIX. If the number of samples (width multiplied by height) of a block is less than MIN_PIX, MTSS cannot be used for this block. Otherwise, if the number of samples of this block is greater than or equal to MIN_PIX, MTSS can be used for this block.
  • MIN_PIX can be set to 32, 64, 256, etc.
  • a relatively small value such as 16
  • a relatively large value such as 256
  • the embodiment of the present application may also use a high level syntax to set a minimum size threshold for applying MTSS technology, such as sps_mtss_min_size.
  • a high level syntax to set a minimum size threshold for applying MTSS technology, such as sps_mtss_min_size.
  • the decoder parses sps_mtss_min_size to determine the minimum size threshold for applying MTSS technology.
  • sps_mtss_min_size a relatively small value for sps_mtss_min_size. If you pay special attention to encoding time and can sacrifice a certain amount of compression efficiency, you can set a relatively large value for sps_mtss_min_size, such as 16.
  • FIG35 is a schematic diagram of the composition structure of an encoder provided in an embodiment of the present application.
  • the encoder 350 may include a first determination unit 3501, a first transformation unit 3502, and an encoding unit 3503, wherein:
  • a first transform unit 3502 configured to determine a residual block of a current block, and transform the residual block of the current block according to a transform kernel to determine a transform coefficient of the current block;
  • the encoding unit 3503 is configured to perform encoding processing on the transform coefficients of the current block and write the obtained encoding bits into the bitstream.
  • the first determination unit 3501 is further configured to perform encoding cost calculation on the current block based on at least two candidate transform core groups, and determine the cost results corresponding to each of the at least two candidate transform core groups; determine the minimum cost result among the cost results corresponding to each of the at least two candidate transform core groups, and determine the candidate transform core group corresponding to the minimum cost result as the transform core group of the current block.
  • the first determination unit 3501 is further configured to determine the transform core group index of the current block; wherein the transform core group index is used to indicate the number of the transform core group of the current block in at least two candidate transform core groups; the encoding unit 3503 is further configured to encode the transform core group index of the current block and write the obtained encoded bits into the bitstream.
  • the first determining unit 3501 is further configured to determine the number of samples of the current block; when the number of samples of the current block is greater than or equal to the minimum sample threshold, perform the step of determining the transformation core group of the current block.
  • the first prediction mode set includes at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, the intra-frame prediction mode including at least one of the following: DC mode, PLANAR mode and angular prediction mode; a mode for prediction by copying intra-frame blocks; a mode for prediction using an interpolation filter; a mode for prediction using matrix operations.
  • the first determination unit 3501 is further configured to determine the transform core group of the current block based on the i-th intra-frame prediction mode derived from the prediction mode of the current block if the transform core group index is the i-th value, where i is a positive integer; wherein the second prediction mode set includes at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, and the intra-frame prediction mode includes at least one of the following: DC mode, PLANAR mode, and angular prediction mode.
  • the first determination unit 3501 is further configured to determine the transform core group of the current block based on the first intra-frame prediction mode derived from the prediction mode of the current block if the transform core group index is a first value; and determine the transform core group of the current block based on the PLANAR mode if the transform core group index is a second value; wherein the third prediction mode set includes at least one of the following prediction modes: a mode for prediction using an interpolation filter and a mode for prediction using matrix operations.
  • the first determining unit 3501 is further configured to determine a first candidate list for the current block, the first candidate list indicating at least two candidate transform core groups; and determine the transform core group of the current block according to the first candidate list and the transform core group index.
  • the first determining unit 3501 is further configured to determine one or more intra prediction modes derived based on the prediction mode of the current block; and determine one or more candidate transform core groups according to the one or more intra prediction modes, and assign the one or more candidate transform core groups to the target frame.
  • the core replacement group is added to the first candidate list.
  • the first determination unit 3501 is further configured to determine a preset texture feature index of the current block when the first candidate list is not filled; and determine one or more candidate transform core groups based on the preset texture feature index, and add the one or more candidate transform core groups to the first candidate list.
  • the first determination unit 3501 is further configured to, when the prediction mode of the current block is one of the items in the second prediction mode set, determine multiple intra-frame prediction modes derived based on the prediction mode of the current block; and determine multiple candidate transform core groups based on the multiple intra-frame prediction modes, and add the multiple candidate transform core groups to the first candidate list; wherein the second prediction mode set includes at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, and the intra-frame prediction mode includes at least one of the following: DC mode, PLANAR mode and angular prediction mode.
  • the first determining unit 3501 is further configured to determine a plurality of candidate transform core groups according to a plurality of intra prediction modes and a PLANAR mode, and add the plurality of candidate transform core groups to the first candidate list.
  • the first determination unit 3501 is further configured to determine an intra-frame prediction mode derived based on the prediction mode of the current block when the prediction mode of the current block is one of the items in the third prediction mode set; and determine two candidate transform core groups based on the intra-frame prediction mode and the PLANAR mode, and add the two candidate transform core groups to the first candidate list; wherein the third prediction mode set includes at least one of the following prediction modes: a mode for prediction using an interpolation filter and a mode for prediction using matrix operations.
  • the first determination unit 3501 is further configured to determine candidate samples for deriving texture feature indexes when the first candidate list is not filled; determine one or more candidate texture feature indexes of the current block based on the candidate samples; and determine one or more candidate transform core groups based on the one or more candidate texture feature indexes, and add the one or more candidate transform core groups to the first candidate list.
  • the first determination unit 3501 is further configured to determine at least one texture feature index and at least one corresponding gradient intensity value when the number of candidate samples is at least one; determine at least one reference texture feature index with different characteristics based on the at least one texture feature index, and perform cumulative calculation on the gradient intensity values belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the gradient intensity cumulative value corresponding to the at least one reference texture feature index; and determine a texture feature statistics table based on the at least one reference texture feature index and the gradient intensity cumulative value corresponding to the at least one reference texture feature index.
  • the first determining unit 3501 is further configured to prune the N reference texture feature indexes to determine one or more candidate texture feature indexes for the current block.
  • the first determination unit 3501 is further configured to determine at least two candidate transformation cores indicated by the first candidate list; perform encoding cost calculation on the current block based on the at least two candidate transformation cores, and determine the cost results corresponding to each of the at least two candidate transformation cores; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transformation cores, and determine the candidate transformation core corresponding to the minimum cost result as the transformation core of the current block.
  • the first determination unit 3501 is further configured to determine at least two candidate transform cores included in the transform core group; perform encoding cost calculation on the current block based on the at least two candidate transform cores, and determine the cost results corresponding to each of the at least two candidate transform cores; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transform cores, and determine the candidate transform core corresponding to the minimum cost result as the transform core of the current block.
  • the first determination unit 3501 is further configured to determine the transform core index of the current block; wherein the transform core index is used to indicate the number of the transform core of the current block in the first candidate list or the transform core group of the current block; the encoding unit 3503 is further configured to encode the transform core index of the current block and write the obtained encoded bits into the bitstream.
  • the first determination unit 3501 is further configured to determine the value of a second syntax element; wherein the second syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index used, and the transform core index is used to indicate the number of the transform core of the current block in the first candidate list or the transform core group of the current block; the encoding unit 3503 is further configured to encode the value of the first syntax element and write the obtained encoded bits into the bitstream.
  • the first determination unit 3501 is further configured to perform intra-frame prediction on the current block to determine the prediction block of the current block; and determine the residual block of the current block based on the initial block of the current block and the prediction block of the current block.
  • the first transformation unit 3502 is further configured to perform an inseparable basic transform on the residual block of the current block according to the transform kernel to determine the transform coefficient of the current block; or, perform a discrete cosine transform on the residual block of the current block to determine the transform block of the current block, and perform a low-frequency inseparable transform on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.
  • the first determination unit 3501 is further configured to quantize the transform coefficients of the current block to determine the quantization coefficients of the current block; the encoding unit 3503 is further configured to encode the quantization coefficients of the current block and write the obtained encoding bits into the bitstream.
  • the first determination unit 3501 is further configured to determine the value of the third syntax element and the value of the fourth syntax element; wherein the third syntax element is used to indicate whether the current sequence allows the use of multi-transform kernel group selection technology, and the fourth syntax element is used to indicate whether the current image allows the use of multi-transform kernel group selection technology, the current sequence includes the current image, and the current image includes the current block; the encoding unit 3503 is further configured to encode the value of the third syntax element and the value of the fourth syntax element, and write the obtained encoded bits into the bitstream.
  • the first determination unit 3501 is further configured to determine the value of the third syntax element and the value of the fifth syntax element; wherein the third syntax element is used to indicate whether the current sequence allows the use of multi-transform core group selection technology, and the fifth syntax element is used to indicate whether the current slice allows the use of multi-transform core group selection technology, the current sequence includes the current slice, and the current slice includes the current block; the encoding unit 3503 is further configured to encode the value of the third syntax element and the value of the fifth syntax element, and write the obtained encoded bits into the bitstream.
  • the encoding unit 3503 is further configured to perform encoding processing on the minimum sample threshold and write the obtained encoded bits into the bitstream.
  • the first determination unit 3501 is further configured to determine the value of the sixth syntax element when the current sequence allows the use of multi-transform core group selection technology; wherein the sixth syntax element is used to indicate the minimum sample threshold; the encoding unit 3503 is further configured to encode the value of the sixth syntax element and write the obtained coded bits into the bitstream.
  • the encoding unit 3503 is further configured to perform encoding processing on the minimum size threshold and write the obtained encoded bits into the bitstream.
  • the first determination unit 3501 is further configured to determine the value of the seventh syntax element when the current sequence allows the use of multi-transform core group selection technology; wherein the seventh syntax element is used to indicate the minimum size threshold; the encoding unit 3503 is further configured to encode the value of the seventh syntax element and write the obtained coded bits into the bitstream.
  • a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular.
  • the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
  • the above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
  • Figure 36 is a schematic diagram of the specific hardware structure of an encoder provided in an embodiment of the present application.
  • the encoder 350 may include: a first communication interface 3601, a first memory 3602 and a first processor 3603; each component is coupled together through a first bus system 3604.
  • the first bus system 3604 is used to achieve connection and communication between these components.
  • the first bus system 3604 also includes a power bus, a control bus and a status signal bus.
  • various buses are labeled as the first bus system 3604 in Figure 36. Among them,
  • a first memory 3602 is used to store computer programs that can be run on the first processor 3603;
  • the first processor 3603 is configured to, when running the computer program, execute:
  • Determine a transform kernel group for the current block determine a transform kernel for the current block based on the transform kernel group; determine a residual block for the current block, and transform the residual block for the current block based on the transform kernel to determine transform coefficients for the current block; encode the transform coefficients for the current block, and write the resulting coded bits into a bitstream.
  • the first memory 3602 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
  • the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
  • the volatile memory can be a random access memory (RAM), which is used as an external cache.
  • RAM static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDR SDRAM double data rate synchronous DRAM
  • ESDRAM enhanced synchronous DRAM
  • SLDRAM synchronous link DRAM
  • DRRAM direct RAM
  • the first processor 3603 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the first processor 3603.
  • the processor 3603 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.
  • the general-purpose processor can be a microprocessor or any conventional processor.
  • the steps of the methods disclosed in the embodiments of this application can be directly implemented as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor.
  • the software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc.
  • the storage medium is located in the first memory 3602, and the first processor 3603 reads the information in the first memory 3602 and completes the steps of the above method in combination with its hardware.
  • the embodiments described in this application can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof.
  • the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.
  • ASICs application specific integrated circuits
  • DSPs digital signal processors
  • DSPDs digital signal processing devices
  • PLDs programmable logic devices
  • FPGAs field programmable gate arrays
  • the technology described in this application can be implemented by modules (such as processes, functions, etc.) that perform the functions described in this application.
  • the software code can be stored in a memory and executed by a processor.
  • the memory can be implemented in the processor or outside the processor.
  • the first processor 3603 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.
  • This embodiment provides an encoder in which, for a current block predicted using certain intra-frame prediction modes, when determining a transform kernel for the current block, multiple candidate texture features derived using a multi-transform kernel group selection technique can be used to guide the transform, thereby improving the accuracy of the transform prediction and thereby improving compression efficiency and, in turn, encoding and decoding performance.
  • FIG37 is a schematic diagram of the composition structure of a decoder provided in an embodiment of the present application.
  • the decoder 370 may include a second determination unit 3701 and a second transformation unit 3702, wherein:
  • the second transform unit 3702 is configured to determine a transform coefficient of the current block, and transform the transform coefficient of the current block according to the transform kernel to determine a residual block of the current block.
  • the decoding unit 3703 is further configured to decode the code stream and determine the value of the first syntax element; the second determination unit 3701 is further configured to determine the transform core group index of the current block according to the value of the first syntax element.
  • the second determining unit 3701 is further configured to determine the number of samples of the current block; and when the number of samples of the current block is greater than or equal to the minimum sample threshold, perform the step of determining the transform core group index of the current block.
  • the second determination unit 3701 is further configured to determine the size of the current block, wherein the size of the current block includes height and width; when the height and width of the current block are both greater than or equal to the minimum size threshold, the step of determining the transform core group index of the current block is performed.
  • the second determination unit 3701 is further configured to determine a prediction mode of the current block; when the prediction mode of the current block is one of the items in the first prediction mode set, perform the step of determining the transform core group index of the current block; wherein the first prediction mode set includes at least one of the following prediction modes: DC mode, PLANAR mode and other prediction modes other than the angle prediction mode.
  • the first prediction mode set includes at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, the intra-frame prediction mode including at least one of the following: DC mode, PLANAR mode and angle prediction mode; a mode for prediction by copying intra-frame blocks; a mode for prediction using an interpolation filter; a mode for prediction using matrix operations.
  • the first prediction mode set includes at least one of the following prediction modes: DIMD mode, TIMD mode, SGPM mode, MIP mode, EIP mode, ITMP mode, and IBC mode.
  • the second determination unit 3701 is further configured to determine the transform core group of the current block based on the i-th intra-frame prediction mode derived from the prediction mode of the current block if the transform core group index is the i-th value, where i is a positive integer; wherein the second prediction mode set includes at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, and the intra-frame prediction mode includes at least one of the following: DC mode, PLANAR mode, and angular prediction mode.
  • the second determining unit 3701 is further configured to determine the first intra prediction mode derived based on the prediction mode of the current block if the transform core group index is the first value. transform core group; if the transform core group index is the second value, determining the transform core group of the current block based on the PLANAR mode; wherein the third prediction mode set includes at least one of the following prediction modes: a mode for prediction using an extrapolation filter and a mode for prediction using a matrix operation.
  • the second determining unit 3701 is further configured to determine a first candidate list for the current block, where the first candidate list indicates at least two candidate transform core groups; and determine the transform core group of the current block according to the first candidate list and the transform core group index.
  • the second determination unit 3701 is further configured to determine one or more intra-frame prediction modes derived based on the prediction mode of the current block; and determine one or more candidate transform core groups based on the one or more intra-frame prediction modes, and add the one or more candidate transform core groups to the first candidate list.
  • the second determination unit 3701 is further configured to determine a preset texture feature index of the current block when the first candidate list is not filled; and determine one or more candidate transform core groups based on the preset texture feature index, and add the one or more candidate transform core groups to the first candidate list.
  • the second determination unit 3701 is further configured to determine a plurality of intra-frame prediction modes derived based on the prediction mode of the current block when the prediction mode of the current block is one of the items in the second prediction mode set; and determine a plurality of candidate transform core groups based on the plurality of intra-frame prediction modes, and add the plurality of candidate transform core groups to the first candidate list; wherein the second prediction mode set includes at least one of the following prediction modes: a mode for combined prediction using at least two intra-frame prediction modes, and the intra-frame prediction mode includes at least one of the following: DC mode, PLANAR mode and angular prediction mode.
  • the second determining unit 3701 is further configured to determine a plurality of candidate transform core groups according to a plurality of intra prediction modes and a PLANAR mode, and add the plurality of candidate transform core groups to the first candidate list.
  • the second determination unit 3701 is further configured to determine candidate samples for deriving texture feature indexes when the first candidate list is not filled; determine one or more candidate texture feature indexes of the current block based on the candidate samples; and determine one or more candidate transform core groups based on the one or more candidate texture feature indexes, and add the one or more candidate transform core groups to the first candidate list.
  • the second determination unit 3701 is further configured to determine the horizontal gradient value and the vertical gradient value of the candidate sample; determine the texture feature index and the gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and the vertical gradient value of the candidate sample; determine the texture feature statistics table based on the texture feature index and the gradient intensity value corresponding to the candidate sample; and determine one or more candidate texture feature indexes of the current block based on the texture feature statistics table.
  • the second determination unit 3701 is further configured to determine at least one texture feature index and at least one corresponding gradient intensity value when the number of candidate samples is at least one; determine at least one reference texture feature index with different characteristics based on the at least one texture feature index, and perform cumulative calculation on the gradient intensity values belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the gradient intensity cumulative value corresponding to the at least one reference texture feature index; and determine a texture feature statistical table based on the at least one reference texture feature index and the gradient intensity cumulative value corresponding to the at least one reference texture feature index.
  • the second determination unit 3701 is further configured to sort the texture feature statistics table from high to low according to the gradient intensity accumulation value, determine the reference texture feature indexes corresponding to the top N gradient intensity accumulation values; where N is a positive integer; and determine the N reference texture feature indexes as one or more candidate texture feature indexes for the current block.
  • the second determining unit 3701 is further configured to prune the N reference texture feature indexes to determine one or more candidate texture feature indexes for the current block.
  • the second determination unit 3701 is further configured to determine the transform core index of the current block; and determine the transform core of the current block based on the first candidate list and the transform core index.
  • the second determining unit 3701 is further configured to determine a transform core index of the current block; and determine the transform core of the current block according to the transform core group and the transform core index.
  • the decoding unit 3703 is further configured to decode the code stream and determine the transform kernel index of the current block.
  • the decoding unit 3703 is further configured to decode the code stream to determine the quantization coefficient of the current block; the second determination unit 3701 is further configured to dequantize the quantization coefficient of the current block to determine the transformation coefficient of the current block.
  • the second transformation unit 3702 is further configured to perform an inseparable basic transform on the transformation coefficients of the current block according to the transformation kernel to determine the residual block of the current block; or, perform a low-frequency inseparable transform on the transformation coefficients of the current block according to the transformation kernel to determine the transformation block of the current block, and perform a discrete cosine transform on the transformation block of the current block to determine the residual block of the current block.
  • the second determining unit 3701 is further configured to perform intra-frame prediction on the current block to determine a prediction block of the current block; and determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block.
  • the decoding unit 3703 is further configured to decode the code stream and determine the value of the third syntax element; when the third syntax element indicates that the current sequence allows the use of multiple transform core group selection technology, perform the step of determining the transform core group index of the current block; wherein the current sequence includes the current block.
  • the decoding unit 3703 is further configured to decode the code stream and determine the value of the third syntax element; and when the third syntax element indicates that the current sequence allows the use of a multi-transform core group selection mode, decode the code stream and determine the value of the fourth syntax element; the second determination unit 3701 is further configured to execute the step of determining the transform core group index of the current block when the fourth syntax element indicates that the current image allows the use of a multi-transform core group selection technology; wherein the current sequence includes the current image, and the current image includes the current block.
  • the decoding unit 3703 is further configured to decode the code stream and determine the value of the third syntax element; and when the third syntax element indicates that the current sequence allows the use of the multi-transform core group selection mode, decode the code stream and determine the value of the fifth syntax element; the second determination unit 3701 is further configured to perform the step of determining the transform core group index of the current block when the fifth syntax element indicates that the current slice allows the use of the multi-transform core group selection technology; wherein the current sequence includes the current slice, and the current slice includes the current block.
  • the decoding unit 3703 is further configured to decode the code stream and determine a minimum sample threshold.
  • the decoding unit 3703 is further configured to decode the code stream and determine the value of the third syntax element; the second determination unit 3701 is further configured to decode the code stream and determine the value of the sixth syntax element when the third syntax element indicates that the current sequence allows the use of the multi-transform core group selection mode; and determine the minimum sample threshold based on the value of the sixth syntax element.
  • the decoding unit 3703 is further configured to decode the code stream and determine a minimum size threshold.
  • the decoding unit 3703 is further configured to decode the code stream and determine the value of the third syntax element; the second determination unit 3701 is further configured to decode the code stream and determine the value of the seventh syntax element when the third syntax element indicates that the current sequence allows the use of the multi-transform core group selection mode; and determine the minimum size threshold based on the value of the seventh syntax element.
  • a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system.
  • the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
  • the aforementioned integrated units can be implemented in the form of hardware or software functional modules.
  • FIG38 is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of the present application.
  • the decoder 370 may include: a second communication interface 3801, a second memory 3802, and a second processor 3803; each component is coupled together through a second bus system 3804.
  • the second bus system 3804 is used to achieve connection and communication between these components.
  • the second bus system 3804 also includes a power bus, a control bus, and a status signal bus.
  • various buses are labeled as the second bus system 3804 in FIG38. Among them,
  • the second communication interface 3801 is used to receive and send signals when sending and receiving information with other external network elements
  • the second memory 3802 is used to store computer programs that can be run on the second processor 3803;
  • the second processor 3803 is configured to, when running the computer program, execute:
  • the second processor 3803 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.
  • the hardware functions of the second memory 3802 are similar to those of the first memory 3602, and the hardware functions of the second processor 3803 are similar to those of the first processor 3603; they will not be described in detail here.
  • This embodiment provides a decoder in which, for a current block predicted using certain intra-frame prediction modes, when determining the transform kernel of the current block, multiple candidate texture features derived from a multi-transform kernel group selection technique can be used to guide the transform, thereby improving the accuracy of the transform prediction, thereby improving compression efficiency, and further improving encoding and decoding performance.
  • FIG39 is a schematic diagram of the structure of a coding and decoding system provided by an embodiment of the present application.
  • the coding and decoding system 390 may include an encoder 3901 and a decoder 3902.
  • the encoder 3901 may be the encoder described in any one of the aforementioned embodiments
  • the decoder 3902 may be the decoder described in any one of the aforementioned embodiments.
  • the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor (eg, the first processor or the second processor), implements the method as described in any of the aforementioned embodiments.
  • a processor eg, the first processor or the second processor
  • embodiments of the present application further provide a computer program product, including a computer program or instructions.
  • a processor e.g., a first processor or a second processor
  • the method described in any one of the aforementioned embodiments is implemented.
  • the embodiments of the present application further provide a computer program, which, when executed by a processor (eg, a first processor or a second processor), implements the method as described in any one of the aforementioned embodiments.
  • a processor eg, a first processor or a second processor
  • the disclosed devices and methods can be implemented in other ways.
  • the device embodiments described above are merely schematic.
  • the division of the units is merely a logical function division.
  • Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
  • the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
  • each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
  • the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
  • the computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
  • the aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.
  • a transform core group for the current block is determined; based on the transform core group, a transform core for the current block is determined; a residual block for the current block is determined, and the residual block for the current block is transformed based on the transform core to determine the transform coefficients for the current block; the transform coefficients for the current block are encoded, and the resulting coded bits are written into the bitstream.
  • a transform core group index for the current block is determined, and based on the transform core group index for the current block, a transform core group for the current block is determined; based on the transform core group, a transform core for the current block is determined; the transform coefficients for the current block are determined, and the transform coefficients for the current block are transformed based on the transform core to determine the residual block for the current block.
  • both the encoding end and the decoding end determine the transform core group for the current block based on the multi-transform core group selection technique, and then determine the transform core for the current block from there.
  • multiple candidate texture features derived from the multi-transform core group selection technique can be used to guide the transform, thereby improving the accuracy of the transform prediction, thereby improving compression efficiency, and further enhancing encoding and decoding performance.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Discrete Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

本申请公开了一种编解码方法、码流、编码器、解码器以及存储介质,该方法包括:确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。这样,可以提高压缩效率,进而提升编解码性能。

Description

编解码方法、码流、编码器、解码器以及存储介质 技术领域
本申请涉及视频编解码技术领域,尤其涉及一种编解码方法、码流、编码器、解码器以及存储介质。
背景技术
随着人们对视频显示质量要求的提高,高清和超高清等高分辨率视频应运而生。然而,高分辨率视频通常具有更多信息,因此需要更多带宽。为降低带宽要求,已经引入了涉及视频压缩的视频编码标准。
在视频编码标准中,每种帧内预测模式对应一种变换核组,故一个块可以推导出一个用于确定变换核组的纹理特征索引。但是对于比较复杂的帧内预测模式来说,这些帧内预测模式都可以对两种或以上的帧内预测模式的预测值进行加权计算。对于这种需要两种或以上的帧内预测模式的预测值进行加权的块,目前的变换过程考虑不全面,导致压缩效率低。
发明内容
本申请提供一种编解码方法、码流、编码器、解码器以及存储介质,可以提高压缩效率,进而提升编解码性能。
本申请的技术方案可以如下实现:
第一方面,本申请实施例提供了一种解码方法,应用于解码器,该方法包括:
确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;
根据变换核组,确定当前块的变换核;
确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
第二方面,本申请实施例提供了一种编码方法,应用于编码器,该方法包括:
确定当前块的变换核组;
根据变换核组,确定当前块的变换核;
确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;
对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
第三方面,本申请实施例提供了一种码流,该码流是根据待编码信息进行比特编码生成的;其中,待编码信息包括下述至少一项:当前块的量化系数、所述当前块的变换核索引、所述当前块的变换核组索引、最小样本阈值、最小尺寸阈值、第一语法元素的取值、第二语法元素的取值、第三语法元素的取值、第四语法元素的取值、第五语法元素的取值、第六语法元素的取值和第七语法元素的取值;
其中,第一语法元素用于指示当前块的变换核组索引,第二语法元素用于指示当前块是否使用第一变换模式以及对应使用的变换核索引,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,第四语法元素用于指示当前图像是否允许使用多变换核组选择技术,第五语法元素用于指示当前片是否允许使用多变换核组选择技术,第六语法元素用于指示最小样本阈值,第七语法元素用于指示最小尺寸阈值。
第四方面,本申请实施例提供了一种编码器,该编码器包括第一确定单元、第一变换单元和编码单元,其中:
第一确定单元,配置为确定当前块的变换核组;以及还配置为根据变换核组,确定当前块的变换核;
第一变换单元,配置为确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;
编码单元,配置为对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
第五方面,本申请实施例提供了一种编码器,该编码器包括第一存储器和第一处理器;其中,
第一存储器,用于存储能够在第一处理器上运行的计算机程序;
第一处理器,用于在运行所述计算机程序时,执行如第一方面所述的方法。
第六方面,本申请实施例提供了一种解码器,该解码器包括第二确定单元和第二变换单元,其中:
第二确定单元,配置为确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;以及还配置为根据变换核组,确定当前块的变换核;
第二变换单元,配置为确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
第七方面,本申请实施例提供了一种解码器,该解码器包括第二存储器和第二处理器;其中,
第二存储器,用于存储能够在第二处理器上运行的计算机程序;
第二处理器,用于在运行所述计算机程序时,执行如第二方面所述的方法。
第八方面,本申请实施例提供了一种计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如第一方面所述的方法、或者实现如第二方面所述的方法。
第九方面,本申请实施例提供了一种计算机程序产品,包括计算机程序或指令,该计算机程序或指令被处理器执行时实现如第一方面所述的方法、或者实现如第二方面所述的方法。
本申请实施例提供了一种编解码方法、码流、编码器、解码器以及存储介质,在编码端,确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。在解码端,确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。这样,无论是编码端还是解码端,都是首先根据多变换核组选择技术确定当前块的变换核组,然后从中确定当前块的变换核。也就是说,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。
附图说明
图1为一种混合编码框架的流程框图示意图;
图2为一种当前块的模板匹配示意图;
图3为一种当前块的参考样本示意图;
图4为一种当前块的多参考行示意图;
图5为一种帧内预测对应的多种预测模式示意图一;
图6为一种帧内预测对应的多种预测模式示意图二;
图7为一种帧内预测对应的多种预测模式示意图三;
图8为一种帧内预测对应的多种预测模式示意图四;
图9为一种屏幕内容的编码示意图;
图10为一种MIP模式的预测流程示意图;
图11为一种当前块的模板与模板参考区域示意图;
图12为一种梯度与帧内预测模式的柱状图示意图;
图13为一种三种帧内预测模式的加权融合示意图;
图14为ITMP模式的搜索范围示意图;
图15为一种GPM模式下的多种模式权重示意图;
图16A为一种典型的滤波器示意图一;
图16B为一种典型的滤波器示意图二;
图16C为一种典型的滤波器示意图三;
图17A为一种用于训练滤波器的重建样本区域示意图一;
图17B为一种用于训练滤波器的重建样本区域示意图二;
图17C为一种用于训练滤波器的重建样本区域示意图三;
图18为一种DCT变换示意图;
图19为一种DCT变换的基图像示意图;
图20为一种无LFNST变换的流程示意图;
图21为一种有LFNST变换的流程示意图;
图22为一种有LFNST变换的详细流程示意图;
图23为一种多个变换核组的基图像示意图;
图24为一种NSPT变换的基图像示意图;
图25为本申请实施例提供的一种视频编解码的网络架构示意图;
图26为本申请实施例提供的一种编码器的系统组成框图示意图;
图27为本申请实施例提供的一种解码器的系统组成框图示意图;
图28为本申请实施例提供的一种解码方法的流程示意图一;
图29为本申请实施例提供的一种解码方法的流程示意图二;
图30为本申请实施例提供的一种解码方法的流程示意图三;
图31为本申请实施例提供的一种解码方法的流程示意图四;
图32为本申请实施例提供的一种解码方法的流程示意图五;
图33为本申请实施例提供的一种编码方法的流程示意图一;
图34为本申请实施例提供的一种编码方法的流程示意图二;
图35为本申请实施例提供的一种编码器的组成结构示意图;
图36为本申请实施例提供的一种编码器的具体硬件结构示意图;
图37为本申请实施例提供的一种解码器的组成结构示意图;
图38为本申请实施例提供的一种解码器的具体硬件结构示意图;
图39为本申请实施例提供的一种编解码系统的组成结构示意图。
具体实施方式
为了能够更加详尽地了解本申请实施例的特点与技术内容,下面结合附图对本申请实施例的实现进行详细阐述,所附附图仅供参考说明之用,并非用来限定本申请实施例。
除非另有定义,本文所使用的所有的技术和科学术语与属于本申请的技术领域的技术人员通常理解的含义相同。本文中所使用的术语只是为了描述本申请实施例的目的,不是旨在限制本申请。
在以下的描述中,涉及到“一些实施例”,其描述了所有可能实施例的子集,但是可以理解,“一些实施例”可以是所有可能实施例的相同子集或不同子集,并且可以在不冲突的情况下相互结合。
还需要指出,本申请实施例所涉及的术语“第一\第二\第三”仅是用于区别类似的对象,不代表针对对象的特定排序,可以理解地,“第一\第二\第三”在允许的情况下可以互换特定的顺序或先后次序,以使这里描述的本申请实施例能够以除了在这里图示或描述的以外的顺序实施。
在视频图像中,一般采用第一颜色分量、第二颜色分量和第三颜色分量来表征编码块(Coding Block,CB)。其中,这三个颜色分量分别为一个亮度分量、一个蓝色色度分量和一个红色色度分量,具体地,亮度分量通常使用符号Y表示,蓝色色度分量通常使用符号Cb或者U表示,红色色度分量通常使用符号Cr或者V表示;这样,视频图像可以用YCbCr格式表示,也可以用YUV格式表示。
对本申请实施例进行进一步详细说明之前,先对本申请实施例中涉及的名词和术语进行说明,本申请实施例中涉及的名词和术语适用于如下的解释:
H.265/高效视频编码(High Efficiency Video Coding,HEVC);
H.266/多功能视频编码(Versatile Video Coding,VVC);
VVC的参考软件测试平台(VVC Test Model,VTM);
VVC之后提升压缩性能的平台(Enhanced Compression Model,ECM);
联合视频专家组(Joint Video Experts Team,JVET);
编码单元(Coding Unit,CU);
编码树单元(Coding Tree Unit,CTU);
最大编码单元(Largest Coding Unit,LCU);
运动矢量(Motion Vector,MV);
预测单元(Prediction Unit,PU);
变换单元(Transform Unit,TU);
融合技术(Merge);
跳过技术(Skip);
量化参数(Quantization Parameter,QP);
使用运动矢量差的融合技术(Merge with Motion Vector Difference,MMVD);
运动矢量预测(Motion Vector Prediction,MVP);
时域运动矢量预测(Temporal Motion Vector Prediction,TMVP);
离散余弦变换(Discrete Cosine Transform,DCT);
离散正弦变换(Discrete Sine Transform,DST);
多变换选择(Multiple Transform Selection,MTS);
低频不可分离的变换(Low Frequency Non-Separable Transform,LFNST);
不可分离的基础变换(Non-Separable Primary Transform,NSPT);
基于上下文的自适应二进制算术编码(Context-based Adaptive Binary Arithmetic Coding,CABAC)。
目前,通用的视频编解码标准都采用基于块的混合编码框架。视频中的每一个图像或子图像或一帧(frame)被分割成相同大小(如256×256,128×128,64×64等)的正方形的最大编码单元或编码树单元。每个最大编码单元或编码树单元可根据规则划分成矩形的编码单元。编码单元可能还会划分为预测单元、变换单元等。具体地,如图1所示,混合编码框架包括有预测(Prediction)、变换(Transform)、量化(Quantization)、熵编码(Entropy Coding)、反量化(Inv.quantization)、反变换(Inv.transform)、环路滤波(In Loop Filter)等模块。其中,预测模块可以包括帧内预测(Intra Prediction)和帧间预测(Inter Prediction),帧间预测可以包括运动估计(Motion Estimation)和运动补偿(Motion Compensation)。由于视频的一个图像中的相邻样本之间存在很强的相关性,在视频编解码技术中使用帧内预测的方法消除相邻样本之间的空间冗余。另外,由于视频中的相邻图像之间存在着很强的相似性,在视频编解码技术中使用图像间预测方法消除相邻图像之间的时间冗余,从而提高编码效率。
样本也可以称作像素,样本有位置信息和值。
视频编解码器的基本流程如下:在编码端,将一个图像划分成块,对当前块使用帧内预测或帧间预测产生当前块的预测块,当前块的初始块减去预测块得到残差块,对残差块进行变换、量化得到量化系数矩阵,对量化系数矩阵进行熵编码输出到码流中。在解码端,对当前块使用帧内预测或帧间预测产生当前块的预测块,另一方面解析码流得到量化系数矩阵,对量化系数矩阵进行反量化与反变换得到残差块,将预测块和残差块相加得到重建块。重建块组成重建图像,基于图像或基于块对重建图像进行环路滤波得到解码图像。编码端同样需要和解码端类似的操作获得解码图像。解码图像可以为后续的图像作为帧间预测的参考图像。编码端确定的块划分信息,预测、变换、量化、熵编码、环路滤波等模式信息或者参数信息如果有必要需要在输出到码流中。解码端通过解析码流及根据已有信息进行分析确定与编码端相同的块划分信息,预测、变换、量化、熵编码、环路滤波等模式信息或者参数信息,从而保证编码端获得的解码图像和解码端获得的解码图像相同。编码端获得的解码图像通常也叫做重建图像。在预测时可以将当前块划分成预测单元,在变换时可以将当前块划分成变换单元,预测单元和变换单元的划分可以不同。上述是基于块的混合编码框架下的视频编解码器的基本流程,随着技术的发展,该框架或流程的一些模块或步骤可能会被优化,本申请实施例适用于该基于块的混合编码框架下的视频编解码器的基本流程,但不限于该框架及流程。
另外,在本申请实施例中,当前块(Current Block,CB)可以是当前编码单元、当前预测单元、或者当前变换单元等。由于并行处理的需要,图像可以被划分成片(slice)等,同一个图像中的slice可以并行处理,也就是说它们之间没有数据依赖。而“帧”是一种常用的说法,一般可以理解为一帧是一个图像。在本申请实施例中所述的帧也可以替换为图像或slice等。
下面分别针对预测技术的相关方案进行详细介绍。
(一)帧间预测。
模板匹配(Template Matching,TM)的方法最早用在帧间预测中,它利用相邻样本之间的相关性,把当前块周边的一些区域作为模板。在当前块进行编解码时,按照编码顺序其左侧及上侧已经编解码完成。当然在现有的硬件解码器实现时,不一定能保证当前块开始解码时,其左侧和上侧已经解码完成,当然这里说的是帧间块,比如在HEVC中帧间编码的块产生预测块时是不需要周边的重建样本的,因而帧间块的预测过程可以并行进行。但是帧内编码的块是一定需要左侧和上侧的重建样本作为参考样本的。理论上左侧和上侧是可得的,也就是说硬件设计做相应的调整是可以实现的。相对来说右侧和下侧在现在标准如VVC的编码顺序下是不可得的。
如图2所示,把当前块的左侧和上侧的矩形区域设为模板,左侧的模板部分的高度一般和当前块的高度相同,上侧的模板的部分的宽度一般和当前块的宽度相同,当然也可以不同。在参考图像中寻找模板的最佳匹配位置从而确定当前块的运动信息或者说运动矢量。这个过程大致可以描述为,在某一个参考图像(Ref0)中,从一个起始位置开始,在周边一定范围内进行搜索。可以预先设定好搜索的规则,如搜索范围搜索步长等。每移动一个到位置,计算该位置对应的模板和当前块周边的模板的匹配程度,所谓匹配程度可以用一些失真代价来衡量,比如说绝对差之和(Sum of Absolute Difference,SAD)、绝对变换差之和(Sum of Absolute Transformed Difference,SATD)、均方误差(Mean-Square Error,MSE)等,一般SATD使用的变换是哈达玛(Hadamard)变换,另外SAD、SATD、MSE等的值越小代表匹配程度越高。用该位置对应的模板的预测块和当前块周边的模板的重建块计算代价。除了整样本位置的搜索还可以进行分样本位置的搜索,根据搜索到的匹配程度最高的位置来确定当前块的运动信息。利用相邻样本之间的相关性,对模板合适的运动信息可能也是当前块合适的运动信息。当然模板匹配的方法可能并不一定对所有的块都适用,因而可以使用一些方法确定当前块是否使用上述模板匹配的方法,比如在当前块用一个控制开关表示是否使用模板匹配的方法。这种模板匹配的方法的一个名字叫解码端运 动矢量推导(Decoder side Motion Vector Derivation,DMVD)。编码器和解码器都可以利用模板进行搜索从而导出运动信息或者在原有的运动信息的基础上找到更好的运动信息。而它不需要传输具体的运动矢量或运动矢量差,而是由编码器和解码器都进行同样规则的搜索从而保证编码和解码的一致。模板匹配的方法可以提高压缩性能,但是它需要在解码端也进行“搜索”,从而带来了一定的解码端复杂度。
(二)帧内预测。
可以理解,图像内部相邻的部分或相邻样本之间存在着很强的空间相关性,帧内预测就是利用当前块周边已编解码的样本和当前块内部的样本的空间相关性的预测方法。示例性地,如图3所示,4×4的白色填充样本是当前块,当前块左侧一列和上侧一行的网格填充样本为当前块的参考样本,帧内预测使用这些参考样本对当前块进行预测。这些参考样本可能已经全部可得,即全部已经编解码。也可能有部分不可得,比如当前块是整帧的最左边,那么当前块的左侧的参考样本不可得。或者编解码当前块时,当前块左下方的部分还没有编解码,那么左下方的参考样本也不可得。对于参考样本不可得的情况,可以使用可得的参考样本或某些值或某些方法进行填充,或者不进行填充。
多参考行(Multiple reference line,MRL)的帧内预测方法可以使用更多的参考样本从而提高编码效率。如图4所示,这里为使用四个参考行/列的示意图。
帧内预测有多种预测模式,如图5所示,这里是H.264中对4×4的块进行帧内预测的9种模式。其中,模式0(垂直模式)是将当前块上面的样本的值按竖直方向复制到当前块作为预测值,模式1(水平模式)是将左侧的参考样本的值按水平方向复制到当前块作为预测值,模式2(DC模式)是将A~D和I~L这8个点的平均值作为所有点的预测值,模式3~8分别按某一个角度将参考样本的值复制到当前块的对应位置,因为当前块某些位置不能正好对应到参考样本,可能需要使用参考样本的加权平均值,或者说是插值的参考样本的分样本。
除此之外,还有PLANE,PLANAR等模式,而随着技术的发展以及块的扩大,角度预测模式也越来越多。如HEVC使用的帧内预测模式有PLANAR、DC和33种角度模式共35种预测模式,详见图6。VVC使用的帧内模式有PLANAR、DC和65种角度模式共67种预测模式,详见图7。当然,除了上述67种模式,VVC对一些长和宽差距较大的长方形的块还提供了宽角度模式,如图8中的虚线所指的模式即-14~-1和67~80两个区间,它们会替换掉一些常规的模式,详见图8。
(三)帧内块复制(Intra Block Copy,IBC)。
IBC能够明显提升屏幕内容编码(Screen Content Coding,SCC)的压缩效率,因而IBC从HEVC到VVC都被用于屏幕内容编码。屏幕内容区别于相机采集的内容(camera captured content),它是由计算机生成的,屏幕内容没有噪声,包含文字、计算机图形等,边界清晰。屏幕内容中存在大量重复的内容,如图9所示。
在本申请实施例中,可以认为IBC是把帧间预测的方法用到了帧内预测。其中,帧间预测把参考图像上的参考块用来生成当前块的预测块,参考图像不是当前图像。而IBC则是从当前图像中已编解码的部分或者叫已重建部分找参考块用来生成当前块的预测块。IBC也可称为帧内块补偿(intra picture block compensation)或当前图像参考(Current Picture Referencing,CPR)。
IBC可以用块矢量(block vector,BV)来表示当前块和参考块之间的位置差别,这和帧间预测的MV是类似的。编码器在搜索范围内通过块匹配的方法确定当前块的最佳的匹配块,并对BV进行编码,对BV的编码有多种方法,比如可以用merge模式,和帧间预测有类似之处,这里不再赘述。
IBC可以认为是一种帧内预测方法,也可以认为是独立于帧内预测和帧间预测的另一类预测方法。IBC对屏幕内容编码有很高的效率,在相机采集的自然序列中同样能提升压缩效率。
(四)矩阵加权帧内预测(Matrix-based Intra Prediction,MIP)。
对于MIP来说,这是一种特殊的帧内预测模式,某些地方也可以称为Matrix weighted Intra Prediction。
如图10所示,为了对一个宽度为W、高度为H的块进行预测,MIP需要当前块左侧一列的H个重建样本和当前块上侧一行的W个重建样本作为输入。MIP按如下三个步骤生成预测块:(a)参考样本平均(Averaging)、(b)矩阵乘法(Matrix Vector Multiplication)和(c)插值(Interpolation)。这里认为MIP的核心是矩阵乘法。它可以认为是用一种矩阵乘法的方式用输入样本(参考样本)生成预测块的过程。MIP提供了多种矩阵,预测方式的不同就体现在矩阵的不同上,相同的输入样本使用不同的矩阵会得到不同的结果。而参考样本平均和插值的过程是一种性能和复杂度折中的设计。对于尺寸较大的块,可以通过参考样本平均来实现一种近似于降采样的效果,使输入能适配到比较小的矩阵,而插值则实现一种上采样的效果。这样就不需要对每一种尺寸的块都提供MIP的矩阵,而是只提供一种或几种特定的尺寸的矩阵即可。随着对压缩性能的需求的提高,以及硬件能力的提高,下一代的标准中也许会出现复杂度更高的MIP。
MIP有些类似于PLANAR,但显然MIP比PLANAR更复杂,灵活性也更强。
(五)基于模板的帧内模式推导(Template-based Intra Mode Derivation,TIMD)。
如图11所示,对当前块,把它左侧和上侧的一个区域作为模板。除了边界情况,在编解码当前块时,当前块的左侧和上侧理论上是可以得到重建值的。这也是众多模板适配方法的基础。TIMD把如图11所示的斜线填充区域作为模板,而图11中模板参考区域(Reference of the template)就是模板的参考样本(网格填充区域)。解码器可以使用某一个帧内预测模式在模板上进行预测,并且将预测值和重建值进行比较,得到该帧内预测模式在模板上的代价。比如说SAD、SATD、SSE等。由于模板和当前块是相邻的,它们有相关性,所以这里可以用一个预测模式在模板上的表现来估计它在当前块上的表现。TIMD将一些候选的帧内预测模式在模板上进行预测,得到它们在模板上的代价,选取代价最低的一个或2个帧内预测模式作为当前块的帧内预测模式。
研究发现如果2个帧内预测模式在模板上的代价差距不大,将2个帧内预测模式的预测值进行加权平均可以得到压缩性能的提升。2个预测模式的预测值的权重跟上述的代价有关,现在的版本中这个权重跟代价成反比。
总结来说,TIMD利用帧内预测模式在模板上的预测效果来筛选帧内预测模式,而且可以将2个帧内预测模式根据模板上的代价进行加权。TIMD的好处在于如果当前块选择了TIMD模式,那么它不需要再去指示具体使用了哪种帧内预测模式,而是由解码器自己通过上述流程导出,一定程度上节省开销。
(六)解码端帧内预测模式导出(Decoder-side Intra Mode Derivation,DIMD)。
DIMD利用当前块左侧和上侧的重建样本导出预测模式,但是它不是在模板上进行预测,而是分析重建样本的梯度。
如图12所示,DIMD分析黑色点的梯度,如水平梯度和竖直梯度,根据它的梯度适配一种帧内预测模式,对所有需要检查的点分析可以得到一个类似于下面的柱状图的结果。即每种帧内预测模式匹配的点的数量的统计。当然所谓柱状图只是帮助理解,具体实现时可以用多种简单的形式实现。现在的DIMD选出柱状图里最高的2个帧内预测模式,再加上PLANAR模式,共3个帧内预测模式的预测值进行加权,权重和分析的结果有关。示例性地,如图13所示,3个帧内预测模式包括M1模式、M2模式和PLANAR模式。对于这3个帧内预测模式所得到的预测值分别设置为Pred1、Pred2、Pred3,这3个帧内预测模式的权重值分别设置为w1、w2、w3,具体计算公式如下:


最终的预测块,可以如下述所示:
总结来说,DIMD利用重建样本的梯度分析来筛选帧内预测模式,而且可以将2个帧内预测模式再加上planar根据分析结果进行加权。DIMD的好处在于如果当前块选择了DIMD模式,那么它不需要再去指示具体使用了哪种帧内预测模式,而是由解码器自己通过上述流程导出,一定程度上节省了开销。
(七)模板匹配(Template Matching,TM)。
模板匹配的方法最早用在帧间预测中,它利用相邻样本之间的相关性,把当前块周边的一些区域作为模板。在当前块进行编解码时,按照编码顺序其左侧及上侧已经编解码完成。当然在现有的硬件解码器实现时,不一定能保证当前块开始解码时,其左侧和上侧已经解码完成,当然这里说的是帧间块,比如在HEVC中帧间编码的块产生预测块时是不需要周边的重建样本的,因而帧间块的预测过程可以并行进行。但是帧内编码的块是一定需要左侧和上侧的重建样本作为参考样本的。理论上左侧和上侧是可得的,也就是说硬件设计做相应的调整是可以实现的。相对来说右侧和下侧在已有标准如VVC的编码顺序下是不可得的。
如前述图2所示,把当前块的左侧和上侧的矩形区域设为模板,左侧的模板部分的高度一般和当前块的高度相同,上侧的模板的部分的宽度一般和当前块的宽度相同,当然也可以不同。在参考图像中寻找模板的最佳匹配位置从而确定当前块的运动信息或者说运动矢量。这个过程大致可以描述为,在某一个参考图像(Ref0)中,从一个起始位置开始,在周边一定范围内进行搜索。可以预先设定好搜索的规则,如搜索范围搜索步长等。每移动一个到位置,计算该位置对应的模板和当前块周边的模板的匹配程度,所谓匹配程度可以用一些失真代价来衡量,比如说SAD,SATD,一般SATD使用的变换是Hadamard变换,MSE等,SAD,SATD,MSE等的值越小代表匹配程度越高。用该位置对应的模板的预测块和当前块周边的模板的重建块计算代价。除了整样本位置的搜索还可以进行分样本位置的搜索,根据搜索到的匹配程度最高的位置来确定当前块的运动信息。利用相邻样本之间的相关性,对模板合适的运动信息可能也是当前块合适的运动信息。当然模板匹配的方法可能并不一定对所有的块都适用,因而可以使 用一些方法确定当前块是否使用上述模板匹配的方法,比如在当前块用一个控制开关表示是否使用模板匹配的方法。一个经典的模板匹配的技术叫DMVD(Decoder side Motion Vector Derivation)。编码器和解码器都可以利用模板进行搜索从而导出运动信息或者在原有的运动信息的基础上找到更好的运动信息。而它不需要传输具体的运动矢量或运动矢量差,而是由编码器和解码器都进行同样规则的搜索从而保证编码和解码的一致。模板匹配的方法可以提高压缩性能,但是它需要在解码器中也进行“搜索”,从而带来了一定的解码器复杂度。
(八)帧内模板匹配预测(Intra Template Matching Prediction,ITMP)。
ITMP可以认为是在IBC和TM结合的一种技术。上面已经提到TM应用到帧间可以减少编码MV的开销,类似地,TM用在IBC上可以减少编码BV的开销。一个示例是不需要编码BV,直接把TM找到的匹配块作为当前块的ITMP模式的预测块。
一个ITMP的示例如图14所示,当前块左上角的倒L形区域作为模板,在点填充的搜索范围内进行搜索,搜索范围为已重建区域,图14所示点填充区域包括R1当前CTU,R2左上侧的CTU,R3上侧的CTU,R4为左侧CTU。这是一个示例,实际应用时搜索范围不一样。该示例在R2中找到最佳匹配块。
(九)空域几何划分模式(Spatial Geometric Partitioning Mode,SGPM)。
在VVC视频编解码标准中,有一个叫做几何划分模式(Geometric Partitioning Mode,GPM)的帧间预测模式。在AVS3视频编解码标准中,有一个叫做角度加权预测(Angular Weighted Prediction,AWP)的帧间预测模式。这两种模式虽然名称不同、具体的实现形式不同、但原理上有共通之处。
传统的单向预测只找一个与当前块大小相同的参考块,传统的双向预测使用两个与当前块大小相同的参考块,且预测块每个点的样本值为两个参考块对应位置的平均值,即每一个参考块的所有点都占50%的比例。双向加权预测使得两个参考块的比例可以不同,如第一个参考块中所有点都占75%的比例,第二个参考块中所有点都占25%的比例。但同一个参考块中的所有点的比例都相同。其他一些优化方法如解码端运动矢量细化(Decoder side Motion Vector Refinement,DMVR)、双向光流(Bi-directional Optical Flow,BIO)会使参考样本或预测样本产生一些变化,但与上面所说的原理无关。BIO也可以简写为BDOF。而GPM或AWP也使用两个与当前块大小相同的参考块,但某些样本位置100%使用第一个参考块对应位置的样本值,某些样本位置100%使用第二个参考块对应位置的样本值,而在交界区域或者称过渡区域,按一定比例使用这两个参考块对应位置的样本值。交界区域的权重也是逐渐过渡的。具体这些权重如何分配,由GPM或AWP的模式决定。根据GPM或AWP的模式确定每个样本位置的权重。当然在某些情况下,比如说块尺寸很小的情况,可能某些GPM或AWP的模式下不能保证一定有某些样本位置100%使用第一个参考块对应位置的样本值,某些样本位置100%使用第二个参考块对应位置的样本值。也可以认为GPM或AWP使用两个与当前块大小不相同的参考块,即各取所需的一部分作为参考块。即将权重不为0的部分作为参考块,而将权重为0的部分剔除出来。
如图15所示,是VVC中的GPM在正方形的块上的64种模式的权重图。黑色表示第一个参考块对应位置的权重值为0%,白色表示第一个参考块对应位置的权重值为100%,灰色区域则按颜色深浅的不同表示第一个参考块对应位置的权重值为大于0%且小于100%的某一个权重值。第二个参考块对应位置的权重值则为100%减去第一个参考块对应位置的权重值。
GPM和AWP的权重导出方法不同。GPM根据每种模式确定角度及偏移量,而后计算出每个模式的权重矩阵。AWP首先做出一维的权重的线,然后使用类似于帧内角度预测的方法将一维的权重的线铺满整个矩阵。
需要说明的是,早先的编解码标准中只存在矩形的划分方式,无论是CU、PU还是TU的划分。而GPM和AWP在没有划分的情况下实现了预测的非矩形的划分效果。GPM和AWP使用了2个参考块的权重的蒙版(mask),即上述的权重图或者叫权重矩阵。这个蒙版确定了两个参考块在产生预测块时的权重,或者可以简单地理解为预测块的一部分位置来自于第一个参考块一部分位置来自于第二个参考块,而过渡区域(blending area)用两个参考块的对应位置加权得到,从而使过渡更平滑。GPM和AWP没有按划分线把当前块划分成两个CU或PU,于是在预测之后的残差的变换、量化、反变换、反量化等也都是将当前块作为一个整体来处理。
还需要说明的是,GPM在VVC中是一种帧间技术,但是它同样也可以使用帧内预测。GPM的两个预测模式可以都是帧间预测模式,可以一个是帧间预测模式一个是帧内预测模式,也可以两个都是帧内预测模式。
ECM中的SGPM模式将这种权重的蒙版或者说权重矩阵用在了帧内预测中,它用权重矩阵将2个帧内预测模式的预测值组合成SGPM的预测值,这样相比于单一的预测模式可以产生出更复杂的纹理。
因为要用到1个“划分”模式和2个帧内预测模式,一般的逻辑是需要像表1这样在码流中分别写 入指示这3个模式的语法元素,即如表1所示的partition_mode_idx、intra_pred_mode0_idx、intra_pred_mode1_idx。但是为了减少这些信息的开销,SGPM利用当前块周边的模板对3个模式的组合进行排序,得到一个组合模式的候选列表,而在码流中只需要写入SGPM的候选索引即可,在解码端SGPM可以构建出组合模式的候选列表,根据SGPM的候选索引,如表2中的sgpm_cand_idx导出1个“划分”模式和2个帧内预测模式partition_mode_idx、intra_pred_mode0_idx、intra_pred_mode1_idx。
表1
表2
在这里,对于sgpm_cand_idx来说,
(十)基于外插滤波器的帧内预测(Extrapolation filter based Intra Prediction,EIP)。
ECM中有一个EIP模式,图16A、图16B和图16C是EIP模式的几种典型的滤波器。其中,网格填充部分的样本位置表示输入的样本(input of EIP),白色填充部分的样本位置表示输出的样本(output of EIP)。以第一个正方形的滤波器为例,通常要求右下角的那一个样本的预测值,它左侧、左上、上侧的15个样本的值是已知的,滤波器的每一个网格填充位置有预设的系数。将输入的每一个位置的样本值乘以对应位置的系数,将所有的乘法的结果累加起来再进行归一化,就可得到待预测位置的预测值。
EIP模式提供的可能是已经预先训练好的滤波器,也可能是根据当前块的周边已重建样本区域训练得到的滤波器。如图17A、图17B和图17C表示EIP模式可以选择用来训练滤波器的3种重建样本区域。假设当前块和用来训练滤波器系数的重建样本区域遵循相似的规律,所以用重建样本区域训练的滤波器可以应用到当前块。
进一步地,下面针对变换技术进行相关介绍。
编码时,现在通用的混合编码框架会先进行预测,预测利用空间或者时间上的相关性能得到一个跟当前块相同或相似的图像。对一个块来说,预测块和当前块是完全相同的情况是有可能出现的,但是很难保证一个视频中的所有块都如此,特别是对自然视频,或者说相机拍摄的视频。视频中不规则的运动,扭曲形变,遮挡,亮度等的变化,很难被完全预测。所以混合编码框架会将当前块的原始图像减去预测图像得到残差图像,或者说当前块减去预测块得到残差块。残差块通常要比原始图像简单很多,因而预测可以显著提升压缩效率。对残差块也不是直接进行编码,而是通常先进行变换。变换是把残差图像从空间域变换到频率域,去除残差图像的相关性。残差图像变换到频率域以后,由于能量大多集中在低频区域,变换后的非零系数大多集中在左上角。接下来利用量化来进一步压缩。而且由于人眼对高频不敏感,高频区域可以使用更大的量化步长。
图18为一个DCT变换的示意图。如图18所示,原始图像经过DCT变换以后只有左上角区域存在非零系数。当然这个示例是对整幅图像做了DCT变换,而在视频编解码中,图像是分割成块来处理的,因而变换也是基于块来进行的。
变换在通常的视频的压缩中非常有用,但是也并非所有的块都必须要做变换,有些情况下变换反而不如不变换压缩效果好,因而在某些标准如VVC中,编码器可以选择当前块是否使用变换。其中,DCT2型(DCT-II)是视频压缩标准中最常用的变换,其变换的基图像如图19所示。
另外,VVC中还可以使用DCT8型(DCT-VIII)和DST7型(DST-VII)。这些变换的基本公式如表3所示,这里示出了N个点输入的DCT2,DCT8和DST7的基本变换公式。
表3

由于图像都是二维的,而直接进行二维的变换运算量和内存开销都是当时硬件条件所不能接受的,因而在标准中使用的上述DCT2、DCT8、DST7变换都是拆分成水平方向和竖直方向的一维变换分成两步进行的。如先进行水平方向的变换再进行竖直方向的变换,或者先进行竖直方向的变换再进行水平方向的变换。
(一)多变换选择MTS。
VVC支持DCT2,DCT8,DST7等变换核,VVC中使用的DCT2,DCT8,DST7等变换核是水平和数值可分离的变换核,可以单独应用于水平方向或竖直方向的变换。对一个块,编码器可以选择合适的变换核并将索引传输到码流中,解码器根据索引确定反变换的变换核。水平方向和竖直方向可以选择不同的变换核,如水平方向用DCT8,竖直方向用DST7。这个技术一般称为MTS。
VVC用一个语法元素mts_idx来确定基础变换的变换核。如下表4所示,其中,trTypeHor表示水平变换的变换核,trTypeVer表示竖直变换的变换核,trTypeHor和trTypeVer的0表示DCT2型变换,1表示DST7型变换,2表示DCT8型变换。如果mts_idx不存在,推断mts_idx的值为0。
表4
(二)低频不可分离的变换LFNST。
上述变换方法对水平方向和竖直方向的纹理比较有效,但是对斜向的纹理效果就会差一些。确实水平和竖直方向的纹理是最常见的,因而上述的变换方法对提升压缩效率是非常有用的。随着对压缩效率需求的不断提高,如果斜向的纹理能够更有效地处理,可以进一步提升压缩效率。
为了更有效地处理斜向纹理的残差,VVC中使用了LFNST变换。将上述变换诸如DCT2、DCT8、DST7称为基础变换(Primary Transform)。在VVC的编码端,LFNST用于DCT2变换之后量化之前。在VVC的解码端,LFNST用于反量化之后,反DCT2变换之前。因为是在DCT2(基础变换)的基础上再做变换,LFNST是一种二次变换。图20为无LFNST(二次变换)的编解码流程示意图,图21为有LFNST(二次变换)的编解码流程示意图。当然编码端可以不通过熵解码而是直接对保存的量化系数进行反量化,因为熵编码是无损的。
图22为一种存在LFNST(二次变换)的详细编解码流程示意图。在编码端,LFNST对基础变换后的左上角的低频系数进行二次变换。基础变换通过对图像进行去相关性,把能量集中到左上角。而二次变换对基础变换的低频系数再去相关性,结果直观地来说如图22所示。在编码端,16个系数输入到4×4的LFNST,输出是8个系数;48个系数输入到8×8的LFNST,对8×8的块输出8个系数,对其他块输出是16个系数。在解码端,8个系数输入到4×4的反LFNST,输出是16个系数;对8×8的块输入8个系数,对其他块输入是16个系数,把系数输入到8×8的反LFNST,输出是48个系数。
图23为VVC中LFNST的一些基图像。在图23中只展示了每个变换核组中每个变换核最低频的2个基图像。可以看出一些明显的斜向纹理。LFNST除了有针对某些斜向纹理优化的变换核,也有针对平坦渐变纹理优化的变换核,如VVC中的LFNST的变换核组0。
LFNST只应用于帧内编码的块。角度预测按照指定的角度将参考样本的值平铺到当前块作为预测值,这意味着预测块会有明显的方向纹理,而当前块经过角度预测后的残差在统计上也会体现出明显的角度特性。因而LFNST所选用的变换核可以跟帧内预测模式进行绑定,即确定了帧内预测模式以后,LFNST只能使用帧内预测模式对应的一组(set)变换核。
具体地,VVC中的LFNST总共有4组变换核,每组可选2个变换核。表5给出了帧内预测模式和 变换核组的对应关系。注意到色度帧内预测使用的跨分量预测模式为81到83,亮度帧内预测并没有这几种模式。LFNST的变换核可以通过转置来用一个变换核组对应处理更多的角度,示例性地,13到23和45到55的模式都对应变换核组2,但是13到23明显是接近于水平的模式而45到55明显是接近于竖直的模式。
表5
VVC的LFNST共有4组变换核,根据帧内预测模式指定LFNST使用哪一组。这样做利用了帧内预测模式和LFNST的变换核之间的相关性,从而减少了选择LFNST的变换核在码流中的传输。而当前块是否会使用LFNST,以及如果使用LFNST,是使用一个组中的第一个还是第二个,是需要通过码流和一些条件来确定的。
在后续的ECM技术演进中,LFNST进一步扩展。LFNST有更多的变换核组,在ECM中是35组,变换核组索引(LFNST set index)和帧内预测模式(Intra pred.mode)的对应关系如表6。每个变换核组对对应角度的纹理更高效。在这里,每个变换核组可以选择3个变换核。
表6
(三)不可分离的基础变换NSPT。
LFNST是水平竖直不可分离的变换,因为有二次变换,所以DCT2可以被称为基础变换。这样先经过DCT2,再进行LFNST,可以说是一种性能和复杂度折衷的方案,因为直接进行不可分离的基础变换效率更高,但是有更高的复杂度,比如说计算量和变换核的存储空间都更高。
在ECM10中,一些小块可以使用NSPT,而大块仍然使用DCT2+LFNST。小块的尺寸如4×4,4×8,8×4,8×8,4×16,16×4,8×16,16×8,8×32,32×8。在ECM10中,NSPT同样根据帧内预测模式匹配变换核组,匹配方法可以参照LFNST的方法,每个变换核组有3个变换核可选。示例性地,ECM10中的一个NSPT的8×8基图像如图24所示,图24是对应帧间角度预测模式7的,可见其处理对应角度的纹理更优。需要注意的是,NSPT只应用于帧内编码的块。
综上可知,NSPT和LFNST都是处理各种角度的纹理的变换,它们可能有多个变换核,一个变换核可能是专门为某种特定的角度纹理优化的。当然除了角度的纹理NSPT和LFNST也包括处理渐变纹理的变换核。其实这些变换核也可以说是训练好的KLT(Karhunen-Loeve Transform)。也可以总结为NSPT和LFNST都有多个变换核,每个变换核为特定的纹理设计,特定的纹理包括角度纹理,渐变纹理等。当然渐变纹理可以进一步扩展如水平渐变纹理,竖直渐变纹理,斜向渐变纹理等。它们有多个变换核组,每种帧内预测模式可以对应一种变换核组。每种帧内预测模式其实也代表了一种纹理特征。所以帧内预测模式索引也是一种纹理特征索引。
在VVC以及目前的ECM中,一个块(CU或TU)只能导出一个用于推导LFNST/NSPT变换核组的纹理特征索引,从而导出唯一的LFNST/NSPT变换核组。这个纹理特征可以说是使用LFNST/NSPT的一个纽带。对使用普通帧内预测模式,即DC、PLANAR以及各种角度模式的块来说,这个纹理特征索引就是当前块使用的帧内预测模式。但是对于DIMD、TIMD、MIP、EIP、SGPM、ITMP和IBC等特殊的帧内预测模式来说,就不是那么直接了。其中,DIMD和TIMD都可以对2种或以上的帧内预测模式的预测值进行加权,SGPM是对2种帧内预测模式的预测值用权重矩阵进行加权,MIP是根据一个矩阵运算进行预测,EIP是根据一个滤波器进行预测,ITMP和IBC是基于复制一个已重建的块来进行预测,它们不是像普通帧内预测模式那样简单的纹理特征。比如用2种或以上帧内预测模式的预测值进行加权的块,就说明它的纹理中可能包含2种或以上的纹理特征,它经过预测后的残差可能表现出其中第一种预测模式的纹理特征,也可能表现出第二个预测模式的纹理特征。由此可见,对于这种需要两种 或以上的帧内预测模式的预测值进行加权的块,目前的变换过程考虑不全面,导致压缩效率低。
基于此,本申请实施例提供了一种编码方法,确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。本申请实施例还提供了一种解码方法,确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
这样,无论是编码端还是解码端,可以根据多变换核组选择技术确定当前块的变换核组,然后从中确定当前块的变换核。也就是说,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。
下面将结合附图对本申请各实施例进行详细说明。
图25为本申请实施例提供的一种视频编解码的网络架构示意图。如图25所示,该网络架构包括一个或多个电子设备13至1N和通信网络01,其中,电子设备13至1N可以通过通信网络01进行视频交互。电子设备在实施的过程中可以为各种类型的具有视频编解码功能的设备,例如,所述电子设备可以包括手机、平板电脑、个人计算机、个人数字助理、导航仪、数字电话、视频电话、电视机、传感设备、服务器等,本申请实施例不作限定。
在本申请实施例中,这里提供了一种包含解码方法和编码方法的视频编解码系统的网络架构。其中,本申请实施例中的解码器或编码器就可以为上述电子设备。也就是说,本申请实施例中的电子设备具有视频编解码功能,一般包括视频编码器(即编码器)和视频解码器(即解码器)。
图26为本申请实施例提供的一种编码器的系统组成框图示意图。如图26所示,编码器100可以包括:分割单元101、预测单元102、第一加法器107、变换单元108、量化单元109、反量化单元110、反变换单元111、第二加法器112、滤波单元113、解码图片缓存(Decoded Picture Buffer,DPB)单元114和熵编码单元115。这里,编码器100的输入可以是由一系列图片或者一张静态图片组成的视频,编码器100的输出可以是用于表示输入视频的压缩版本的比特流(也可以称为“码流”)。
其中,分割单元101将输入视频中的图片分割成一个或多个编码树单元(Coding Tree Units,CTUs)。分割单元101将图片分成多个图块(或称为瓦片,tiles),还可以进一步将一个tile分成一个或多个砖块(bricks),这里,一个tile或者一个brick中可以包括一个或多个完整的和/或部分的CTUs。另外,分割单元101可以形成一个或多个切片(slices),其中一个slice可以包括图片中按照栅格顺序排列的一个或多个tiles,或者覆盖图片中矩形区域的一个或多个tiles。分割单元101还可形成一个或多个子图片,其中,一个子图片可以包括一个或多个slices、tiles或bricks。
在编码器100的编码过程中,分割单元101将CTU传送到预测单元102。通常,预测单元102可以由块分割单元103、运动估计(Motion Estimation,ME)单元104、运动补偿(Motion Compensation,MC)单元105和帧内预测单元106组成。具体地,块分割单元103迭代地使用四叉树分割、二叉树分割和三叉树分割而进一步将输入CTU划分成更小的编码单元(Coding Units,CUs)。预测单元102可使用ME单元104和MC单元105获取CU的帧间预测块。帧内预测单元106可使用包括MIP模式的各种帧内预测模式获取CU的帧内预测块。在示例中,率失真优化的运动估计方式可被ME单元104和MC单元105调用以获取帧间预测块,以及率失真优化的模式确定方式可被帧内预测单元106调用以获取帧内预测块。预测单元102输出CU的预测块,第一加法器107计算分割单元101的输出中的CU和CU的预测块之间的差值,即残差CU。变换单元108读取残差CU并对残差CU执行一个或多个变换操作以获取系数。量化单元109对系数进行量化并输出量化系数(即levels)。反量化单元110对量化系数执行缩放操作以输出重构系数。反变换单元111执行对应于变换单元108中的变换的一个或多个反变换并输出重构残差。第二加法器112通过使重构残差和来自预测单元102的CU的预测块相加而计算出重构CU。第二加法器112还将其输出发送到预测单元102以用作帧内预测参考。在图片或子图片中的所有CU被重构之后,滤波单元113对重构图片或子图片执行环路滤波。这里,滤波单元113包含一个或多个滤波器,例如去方块滤波器、采样自适应偏移(Sample Adaptive Offset,SAO)滤波器、自适应环路滤波器(Adaptive Loop Filter,ALF)、亮度映射和色度缩放(Luma Mapping with Chroma Scaling,LMCS)滤波器以及基于神经网络的滤波器等。或者,当滤波单元113确定CU不用作其它CU编码时的参考时,滤波单元113对CU中的一个或多个目标样本执行环路滤波。滤波单元113的输出是解码图片或子图片,这些解码图片或子图片缓存至DPB单元114。DPB单元114根据时序和控制信息输出解码图片或子图片。这里,存储在DPB单元114中的图片还可用作预测单元102执行帧间预测或帧内预测的参考。最后熵编码单元115将来自编码器100中解码图片所必需的参数(比如控制参数和补充信息 等)转换成二进制形式,并根据每个数据单元的语法结构将这样的二进制形式写入码流中,即编码器100最终输出码流。
进一步地,编码器100可以是具有第一处理器和记录计算机程序的第一存储器。当第一处理器读取并运行计算机程序时,编码器100读取输入视频并生成对应的码流。另外,编码器100还可以是具有一个或多个芯片的计算设备。在芯片上实现为集成电路的这些单元具有与图26中相应单元类似的连接和数据交换功能。
图27为本申请实施例提供的一种解码器的系统组成框图示意图。如图27所示,该解码器200可以包括:解析单元201、预测单元202、反量化单元205、反变换单元206、加法器207、滤波单元208和解码图片缓存单元209。这里,解码器200的输入是用于表示视频或者一张静态图片的压缩版本的比特流,解码器200的输出可以是由一系列图片组成的解码视频或者一张解码的静态图片。
其中,解码器200的输入码流可以是编码器100所生成的码流。解析单元201对输入码流进行解析并从输入码流获取语法元素的值。解析单元201将语法元素的二进制表示转换成数字值并将数字值发送到解码器200中的单元以获取一个或多个解码图片。解析单元201还可从输入码流解析一个或多个语法元素以显示解码图片。
在解码器200的解码过程中,解析单元201将语法元素的值以及根据语法元素的值设置或确定的、用于获取一个或多个解码图片的一个或多个变量发送到解码器200中的单元。预测单元202确定当前解码块(例如CU)的预测块。这里,预测单元202可以包括运动补偿单元203和帧内预测单元204。具体地,当指示帧间解码模式用于对当前解码块进行解码时,预测单元202将来自解析单元201的相关参数传递到运动补偿单元203以获取帧间预测块;当指示帧内预测模式(包括基于MIP模式索引值指示的MIP模式)用于对当前解码块进行解码时,预测单元202将来自解析单元201的相关参数传送到帧内预测单元204以获取帧内预测块。反量化单元205具有与编码器100中的反量化单元110相同的功能。反量化单元205对来自解析单元201的量化系数(即levels)执行缩放操作以获取重构系数。反变换单元206具有与编码器100中的反变换单元111相同的功能。反变换单元206执行一个或多个变换操作(即通过编码器100中的反变换单元111执行的一个或多个变换操作的反操作)以获取重构残差。加法器207对其输入(来自预测单元202的预测块和来自反变换单元206的重构残差)执行相加操作以获取当前解码块的重构块。重构块还发送到预测单元202以用作在帧内预测模式下编码的其它块的参考。
在图片或子图片中的所有CU被重构之后,滤波单元208对重构图片或子图片执行环路滤波。滤波单元208包含一个或多个滤波器,例如去方块滤波器、采样自适应补偿滤波器、自适应环路滤波器、亮度映射和色度缩放滤波器以及基于神经网络的滤波器等。或者,当滤波单元208确定重构块不用作对其它块解码时的参考时,滤波单元208对重构块中的一个或多个目标样本执行环路滤波。这里,滤波单元208的输出是解码图片或子图片,解码图片或子图片缓存至DPB单元209。DPB单元209根据时序和控制信息输出解码图片或子图片。存储在DPB单元209中的图片还可用作通过预测单元202执行帧间预测或帧内预测的参考。
进一步地,解码器200可以是具有第二处理器和记录计算机程序的第二存储器。当第一处理器读取并运行计算机程序时,解码器200读取输入码流并生成对应的解码视频。另外,解码器200还可以是具有一个或多个芯片的计算设备。在芯片上实现为集成电路的这些单元具有与图27中相应单元类似的连接和数据交换功能。
还需要说明的是,当本申请实施例应用于编码器100时,“当前块”具体是指视频图像中的当前待编码的块(也可以简称为“编码块”);当本申请实施例应用于解码器200时,“当前块”具体是指视频图像中的当前待解码的块(也可以简称为“解码块”)。
在本申请的一实施例中,图28为本申请实施例提供的一种解码方法的流程示意图一。如图28所示,该方法可以包括:
S2801,确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组。
需要说明的是,在本申请实施例中,该方法应用于解码器。具体来说,基于图27所示解码器200的组成结构,本申请实施例的解码方法主要应用于帧内预测的块。其中,在当前块采用帧内预测模式时,这里主要是针对帧内预测模式下的NSPT和LFNST变换所提出的优化方案,以提高压缩效率。
还需要说明的是,在本申请实施例中,该方法适用于多变换组选择(Multiple Transform Set Selection,MTSS)方式。其中,NSPT和LFNST都是处理各种角度的纹理的变换,它们可能有多个变换核,一个变换核可能是专门为某种特定的角度纹理优化的。当然除了角度的纹理,NSPT和LFNST也包括处理渐变纹理的变换核。其实这些变换核也可以说是训练好的KL变换(Karhunen-Loeve Transform,KLT)。 也就是说,NSPT和LFNST都有多个变换核,每个变换核为特定的纹理设计,特定的纹理包括角度纹理,渐变纹理等。另外,渐变纹理可以进一步扩展如水平渐变纹理,竖直渐变纹理,斜向渐变纹理等。进一步地,MTSS方式不局限于只能用于NSPT和LFNST这样的不可分离变换,针对为特定纹理优化的可分离变换也可以应用MTSS方式。
还需要说明的是,在本申请实施例中,帧内预测的块可以有多个变换核组,每种帧内预测模式可以对应一种变换核组。换句话说,每种帧内预测模式其实代表了一种纹理特征。所以帧内预测模式索引也是一种纹理特征索引。如DC和PLANAR对应渐变的纹理特征,某一种角度预测模式对应这种角度的纹理特征。纹理特征索引一方面可以避免在“帧间”出现帧内预测模式,另一方面它也更利于可能的扩展,例如一种帧内预测模式可以对应多种纹理特征,比如DC模式可以对应水平渐变纹理,竖直渐变纹理,斜向渐变纹理等。
在本申请实施例中,考虑到一些帧内预测模式并非是简单的纹理特征,而是有可能包含两种或以上的纹理特征。因此,这里可以使用多变换核组选择(Multiple Transform Set Selection,MTSS)技术。在MTSS技术中,如果当前块的帧内预测模式是某种特殊的帧内预测模式,那么可选的变换核组不止一个。
在本申请实施例中,对于当前块是否使用MTSS技术,可以增加一些判断条件,示例性地,这些判断条件可以包括:当前块的样本个数、当前块的尺寸和当前块的预测模式等。也就是说,在本申请实施例中,可以选择当前块的样本个数来判断当前块是否使用MTSS技术,和/或,也可以选择当前块的尺寸来判断当前块是否使用MTSS技术,和/或,还可以选择当前块的预测模式来判断当前块是否使用MTSS技术,这里不作任何限定。
在本申请实施例中,在当前块使用MTSS技术时,这时候可以根据当前块的变换核组索引来确定当前块的变换核组。
其中,MTSS技术对解码器的复杂度增加体现在解析语法元素时的不同,一般认为对解码器的复杂度影响不大。但是对编码器而言,由于增加了变换核组的候选,这会增加编码器的复杂度。具体来说,一个块可能从2个或多个LFNST/NSPT的变换核组中选择一个。而VVC中的一个块只有一个可用的变换核组。因而一种可能的实现方式是限制MTSS技术所适用的块的尺寸,从而可以在MTSS技术需要增加很多计算而对压缩效率提升不明显的块尺寸禁用MTSS技术。简单来说,本申请实施例可以第MTSS技术所适用的块尺寸进行限制。
在一种可能的实施例中,该方法还可以包括:确定当前块的样本个数;在当前块的样本个数大于或等于最小样本阈值时,执行确定当前块的变换核组索引的步骤。
示例性地,可以设置一个应用MTSS技术的最小的样本数的阈值(即“最小样本阈值”),用MIN_PIX表示。如果一个块的样本数(宽度×高度)小于MIN_PIX,则这个块不能使用MTSS技术。否则,若这个块的样本数大于或等于MIN_PIX,则这个块可以使用MTSS技术。在本申请实施例中,MIN_PIX的取值可能是32,64,256等。示例性地,MIN_PIX的取值可以等于256。
在另一种可能的实施例中,该方法还可以包括:确定当前块的尺寸,其中,当前块的尺寸包括高度和宽度;在当前块的高度和宽度均大于或等于最小尺寸阈值时,执行确定当前块的变换核组索引的步骤。
示例性地,可以设置一个应用MTSS技术的最小的尺寸的阈值(即“最小尺寸阈值”),用MIN_SIZE表示。如果一个块的宽度或高度小于MIN_SIZE,则这个块不能使用MTSS技术。否则,若这个块的宽度和高度均大于或等于MIN_SIZE,则这个块可以使用MTSS技术。在本申请实施例中,MIN_SIZE的值可能是8,16等。示例性地,MIN_PIX的取值可以等于8。
在又一种可能的实施例中,该方法还可以包括:确定当前块的预测模式;在当前块的预测模式为第一预测模式集合中的其中一项时,执行确定当前块的变换核组索引的步骤。
在一些实施例中,第一预测模式集合可以包括以下预测模式中至少之一:DC模式、PLANAR模式和角度预测模式之外的其他预测模式。
在一些实施例中,第一预测模式集合可以包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式;复制帧内块进行预测的模式;使用外插滤波器进行预测的模式;使用矩阵运算进行预测的模式。
在一些实施例中,第一预测模式集合可以包括以下预测模式中至少之一:DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式等。
在本申请实施例中,对于将至少两种帧内预测模式的预测值进行组合预测的模式,具体可以是DIMD模式、TIMD模式和SGPM模式;其中,帧内预测模式包括但不限于DC模式、PLANAR模式和角度预测模式等。
在本申请实施例中,对于复制帧内块进行预测的模式,具体可以是ITMP模式和IBC模式。
在本申请实施例中,对于使用外插滤波器进行预测的模式,具体可以是EIP模式。
在本申请实施例中,对于使用矩阵运算进行预测的模式,具体可以是MIP模式。
也就是说,第一预测模式集合中的模式为比较复杂的预测模式,其对应的纹理中可能包含两种或以上的纹理特征。示例性地,DIMD模式和TIMD模式都可以对两种或以上的帧内预测模式的预测值进行加权,SGPM模式是对两种帧内预测模式的预测值用权重矩阵进行加权,MIP模式是根据一个矩阵运算进行预测,EIP模式是根据一个外插滤波器(Extrapolation filter)进行预测,ITMP模式和IBC模式则是基于复制一个已重建的参考块来进行预测,它们的纹理特征并不像DC模式、PLANAR模式等那样简单的纹理特征。也就是说,在当前块的预测模式为第一预测模式集合中的任意一项时,这时候当前块可以使用MTSS技术。
可以理解地,在一些实施例中,参见图29,对于步骤S2801来说,该步骤可以包括:
S2901,解码码流,确定当前块的变换核组索引。
S2902,根据当前块的变换核组索引确定当前块的变换核组。
需要说明的是,在本申请实施例中,变换核组索引用于表示当前块的变换核组在至少两个候选变换核组中的编号,该变换核组索引可以用lfnst_feature_idx表示,或者也可以lfnst_set_idx表示。其中,当前块的变换核组索引可以为大于或等于零的整数,例如0,1,2等。示例性地,若lfnst_feature_idx的取值等于0,则指示选择第一个候选变换核组作为当前块的变换核组;若lfnst_feature_idx的取值等于1,则指示选择第二个候选变换核组作为当前块的变换核组。
在一些实施例中,对于确定当前块的变换核组索引,该方法还可以包括:解码码流,确定第一语法元素的取值;根据第一语法元素的取值,确定当前块的变换核组索引。
需要说明的是,在本申请实施例中,对于当前块的变换核组索引可以是直接解码码流确定,或者也可以是通过解码第一语法元素的取值确定。
还需要说明的是,在本申请实施例中,第一语法元素可以用lfnst_feature_idx或者lfnst_set_idx表示。第一语法元素可以用于指示当前块的变换核组索引,具体是当前块的变换核组在至少两个候选变换核组中的编号。其中,第一语法元素的取值可以为大于或等于零的整数,例如0,1,2等。示例性地,若第一语法元素的取值等于0,则指示选择第一个候选变换核组作为当前块的变换核组;若第一语法元素的取值等于1,则指示选择第二个候选变换核组作为当前块的变换核组;若第一语法元素的取值等于2,则指示选择第三个候选变换核组作为当前块的变换核组。
在一种具体的实现方式中,在当前块的预测模式为第二预测模式集合中的其中一项时,根据当前块的变换核组索引确定当前块的变换核组,可以包括:若变换核组索引为第i值,则基于当前块的预测模式推导的第i个帧内预测模式确定当前块的变换核组,i为正整数。
其中,第二预测模式集合可以包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
在一种具体的实施例中,第二预测模式集合可以包括以下预测模式中至少之一:DIMD模式、TIMD模式和SGPM模式。
在本申请实施例中,MTSS可以应用于DIMD模式、TIMD模式和SGPM模式。若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则基于当前块的预测模式推导的第二个帧内预测模式确定当前块的变换核组;若变换核组索引为第三值,则基于当前块的预测模式推导的第三个帧内预测模式确定当前块的变换核组等等。
在这里,对于变换核组索引来说,第一值可以0,第二值可以为1,第三值可以为2。也就是说,针对第二预测模式集合中的预测模式,在确定出当前块的变换核组索引之后,可以根据变换核组索引直接来确定对应的变换核组。
示例性地,对于DIMD模式,由于DIMD模式本身就使用梯度推导纹理特征,而且它可以推导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据当前块的变换核组索引可以确定对应的变换核组。
示例性地,对于TIMD模式,由于TIMD模式本身可以导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据当前块的变换核组索引可以确定对应的变换核组。
示例性地,对于SGPM模式,由于SGPM模式本身可以导出一个“划分”模式和2个帧内预测模式,因而也不需要MTSS技术额外使用梯度推导纹理特征,根据当前块的变换核组索引可以确定对应的变换核组。
更具体地,在一种可能的实施例中,如下:
示例性地,对于DIMD模式,假设DIMD模式推导的第一个帧内预测模式对应的变换核组索引为0,记作candFeature0;DIMD模式推导的第二个帧内预测模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于DIMD模式推导的第一个帧内预测模式确定当前块的变换核组;若 变换核组索引为1,则基于DIMD模式推导的第二个帧内预测模式确定当前块的变换核组。
示例性地,对于TIMD模式,假设TIMD模式推导的第一个帧内预测模式对应的变换核组索引为0,记作candFeature0;TIMD模式推导的第二个帧内预测模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于TIMD模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为1,则基于TIMD模式推导的第二个帧内预测模式确定当前块的变换核组。
示例性地,对于SGPM模式,假设SGPM模式的划分模式对应的变换核组索引为0,记作candFeature0;SGPM模式推导的第一个帧内预测模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于SGPM模式的划分模式确定当前块的变换核组;若变换核组索引为1,则基于SGPM模式推导的第一个帧内预测模式确定当前块的变换核组。
在另一种具体的实现方式中,在当前块的预测模式为第三预测模式集合中的其中一项时,根据当前块的变换核组索引确定当前块的变换核组,可以包括:若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则基于PLANAR模式确定当前块的变换核组。
在又一种具体的实现方式中,在当前块的预测模式为第三预测模式集合中的其中一项时,根据当前块的变换核组索引确定当前块的变换核组,可以包括:若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则当前块的预测模式推导的第二个帧内预测模式确定当前块的变换核组。
其中,第三预测模式集合可以包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
在一种具体的实施例中,第三预测模式集合可以包括以下预测模式中至少之一:MIP模式和EIP模式。
在本申请实施例中,MTSS可以应用于MIP模式和EIP模式。由于MIP模式和EIP模式可以大量应用于纹理渐变的块,其特征和PLANAR模式是相似的。因而对MIP模式和EIP模式也可以将PLANAR模式作为其候选纹理特征。
在这里,对于变换核组索引来说,第一值可以0,第二值可以为1。也就是说,针对第三预测模式集合中的预测模式,在确定出当前块的变换核组索引之后,也可以根据变换核组索引直接来确定对应的变换核组。
示例性地,在当前块的预测模式为MIP模式或者EIP模式时,假设MIP模式或者EIP模式推导的第一个帧内预测模式对应的变换核组索引为0,记作candFeature0;PLANAR模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为1,则基于PLANAR模式确定当前块的变换核组。
还可以理解地,在本申请实施例中,对于MTSS技术的至少两个候选变换核组来说,解码端还可以构建第一候选列表。在一些实施例中,参见图30,对于步骤S2801来说,该步骤可以包括:
S3001,确定当前块的第一候选列表,第一候选列表指示至少两个候选变换核组。
S3002,根据第一候选列表和变换核组索引,确定当前块的变换核组。
需要说明的是,在本申请实施例中,在当前块的预测模式为DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式等中的任意一项时,第一候选列表可以包括至少两个候选纹理特征索引,或者第一候选列表可以包括至少两个变换核组。在这里,每一个候选纹理特征索引对应一个变换核组。因此,可以说是:第一候选列表指示至少两个候选变换核组。
在一些实施例中,确定当前块的第一候选列表,可以包括:确定基于当前块的预测模式推导的一个或多个帧内预测模式;根据一个或多个帧内预测模式确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一种可能的实施例中,确定当前块的第一候选列表,具体可以包括:在当前块的预测模式为第二预测模式集合中的其中一项时,确定基于当前块的预测模式推导的多个帧内预测模式;根据多个帧内预测模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表。
在本申请实施例中,第二预测模式集合可以包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式;示例性地,第二预测模式集合可以包括以下预测模式中至少之一:DIMD模式、TIMD模式和SGPM模式。也就是说,MTSS技术可以应用于DIMD模式、TIMD模式和SGPM模式,这时候可以不额外使用梯度推导纹理特征,仅根据该预测模式自身所推导出的一个或多个帧内预测模式来构建第一候选列表。
示例性地,对于DIMD模式,由于DIMD模式本身就使用梯度推导纹理特征,而且它可以推导出 多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据DIMD模式本身推导的多个帧内预测模式来构建第一候选列表。
示例性地,对于TIMD模式,由于TIMD模式本身可以导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据TIMD模式本身推导的多个帧内预测模式来构建第一候选列表。
示例性地,对于SGPM模式,由于SGPM模式本身可以导出一个“划分”模式和2个帧内预测模式,因而也不需要MTSS技术额外使用梯度推导纹理特征,根据SGPM模式本身推导的一个“划分”模式和2个帧内预测模式来构建第一候选列表。
在一些实施例中,确定当前块的第一候选列表,还可以包括:在第一候选列表未填满时,确定当前块的预设纹理特征索引;根据预设纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在本申请实施例中,预设纹理特征索引可以为PLANAR模式。在一种具体的实施例中,该方法还可以包括:根据多个帧内预测模式和PLANAR模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表。
示例性地,如果DIMD模式、TIMD模式、SGPM模式所推导出的帧内预测模式中包括ITMP或IBC,或EIP或MIP的模式,这时候也可以使用PLANAR模式对应的纹理特征作为候选纹理特征。由于DIDM模式、TIMD模式、SGPM模式可以使用2个以上的帧内预测模式进行加权或组合,所述帧内预测模式可能是ITMP或IBC,或EIP或MIP的模式。在这种情况下,也可以使用PLANAR模式对应的纹理特征作为候选纹理特征。简单来说,可以根据DIMD模式、TIMD模式、SGPM模式所推导出的多个帧内预测模式和PLANAR模式共同来构建第一候选列表。
在另一种可能的实施例中,确定当前块的第一候选列表,具体可以包括:在当前块的预测模式为第三预测模式集合中的其中一项时,确定基于当前块的预测模式推导的帧内预测模式;根据帧内预测模式和PLANAR模式确定两个候选变换核组,并将两个候选变换核组添加到第一候选列表。
在本申请实施例中,第三预测模式集合可以包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。示例性地,第三预测模式集合可以包括以下预测模式中至少之一:MIP模式和EIP模式。也就是说,MTSS技术可以应用于MIP模式和EIP模式,这时候也可以不额外使用梯度推导纹理特征,仅根据该预测模式自身所推导出的第一个帧内预测模式和PLANAR模式共同来构建第一候选列表。
在一些实施例中,确定当前块的第一候选列表,还可以包括:在第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;根据所述候选样本,确定所述当前块的一个或多个候选纹理特征索引;根据一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在本申请实施例中,在构建第一候选列表时,MTSS技术也可以额外使用梯度推导纹理特征。在这里,根据候选样本确定当前块的一个或多个候选变换核组时,可以是先根据候选样本确定包含一个或多个候选纹理特征索引的一个或多个候选纹理特征索引,然后根据这一个或多个候选纹理特征索引来确定当前块的一个或多个候选变换核组。
通常情况下,一个候选纹理特征索引对应一个候选变换核组;但是某些情况下,针对多个相近的候选纹理特征索引可以对应同一个变换核组。例如多个类似角度的帧内预测模式对应同一个变换核组。在本申请实施例中,如果多个相邻的帧内预测模式(或候选纹理特征索引)对应一个变换核组,那么在确定候选变换核组时,需要保证各个候选纹理特征索引不会确定为相同的变换核组。
在一种可能的实现方式中,对于候选样本来说,可以是确定当前块的预测块;将预测块中的至少部分样本作为候选样本。
在本申请实施例中,如果预测块存在某种纹理,可以认为残差块中存在相同特征的纹理。这样,关于帧间推导候选纹理特征索引所使用的候选样本,可以是预测块中的全部样本或者预测块中的部分样本。
在另一种可能的实现方式中,对于候选样本来说,可以是确定当前块的已重建区域的相邻样本;将已重建区域的相邻样本作为候选样本。
在本申请实施例中,关于帧间推导候选纹理特征索引所使用的候选样本,可以是使用当前块的已重建区域的相邻样本,例如当前块左侧和右侧的已重建区域。因为左侧和上侧的已重建区域虽然不是当前块但是和当前块相邻,比如说纹理有相连的情况,因此一定程度上可以用于估计当前块的纹理。
在又一种可能的实现方式中,考虑使用更多的样本,对于候选样本来说,可以是将已重建区域的相邻样本和预测块中的至少部分样本共同作为候选样本。
在本申请实施例中,关于帧间推导候选纹理特征索引所使用的候选样本,也可以是同时使用当前块的预测块和当前块左侧和上侧的已重建区域,这样用于推导候选纹理特征索引的样本更多,使得所推导 的候选纹理特征索引更准确。
还需要说明的是,在本申请实施例中,用于推导候选纹理特征索引的候选样本的数量可以为至少一个,例如1个、2个、3个或更多个。
在一些实施例中,该方法还可以包括:根据当前块的尺寸参数,确定候选样本的数量。
也就是说,在根据候选样本推导一个或多个候选纹理特征索引时,使用多少个候选样本可以是由当前块的尺寸参数确定。示例性地,如果当前块的尺寸小,那么可以对所有可用的样本进行统计;如果当前块的尺寸大,那么可以对当前块降采样进行统计,比如水平方向和/或竖直方向每2个、或4个、或8个样本中的一个样本进行统计。或者,如果当前块的水平或竖直某一个方向的尺寸小于或等于8,那么对该方向上的所有可用的样本进行统计;否则,如果当前块的水平或竖直某一个方向的尺寸小于或等于16,那么对该方向上的每2个样本中的一个样本进行统计;否则,对该方向上的每4个样本中的一个样本进行统计,这里不作具体限定。
在一些实施例中,根据候选样本,确定当前块的一个或多个候选纹理特征索引,可以包括:确定候选样本的水平梯度值和竖直梯度值;根据候选样本的水平梯度值和竖直梯度值,确定候选样本对应的纹理特征索引以及梯度强度值;根据候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表;根据纹理特征统计表,确定当前块的一个或多个候选纹理特征索引。
需要说明的是,在本申请实施例中,在根据候选样本的水平梯度值和竖直梯度值,确定候选样本对应的纹理特征索引以及梯度强度值时,可以包括:根据候选样本的水平梯度值和竖直梯度值进行角度映射,确定候选样本对应的纹理特征索引;以及根据候选样本的水平梯度值和竖直梯度值进行梯度强度计算,确定候选样本对应的梯度强度值。
在一种具体的实施例中,根据候选样本的水平梯度值和竖直梯度值进行角度映射,确定候选样本对应的纹理特征索引,可以包括:根据候选样本的水平梯度值和竖直梯度值,利用预设查找表确定候选样本对应的纹理特征索引。
在本申请实施例中,候选样本的水平梯度值可以用gradx表示,候选样本的竖直梯度值可以用grady表示。这样,根据gradx和grady导出纹理特征索引(或称为“虚拟的帧内预测模式”)可以通过查表来实现。
示例性地,如果abs(gradx)等于0且abs(grady)不等于0,那么有水平方向纹理,对应VVC中的帧内预测模式18。如果abs(grady)等于0且abs(gradx)不等于0,那么有竖直方向纹理,对应VVC中的帧内预测模式50。在abs(gradx)和abs(grady)都不等于0的情况下,如果abs(gradx)等于abs(grady),且gradx和grady符号相同,对应VVC中的帧内预测模式34。如果abs(gradx)等于2倍的abs(grady),且gradx和grady符号相同,对应VVC中的帧内预测模式40。另外,其他情况可以按照相同的原则查表确定。
在一种具体的实施例中,根据候选样本的水平梯度值和竖直梯度值进行梯度强度计算,确定候选样本对应的梯度强度值,可以包括:对水平梯度值的绝对值和竖直梯度值的绝对值进行加法运算,确定候选样本对应的梯度强度值。
在这里,候选样本对应的梯度强度值可以记为amp。示例性地,amp=abs(gradx)+abs(grady)。
需要说明的是,在本申请实施例中,针对候选样本的水平梯度值和竖直梯度值可以使用索贝尔算子(sobel)计算得到。示例性地,对于索贝尔算子来说,具体如下所示:
水平梯度值的算子:
竖直梯度值的算子:
如此,假设样本位置为(x,y)的样本值为Px,y,则水平梯度值gradx和竖直梯度值grady的计算如下所示:
gradx=Px+1,y-1+2*Px+1,y+Px+1,y+1-Px-1,y-1-2*Px-1,y-Px-1,y+1             (5)
grady=Px-1,y+1+2*Px,y+1+Px+1,y+1-Px-1,y-1-2*Px,y-1-Px+1,y-1             (6)
还需要说明的是,在本申请实施例中,还可以设置候选样本不包括当前块最外面上下左右一行一列的样本。其中,考虑到索贝尔算子要使用到当前样本的上下左右一行一列的样本,所以本申请实施例可以设置不计算当前块最外面上下左右一行一列的样本的梯度。
在一些实施例中,根据候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表,可以包括:在候选样本的数量为至少一个时,确定至少一个纹理特征索引以及对应的至少一个梯度强度值;根据至少一个纹理特征索引确定具有互异特性的至少一种参考纹理特征索引,以及根据至少一个梯度强度值,对归属于同一参考纹理特征索引的梯度强度值进行累加计算,确定至少一种参考纹理特征索引对应的梯度强度累加值;根据至少一种参考纹理特征索引和至少一种参考纹理特征索引对应的梯度强度累加值,确定纹理特征统计表。
也就是说,在本申请实施例中,以预测块中的至少部分样本作为候选样本为例,对预测块中的全部或部分样本计算其梯度,一般来说可以计算水平梯度值和竖直梯度值,这里可以使用索贝尔算子来计算梯度值。对某一个样本,根据其水平梯度值和竖直梯度值可以推测该样本的纹理方向,比如水平梯度值非零,竖直梯度值为零,那该样本的纹理是竖直方向。反之,水平梯度值为零,竖直梯度值非零,那该样本的纹理是水平方向。比如水平梯度值和竖直梯度值相等且不为零,那该样本的纹理是45度。当然,本申请实施例还有很多其他情况是水平梯度值和竖直梯度值都不为零,可以根据它们的比值确定该样本的纹理的方向。这样可以根据每一个样本的梯度强度值可以对应到对应的纹理特征索引。从而构建一个纹理特征统计表,将每一个计算的样本的梯度强度值累加到统计表中对应的纹理特征索引项中,以得到最终的纹理特征统计表,进而根据纹理特征统计表可以确定出当前块的一个或多个候选纹理特征索引。
在一些实施例中,根据纹理特征统计表,确定当前块的一个或多个候选纹理特征索引,该方法可以包括:将纹理特征统计表按照梯度强度累加值由高向低进行排序,确定排序靠前的N个梯度强度累加值对应的参考纹理特征索引;将N个参考纹理特征索引确定为当前块的一个或多个候选纹理特征索引。其中,N为正整数。
也就是说,在本申请实施例中,在纹理特征统计表构建完成之后,为了降低复杂度,还可以按照累计得到的梯度强度累加值从大到小的顺序对纹理特征统计表进行排序,然后只选取排序靠前的N个参考纹理特征索引。在这里,为了复杂度的考虑,这里还可以只维护一个长度为N的候选纹理特征列表,其包含一个或多个候选纹理特征索引,而且该候选纹理特征列表是按照梯度强度累加值由高向低进行排序的。另外,在本申请实施例中,N的取值可以为2、3、4、5、…、10等,这里对于N的取值不作具体限定。
可以理解地,在本申请实施例中,考虑到有些相邻的帧内预测模式的角度是非常接近的,为了排除过于接近的帧内预测模式,该方法还可以包括:对候选纹理特征列表中的N个参考纹理特征索引进行剪枝,确定当前块的一个或多个候选纹理特征索引。
在一些实施例中,对N个参考纹理特征索引进行剪枝,确定当前块的一个或多个候选纹理特征索引,可以包括:根据N个参考纹理特征索引中处于第一位置的参考纹理特征索引确定第一个候选纹理特征索引;在N个参考纹理特征索引中除第一位置之外的其他参考纹理特征索引与第i个候选纹理特征索引之间满足第二条件时,根据其他参考纹理特征索引确定第i+1个候选纹理特征索引,以确定当前块的一个或多个候选纹理特征索引;其中,i为大于零且小于N的整数。
需要说明的是,在本申请实施例中,第二条件可以包括:第i个候选纹理特征索引与第i+1个候选纹理特征索引之间的差距满足预设阈值,即相邻两个候选纹理特征索引之间具有一定的差距。或者在每一个候选纹理特征索引对应一个候选变换核组的情况下,也可以说,第i个候选变换核组与第i+1个候选变换核组之间的差距满足预设阈值。
还需要说明的是,在本申请实施例中,N个参考纹理特征索引是按照梯度强度累加值由高向低进行排序的。在这种情况下,根据N个参考纹理特征索引中处于第一位置的参考纹理特征索引确定第一个候选纹理特征索引,可以包括:根据N个参考纹理特征索引中梯度强度累加值最大的参考纹理特征索引确定第一个候选纹理特征索引。
还需要说明的是,在本申请实施例中,预设阈值可以用THR表示,第i个候选纹理特征索引可以用candFeature(i)表示,第i+1个候选纹理特征索引可以用candFeature(i+1)表示。在一种具体的实施例中,第i个候选纹理特征索引与第i+1个候选纹理特征索引之间的差距满足预设阈值,可以包括:candFeature(i+1)+THR<candFeature(i)||candFeature(i+1)-THR>candFeature(i)。
还需要说明的是,在本申请实施例中,THR的取值可能是3、4、5、6等。这样,对于多个相邻的角度累计值都很高的情况,这种剪枝方法会优先选择具有一定区分度的候选纹理特征索引。在一种可能的实施例中,THR等于0或者说不设置THR,也就是只需要这些候选纹理特征索引不同就可以。在另一种可能的实施例中,这里也可以不执行剪枝操作,即剪枝并不是必须的操作步骤。
还可以理解地,在本申请实施例中,上述方法没有考虑到DIMD、TIMD、SGPM等模式本身所导出的一些帧内预测模式。因此,在一些实施例中,该方法还可以包括:确定当前块的预测模式推导的一个或多个帧内预测模式;根据一个或多个帧内预测模式确定一个或多个候选变换核组,并将一个或多个 候选变换核组添加到第一候选列表。
在一些实施例中,该方法还包括:在第一候选列表未填满时,确定当前块的预设纹理特征索引;根据预设纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,该方法还包括:在第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;根据候选样本,确定当前块的一个或多个候选纹理特征索引;根据一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
也就是说,在本申请实施例中,考虑到DIMD、TIMD、SGPM等模式本身所导出的一些模式,比如DIMD本身会导出一个或几个帧内预测模式用来加权,TIMD本身也会导出一个或几个帧内预测模式用来加权,SGPM不仅有2个帧内预测模式,还有一个“划分”模式也可以找到对应的帧内预测模式,而且残差往往出现在“划分”的边界区域。因此,一种可能的实现方式是根据DIMD、TIMD、SGPM等模式本身所导出的帧内预测模式和上述方法导出的候选纹理特征索引来确定第一候选列表;或者另一种可能的实现方式是根据DIMD、TIMD、SGPM等模式本身所导出的帧内预测模式和默认的纹理特征索引来确定第一候选列表。
另外,在本申请实施例中,由于DIMD、TIMD、SGPM都可以导出不只一个帧内预测模式,因而对这几个模式也可以优先使用各模式自己导出的多个模式确定候选纹理特征索引,在各模式自己导出的模式不能填满所有的候选纹理特征索引时,一种可能的方法是加入默认的纹理特征索引,一种可能的方法是使用上述长度为N的排序的候选纹理特征列表确定候选纹理特征索引。
示例性地,在当前块的预测模式为DIMD模式时,如果它在预测时使用多个帧内预测模式的预测值加权,那么可以把用来加权多个帧内预测模式依次尝试确定为候选纹理特征索引。
示例性地,在当前块的预测模式为TIMD模式时,如果它在预测时使用多个帧内预测模式的预测值加权,那么可以把用来加权多个帧内预测模式依次尝试确定为候选纹理特征索引。
示例性地,在当前块的预测模式为SGPM模式时,可以将它所导出的“划分”模式对应的帧内预测模式,以及2个用来预测的帧内预测模式依次尝试确定为候选纹理特征索引。
在又一种可能的实现方式中,对于确定当前块的第一候选列表,该方法可以包括:根据候选纹理特征列表中处于第一位置的参考纹理特征索引确定第一个候选变换核组;在候选纹理特征列表中除第一位置之外的所有参考纹理特征索引与第一个候选变换核组之间均不满足第二条件时,根据候选纹理特征列表中处于第二位置的参考纹理特征索引确定第二个候选变换核组;根据第一个候选变换核组和第二个候选变换核组,确定第一候选列表。
需要说明的是,在本申请实施例中,根据候选纹理特征列表中处于第一位置的参考纹理特征索引确定第一个候选变换核组,具体可以是:根据候选纹理特征列表中梯度强度累加值最大的参考纹理特征索引确定第一个候选变换核组。
还需要说明的是,在本申请实施例中,以第一候选列表指示两个候选变换核组为例,第一个候选变换核组可以用candFeature0表示,第二个候选变换核组可以用candFeature1表示。示例性地,以梯度强度累计值最大的帧内预测模式为第一个候选变换核组candFeature0,那么在选择第二个候选变换核组candFeature1时,就需要candFeature1和candFeature0具有一定差距,例如candFeature1+THR<candFeature0||candFeature1-THR>candFeature0。如果检查完所有的N-1个参考纹理特征索引中仍然没有符合要求的,那么可以根据N个参考纹理特征索引中处于第二位置的参考纹理特征索引确定第二个候选变换核组candFeature1。也就是说,这时候可以将梯度强度累加值最大的前两个参考纹理特征索引对应的候选变换核组添加到第一候选列表。
在又一种可能的实现方式中,对于确定当前块的第一候选列表,该方法可以包括:根据候选纹理特征列表中处于第一位置的参考纹理特征索引确定第一个候选变换核组;在候选纹理特征列表中除第一位置之外的所有参考纹理特征索引与第一个候选变换核组之间均不满足第二条件时,根据预设纹理特征索引确定第二个候选变换核组;根据第一个候选变换核组和第二个候选变换核组,确定第一候选列表。
需要说明的是,在本申请实施例中,如果检查完所有的N-1个参考纹理特征索引中仍然没有符合要求的,那么也可以确定当前块的默认的纹理特征索引,然后将梯度强度累加值最大的参考纹理特征索引和默认的纹理特征索引对应的候选变换核组添加到第一候选列表。
在又一种可能的实现方式中,对于确定当前块的第一候选列表,该方法可以包括:根据当前块的预测模式推导的第一个帧内预测模式确定第一个候选变换核组;在候选纹理特征列表中的第一参考纹理特征索引与第一个候选变换核组之间满足第二条件时,根据第一参考纹理特征索引确定第二个候选变换核组;根据第一个候选变换核组和第二个候选变换核组,确定第一候选列表。
需要说明的是,在本申请实施例中,第一候选列表不仅考虑了根据候选样本的水平梯度值和竖直梯度值推导出的候选纹理特征索引,而且考虑了根据当前块的预测模式本身所推导出的帧内预测模式。下 面以几种示例进行详细说明。
示例性地,在当前块的预测模式为DIMD模式时,将DIMD模式导出的第一个帧内预测模式作为candFeature0。
示例性地,在当前块的预测模式为TIMD模式时,将TIMD模式导出的第一个帧内预测模式作为candFeature0。
示例性地,在当前块的预测模式为SGPM模式时,将SGPM“划分”模式对应的帧内预测模式作为candFeature0。
然后再按照上述方法确定candFeature1。例如尝试将N个参考纹理特征索引中的参考纹理特征索引,从第一位置开始如果某一个参考纹理特征索引符合THR的限制,那么可以将它作为candFeature1。
还需要说明的是,在本申请实施例中,对于IBC、ITMP等模式,可以使用预测块或者预测块加模板来推导候选纹理特征索引(或帧内预测模式)。
还需要说明的是,在本申请实施例中,如果MTSS技术推导N个参考纹理特征索引所使用的候选样本和DIMD模式使用的候选样本相同,那么DIMD推导出来的第一个帧内预测模式和MTSS技术推导出来的第一个参考纹理特征索引是相同的。
S2802,根据变换核组,确定当前块的变换核。
需要说明的是,在确定出当前块的变换核组之后,可以进一步确定出当前块的变换核。在一些实施例中,根据变换核组,确定当前块的变换核,可以包括:确定当前块的变换核索引;根据变换核组以及变换核索引,确定当前块的变换核。
还需要说明的是,在第一候选列表指示至少两个变换核组包括的变换核时,该方法还可以包括:确定当前块的变换核索引;根据第一候选列表以及变换核索引,确定当前块的变换核。
在本申请实施例中,对于第一候选列表来说,假设第一候选列表指示两个候选变换核组包括的变换核,若第一个候选变换核组包括3个候选变换核,第二个候选变换核组包括2个候选变换核,则第一候选列表可以指示5个候选变换核;若第一个候选变换核组包括3个候选变换核,第二个候选变换核组包括3个候选变换核,则第一候选列表可以指示6个候选变换核。在这种情况下,确定出当前块的变换核索引之后,可以根据变换核索引在第一候选列表中确定当前块的变换核。
还需要说明的是,在本申请实施例中,当前块的变换核索引可以为正整数,例如1,2,3,4,5,6等等。对于当前块的变换核索引可以是直接解码码流确定,或者可以是通过解码第二语法元素的取值确定。
示例性地,一种可能的实现方式为:解码码流,确定当前块的变换核索引。或者,另一种可能的实现方式为:解码码流,确定第二语法元素的取值;在第二语法元素指示当前块使用第一变换模式时,根据第二语法元素的取值,确定当前块的变换核索引。
还需要说明的是,在本申请实施例中,第一变换模式可以为LFNST/NSPT,第二语法元素可以用lfnst_idx表示。其中,第二语法元素可以用于指示当前块是否使用第一变换模式,以及在当前块使用第一变换模式时对应的变换核索引。
还需要说明的是,在本申请实施例中,若第二语法元素的取值为第三值,则确定当前块不使用第一变换模式;若第二语法元素的取值为第四值,则确定当前块使用第一变换模式以及对应的变换核索引。其中,第三值可以设置为0,第四值可以设置为非0,例如1,2,3,4,5,6等等。
也就是说,在本申请实施例中,对于LFNST/NSPT,当前块的变换核索引也可以用lfnst_idx表示。其中,lfnst_idx为0表示当前块不使用LFNST/NSPT,ECM中LFNST/NSPT的每个变换核组有3个变换核,所以lfnst_idx的值为1或2或3即表示当前块使用LFNST/NSPT的所选变换核组的第1个变换核或第2个变换核或第3个变换核。
在本申请实施例中,如果当前块的预测模式是某种特殊的帧内预测模式,那么它可选的变换核组不止一个。示例性地,它可选的变换核组是2个,那么lfnst_idx可能的值是0,1,2,3,4,5,6。其中1,2,3对应第一个变换核组的3个变换核,4,5,6对应第二个变换核组的3个变换核。
在一种具体的实施例中,lfnst_idx的二元符号对应表如表7所示。
表7

其中,第三个二元符号,即BinIdx为2的二元符号,也可以理解为选择第一个候选变换核组或第二个候选变换核组。
可以理解地,在本申请实施例中,这里的变换核组可以为第一候选列表指示的至少两个候选变换核组中的其中一个。另外,在本申请实施例中,当前块的变换核组索引可以用lfnst_feature_idx表示,或者用lfnst_set_idx表示。其中,当前块的变换核组索引可以为大于或等于零的整数,例如0,1,2等。示例性地,若lfnst_feature_idx的取值等于0,则指示选择第一个候选变换核组作为当前块的变换核组;若lfnst_feature_idx的取值等于1,则指示选择第二个候选变换核组作为当前块的变换核组。
需要说明的是,在确定当前块的变换核组之后,可以根据变换核组,确定当前块的变换核。具体可以为:先确定当前块的变换核索引,然后根据变换核组以及变换核索引,确定当前块的变换核。
示例性地,一种可能的实现方式为:解码码流,确定当前块的变换核索引。或者,另一种可能的实现方式为:解码码流,确定第二语法元素的取值;在第二语法元素指示当前块使用第一变换模式时,根据第一语法元素的取值,确定当前块的变换核索引。
还需要说明的是,在本申请实施例中,第一变换模式可以为LFNST/NSPT,第一语法元素可以用lfnst_idx表示。其中,第一语法元素可以用于指示当前块是否使用第一变换模式,以及在当前块使用第一变换模式时对应的变换核索引。示例性地,在这种情况下,对于LFNST/NSPT,当前块的变换核索引也可以用lfnst_idx表示。其中,VVC中,lfnst_idx可以有3个值即0,1,2。其中lfnst_idx为0表示当前块不使用LFNST,VVC中LFNST的每个变换核组有2个变换核,所以lfnst_idx的值为1或2即表示当前块使用LFNST的所选变换核组的第1个变换核或第2个变换核。在已有的ECM中,lfnst_idx可以有4个值即0,1,2,3。其中lfnst_idx为0表示当前块不使用LFNST/NSPT,ECM中LFNST/NSPT的每个变换核组有3个变换核,所以lfnst_idx的值为1或2或3即表示当前块使用LFNST/NSPT的所选变换核组的第1个变换核或第2个变换核或第3个变换核。
在一种具体的实施例中,lfnst_idx的二元符号对应表如表8所示。
表8
还需要说明的是,在本申请实施例中,在解码码流,确定当前块的变换核组索引时,该方法还可以包括:在当前块使用第一变换模式时,解码码流,确定当前块的变换核组索引。也就是说,在当前块使用MTSS技术,且lfnst_idx>0时,则执行解码码流,确定当前块的变换核组索引的步骤。
S2803,确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
需要说明的是,在本申请实施例中,对于确定当前块的变换系数,可以包括:解码码流,确定当前块的量化系数;对当前块的量化系数进行反量化,确定当前块的变换系数。
还需要说明的是,在本申请实施例中,在根据变换核对当前块的变换系数进行变换,确定当前块的残差块时,可以包括:根据变换核对当前块的变换系数进行不可分离基础变换,确定当前块的残差块;或者,根据变换核对当前块的变换系数进行低频不可分离变换,确定当前块的变换块,并对当前块的变换块进行离散余弦变换,确定当前块的残差块。
在一种具体的实施例中,若当前块的尺寸参数满足第一条件,则根据变换核对当前块的变换系数进行不可分离基础变换,确定当前块的残差块;若当前块的尺寸参数满足第二条件,则根据变换核对当前块的变换系数进行低频不可分离变换,确定当前块的变换块,并对当前块的变换块进行离散余弦变换,确定当前块的残差块。
在这里,当前块的尺寸参数满足第一条件,包括:当前块的尺寸参数较小,例如当前块的尺寸参数小于某一阈值。也就是说,针对较小尺寸的块,这里使用NSPT的变换核,即根据变换核对当前块的变换系数进行反NSPT变换,确定当前块的残差块。
在这里,当前块的尺寸参数满足第二条件,包括:当前块的尺寸参数较大,例如当前块的尺寸参数大于某一阈值。也就是说,针对较大尺寸的块,这里使用LFNST的变换核,即根据变换核对当前块的变换系数进行反LFNST变换,确定当前块的变换块;以及对当前块的变换块进行反DCT2变换,确定当前块的残差块。
还需要说明的是,在本申请实施例中,解码端对变换系数的“反变换”,在标准文本中也可称为“变换”。本文中的“变换”和“反变换”对应的是两个相反的过程,如“变换”将空间域的数值转换到频率域的系数,那么“反变换”将频率域的系数转换到空间域的数值。“反”是相对于“正”而言,它们实质上都是变换。需要注意的是,标准如果只规定解码,那么标准文本中的“变换”是解码的部分,具体指的是本文中的“反变换”。解码端对变换系数的“反变换”,在标准文本中也可称为“变换”。
在一些实施例中,参见图31,在步骤S2803之后,该方法还可以包括:
S3101,对当前块进行帧内预测,确定当前块的预测块。
S3102,根据当前块的预测块和当前块的残差块,确定当前块的重建块。
需要说明的是,在本申请实施例中,对于步骤S3101来说,该步骤可以是与步骤S2801~S2803并行操作,或者还可以是在步骤S2801~S2803之前执行,或者也可以是在步骤S2801~S2803之后执行,这里对步骤的先后顺序并不作具体限定。
还需要说明的是,在本申请实施例中,在确定当前块的预测块之后,可以对当前块的预测块和当前块的残差块进行加法运算,从而确定出当前块的重建块。
简单来说,解码器在确定出预测块后,根据预测块推导出候选纹理特征索引,然后根据候选纹理特征索引确定NSPT/LFNST的变换核组。如果一个变换核组有多个变换核可选,解码器通过在码流中解码lfnst_nspt_idx来确定变换核。在码流中解码lfnst_nspt_idx的过程和确定变换核组的过程没有依赖关系。通过熵解码从码流中解码得到量化系数。
如果是NSPT变换,对量化系数进行反量化得到解码的变换系数,对解码的变换系数进行反NSPT变换得到解码的残差块,最后根据解码的残差块和预测块得到重建块。
如果是LFNST变换,对量化系数进行反量化得到解码的变换系数,对解码的变换系数进行反LFNST变换,再进行反DCT2变换得到解码的残差块,最后根据解码的残差块和预测块得到重建块。
除此之外,在一些实施例中,参见图32,在当前块不使用MTSS技术时,该方法还可以包括:
S3201,确定当前块的纹理特征索引。
S3202,根据纹理特征索引,确定当前块的变换核。
S3203,确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
需要说明的是,在本申请实施例中,根据纹理特征索引,确定当前块的变换核,可以包括:根据纹理特征索引,确定当前块的变换核组;解码码流,确定当前块的变换核索引;根据变换核组以及变换核索引,确定当前块的变换核。
在一些实施例中,该方法还包括:在当前块的样本个数小于最小样本阈值时,执行确定当前块的纹理特征索引的步骤。
在一些实施例中,当前块的尺寸包括高度和宽度;该方法还包括:在当前块的高度或宽度小于最小尺寸阈值时,执行确定当前块的纹理特征索引的步骤。
在一些实施例中,该方法还包括:在当前块的预测模式为第一预测模式集合之外的预测模式时,执行确定当前块的纹理特征索引的步骤。
需要说明的是,在本申请实施例中,在当前块的预测模式为第一预测模式集合之外的预测模式时,当前块的预测模式可以为第四预测模式集合中的其中一项。其中,第四预测模式集合至少包括:DC模式、PLANAR模式和角度预测模式。
还需要说明的是,在本申请实施例中,第四预测模式集合中的模式为比较简单的预测模式。示例性地,可以做一个归类,DC模式、PLANAR模式和各种角度预测模式归为第四预测模式集合,DIMD模式、TIMD模式、MIP模式、EIP模式、SGPM模式、ITMP模式和IBC模式等归为第一预测模式集合。
示例性地,第一预测模式集合也可能只包含DIMD模式、TIMD模式、MIP模式、EIP模式、SGPM模式、ITMP模式和IBC模式中的一个或几个。如第一预测模式集合包含DIMD模式、TIMD模式和SGPM模式。其他模式都属于第四预测模式集合。所述第四预测模式集合的lfnst_idx的编码方法和相关技术的方法相同,而第一预测模式集合的lfnst_idx编码方法和相关技术的方法不同。
示例性地,LFNST/NSPT的每个变换核组有3个变换核,使用第四预测模式集合的块只可能选择一个变换核组,使用第一预测模式集合的块可以选择2个变换核组,每个变换核组有3个变换核。如果当前块的预测模式属于第四预测模式集合,它只可能选择一种纹理特征索引,那么lfnst_idx可能的值是0, 1,2,3。如果当前块的预测模式属于第一预测模式集合,那么lfnst_idx可能的值是0,1,2,3,4,5,6。其中1,2,3对应第一个纹理特征索引对应的变换核组的3个变换核,4,5,6对应第二个纹理特征索引对应的变换核组的3个变换核。在这里,使用第四预测模式集合的块的lfnst_idx的二元符号对应表如表8所示,使用第一预测模式集合的块的lfnst_idx的二元符号对应表如表7所示。
还可以理解地,本申请实施例可以使用一个高层语法(high level syntax)控制本技术方案的开关。在一些实施例中,该方法还包括:解码码流,确定第三语法元素的取值;在第三语法元素指示当前序列允许使用多变换核组选择技术时,执行确定当前块的变换核组索引的步骤;其中,当前序列包括当前块。
在本申请实施例中,第三语法元素可以用sps_mtss_enabled_flag表示,而且第三语法元素为序列参数集(Sequence Parameter Set,SPS)中的语法元素。
还需要说明的是,在本申请实施例中,如果第三语法元素的取值为第五值,则确定第三语法元素指示当前序列允许使用MTSS技术;如果第三语法元素的取值为第六值,则确定第三语法元素指示当前序列不允许使用MTSS技术。
在本申请实施例中,第五值与第六值不同,而且第五值和第六值可以是参数形式,也可以是数字形式。具体地,第三语法元素可以是写入在概述(profile)中的参数,也可以是一个标志(flag)的取值,这里对此不作具体限定。示例性地,第五值可以为1,第六值可以为0;或者,第五值可以为0,第六值可以为1;或者,第五值可以为true,第六值可以为false;或者,第五值可以为false,第六值可以为true。在一种具体的实施例中,第五值为1,第六值为0。
也就是说,本申请实施例可以使用一个高层语法(high level syntax)控制本技术方案的开关。比如使用一个序列级的flag,如在序列参数集中增加语法元素sps_mtss_enabled_flag。如果sps_mtss_enabled_flag的取值为1,当前序列允许使用MTSS技术;如果sps_mtss_enabled_flag的取值为0,当前序列不允许使用MTSS技术。其中,在允许使用MTSS技术的情况下,更具体地,允许在块级(CU或TU)解码本技术方案所述的纹理特征索引、LFNST/NSPT变换核索引的方法;如果当前序列不允许使用MTSS技术,更具体地,在块级(CU或TU)不会使用本技术方案所述的纹理特征索引、LFNST/NSPT变换核索引的方法。
在一些实施例中,该方法还包括:解码码流,确定第三语法元素的取值;在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第四语法元素的取值;在第四语法元素指示当前图像允许使用多变换核组选择技术时,执行确定当前块的变换核组索引的步骤。
在本申请实施例中,当前序列可以包括当前图像,且当前图像包括当前块。其中,第四语法元素可以用ph_inter_lfnst_nspt_enabled_flag表示,而且第四语法元素为图像级别的语法元素。
还需要说明的是,在本申请实施例中,如果第四语法元素的取值为第五值,则确定第四语法元素指示当前图像允许使用MTSS技术;如果第四语法元素的取值为第六值,则确定第四语法元素指示当前图像不允许使用MTSS技术。
在一些实施例中,该方法还包括:解码码流,确定第三语法元素的取值;在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第五语法元素的取值;在第五语法元素指示当前片允许使用多变换核组选择技术时,执行确定当前块的变换核组索引的步骤;其中,当前序列包括当前片,当前片包括当前块。
在本申请实施例中,当前序列可以包括当前片,且当前片包括当前块。其中,第五语法元素可以用sh_inter_lfnst_nspt_enabled_flag表示,而且第五语法元素为片(slice)级别的语法元素。
还需要说明的是,在本申请实施例中,如果第五语法元素的取值为第五值,则确定第五语法元素指示当前片允许使用MTSS技术;如果第五语法元素的取值为第六值,则确定第五语法元素指示当前片不允许使用MTSS技术。
在本申请实施例中,第五值与第六值不同,而且第五值和第六值可以是参数形式,也可以是数字形式。具体地,第四语法元素或第五语法元素可以是写入在概述(profile)中的参数,也可以是一个标志(flag)的取值,这里对此不作具体限定。示例性地,第五值可以为1,第六值可以为0;或者,第五值可以为0,第六值可以为1;或者,第五值可以为true,第六值可以为false;或者,第五值可以为false,第六值可以为true。在一种具体的实施例中,第五值为1,第六值为0。
也就是说,在本申请实施例中,除了序列级的语法元素sps_mtss_enabled_flag之外,当然也可以使用其他级别的语法来实现更灵活的控制,如图像参数集(Picture Parameter Set,PPS)的flag,或图像头(picture header)或片头(slice header)的flag等。比如先在SPS确定当前序列是否可以使用MTSS技术,如果当前序列使用MTSS技术,那么设置一个片头(slice header)的sh_inter_lfnst_nspt_enabled_flag确定当前slice是否使用MTSS技术,提供更高的灵活性。
还可以理解地,对于最小样本阈值来说,在一种可能的实现方式中,可以包括:解码码流,确定最 小样本阈值。或者,在另一种可能的实现方式中,可以包括:解码码流,确定第三语法元素的取值;在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第六语法元素的取值;根据第六语法元素的取值,确定最小样本阈值。
在本申请实施例中,最小样本阈值可以是直接写入码流,或者以第六语法元素的形式写入码流。另外,第六语法元素可以用sps_mtss_min_pix表示,其可以为序列级别的语法元素。
在解码码流时,第六语法元素的取值可以与最小样本阈值相等,或者也可以是按照一定的映射规则进行解码。示例性地,若最小样本阈值等于32,则确定第六语法元素的取值等于0;若最小样本阈值等于64,则确定第六语法元素的取值等于1;若最小样本阈值等于256,则确定第六语法元素的取值等于2。这样,在解码获得第六语法元素的取值等于2时,那么可以确定最小样本阈值等于256。
也就是说,在本申请实施例中,可以使用一个高层语法(high level syntax)设置一个应用MTSS技术的最小样本阈值,例如sps_mtss_min_pix。当sps_mtss_enabled_flag的取值为1时,解码器解析sps_mtss_min_pix确定应用MTSS技术的最小样本阈值。用高层语法可以根据需求在编码复杂度和压缩效率之间进行取舍。也就是说,这种方法下解码器需要支持所有sps_mtss_min_pix可能的情况,但是编码器可以配置编码当前码流需要的sps_mtss_min_pix。例如编码一个码流时,如果需要更好的压缩效率而不特别注重编码时间,可以给sps_mtss_min_pix设置一个比较小的值,如16。如果特别注重编码时间而且可以损失一定的压缩效率,可以给sps_mtss_min_pix设置一个比较大的值,如256。
还可以理解地,对于最小尺寸阈值来说,在一种可能的实现方式中,可以包括:解码码流,确定最小尺寸阈值。或者,在另一种可能的实现方式中,可以包括:解码码流,确定第三语法元素的取值;在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第七语法元素的取值;根据第七语法元素的取值,确定最小尺寸阈值。
在本申请实施例中,最小尺寸阈值可以是直接写入码流,或者以第七语法元素的形式写入码流。另外,第七语法元素可以用sps_mtss_min_size表示,其可以为序列级别的语法元素。
在解码码流时,第七语法元素的取值可以与最小尺寸阈值相等,或者也可以是按照一定的映射规则进行解码。示例性地,若最小尺寸阈值等于4,则确定第七语法元素的取值等于0;若最小样本阈值等于16,则确定第七语法元素的取值等于1。这样,在解码获得第七语法元素的取值等于1时,那么可以确定最小尺寸阈值等于16。
也就是说,在本申请实施例中,可以使用一个高层语法(high level syntax)设置一个应用MTSS技术的最小尺寸阈值,例如sps_mtss_min_size。当sps_mtss_enabled_flag的取值为1时,解码器解析sps_mtss_min_size确定应用MTSS技术的最小尺寸阈值。用高层语法可以根据需求在编码复杂度和压缩效率之间进行取舍。也就是说,这种方法下解码器需要支持所有sps_mtss_min_size可能的情况,但是编码器可以配置编码当前码流需要的sps_mtss_min_size。比如编码一个码流时,如果需要更好的压缩效率而不特别注重编码时间,可以给sps_mtss_min_size设置一个比较小的值,如4。如果特别注重编码时间而且可以损失一定的压缩效率,可以给sps_mtss_min_size设置一个比较大的值,如16。
本申请实施例提供了一种解码方法,具体为一种帧内LFNST/NSPT多角度选择方案。首先确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;然后根据变换核组,确定当前块的变换核;再确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。这样,根据多变换核组选择技术确定当前块的变换核组,然后从中确定当前块的变换核。如此,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。
在本申请的另一实施例中,图33为本申请实施例提供的一种编码方法的流程示意图一。如图33所示,该方法可以包括:
S3301,确定当前块的变换核组。
需要说明的是,在本申请实施例中,该方法应用于编码器。具体来说,基于图26所示编码器100的组成结构,本申请实施例的编码方法主要应用于帧内预测的块。其中,在当前块采用帧内预测模式时,这里主要是针对帧内预测模式下的NSPT和LFNST变换所提出的优化方案,以提高压缩效率。
还需要说明的是,在本申请实施例中,该方法适用于多变换组选择(Multiple Transform Set Selection,MTSS)方式。其中,NSPT和LFNST都是处理各种角度的纹理的变换,它们可能有多个变换核,一个变换核可能是专门为某种特定的角度纹理优化的。当然除了角度的纹理,NSPT和LFNST也包括处理渐变纹理的变换核。其实这些变换核也可以说是训练好的KL变换(Karhunen-Loeve Transform,KLT)。也就是说,NSPT和LFNST都有多个变换核,每个变换核为特定的纹理设计,特定的纹理包括角度纹 理,渐变纹理等。另外,渐变纹理可以进一步扩展如水平渐变纹理,竖直渐变纹理,斜向渐变纹理等。进一步地,MTSS方式不局限于只能用于NSPT和LFNST这样的不可分离变换,针对为特定纹理优化的可分离变换也可以应用MTSS方式。
还需要说明的是,在本申请实施例中,帧内预测的块可以有多个变换核组,每种帧内预测模式可以对应一种变换核组。换句话说,每种帧内预测模式其实代表了一种纹理特征。所以帧内预测模式索引也是一种纹理特征索引。如DC和PLANAR对应渐变的纹理特征,某一种角度预测模式对应这种角度的纹理特征。纹理特征索引一方面可以避免在“帧间”出现帧内预测模式,另一方面它也更利于可能的扩展,例如一种帧内预测模式可以对应多种纹理特征,比如DC模式可以对应水平渐变纹理,竖直渐变纹理,斜向渐变纹理等。
在本申请实施例中,考虑到一些帧内预测模式并非是简单的纹理特征,而是有可能包含两种或以上的纹理特征。因此,这里可以使用多变换核组选择(Multiple Transform Set Selection,MTSS)技术。在MTSS技术中,如果当前块的帧内预测模式是某种特殊的帧内预测模式,那么可选的变换核组不止一个。
在本申请实施例中,对于当前块是否使用MTSS技术,可以增加一些判断条件,示例性地,这些判断条件可以包括:当前块的样本个数、当前块的尺寸和当前块的预测模式等。也就是说,在本申请实施例中,可以选择当前块的样本个数来判断当前块是否使用MTSS技术,和/或,也可以选择当前块的尺寸来判断当前块是否使用MTSS技术,和/或,还可以选择当前块的预测模式来判断当前块是否使用MTSS技术,这里不作任何限定。
在本申请实施例中,MTSS技术对编码器而言,由于增加了变换核组的候选,这会增加编码器的复杂度。具体来说,一个块可能从2个或多个LFNST/NSPT的变换核组中选择一个。而VVC中的一个块只有一个可用的变换核组。因而一种可能的实现方式是限制MTSS技术所适用的块的尺寸,从而可以在MTSS技术需要增加很多计算而对压缩效率提升不明显的块尺寸禁用MTSS技术。简单来说,本申请实施例可以第MTSS技术所适用的块尺寸进行限制。
在一种可能的实施例中,该方法还可以包括:确定当前块的样本个数;在当前块的样本个数大于或等于最小样本阈值时,执行确定当前块的变换核组的步骤。
示例性地,可以设置一个应用MTSS技术的最小的样本数的阈值(即“最小样本阈值”),用MIN_PIX表示。如果一个块的样本数(宽度×高度)小于MIN_PIX,则这个块不能使用MTSS技术。否则,若这个块的样本数大于或等于MIN_PIX,则这个块可以使用MTSS技术。在本申请实施例中,MIN_PIX的取值可能是32,64,256等。示例性地,MIN_PIX的取值可以等于256。
在另一种可能的实施例中,该方法还可以包括:确定当前块的尺寸,其中,当前块的尺寸包括高度和宽度;在当前块的高度和宽度均大于或等于最小尺寸阈值时,执行确定当前块的变换核组的步骤。
示例性地,可以设置一个应用MTSS技术的最小的尺寸的阈值(即“最小尺寸阈值”),用MIN_SIZE表示。如果一个块的宽度或高度小于MIN_SIZE,则这个块不能使用MTSS技术。否则,若这个块的宽度和高度均大于或等于MIN_SIZE,则这个块可以使用MTSS技术。在本申请实施例中,MIN_SIZE的值可能是8,16等。示例性地,MIN_PIX的取值可以等于8。
在又一种可能的实施例中,该方法还可以包括:确定当前块的预测模式;在当前块的预测模式为第一预测模式集合中的其中一项时,执行确定当前块的变换核组的步骤。
在一些实施例中,第一预测模式集合可以包括以下预测模式中至少之一:DC模式、PLANAR模式和角度预测模式之外的其他预测模式。
在一些实施例中,第一预测模式集合可以包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式;复制帧内块进行预测的模式;使用外插滤波器进行预测的模式;使用矩阵运算进行预测的模式。
在一些实施例中,第一预测模式集合可以包括以下预测模式中至少之一:DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式等。
在本申请实施例中,对于将至少两种帧内预测模式的预测值进行组合预测的模式,具体可以是DIMD模式、TIMD模式和SGPM模式;其中,帧内预测模式包括但不限于DC模式、PLANAR模式和角度预测模式等。
在本申请实施例中,对于复制帧内块进行预测的模式,具体可以是ITMP模式和IBC模式。
在本申请实施例中,对于使用外插滤波器进行预测的模式,具体可以是EIP模式。
在本申请实施例中,对于使用矩阵运算进行预测的模式,具体可以是MIP模式。
也就是说,第一预测模式集合中的模式为比较复杂的预测模式,其对应的纹理中可能包含两种或以上的纹理特征。示例性地,DIMD模式和TIMD模式都可以对两种或以上的帧内预测模式的预测值进行加权,SGPM模式是对两种帧内预测模式的预测值用权重矩阵进行加权,MIP模式是根据一个矩阵运算 进行预测,EIP模式是根据一个外插滤波器(Extrapolation filter)进行预测,ITMP模式和IBC模式则是基于复制一个已重建的参考块来进行预测,它们的纹理特征并不像DC模式、PLANAR模式等那样简单的纹理特征。也就是说,在当前块的预测模式为第一预测模式集合中的任意一项时,这时候当前块可以使用MTSS技术。
可以理解地,在本申请实施例中,在确定当前块使用MTSS技术时,这时候首先需要确定当前块的变换核组。在一些实施例中,确定当前块的变换核组,可以包括:根据至少两个候选变换核组对当前块进行编码代价计算,确定至少两个候选变换核组各自对应的代价结果;在至少两个候选变换核组各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核组确定为当前块的变换核组。
在一种可能的实施例中,该方法还包括:确定当前块的变换核组索引;对当前块的变换核组索引进行编码处理,将所得到的编码比特写入码流。
在另一种可能的实施例中,该方法还包括:确定当前块的变换核组索引;根据当前块的变换核组索引,确定第一语法元素的取值;对第一语法元素的取值进行编码处理,将所得到的编码比特写入码流。
需要说明的是,在本申请实施例中,变换核组索引可以用于指示当前块的变换核组在至少两个候选变换核组中的编号,该变换核组索引可以用lfnst_feature_idx表示,或者也可以lfnst_set_idx表示。其中,当前块的变换核组索引可以为大于或等于零的整数,例如0,1,2等。示例性地,若lfnst_feature_idx的取值等于0,则指示选择第一个候选变换核组作为当前块的变换核组;若lfnst_feature_idx的取值等于1,则指示选择第二个候选变换核组作为当前块的变换核组。
还需要说明的是,在本申请实施例中,对于当前块的变换核组索引可以是直接写入码流,或者也可以是用第一语法元素表示,然后将第一语法元素的取值写入码流。
还需要说明的是,在本申请实施例中,第一语法元素可以用lfnst_feature_idx或者lfnst_set_idx表示。第一语法元素可以用于指示当前块的变换核组索引,具体是当前块的变换核组在至少两个候选变换核组中的编号。其中,第一语法元素的取值可以为大于或等于零的整数,例如0,1,2等。示例性地,若第一语法元素的取值等于0,则指示选择第一个候选变换核组作为当前块的变换核组;若第一语法元素的取值等于1,则指示选择第二个候选变换核组作为当前块的变换核组;若第一语法元素的取值等于2,则指示选择第三个候选变换核组作为当前块的变换核组。
在一种具体的实现方式中,在当前块的预测模式为第二预测模式集合中的其中一项时,根据当前块的变换核组索引确定当前块的变换核组,可以包括:若变换核组索引为第i值,则基于当前块的预测模式推导的第i个帧内预测模式确定当前块的变换核组,i为正整数。
其中,第二预测模式集合可以包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
在一种具体的实施例中,第二预测模式集合可以包括以下预测模式中至少之一:DIMD模式、TIMD模式和SGPM模式。
在本申请实施例中,MTSS可以应用于DIMD模式、TIMD模式和SGPM模式。若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则基于当前块的预测模式推导的第二个帧内预测模式确定当前块的变换核组;若变换核组索引为第三值,则基于当前块的预测模式推导的第三个帧内预测模式确定当前块的变换核组等等。
在这里,对于变换核组索引来说,第一值可以0,第二值可以为1,第三值可以为2。也就是说,针对第二预测模式集合中的预测模式,在确定出当前块的变换核组索引之后,可以根据变换核组索引直接来确定对应的变换核组。
示例性地,对于DIMD模式,由于DIMD模式本身就使用梯度推导纹理特征,而且它可以推导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据当前块的变换核组索引可以确定对应的变换核组。
示例性地,对于TIMD模式,由于TIMD模式本身可以导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据当前块的变换核组索引可以确定对应的变换核组。
示例性地,对于SGPM模式,由于SGPM模式本身可以导出一个“划分”模式和2个帧内预测模式,因而也不需要MTSS技术额外使用梯度推导纹理特征,根据当前块的变换核组索引可以确定对应的变换核组。
更具体地,在一种可能的实施例中,如下:
示例性地,对于DIMD模式,假设DIMD模式推导的第一个帧内预测模式对应的变换核组索引为0,记作candFeature0;DIMD模式推导的第二个帧内预测模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于DIMD模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为1,则基于DIMD模式推导的第二个帧内预测模式确定当前块的变换核组。
示例性地,对于TIMD模式,假设TIMD模式推导的第一个帧内预测模式对应的变换核组索引为0,记作candFeature0;TIMD模式推导的第二个帧内预测模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于TIMD模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为1,则基于TIMD模式推导的第二个帧内预测模式确定当前块的变换核组。
示例性地,对于SGPM模式,假设SGPM模式的划分模式对应的变换核组索引为0,记作candFeature0;SGPM模式推导的第一个帧内预测模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于SGPM模式的划分模式确定当前块的变换核组;若变换核组索引为1,则基于SGPM模式推导的第一个帧内预测模式确定当前块的变换核组。
在另一种具体的实现方式中,在当前块的预测模式为第三预测模式集合中的其中一项时,根据当前块的变换核组索引确定当前块的变换核组,可以包括:若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则基于PLANAR模式确定当前块的变换核组。
在又一种具体的实现方式中,在当前块的预测模式为第三预测模式集合中的其中一项时,根据当前块的变换核组索引确定当前块的变换核组,可以包括:若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则当前块的预测模式推导的第二个帧内预测模式确定当前块的变换核组。
其中,第三预测模式集合可以包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
在一种具体的实施例中,第三预测模式集合可以包括以下预测模式中至少之一:MIP模式和EIP模式。
在本申请实施例中,MTSS可以应用于MIP模式和EIP模式。由于MIP模式和EIP模式可以大量应用于纹理渐变的块,其特征和PLANAR模式是相似的。因而对MIP模式和EIP模式也可以将PLANAR模式作为其候选纹理特征。
在这里,对于变换核组索引来说,第一值可以0,第二值可以为1。也就是说,针对第三预测模式集合中的预测模式,在确定出当前块的变换核组索引之后,也可以根据变换核组索引直接来确定对应的变换核组。
示例性地,在当前块的预测模式为MIP模式或者EIP模式时,假设MIP模式或者EIP模式推导的第一个帧内预测模式对应的变换核组索引为0,记作candFeature0;PLANAR模式对应的变换核组索引为1,记作candFeature1。这样,若变换核组索引为0,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为1,则基于PLANAR模式确定当前块的变换核组。
还可以理解地,在本申请实施例中,对于MTSS技术的至少两个候选变换核组来说,编码端还可以构建第一候选列表。在一些实施例中,对于步骤S3301来说,该步骤还可以包括:确定当前块的第一候选列表,第一候选列表指示至少两个候选变换核组;根据第一候选列表和变换核组索引,确定当前块的变换核组。
需要说明的是,在本申请实施例中,在当前块的预测模式为DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式等中的任意一项时,第一候选列表可以包括至少两个候选纹理特征索引,或者第一候选列表可以包括至少两个变换核组。在这里,每一个候选纹理特征索引对应一个变换核组。因此,可以说是:第一候选列表指示至少两个候选变换核组。
在一些实施例中,确定当前块的第一候选列表,可以包括:确定基于当前块的预测模式推导的一个或多个帧内预测模式;根据一个或多个帧内预测模式确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一种可能的实施例中,确定当前块的第一候选列表,具体可以包括:在当前块的预测模式为第二预测模式集合中的其中一项时,确定基于当前块的预测模式推导的多个帧内预测模式;根据多个帧内预测模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表。
在本申请实施例中,第二预测模式集合可以包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式;示例性地,第二预测模式集合可以包括以下预测模式中至少之一:DIMD模式、TIMD模式和SGPM模式。也就是说,MTSS技术可以应用于DIMD模式、TIMD模式和SGPM模式,这时候可以不额外使用梯度推导纹理特征,仅根据该预测模式自身所推导出的一个或多个帧内预测模式来构建第一候选列表。
示例性地,对于DIMD模式,由于DIMD模式本身就使用梯度推导纹理特征,而且它可以推导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据DIMD模式本身推导的 多个帧内预测模式来构建第一候选列表。
示例性地,对于TIMD模式,由于TIMD模式本身可以导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征,根据TIMD模式本身推导的多个帧内预测模式来构建第一候选列表。
示例性地,对于SGPM模式,由于SGPM模式本身可以导出一个“划分”模式和2个帧内预测模式,因而也不需要MTSS技术额外使用梯度推导纹理特征,根据SGPM模式本身推导的一个“划分”模式和2个帧内预测模式来构建第一候选列表。
在一些实施例中,确定当前块的第一候选列表,还可以包括:在第一候选列表未填满时,确定当前块的预设纹理特征索引;根据预设纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在本申请实施例中,预设纹理特征索引可以为PLANAR模式。在一种具体的实施例中,该方法还可以包括:根据多个帧内预测模式和PLANAR模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表。
示例性地,如果DIMD模式、TIMD模式、SGPM模式所推导出的帧内预测模式中包括ITMP或IBC,或EIP或MIP的模式,这时候也可以使用PLANAR模式对应的纹理特征作为候选纹理特征。由于DIDM模式、TIMD模式、SGPM模式可以使用2个以上的帧内预测模式进行加权或组合,所述帧内预测模式可能是ITMP或IBC,或EIP或MIP的模式。在这种情况下,也可以使用PLANAR模式对应的纹理特征作为候选纹理特征。简单来说,可以根据DIMD模式、TIMD模式、SGPM模式所推导出的多个帧内预测模式和PLANAR模式共同来构建第一候选列表。
在另一种可能的实施例中,确定当前块的第一候选列表,具体可以包括:在当前块的预测模式为第三预测模式集合中的其中一项时,确定基于当前块的预测模式推导的帧内预测模式;根据帧内预测模式和PLANAR模式确定两个候选变换核组,并将两个候选变换核组添加到第一候选列表。
在本申请实施例中,第三预测模式集合可以包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。示例性地,第三预测模式集合可以包括以下预测模式中至少之一:MIP模式和EIP模式。也就是说,MTSS技术可以应用于MIP模式和EIP模式,这时候也可以不额外使用梯度推导纹理特征,仅根据该预测模式自身所推导出的第一个帧内预测模式和PLANAR模式共同来构建第一候选列表。
在一些实施例中,确定当前块的第一候选列表,还可以包括:在第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;根据所述候选样本,确定所述当前块的一个或多个候选纹理特征索引;根据一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在本申请实施例中,在构建第一候选列表时,MTSS技术也可以额外使用梯度推导纹理特征。在这里,根据候选样本确定当前块的一个或多个候选变换核组时,可以是先根据候选样本确定包含一个或多个候选纹理特征索引的一个或多个候选纹理特征索引,然后根据这一个或多个候选纹理特征索引来确定当前块的一个或多个候选变换核组。
通常情况下,一个候选纹理特征索引对应一个候选变换核组;但是某些情况下,针对多个相近的候选纹理特征索引可以对应同一个变换核组。例如多个类似角度的帧内预测模式对应同一个变换核组。在本申请实施例中,如果多个相邻的帧内预测模式(或候选纹理特征索引)对应一个变换核组,那么在确定候选变换核组时,需要保证各个候选纹理特征索引不会确定为相同的变换核组。
在一种可能的实现方式中,对于候选样本来说,可以是确定当前块的预测块;将预测块中的至少部分样本作为候选样本。
在本申请实施例中,如果预测块存在某种纹理,可以认为残差块中存在相同特征的纹理。这样,关于帧间推导候选纹理特征索引所使用的候选样本,可以是预测块中的全部样本或者预测块中的部分样本。
在另一种可能的实现方式中,对于候选样本来说,可以是确定当前块的已重建区域的相邻样本;将已重建区域的相邻样本作为候选样本。
在本申请实施例中,关于帧间推导候选纹理特征索引所使用的候选样本,可以是使用当前块的已重建区域的相邻样本,例如当前块左侧和右侧的已重建区域。因为左侧和上侧的已重建区域虽然不是当前块但是和当前块相邻,比如说纹理有相连的情况,因此一定程度上可以用于估计当前块的纹理。
在又一种可能的实现方式中,考虑使用更多的样本,对于候选样本来说,可以是将已重建区域的相邻样本和预测块中的至少部分样本共同作为候选样本。
在本申请实施例中,关于帧间推导候选纹理特征索引所使用的候选样本,也可以是同时使用当前块的预测块和当前块左侧和上侧的已重建区域,这样用于推导候选纹理特征索引的样本更多,使得所推导的候选纹理特征索引更准确。
还需要说明的是,在本申请实施例中,用于推导候选纹理特征索引的候选样本的数量可以为至少一个,例如1个、2个、3个或更多个。
在一些实施例中,该方法还可以包括:根据当前块的尺寸参数,确定候选样本的数量。
也就是说,在根据候选样本推导一个或多个候选纹理特征索引时,使用多少个候选样本可以是由当前块的尺寸参数确定。示例性地,如果当前块的尺寸小,那么可以对所有可用的样本进行统计;如果当前块的尺寸大,那么可以对当前块降采样进行统计,比如水平方向和/或竖直方向每2个、或4个、或8个样本中的一个样本进行统计。或者,如果当前块的水平或竖直某一个方向的尺寸小于或等于8,那么对该方向上的所有可用的样本进行统计;否则,如果当前块的水平或竖直某一个方向的尺寸小于或等于16,那么对该方向上的每2个样本中的一个样本进行统计;否则,对该方向上的每4个样本中的一个样本进行统计,这里不作具体限定。
在一些实施例中,根据候选样本,确定当前块的一个或多个候选纹理特征索引,可以包括:确定候选样本的水平梯度值和竖直梯度值;根据候选样本的水平梯度值和竖直梯度值,确定候选样本对应的纹理特征索引以及梯度强度值;根据候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表;根据纹理特征统计表,确定当前块的一个或多个候选纹理特征索引。
需要说明的是,在本申请实施例中,在根据候选样本的水平梯度值和竖直梯度值,确定候选样本对应的纹理特征索引以及梯度强度值时,可以包括:根据候选样本的水平梯度值和竖直梯度值进行角度映射,确定候选样本对应的纹理特征索引;以及根据候选样本的水平梯度值和竖直梯度值进行梯度强度计算,确定候选样本对应的梯度强度值。
在一种具体的实施例中,根据候选样本的水平梯度值和竖直梯度值进行角度映射,确定候选样本对应的纹理特征索引,可以包括:根据候选样本的水平梯度值和竖直梯度值,利用预设查找表确定候选样本对应的纹理特征索引。
在本申请实施例中,候选样本的水平梯度值可以用gradx表示,候选样本的竖直梯度值可以用grady表示。这样,根据gradx和grady导出纹理特征索引(或称为“虚拟的帧内预测模式”)可以通过查表来实现。
示例性地,如果abs(gradx)等于0且abs(grady)不等于0,那么有水平方向纹理,对应VVC中的帧内预测模式18。如果abs(grady)等于0且abs(gradx)不等于0,那么有竖直方向纹理,对应VVC中的帧内预测模式50。在abs(gradx)和abs(grady)都不等于0的情况下,如果abs(gradx)等于abs(grady),且gradx和grady符号相同,对应VVC中的帧内预测模式34。如果abs(gradx)等于2倍的abs(grady),且gradx和grady符号相同,对应VVC中的帧内预测模式40。另外,其他情况可以按照相同的原则查表确定。
在一种具体的实施例中,根据候选样本的水平梯度值和竖直梯度值进行梯度强度计算,确定候选样本对应的梯度强度值,可以包括:对水平梯度值的绝对值和竖直梯度值的绝对值进行加法运算,确定候选样本对应的梯度强度值。
在这里,候选样本对应的梯度强度值可以记为amp。示例性地,amp=abs(gradx)+abs(grady)。
需要说明的是,在本申请实施例中,针对候选样本的水平梯度值和竖直梯度值可以使用索贝尔算子(sobel)计算得到。示例性地,对于索贝尔算子来说,具体如下所示:
水平梯度值的算子:
竖直梯度值的算子:
如此,假设样本位置为(x,y)的样本值为Px,y,则水平梯度值gradx和竖直梯度值grady的计算如下所示:
gradx=Px+1,y-1+2*Px+1,y+Px+1,y+1-Px-1,y-1-2*Px-1,y-Px-1,y+1             (7)
grady=Px-1,y+1+2*Px,y+1+Px+1,y+1-Px-1,y-1-2*Px,y-1-Px+1,y-1             (8)
还需要说明的是,在本申请实施例中,还可以设置候选样本不包括当前块最外面上下左右一行一列的样本。其中,考虑到索贝尔算子要使用到当前样本的上下左右一行一列的样本,所以本申请实施例可以设置不计算当前块最外面上下左右一行一列的样本的梯度。
在一些实施例中,根据候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表,可以 包括:在候选样本的数量为至少一个时,确定至少一个纹理特征索引以及对应的至少一个梯度强度值;根据至少一个纹理特征索引确定具有互异特性的至少一种参考纹理特征索引,以及根据至少一个梯度强度值,对归属于同一参考纹理特征索引的梯度强度值进行累加计算,确定至少一种参考纹理特征索引对应的梯度强度累加值;根据至少一种参考纹理特征索引和至少一种参考纹理特征索引对应的梯度强度累加值,确定纹理特征统计表。
也就是说,在本申请实施例中,以预测块中的至少部分样本作为候选样本为例,对预测块中的全部或部分样本计算其梯度,一般来说可以计算水平梯度值和竖直梯度值,这里可以使用索贝尔算子来计算梯度值。对某一个样本,根据其水平梯度值和竖直梯度值可以推测该样本的纹理方向,比如水平梯度值非零,竖直梯度值为零,那该样本的纹理是竖直方向。反之,水平梯度值为零,竖直梯度值非零,那该样本的纹理是水平方向。比如水平梯度值和竖直梯度值相等且不为零,那该样本的纹理是45度。当然,本申请实施例还有很多其他情况是水平梯度值和竖直梯度值都不为零,可以根据它们的比值确定该样本的纹理的方向。这样可以根据每一个样本的梯度强度值可以对应到对应的纹理特征索引。从而构建一个纹理特征统计表,将每一个计算的样本的梯度强度值累加到统计表中对应的纹理特征索引项中,以得到最终的纹理特征统计表,进而根据纹理特征统计表可以确定出当前块的一个或多个候选纹理特征索引。
在一些实施例中,根据纹理特征统计表,确定当前块的一个或多个候选纹理特征索引,该方法可以包括:将纹理特征统计表按照梯度强度累加值由高向低进行排序,确定排序靠前的N个梯度强度累加值对应的参考纹理特征索引;将N个参考纹理特征索引确定为当前块的一个或多个候选纹理特征索引。其中,N为正整数。
也就是说,在本申请实施例中,在纹理特征统计表构建完成之后,为了降低复杂度,还可以按照累计得到的梯度强度累加值从大到小的顺序对纹理特征统计表进行排序,然后只选取排序靠前的N个参考纹理特征索引。在这里,为了复杂度的考虑,这里还可以只维护一个长度为N的候选纹理特征列表,其包含一个或多个候选纹理特征索引,而且该候选纹理特征列表是按照梯度强度累加值由高向低进行排序的。另外,在本申请实施例中,N的取值可以为2、3、4、5、…、10等,这里对于N的取值不作具体限定。
可以理解地,在本申请实施例中,考虑到有些相邻的帧内预测模式的角度是非常接近的,为了排除过于接近的帧内预测模式,该方法还可以包括:对候选纹理特征列表中的N个参考纹理特征索引进行剪枝,确定当前块的一个或多个候选纹理特征索引。
在一些实施例中,对N个参考纹理特征索引进行剪枝,确定当前块的一个或多个候选纹理特征索引,可以包括:根据N个参考纹理特征索引中处于第一位置的参考纹理特征索引确定第一个候选纹理特征索引;在N个参考纹理特征索引中除第一位置之外的其他参考纹理特征索引与第i个候选纹理特征索引之间满足第二条件时,根据其他参考纹理特征索引确定第i+1个候选纹理特征索引,以确定当前块的一个或多个候选纹理特征索引;其中,i为大于零且小于N的整数。
需要说明的是,在本申请实施例中,第二条件可以包括:第i个候选纹理特征索引与第i+1个候选纹理特征索引之间的差距满足预设阈值,即相邻两个候选纹理特征索引之间具有一定的差距。或者在每一个候选纹理特征索引对应一个候选变换核组的情况下,也可以说,第i个候选变换核组与第i+1个候选变换核组之间的差距满足预设阈值。
还需要说明的是,在本申请实施例中,N个参考纹理特征索引是按照梯度强度累加值由高向低进行排序的。在这种情况下,根据N个参考纹理特征索引中处于第一位置的参考纹理特征索引确定第一个候选纹理特征索引,可以包括:根据N个参考纹理特征索引中梯度强度累加值最大的参考纹理特征索引确定第一个候选纹理特征索引。
还需要说明的是,在本申请实施例中,预设阈值可以用THR表示,第i个候选纹理特征索引可以用candFeature(i)表示,第i+1个候选纹理特征索引可以用candFeature(i+1)表示。在一种具体的实施例中,第i个候选纹理特征索引与第i+1个候选纹理特征索引之间的差距满足预设阈值,可以包括:candFeature(i+1)+THR<candFeature(i)||candFeature(i+1)-THR>candFeature(i)。
还需要说明的是,在本申请实施例中,THR的取值可能是3、4、5、6等。这样,对于多个相邻的角度累计值都很高的情况,这种剪枝方法会优先选择具有一定区分度的候选纹理特征索引。在一种可能的实施例中,THR等于0或者说不设置THR,也就是只需要这些候选纹理特征索引不同就可以。在另一种可能的实施例中,这里也可以不执行剪枝操作,即剪枝并不是必须的操作步骤。
还可以理解地,在本申请实施例中,上述方法没有考虑到DIMD、TIMD、SGPM等模式本身所导出的一些帧内预测模式。因此,在一些实施例中,该方法还可以包括:确定当前块的预测模式推导的一个或多个帧内预测模式;根据一个或多个帧内预测模式确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,该方法还包括:在第一候选列表未填满时,确定当前块的预设纹理特征索引;根据预设纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,该方法还包括:在第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;根据候选样本,确定当前块的一个或多个候选纹理特征索引;根据一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
也就是说,在本申请实施例中,考虑到DIMD、TIMD、SGPM等模式本身所导出的一些模式,比如DIMD本身会导出一个或几个帧内预测模式用来加权,TIMD本身也会导出一个或几个帧内预测模式用来加权,SGPM不仅有2个帧内预测模式,还有一个“划分”模式也可以找到对应的帧内预测模式,而且残差往往出现在“划分”的边界区域。因此,一种可能的实现方式是根据DIMD、TIMD、SGPM等模式本身所导出的帧内预测模式和上述方法导出的候选纹理特征索引来确定第一候选列表;或者另一种可能的实现方式是根据DIMD、TIMD、SGPM等模式本身所导出的帧内预测模式和默认的纹理特征索引来确定第一候选列表。
另外,在本申请实施例中,由于DIMD、TIMD、SGPM都可以导出不只一个帧内预测模式,因而对这几个模式也可以优先使用各模式自己导出的多个模式确定候选纹理特征索引,在各模式自己导出的模式不能填满所有的候选纹理特征索引时,一种可能的方法是加入默认的纹理特征索引,一种可能的方法是使用上述长度为N的排序的候选纹理特征列表确定候选纹理特征索引。
示例性地,在当前块的预测模式为DIMD模式时,如果它在预测时使用多个帧内预测模式的预测值加权,那么可以把用来加权多个帧内预测模式依次尝试确定为候选纹理特征索引。
示例性地,在当前块的预测模式为TIMD模式时,如果它在预测时使用多个帧内预测模式的预测值加权,那么可以把用来加权多个帧内预测模式依次尝试确定为候选纹理特征索引。
示例性地,在当前块的预测模式为SGPM模式时,可以将它所导出的“划分”模式对应的帧内预测模式,以及2个用来预测的帧内预测模式依次尝试确定为候选纹理特征索引。
在又一种可能的实现方式中,对于确定当前块的第一候选列表,该方法可以包括:根据候选纹理特征列表中处于第一位置的参考纹理特征索引确定第一个候选变换核组;在候选纹理特征列表中除第一位置之外的所有参考纹理特征索引与第一个候选变换核组之间均不满足第二条件时,根据候选纹理特征列表中处于第二位置的参考纹理特征索引确定第二个候选变换核组;根据第一个候选变换核组和第二个候选变换核组,确定第一候选列表。
需要说明的是,在本申请实施例中,根据候选纹理特征列表中处于第一位置的参考纹理特征索引确定第一个候选变换核组,具体可以是:根据候选纹理特征列表中梯度强度累加值最大的参考纹理特征索引确定第一个候选变换核组。
还需要说明的是,在本申请实施例中,以第一候选列表指示两个候选变换核组为例,第一个候选变换核组可以用candFeature0表示,第二个候选变换核组可以用candFeature1表示。示例性地,以梯度强度累计值最大的帧内预测模式为第一个候选变换核组candFeature0,那么在选择第二个候选变换核组candFeature1时,就需要candFeature1和candFeature0具有一定差距,例如candFeature1+THR<candFeature0||candFeature1-THR>candFeature0。如果检查完所有的N-1个参考纹理特征索引中仍然没有符合要求的,那么可以根据N个参考纹理特征索引中处于第二位置的参考纹理特征索引确定第二个候选变换核组candFeature1。也就是说,这时候可以将梯度强度累加值最大的前两个参考纹理特征索引对应的候选变换核组添加到第一候选列表。
在又一种可能的实现方式中,对于确定当前块的第一候选列表,该方法可以包括:根据候选纹理特征列表中处于第一位置的参考纹理特征索引确定第一个候选变换核组;在候选纹理特征列表中除第一位置之外的所有参考纹理特征索引与第一个候选变换核组之间均不满足第二条件时,根据预设纹理特征索引确定第二个候选变换核组;根据第一个候选变换核组和第二个候选变换核组,确定第一候选列表。
需要说明的是,在本申请实施例中,如果检查完所有的N-1个参考纹理特征索引中仍然没有符合要求的,那么也可以确定当前块的默认的纹理特征索引,然后将梯度强度累加值最大的参考纹理特征索引和默认的纹理特征索引对应的候选变换核组添加到第一候选列表。
在又一种可能的实现方式中,对于确定当前块的第一候选列表,该方法可以包括:根据当前块的预测模式推导的第一个帧内预测模式确定第一个候选变换核组;在候选纹理特征列表中的第一参考纹理特征索引与第一个候选变换核组之间满足第二条件时,根据第一参考纹理特征索引确定第二个候选变换核组;根据第一个候选变换核组和第二个候选变换核组,确定第一候选列表。
需要说明的是,在本申请实施例中,第一候选列表不仅考虑了根据候选样本的水平梯度值和竖直梯度值推导出的候选纹理特征索引,而且考虑了根据当前块的预测模式本身所推导出的帧内预测模式。下面以几种示例进行详细说明。
示例性地,在当前块的预测模式为DIMD模式时,将DIMD模式导出的第一个帧内预测模式作为candFeature0。
示例性地,在当前块的预测模式为TIMD模式时,将TIMD模式导出的第一个帧内预测模式作为candFeature0。
示例性地,在当前块的预测模式为SGPM模式时,将SGPM“划分”模式对应的帧内预测模式作为candFeature0。
然后再按照上述方法确定candFeature1。例如尝试将N个参考纹理特征索引中的参考纹理特征索引,从第一位置开始如果某一个参考纹理特征索引符合THR的限制,那么可以将它作为candFeature1。
还需要说明的是,在本申请实施例中,对于IBC、ITMP等模式,可以使用预测块或者预测块加模板来推导候选纹理特征索引(或帧内预测模式)。
还需要说明的是,在本申请实施例中,如果MTSS技术推导N个参考纹理特征索引所使用的候选样本和DIMD模式使用的候选样本相同,那么DIMD推导出来的第一个帧内预测模式和MTSS技术推导出来的第一个参考纹理特征索引是相同的。
S3302,根据变换核组,确定当前块的变换核。
需要说明的是,在确定出当前块的变换核组之后,可以进一步确定出当前块的变换核。在一些实施例中,该方法可以包括:确定变换核组包括的至少两个候选变换核;根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果;在至少两个候选变换核各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核确定为当前块的变换核。
还需要说明的是,在第一候选列表指示至少两个变换核组包括的变换核时,该方法还可以包括:确定第一候选列表指示的至少两个候选变换核;根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果;在至少两个候选变换核各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核确定为当前块的变换核。
在本申请实施例中,这里的代价计算可以根据率失真优化(Rate Distortion Optimization,RDO)的代价结果进行确定,也可以根据绝对误差和(Sum of Absolute Difference,SAD)的代价结果进行确定,甚至也可以根据绝对变换差和(Sum of Absolute Transformed Difference,SATD)的代价结果进行确定,但是这里并不作任何限定。
在一种具体的实施例中,根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果,可以包括:基于第一候选变换核对当前块的残差块进行变换与量化,确定当前块的第一候选量化系数,并对第一候选量化系数进行熵编码处理,确定第一候选变换核的第一代价值;对第一候选量化系数进行反量化与反变换,确定当前块的第一候选残差块,并根据第一候选残差块确定当前块的第一候选预测块;根据第一候选预测块与当前块的原始图像进行代价计算,确定第一候选变换核的第二代价值;根据第一候选变换核的第一代价值和第二代价值,确定第一候选变换核对应的代价结果;其中,第一候选变换核为至少两个候选变换核中的任意一个。
在本申请实施例中,对于当前块的变换核索引来说,变换核索引可以用于指示当前块的变换核在当前块的变换核组或者第一候选列表中的编号。其中,当前块的变换核索引可以为正整数,例如1,2,3,4,5,6等等。该变换核索引序号可以是直接写入码流,或者也可以是通过第二语法元素的取值写入码流。
在一种可能的实现方式中,确定当前块的变换核索引;对当前块的变换核索引进行编码处理,将所得到的编码比特写入码流。
在另一种可能的实现方式中,确定第二语法元素的取值;其中,第二语法元素用于指示当前块是否使用第一变换模式以及对应使用的变换核索引;对第二语法元素的取值进行编码处理,将所得到的编码比特写入码流。
还需要说明的是,在本申请实施例中,第一变换模式可以为LFNST/NSPT,第二语法元素可以用lfnst_idx表示。其中,第二语法元素可以用于指示当前块是否使用第一变换模式,以及在当前块使用第一变换模式时对应的变换核索引。在这种情况下,第二语法元素的取值可以为0,1,2,3,4,5,6等等。
在一种具体的实施例中,若第二语法元素的取值为第三值,则确定当前块不使用第一变换模式;若第二语法元素的取值为第四值,则确定当前块使用第一变换模式以及对应使用的变换核索引。其中,第三值可以设置为0,第四值可以设置为非0,例如1,2,3,4,5,6等等。
也就是说,在本申请实施例中,对于LFNST/NSPT,当前块的变换核索引也可以用lfnst_idx表示。其中,lfnst_idx为0表示当前块不使用LFNST/NSPT,ECM中LFNST/NSPT的每个变换核组有3个变换核,所以lfnst_idx的值为1或2或3即表示当前块使用LFNST/NSPT的所选变换核组的第1个变换核或第2个变换核或第3个变换核。
在本申请实施例中,如果当前块的预测模式是某种特殊的帧内预测模式,那么它可选的变换核组不 止一个。示例性地,它可选的变换核组是2个,那么lfnst_idx可能的值是0,1,2,3,4,5,6。其中1,2,3对应第一个变换核组的3个变换核,4,5,6对应第二个变换核组的3个变换核。
在一种具体的实施例中,lfnst_idx的二元符号对应表如前述表7所示。其中,第三个二元符号,即BinIdx为2的二元符号,也可以理解为选择第一个候选变换核组或第二个候选变换核组。
可以理解地,在本申请实施例中,这里的变换核组可以为第一候选列表指示的至少两个候选变换核组中的其中一个。另外,在本申请实施例中,当前块的变换核组索引可以用lfnst_feature_idx表示,或者用lfnst_set_idx表示。其中,当前块的变换核组索引可以为大于或等于零的整数,例如0,1,2等。示例性地,若lfnst_feature_idx的取值等于0,则指示选择第一个候选变换核组作为当前块的变换核组;若lfnst_feature_idx的取值等于1,则指示选择第二个候选变换核组作为当前块的变换核组。
需要说明的是,在确定当前块的变换核组之后,对于确定当前块的变换核,可以包括:确定变换核组包括的至少两个候选变换核;根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果;在至少两个候选变换核各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核确定为当前块的变换核。
在一种具体的实施例中,根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果,可以包括:基于第一候选变换核对当前块的残差块进行变换与量化,确定当前块的第一候选量化系数,并对第一候选量化系数进行熵编码处理,确定第一候选变换核的第一代价值;对第一候选量化系数进行反量化与反变换,确定当前块的第一候选残差块,并根据第一候选残差块确定当前块的第一候选预测块;根据第一候选预测块与当前块的原始图像进行代价计算,确定第一候选变换核的第二代价值;根据第一候选变换核的第一代价值和第二代价值,确定第一候选变换核对应的代价结果;其中,第一候选变换核为至少两个候选变换核中的任意一个。
在本申请实施例中,对于当前块的变换核索引来说,变换核索引可以用于指示当前块的变换核在当前块的变换核组中的编号。其中,当前块的变换核索引可以为正整数,例如1,2,3等。该变换核索引可以是直接写入码流,或者也可以是通过第二语法元素的取值写入码流。
在一种可能的实现方式中,确定当前块的变换核索引;对当前块的变换核索引进行编码处理,将所得到的编码比特写入码流。
在另一种可能的实现方式中,确定第二语法元素的取值;其中,第二语法元素用于指示当前块是否使用第一变换模式以及对应使用的变换核索引;对第二语法元素的取值进行编码处理,将所得到的编码比特写入码流。
还需要说明的是,在本申请实施例中,第一变换模式可以为LFNST/NSPT,第二语法元素可以用lfnst_idx表示。其中,第二语法元素可以用于指示当前块是否使用第一变换模式,以及在当前块使用第一变换模式时对应的变换核索引。在这种情况下,第二语法元素的取值可以为0,1,2,3等等。
示例性地,在这种情况下,对于LFNST/NSPT,当前块的变换核索引也可以用lfnst_idx表示。其中,VVC中,lfnst_idx可以有3个值即0,1,2。其中lfnst_idx为0表示当前块不使用LFNST,VVC中LFNST的每个变换核组有2个变换核,所以lfnst_idx的值为1或2即表示当前块使用LFNST的所选变换核组的第1个变换核或第2个变换核。在已有的ECM中,lfnst_idx可以有4个值即0,1,2,3。其中lfnst_idx为0表示当前块不使用LFNST/NSPT,ECM中LFNST/NSPT的每个变换核组有3个变换核,所以lfnst_idx的值为1或2或3即表示当前块使用LFNST/NSPT的所选变换核组的第1个变换核或第2个变换核或第3个变换核。
在一种具体的实施例中,lfnst_idx的二元符号对应表如前述表8所示。
还需要说明的是,在本申请实施例中,在编码当前块的变换核组索引时,该方法还可以包括:在当前块使用第一变换模式时,对当前块的变换核组索引进行编码处理,将所得到的编码比特写入码流。也就是说,在当前块使用MTSS技术,且lfnst_idx>0时,则执行对当前块的变换核组索引进行编码处理,将所得到的编码比特写入码流的步骤。
S3303,确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数。
S3304,对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
需要说明的是,在本申请实施例中,在对当前块的变换系数进行编码处理时,该方法可以包括:对当前块的变换系数进行量化,确定当前块的量化系数;对当前块的量化系数进行编码处理,将所得到的编码比特写入码流。
还需要说明的是,在本申请实施例中,编码端对残差块的“变换”,也可称为“正变换”,具体是指从空间域到频率域的变换,以去除残差的相关性。需要注意的是,标准如果只规定解码,那么标准文本中的“变换”是解码的部分,具体指的是本文中的“反变换”。
还需要说明的是,在本申请实施例中,参见图34,对于步骤S3303来说,该方法可以包括:
S3401,对当前块进行帧内预测,确定当前块的预测块。
S3402,根据当前块的初始块和当前块的预测块,确定当前块的残差块。
S3403,根据变换核对当前块的残差块进行变换,确定当前块的变换系数。
需要说明的是,在本申请实施例中,对于步骤S3401~S3402来说,可以是与步骤S3301~S3302并行操作,也可以是在步骤S3301~S3302之前执行,或者还可以在步骤S3301~S3303之后执行,这里对步骤的先后顺序并不作具体限定。
还需要说明的是,在本申请实施例中,在确定当前块的预测块之后,可以对当前块的初始块和当前块的预测块进行减法运算,从而确定出当前块的残差块。
还需要说明的是,在本申请实施例中,在根据变换核对当前块的残差块进行变换,确定当前块的变换系数时,可以包括:根据变换核对当前块的残差块进行不可分离基础变换,确定当前块的变换系数;或者,对当前块的残差块进行离散余弦变换,确定当前块的变换块;以及根据变换核对当前块的变换块进行低频不可分离变换,确定当前块的变换系数。
在一种具体的实施例中,若当前块的尺寸参数满足第一条件,则根据变换核对当前块的残差块进行不可分离基础变换,确定当前块的变换系数;若当前块的尺寸参数满足第二条件,则对当前块的残差块进行离散余弦变换,确定当前块的变换块;以及根据变换核对当前块的变换块进行低频不可分离变换,确定当前块的变换系数。
在这里,当前块的尺寸参数满足第一条件,包括:当前块的尺寸参数较小,例如当前块的尺寸参数小于某一阈值。也就是说,针对较小尺寸的块,这里使用NSPT的变换核,即根据变换核对当前块的残差块进行NSPT变换,确定当前块的变换系数。
在这里,当前块的尺寸参数满足第二条件,包括:当前块的尺寸参数较大,例如当前块的尺寸参数大于某一阈值。也就是说,针对较大尺寸的块,这里使用LFNST的变换核,即先对当前块的残差块进行DCT2的基础变换,然后根据变换核对当前块的变换块进行LFNST变换,确定当前块的变换系数。
简单来说,编码器在确定出预测块后,根据预测块推导出候选纹理特征索引,然后根据候选纹理特征索引来确定NSPT/LFNST的变换核组。如果一个变换核组有多个变换核可选,编码器对该变换核组中的每一个变换核进行尝试。
如果是NSPT变换,可以使用NSPT对残差块进行正变换得到变换系数,对变换系数进行量化得到量化系数,然后对量化系数熵编码。通过熵编码可以得到该变换核下的码流中的开销的代价。对量化系数进行反量化得到解码的变换系数,对解码的变换系数进行反NSPT变换得到解码的残差块。解码的变换系数和原始的变换系数可能是不同的,因为量化是有损的,同样地,解码的残差块和原始的残差块也可能是不同的。根据解码的残差块和预测块得到重建块。根据重建块和当前块的原始图像可以得到失真的代价。使用当前NSPT的变换核进行编码的代价为开销的代价加上失真的代价。比较几个变换核的代价,选最小者作为当前块的NSPT的最佳选择。
如果是LFNST变换,可以使用DCT2对残差块进行正变换,再用LFNST进行正变换得到变换系数,对变换系数进行量化得到量化系数,然后对量化系数熵编码。通过熵编码可以得到该变换核下的码流中的开销的代价。对量化系数进行反量化得到解码的变换系数,对解码的变换系数进行反LFNST变换,再进行反DCT2变换得到解码的残差块。解码的变换系数和原始的变换系数可能是不同的,因为量化是有损的,同样地,解码的残差块和原始的残差块也可能是不同的。根据解码的残差块和预测块得到重建块。根据重建块和当前块的原始图像可以得到失真的代价。使用当前NSPT的变换核进行编码的代价为开销的代价加上失真的代价。比较几个变换核的代价,选最小者作为当前块的NSPT的最佳选择。
除此之外,在一些实施例中,在当前块不使用MTSS技术时,该方法还可以包括:确定当前块的纹理特征索引;根据纹理特征索引,确定当前块的变换核;确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
需要说明的是,在本申请实施例中,根据纹理特征索引,确定当前块的变换核,可以包括:根据纹理特征索引,确定当前块的变换核组;根据变换核组,确定当前块的变换核。
在一些实施例中,该方法还包括:在当前块的样本个数小于最小样本阈值时,执行确定当前块的纹理特征索引的步骤。
在一些实施例中,当前块的尺寸包括高度和宽度;该方法还包括:在当前块的高度或宽度小于最小尺寸阈值时,执行确定当前块的纹理特征索引的步骤。
在一些实施例中,该方法还包括:在当前块的预测模式为第一预测模式集合之外的预测模式时,执行确定当前块的纹理特征索引的步骤。
需要说明的是,在本申请实施例中,在当前块的预测模式为第一预测模式集合之外的预测模式时,当前块的预测模式可以为第四预测模式集合中的其中一项。其中,第四预测模式集合至少包括:DC模 式、PLANAR模式和角度预测模式。
还需要说明的是,在本申请实施例中,第四预测模式集合中的模式为比较简单的预测模式。示例性地,可以做一个归类,DC模式、PLANAR模式和各种角度预测模式归为第四预测模式集合,DIMD模式、TIMD模式、MIP模式、EIP模式、SGPM模式、ITMP模式和IBC模式等归为第一预测模式集合。
示例性地,第一预测模式集合也可能只包含DIMD模式、TIMD模式、MIP模式、EIP模式、SGPM模式、ITMP模式和IBC模式中的一个或几个。如第一预测模式集合包含DIMD模式、TIMD模式和SGPM模式。其他模式都属于第四预测模式集合。所述第四预测模式集合的lfnst_idx的编码方法和相关技术的方法相同,而第一预测模式集合的lfnst_idx编码方法和相关技术的方法不同。
示例性地,LFNST/NSPT的每个变换核组有3个变换核,使用第四预测模式集合的块只可能选择一个变换核组,使用第一预测模式集合的块可以选择2个变换核组,每个变换核组有3个变换核。如果当前块的预测模式属于第四预测模式集合,它只可能选择一种纹理特征索引,那么lfnst_idx可能的值是0,1,2,3。如果当前块的预测模式属于第一预测模式集合,那么lfnst_idx可能的值是0,1,2,3,4,5,6。其中1,2,3对应第一个纹理特征索引对应的变换核组的3个变换核,4,5,6对应第二个纹理特征索引对应的变换核组的3个变换核。在这里,使用第四预测模式集合的块的lfnst_idx的二元符号对应表如表8所示,使用第一预测模式集合的块的lfnst_idx的二元符号对应表如表7所示。
还可以理解地,本申请实施例可以使用一个高层语法(high level syntax)控制本技术方案的开关。在一些实施例中,该方法还包括:确定第三语法元素的取值;对第三语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术。另外,第三语法元素可以用sps_mtss_enabled_flag表示,而且第三语法元素为序列参数集(Sequence Parameter Set,SPS)中的语法元素。
在本申请实施例中,如果当前序列允许使用MTSS技术,则确定第三语法元素的取值为第五值;如果当前序列不允许使用MTSS技术,则确定第三语法元素的取值为第六值。
进一步地,在一些实施例中,该方法还包括:在当前序列允许使用MTSS技术时,执行确定当前块的变换核组的步骤。其中,当前序列包括当前块。
需要说明的是,在本申请实施例中,第五值与第六值不同,而且第五值和第六值可以是参数形式,也可以是数字形式。具体地,第三语法元素可以是写入在概述(profile)中的参数,也可以是一个标志(flag)的取值,这里对此不作具体限定。示例性地,第五值可以为1,第六值可以为0;或者,第五值可以为0,第六值可以为1;或者,第五值可以为true,第六值可以为false;或者,第五值可以为false,第六值可以为true。在一种具体的实施例中,第五值为1,第六值为0。
也就是说,本申请实施例可以使用一个高层语法(high level syntax)控制本技术方案的开关。比如使用一个序列级的flag,如在序列参数集中增加语法元素sps_mtss_enabled_flag。如果sps_mtss_enabled_flag的取值为1,当前序列允许使用MTSS技术;如果sps_mtss_enabled_flag的取值为0,当前序列不允许使用MTSS技术。其中,在允许使用MTSS技术的情况下,更具体地,允许在块级(CU或TU)解码本技术方案所述的纹理特征索引、LFNST/NSPT变换核索引的方法;如果当前序列不允许使用MTSS技术,更具体地,在块级(CU或TU)不会使用本技术方案所述的纹理特征索引、LFNST/NSPT变换核索引的方法。
在一些实施例中,该方法还包括:确定第三语法元素的取值和第四语法元素的取值;其中,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,第四语法元素用于指示当前图像是否允许使用多变换核组选择技术;对第三语法元素的取值和第四语法元素的取值进行编码处理,将所得到的编码比特写入码流。
进一步地,在一些实施例中,该方法还包括:在当前序列允许使用MTSS技术时,确定当前图像是否允许使用MTSS技术;在当前图像允许使用MTSS技术时,执行确定当前块的变换核组的步骤。
在本申请实施例中,当前序列可以包括当前图像,且当前图像包括当前块。其中,第四语法元素可以用ph_inter_lfnst_nspt_enabled_flag表示,而且第四语法元素为图像级别的语法元素。
在本申请实施例中,如果当前图像允许使用MTSS技术,则确定第四语法元素的取值为第五值;如果当前图像不允许使用MTSS技术,则确定第四语法元素的取值为第六值。
在一些实施例中,该方法还包括:确定第三语法元素的取值和第五语法元素的取值;其中,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,第五语法元素用于指示当前片是否允许使用多变换核组选择技术;对第三语法元素的取值和第五语法元素的取值进行编码处理,将所得到的编码比特写入码流。
进一步地,在一些实施例中,该方法还包括:在当前序列允许使用MTSS技术时,确定当前片是 否允许使用MTSS技术;在当前片允许使用MTSS技术时,执行确定当前块的变换核组的步骤。
在本申请实施例中,当前序列可以包括当前片,且当前片包括当前块。其中,第五语法元素可以用sh_inter_lfnst_nspt_enabled_flag表示,而且第五语法元素为片(slice)级别的语法元素。
还需要说明的是,在本申请实施例中,如果当前片允许使用MTSS技术,则确定第五语法元素的取值为第五值;如果当前片不允许使用MTSS技术,则确定第五语法元素的取值为第六值。
在本申请实施例中,第五值与第六值不同,而且第五值和第六值可以是参数形式,也可以是数字形式。具体地,第四语法元素或第五语法元素可以是写入在概述(profile)中的参数,也可以是一个标志(flag)的取值,这里对此不作具体限定。示例性地,第五值可以为1,第六值可以为0;或者,第五值可以为0,第六值可以为1;或者,第五值可以为true,第六值可以为false;或者,第五值可以为false,第六值可以为true。在一种具体的实施例中,第五值为1,第六值为0。
也就是说,在本申请实施例中,除了序列级的语法元素sps_mtss_enabled_flag之外,当然也可以使用其他级别的语法来实现更灵活的控制,如图像参数集(Picture Parameter Set,PPS)的flag,或图像头(picture header)或片头(slice header)的flag等。比如先在SPS确定当前序列是否可以使用MTSS技术,如果当前序列使用MTSS技术,那么设置一个片头(slice header)的sh_inter_lfnst_nspt_enabled_flag确定当前slice是否使用MTSS技术,提供更高的灵活性。
还可以理解地,对于最小样本阈值来说,在一种可能的实现方式中,可以包括:确定最小样本阈值;对最小样本阈值进行编码处理,将所得到的编码比特写入码流。或者,在另一种可能的实现方式中,可以包括:在当前序列允许使用多变换核组选择技术时,确定第六语法元素的取值;其中,第六语法元素用于指示最小样本阈值;对第六语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,最小样本阈值可以是直接写入码流,或者以第六语法元素的形式写入码流。另外,第六语法元素可以用sps_mtss_min_pix表示,其可以为序列级别的语法元素。
在以第六语法元素的形式写入码流时,第六语法元素的取值可以与最小样本阈值相等,或者也可以是按照一定的映射规则进行编码。示例性地,若最小样本阈值等于32,则确定第六语法元素的取值等于0;若最小样本阈值等于64,则确定第六语法元素的取值等于1;若最小样本阈值等于256,则确定第六语法元素的取值等于2。这样,在将第六语法元素的取值写入码流之后,后续解码端在解码获得第六语法元素的取值等于2时,那么可以确定最小样本阈值等于256。
也就是说,在本申请实施例中,可以使用一个高层语法(high level syntax)设置一个应用MTSS技术的最小样本阈值,例如sps_mtss_min_pix。当sps_mtss_enabled_flag的取值为1时,编码器编码sps_mtss_min_pix以确定应用MTSS技术的最小样本阈值。用高层语法可以根据需求在编码复杂度和压缩效率之间进行取舍。也就是说,这种方法下编码器需要支持所有sps_mtss_min_pix可能的情况,但是编码器可以配置编码当前码流需要的sps_mtss_min_pix。例如编码一个码流时,如果需要更好的压缩效率而不特别注重编码时间,可以给sps_mtss_min_pix设置一个比较小的值,如16。如果特别注重编码时间而且可以损失一定的压缩效率,可以给sps_mtss_min_pix设置一个比较大的值,如256。
还可以理解地,对于最小尺寸阈值来说,在一种可能的实现方式中,可以包括:确定最小尺寸阈值;对最小尺寸阈值进行编码处理,将所得到的编码比特写入码流。或者,在另一种可能的实现方式中,可以包括:在当前序列允许使用多变换核组选择技术时,确定第七语法元素的取值;其中,第七语法元素用于指示最小尺寸阈值;对第七语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,最小尺寸阈值可以是直接写入码流,或者以第七语法元素的形式写入码流。另外,第七语法元素可以用sps_mtss_min_size表示,其可以为序列级别的语法元素。
在以第七语法元素的形式写入码流时,第七语法元素的取值可以与最小尺寸阈值相等,或者也可以是按照一定的映射规则进行编码。示例性地,若最小尺寸阈值等于4,则确定第七语法元素的取值等于0;若最小样本阈值等于16,则确定第七语法元素的取值等于1。这样,在将第七语法元素的取值写入码流之后,后续解码端在解码获得第七语法元素的取值等于1时,那么可以确定最小尺寸阈值等于16。
也就是说,在本申请实施例中,可以使用一个高层语法(high level syntax)设置一个应用MTSS技术的最小尺寸阈值,例如sps_mtss_min_size。当sps_mtss_enabled_flag的取值为1时,编码器编码sps_mtss_min_size以确定应用MTSS技术的最小尺寸阈值。用高层语法可以根据需求在编码复杂度和压缩效率之间进行取舍。也就是说,这种方法下编码器需要支持所有sps_mtss_min_size可能的情况,但是编码器可以配置编码当前码流需要的sps_mtss_min_size。比如编码一个码流时,如果需要更好的压缩效率而不特别注重编码时间,可以给sps_mtss_min_size设置一个比较小的值,如4。如果特别注重编码时间而且可以损失一定的压缩效率,可以给sps_mtss_min_size设置一个比较大的值,如16。
在本申请的又一实施例中,本申请实施例提供了一种码流,其中,码流是根据待编码信息进行比特编码生成的;其中,待编码信息包括下述至少一项:当前块的量化系数、所述当前块的变换核索引、所 述当前块的变换核组索引、最小样本阈值、最小尺寸阈值、第一语法元素的取值、第二语法元素的取值、第三语法元素的取值、第四语法元素的取值、第五语法元素的取值、第六语法元素的取值和第七语法元素的取值。
在本申请实施例中,第一语法元素用于指示当前块的变换核组索引,第二语法元素用于指示当前块是否使用第一变换模式以及对应使用的变换核索引,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,第四语法元素用于指示当前图像是否允许使用多变换核组选择技术,第五语法元素用于指示当前片是否允许使用多变换核组选择技术,第六语法元素用于指示最小样本阈值,第七语法元素用于指示最小尺寸阈值。示例性地,以第二语法元素为例,若第二语法元素的取值为0,则确定当前块不使用第一变换模式;若第二语法元素的取值为非零,则确定当前块使用第一变换模式以及对应使用的变换核索引。例如,若第二语法元素的取值为1,则确定当前块使用第1个变换核;若第二语法元素的取值为2,则确定当前块使用第2个变换核;若第二语法元素的取值为4,则确定当前块使用第4个变换核等,这里不作任何限定。
本申请实施例提供了一种编码方法,具体为一种帧内LFNST/NSPT多角度选择方案。首先确定当前块的变换核组;然后根据变换核组,确定当前块的变换核;再确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;最后对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。这样,根据多变换核组选择技术确定当前块的变换核组,然后从中确定当前块的变换核。如此,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。
在本申请的又一实施例中,基于前述实施例所述的编解码方法,VVC和ECM中,LFNST/NSPT变换核索引可以用语法元素lfnst_idx表示。VVC中,lfnst_idx可以有3个值即0,1,2。其中lfnst_idx为0表示当前块不使用LFNST,VVC中LFNST的每个变换核组有2个变换核,所以lfnst_idx的值为1或2即表示当前块使用LFNST的所选变换核组的第1个变换核或第2个变换核。在现在的ECM中,lfnst_idx可以有4个值即0,1,2,3。其中lfnst_idx为0表示当前块不使用LFNST/NSPT,ECM中LFNST/NSPT的每个变换核组有3个变换核,所以lfnst_idx的值为1或2或3即表示当前块使用LFNST/NSPT的所选变换核组的第1个变换核或第2个变换核或第3个变换核。
在MTSS技术中,如果当前块的帧内预测模式是某种特殊的帧内预测模式,那么它可选的变换核组不止一个。一个例子是它可选的变换核组是2个。这些特殊的帧内预测模式包含但不限于DIMD、TIMD、MIP、EIP、SGPM、ITMP、IBC等模式。
在一个具体的实施例中,这里做一个归类,DC、PLANAR、各种角度预测模式归为第一帧内预测模式集合(相当于前述实施例的“第四预测模式集合”),DIMD、TIMD、MIP、EIP、SGPM、ITMP、IBC等归为第二帧内预测模式集合(相当于前述实施例的“第一预测模式集合”)。
LFNST/NSPT每个变换核组有3个变换核,使用第一帧内预测模式集合的块只可能选择一个变换核组,使用第二帧内预测模式集合的块可以选择2个变换核组,每个变换核组有3个变换核。如果当前块的预测模式属于第一帧内预测模式集合,它只可能选择一种纹理特征,那么lfnst_idx可能的值是0,1,2,3。如果当前块的预测模式属于第二帧内预测模式集合,那么lfnst_idx可能的值是0,1,2,3,4,5,6。其中1,2,3对应第一纹理特征索引对应的变换核组的3个变换核,4,5,6对应第二个纹理特征索引对应的变换核组的3个变换核。
在本申请实施例中,使用第一帧内预测模式集合的块的lfnst_idx的二元符合对应表如前述表8所示,使用第二帧内预测模式集合的块的lfnst_idx的二元符号对应表如前述表7所示。其中,第三个二元符号,即BinIdx为2的二元符号,也可以理解为选择第一个候选纹理特征索引或第二个候选纹理特征索引,或者说选择第一个候选变换核组或第二个候选变换核组。
在另一个具体的实施例中,第二帧内预测模式集合也可能只包含DIMD、TIMD、MIP、EIP、SGPM、ITMP、IBC中的一个或几个。如第二帧内预测模式集合包含DIMD、TIMD和SGPM。其他模式都属于第一帧内预测模式集合。所述第一帧内预测模式集合的lfnst_idx编码方法和相关技术的方法相同,而第二帧内预测模式集合的lfnst_idx编码方法和相关技术的方法不同。
还需要说明的是,在本申请实施例中,当然也可以设置一个专门的语法元素,比如lfnst_feature_idx,lfnst_feature_idx可能的值是0或1,指示选择哪个候选纹理特征,或者说选择哪个候选变换核组。如果当前块的帧内预测模式属于第二帧内预测模式集合,且lfnst_idx>0,则解析lfnst_feature_idx,结果和上述的实施方式是等价的。或者,该语法元素也可以称为lfnst_set_idx,即候选的变换核组索引。
可以理解地,对于候选纹理特征索引的导出来说,MTSS技术可以使用当前块左侧和上侧的已重建 区域来推导候选纹理特征索引,和DIMD已有的做法类似。因为左侧和上侧的已重建区域虽然不是当前块但是和当前块相邻,比如说纹理有相连的情况,因而一定程度上可以用来估计当前块的纹理。另外一种可能是使用当前块的预测块来推导候选纹理特征索引,另外一种可能是同时使用当前块的预测块和当前块左侧和上侧的已重建区域,这样可用来推导候选纹理特征索引的样本更多。
在本申请实施例中,这里的候选纹理特征索引(或者说,候选纹理特征)也可以直接称之为候选变换核组。
一种推导方式是对所选区域中的全部或部分样本计算其梯度,一般来说可以计算水平和竖直的梯度,可以使用索贝尔(sobel)算子来计算梯度。对某一个样本,根据其水平和竖直的梯度推测该样本的纹理方向,比如水平梯度非零,竖直梯度为零,那该点的纹理是竖直方向的。反之,水平梯度为零,竖直梯度非零,那该点的纹理是水平方向的。比如水平梯度和竖直梯度相等且不为零,那该点的纹理是45度的。当然,还有很多其他情况是水平梯度和竖直梯度都不为零,可以根据它们的比值确定该样本的纹理的方向。这样可以根据梯度对应到对应的帧内预测模式。构建一个帧内预测模式统计表,将每一个计算的样本的梯度强度累加到统计表中对应的帧内预测模式项中。梯度统计完成后,按累计的梯度强度从大到小的顺序对帧内预测模式进行排序。为了复杂度的考虑,可以只维护一个长度为N的列表,N的长度可以是2,3,4,5…10等。
在一种可能的实现方式中,一个索贝尔算子的示例如下。
水平梯度的算子:
竖直梯度的算子:
假设预测块的样本位置为(x,y)的样本值为Px,y,则水平梯度gradx竖直梯度grady计算如下:
gradx=Px+1,y-1+2*Px+1,y+Px+1,y+1-Px-1,y-1-2*Px-1,y-Px-1,y+1             (9)
grady=Px-1,y+1+2*Px,y+1+Px+1,y+1-Px-1,y-1-2*Px,y-1-Px+1,y-1            (10)
所述梯度强度记为amp的一个例子是amp=abs(gradx)+abs(grady)。
根据gradx和grady导出虚拟的帧内预测模式可以通过查表来实现。示例性地,如果abs(gradx)等于0且abs(grady)不等于0,有水平方向纹理,对应VVC中的帧内预测模式18。如果abs(grady)等于0且abs(gradx)不等于0,有竖直方向纹理,对应VVC中的帧内预测模式50。如果abs(gradx)和abs(grady)都不等于0:如果abs(gradx)等于abs(grady),且gradx和grady符号相同,对应VVC中的帧内预测模式34;如果abs(gradx)等于2倍的abs(grady),且gradx和grady符号相同,对应VVC中的帧内预测模式40。
其他情况可以按照相同的原则查表确定。DIMD也要做类似的统计导出帧内预测模式,在实现上,可以和DIMD复用部分逻辑。
按上面的方法可以得到一个长度为N且排序的帧内预测模式列表。然后再产生候选纹理特征列表。一个可选的操作是对N个帧内预测模式进行剪枝,或者说排除过于接近的帧内预测模式。以VVC中的帧内预测模式为例,有些相邻的帧内预测模式的角度是非常接近的,如果2个候选纹理特征选了2个非常接近的角度,它们的差别并不明显。所以可以设置一个阈值,这个阈值表征2个帧内预测模式之间的差距。比如以梯度强度累计值最大的帧内预测模式为第一候选纹理特征candFeature0,那么在选择第二候选纹理特征candFeature1时,就需要candFeature1和candFeature0有一定差距,如candFeature1+THR<candFeature0||candFeature1-THR>candFeature0。如果检查完了所有的N-1个帧内预测模式仍然没有符合要求的,一种可能的方法是返回来将长度为N的排序的帧内预测模式列表中的第二个帧内预测模式设为candFeature1。一种可能的方法是加入默认的纹理特征索引。所述THR可能是3,4,5,6等。对于多个相邻的角度累计值都很高的情况,这种方法会优先选择有一定区分度的角度。一种可能的实施例是THR的取值等于0,或者说不设置THR,也就是只要它们不同就可以。
上面这种方法没有考虑DIMD、TIMD、SGPM等模式本身所导出的一些模式。比如DIMD本身会导出一个或几个帧内预测模式用来加权,TIMD本身也会导出一个或几个帧内预测模式用来加权,SGPM不仅有2个帧内预测模式,还有一个“划分”模式也可以找到对应的帧内预测模式,而且残差往往出现在“划分”的边界区域。一种可能的方法是根据DIMD,TIMD,SGPM等模式本身所导出的模式和上述方法导出的模式来确定候选纹理特征索引。
一种具体的实施例中,如下:
对DIMD模式,将DIMD导出的第一个帧内预测模式作为candFeature0。
对TIMD模式,将TIMD导出的第一个帧内预测模式作为candFeature0。
对SGPM模式,将SGPM“划分”模式对应的帧内预测模式作为candFeature0。
然后再按上述方法确定candFeature1。即尝试将长度为N的排序的帧内预测模式列表中的帧内预测模式,从第一个帧内预测模式开始如果某一个帧内预测模式符合THR的限制,则将它作为candFeature1。
对IBC、ITMP等,可以使用预测块,或预测块加模板来推导候选纹理特征索引。
如果MTSS技术推导N个排序的帧内预测模式的列表所使用的样本和DIMD使用的样本相同,那么DIMD推导出来的第一个帧内预测模式和MTSS技术推导出来的第一个帧内预测模式是相同的。
在另一种具体的实施例中,如下:
由于DIMD、TIMD、SGPM都可以导出不止一个帧内预测模式,因而对这几个模式也可以优先使用各模式自己导出的多个模式确定候选纹理特征索引,在各模式自己导出的模式不能填满所有的候选纹理特征索引时,一种可能的方法是加入默认的纹理特征索引,一种可能的方法是使用上述长度为N的排序的帧内预测模式列表确定候选纹理特征索引。
示例性地,对于DIMD模式,如果它在预测时使用多个帧内预测模式的预测值加权,那么可以把用来加权多个帧内预测模式依次尝试确定为候选纹理特征索引。
示例性地,对于TIMD模式,如果它在预测时使用多个帧内预测模式的预测值加权,那么可以把用来加权多个帧内预测模式依次尝试确定为候选纹理特征索引。
示例性地,对于SGPM模式,可以将它所导出的“划分”模式对应的帧内预测模式,以及2个用来预测的帧内预测模式依次尝试确定为候选纹理特征索引。
在又一种具体的实施例中,如下:
对于MIP模式和EIP模式,可以使用MTSS所述的纹理特征的导出方法利用梯度来推导纹理特征,使用的样本可以是预测样本(predicted sample,预测样本的值是预测值),或者当前块周边的重建样本,或者预测样本和重建样本一起使用。为了方便描述,称这种使用梯度推导出来的纹理特征叫推导的纹理特征。
示例性地,对MIP模式和EIP模式,将第一个推导的纹理特征作为它的candFeature0,将第二个推导的纹理特征作为它的candFeature1。
而MIP模式和EIP模式可以大量应用于纹理渐变的块,其特征和PLANAR模式是相似的。因而对MIP模式和EIP模式也可以将PLANAR模式作为其候选纹理特征。
示例性地,对MIP模式和EIP模式,将第一个推导的纹理特征作为它的candFeature0,将PLANAR模式作为它的candFeature1。
当然,MTSS技术也可能不额外使用梯度推导纹理特征。或者说,使用梯度推导纹理特征并不是MTSS多变换核组选择的一个必要条件。
在又一种具体的实施例中,如下:
MTSS技术只应用于DIMD、TIMD、SGPM模式。
对于DIMD模式,由于DIMD模式本身就使用梯度推导纹理特征,而且它可以导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征。
对于TIMD模式,由于TIMD模式本身可以导出多个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征。
对于SGPM模式,由于SGPM模式本身可以导出一个“划分”模式和2个帧内预测模式,因而不需要MTSS技术额外使用梯度推导纹理特征。
更具体地,一个可能的实施例如下:
对DIMD模式,将DIMD模式导出的第一个帧内预测模式作为candFeature0,将DIMD模式导出的第二个帧内预测模式作为candFeature1。
对TIMD模式,将TIMD模式导出的第一个帧内预测模式作为candFeature0,将TIMD模式导出的第二个帧内预测模式作为candFeature0。
对SGPM模式,将SGPM模式的“划分”模式对应的帧内预测模式作为candFeature0,将SGPM模式导出的第一个帧内预测模式作为candFeature0。
当然,如果MTSS可以应用到EIP模式和MIP模式。而EIP模式、MIP模式本身会用梯度推导纹理特征,而MTSS技术可以只增加一个PLANAR模式对应的纹理特征,这样它也不会额外使用梯度推导纹理特征。
另外,如果DIMD、TIMD、SGPM所导出的帧内预测模式中包括ITMP或IBC,或EIP或MIP的 模式,这时候也可以使用PLANAR模式对应的纹理特征作为候选纹理特征。由于DIDM、TIMD、SGPM可以使用2个以上的帧内预测模式进行加权或组合,所述帧内预测模式可能是ITMP或IBC,或EIP或MIP的模式。在遇到这种情况时,可以使用PLANAR模式对应的纹理特征作为候选纹理特征。
还需要说明的是,已有的ECM中LFNST/NSPT都有很多的变换核组,使得相邻的帧内预测模式都对应不同的变换核组。后续标准中的一种折衷方案是适当减少变换核组的数目,使得多个帧内预测模式对应一个变换核组。这类似于VVC中使用4个变换核组,多个类似角度的帧内预测模式对应同一个变换核组。后续标准中的变换核组数量可能是4个或者其他,比如8个,12个等。对于本申请实施例而言,如果多个相邻的帧内预测模式对应一个变换核组,那么在确定候选纹理特征时,要保证各个候选纹理特征不会确定为相同的变换核组。
也就是说,本文中所述纹理特征索引是为了确定变换核组,用纹理特征索引方便理解,也可以直接用变换核组取代纹理特征索引。
还需要说明的是,计算多少个样本的梯度可以根据当前块的大小来确定。因为索贝尔算子要用到当前样本的上下左右一行一列的样本,所以可以设置不计算当前块最外面上下左右一行一列的样本的梯度。如果当前块的尺寸小,可以对所有能计算梯度的样本进行统计。如果当前块的尺寸大,可以对降采样进行统计,比如水平方向和/或竖直方向每2个,或4,或8个样本中的一个样本进行统计。
示例性地,如果当前块的水平或竖直某一个方向的尺寸小于等于8,对该方向上的所有可用的样本进行梯度统计。否则,如果当前块的水平或竖直某一个方向的尺寸小于等于16,对该方向上的每2个样本中的一个进行梯度统计。否则,对该方向上的每4个样本中的一个进行梯度统计。
还需要说明的是,这里还可以对MTSS技术所适用的块的尺寸进行限制。具体地,MTSS技术对解码器的复杂度增加体现在解析语法元素时的不同,这一般认为对解码器的复杂度影响不大。但是对编码器而言,由于增加了变换核组的候选,这会增加编码器的复杂度。具体来说,一个块可能从2个或多个LFNST/NSPT的变换核组中选择一个。而VVC的一个块只有一个可用的变换核组。因而一种解决方案是限制MTSS技术所适用的块的尺寸,从而可以在MTSS技术需要增加很多计算而对压缩效率提升不明显的块尺寸禁用MTSS技术。
示例性地,可以设置一个应用MTSS技术的最小样本阈值,称为MIN_PIX。如果一个块的样本数(宽乘以高)小于MIN_PIX,则这个块不能使用MTSS技术。否则,即这个块的样本数大于或等于MIN_PIX,则这个块可以使用MTSS技术。MIN_PIX的值可能是32,64,256等。
示例性地,可以设置一个应用MTSS技术的最小尺寸阈值,称为MIN_SIZE.如果一个块的宽或高小于MIN_SIZE,则这个块不能使用MTSS技术。否则,即这个块的宽和高都大于或等于MIN_SIZE,则这个块可以使用MTSS技术。MIN_SIZE的值可能是8等。
还需要说明的是,可以使用一个高层语法(high level syntax)控制本技术方案的开关。示例性地,使用一个序列级的语法元素flag,例如在序列参数集(Sequence Parameter Set,SPS)中增加语法元素sps_mtss_enabled_flag。如果sps_mtss_enabled_flag的取值为1,当前序列允许使用MTSS技术,更具体地,允许在块级(CU或TU)编解码本技术方案所述的纹理特征索引、LFNST/NSPT变换核索引的编码方法,如果sps_mtss_enabled_flag的值为0,当前序列不允许使用MTSS技术,更具体地,在块级(CU或TU)不会使用本技术方案所述的纹理特征索引、LFNST/NSPT变换核索引的编码方法。
当然也可以使用其他级别的语法来实现更灵活的控制,如图像参数集(Picture Parameter Set,PPS)的flag,或图像头(picture header)或片头(slice header)的flag等。比如先在SPS确定当前序列是否可以使用本技术方案,如果当前序列使用本技术方案,那么设置一个片头(slice header)的sh_inter_lfnst_nspt_enabled_flag确定当前slice是否使用本技术方案,提供更高的灵活性。
还需要说明的是,本申请实施例可以使用一个高层语法(high level syntax)设置一个应用MTSS技术的最小样本阈值,例如sps_mtss_min_pix,当sps_mtss_enabled_flag的取值为1时,解码器解析sps_mtss_min_pix确定应用MTSS技术的最小样本阈值。用高层语法可以根据需求在编码复杂度和压缩效率之间做出取舍。也就是说这种方法下,解码器需要支持所有sps_mtss_min_pix可能的情况,但是编码器可以配置编码当前码流需要的sps_mtss_min_pix。比如编码一个码流时,如果需要更好的压缩效率而不特别注重编码时间,可以给sps_mtss_min_pix设置一个比较小的值,如16。如果特别注重编码时间而且可以损失一定的压缩效率,可以给sps_mtss_min_pix设置一个比较大的值,如256。
类似地,本申请实施例也可以使用一个高层语法(high level syntax)设置一个应用MTSS技术的最小尺寸阈值,例如sps_mtss_min_size。当sps_mtss_enabled_flag的取值为1时,解码器解析sps_mtss_min_size确定应用MTSS技术的最小尺寸阈值。用高层语法可以根据需求在编码复杂度和压缩效率之间做出取舍。也就是说这种方法下,解码器需要支持所有sps_mtss_min_size可能的情况,但是编码器可以配置编码当前码流需要的sps_mtss_min_size。比如编码一个码流时,如果需要更好的压缩 效率而不特别注重编码时间,可以给sps_mtss_min_size设置一个比较小的值,如4。如果特别注重编码时间而且可以损失一定的压缩效率,可以给sps_mtss_min_size设置一个比较大的值,如16。
在本申请实施例中,通过上述实施例对前述实施例的具体实现进行详细阐述,从中可以看出,根据前述实施例的技术方案,这里实现了对MIP模式和EIP模式的特殊处理;某些情况用PLANAR模式作为候选的纹理特征;MTSS技术也可能不额外使用梯度推导纹理特征;以及对块尺寸的限制。如此,针对帧内编码的块中使用某些特殊帧内预测模式进行预测的块,它们的预测残差纹理特征不像普通帧内预测模式那样明朗,使用MTSS技术导出多个候选纹理特征指导变换,从而可以提高压缩效率,进而提升编解码性能。
在本申请的再一实施例中,基于前述实施例相同的发明构思,图35为本申请实施例提供的一种编码器的组成结构示意图。如图35所示,该编码器350可以包括第一确定单元3501、第一变换单元3502和编码单元3503,其中:
第一确定单元3501,配置为确定当前块的变换核组;以及还配置为根据变换核组,确定当前块的变换核;
第一变换单元3502,配置为确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;
编码单元3503,配置为对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为根据至少两个候选变换核组对当前块进行编码代价计算,确定至少两个候选变换核组各自对应的代价结果;在至少两个候选变换核组各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核组确定为当前块的变换核组。
在一些实施例中,第一确定单元3501,还配置为确定当前块的变换核组索引;其中,变换核组索引用于指示当前块的变换核组在至少两个候选变换核组中的编号;编码单元3503,还配置为对当前块的变换核组索引进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为根据当前块的变换核组索引,确定第一语法元素的取值;编码单元3503,还配置为对第一语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为确定当前块的样本个数;在当前块的样本个数大于或等于最小样本阈值时,执行确定当前块的变换核组的步骤。
在一些实施例中,第一确定单元3501,还配置为确定当前块的尺寸,当前块的尺寸包括高度和宽度;在当前块的高度和宽度均大于或等于最小尺寸阈值时,执行确定当前块的变换核组的步骤。
在一些实施例中,第一确定单元3501,还配置为确定当前块的预测模式;在当前块的预测模式为第一预测模式集合中的其中一项时,执行确定当前块的变换核组的步骤;其中,第一预测模式集合包括以下预测模式中至少之一:DC模式、PLANAR模式和角度预测模式之外的其他预测模式。
在一些实施例中,第一预测模式集合包括以下预测模式中至少之一:使用至少两个帧内预测模式进行组合预测的模式,帧内预测模式包括下述至少之一:DC模式、PLANAR模式和角度预测模式;复制帧内块进行预测的模式;使用外插滤波器进行预测的模式;使用矩阵运算进行预测的模式。
在一些实施例中,第一预测模式集合包括以下预测模式中至少之一:DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式。
在一些实施例中,在当前块的预测模式为第二预测模式集合中的其中一项时,第一确定单元3501,还配置为若变换核组索引为第i值,则基于当前块的预测模式推导的第i个帧内预测模式确定当前块的变换核组,i为正整数;其中,第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
在一些实施例中,在当前块的预测模式为第三预测模式集合中的其中一项时,第一确定单元3501,还配置为若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的变换核组;若变换核组索引为第二值,则基于PLANAR模式确定当前块的变换核组;其中,第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
在一些实施例中,第一确定单元3501,还配置为确定当前块的第一候选列表,第一候选列表指示至少两个候选变换核组;以及根据第一候选列表和变换核组索引,确定当前块的变换核组。
在一些实施例中,第一确定单元3501,还配置为确定基于当前块的预测模式推导的一个或多个帧内预测模式;以及根据一个或多个帧内预测模式确定一个或多个候选变换核组,并将一个或多个候选变 换核组添加到第一候选列表。
在一些实施例中,第一确定单元3501,还配置为在第一候选列表未填满时,确定当前块的预设纹理特征索引;以及根据预设纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,第一确定单元3501,还配置为在当前块的预测模式为第二预测模式集合中的其中一项时,确定基于当前块的预测模式推导的多个帧内预测模式;以及根据多个帧内预测模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表;其中,第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
在一些实施例中,第一确定单元3501,还配置为根据多个帧内预测模式和PLANAR模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表。
在一些实施例中,第一确定单元3501,还配置为在当前块的预测模式为第三预测模式集合中的其中一项时,确定基于当前块的预测模式推导的帧内预测模式;以及根据帧内预测模式和PLANAR模式确定两个候选变换核组,并将两个候选变换核组添加到第一候选列表;其中,第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
在一些实施例中,第一确定单元3501,还配置为在第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;根据候选样本,确定当前块的一个或多个候选纹理特征索引;以及根据一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,第一确定单元3501,还配置为确定候选样本的水平梯度值和竖直梯度值;根据候选样本的水平梯度值和竖直梯度值,确定候选样本对应的纹理特征索引以及梯度强度值;根据候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表;以及根据纹理特征统计表,确定当前块的一个或多个候选纹理特征索引。
在一些实施例中,第一确定单元3501,还配置为在候选样本的数量为至少一个时,确定至少一个纹理特征索引以及对应的至少一个梯度强度值;根据至少一个纹理特征索引确定具有互异特性的至少一种参考纹理特征索引,以及根据至少一个梯度强度值,对归属于同一参考纹理特征索引的梯度强度值进行累加计算,确定至少一种参考纹理特征索引对应的梯度强度累加值;以及根据至少一种参考纹理特征索引和至少一种参考纹理特征索引对应的梯度强度累加值,确定纹理特征统计表。
在一些实施例中,第一确定单元3501,还配置为将纹理特征统计表按照梯度强度累加值由高向低进行排序,确定排序靠前的N个梯度强度累加值对应的参考纹理特征索引;其中,N为正整数;将N个参考纹理特征索引确定为当前块的一个或多个候选纹理特征索引。
在一些实施例中,第一确定单元3501,还配置为对N个参考纹理特征索引进行剪枝,确定当前块的一个或多个候选纹理特征索引。
在一些实施例中,在第一候选列表指示至少两个变换核组包括的变换核时,第一确定单元3501,还配置为确定第一候选列表指示的至少两个候选变换核;根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果;以及在至少两个候选变换核各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核确定为当前块的变换核。
在一些实施例中,第一确定单元3501,还配置为确定变换核组包括的至少两个候选变换核;根据至少两个候选变换核对当前块进行编码代价计算,确定至少两个候选变换核各自对应的代价结果;以及在至少两个候选变换核各自对应的代价结果中确定最小代价结果,将最小代价结果对应的候选变换核确定为当前块的变换核。
在一些实施例中,第一确定单元3501,还配置为确定当前块的变换核索引;其中,变换核索引用于指示当前块的变换核在第一候选列表或者当前块的变换核组中的编号;编码单元3503,还配置为对当前块的变换核索引进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为确定第二语法元素的取值;其中,第二语法元素用于指示当前块是否使用第一变换模式以及对应使用的变换核索引,变换核索引用于指示当前块的变换核在第一候选列表或者当前块的变换核组中的编号;编码单元3503,还配置为对第一语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为对当前块进行帧内预测,确定当前块的预测块;以及根据当前块的初始块和当前块的预测块,确定当前块的残差块。
在一些实施例中,第一变换单元3502,还配置为根据变换核对当前块的残差块进行不可分离基础变换,确定当前块的变换系数;或者,对当前块的残差块进行离散余弦变换,确定当前块的变换块,并根据变换核对当前块的变换块进行低频不可分离变换,确定当前块的变换系数。
在一些实施例中,第一确定单元3501,还配置为对当前块的变换系数进行量化,确定当前块的量化系数;编码单元3503,还配置为对当前块的量化系数进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为确定第三语法元素的取值;其中,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,当前序列包括当前块;编码单元3503,还配置为对第三语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为确定第三语法元素的取值和第四语法元素的取值;其中,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,第四语法元素用于指示当前图像是否允许使用多变换核组选择技术,当前序列包括当前图像,且当前图像包括当前块;编码单元3503,还配置为对第三语法元素的取值和第四语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为确定第三语法元素的取值和第五语法元素的取值;其中,第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,第五语法元素用于指示当前片是否允许使用多变换核组选择技术,当前序列包括当前片,且当前片包括当前块;编码单元3503,还配置为对第三语法元素的取值和第五语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,编码单元3503,还配置为对最小样本阈值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为在当前序列允许使用多变换核组选择技术时,确定第六语法元素的取值;其中,第六语法元素用于指示最小样本阈值;编码单元3503,还配置为对第六语法元素的取值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,编码单元3503,还配置为对最小尺寸阈值进行编码处理,将所得到的编码比特写入码流。
在一些实施例中,第一确定单元3501,还配置为在当前序列允许使用多变换核组选择技术时,确定第七语法元素的取值;其中,第七语法元素用于指示最小尺寸阈值;编码单元3503,还配置为对第七语法元素的取值进行编码处理,将所得到的编码比特写入码流。
可以理解地,在本申请实施例中,“单元”可以是部分电路、部分处理器、部分程序或软件等等,当然也可以是模块,还可以是非模块化的。而且在本实施例中的各组成部分可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。
在本申请的再一实施例中,图36为本申请实施例提供的一种编码器的具体硬件结构示意图。如图36所示,编码器350可以包括:第一通信接口3601、第一存储器3602和第一处理器3603;各个组件通过第一总线系统3604耦合在一起。可理解,第一总线系统3604用于实现这些组件之间的连接通信。第一总线系统3604除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图36中将各种总线都标为第一总线系统3604。其中,
第一通信接口3601,用于在与其他外部网元之间进行收发信息过程中,信号的接收和发送;
第一存储器3602,用于存储能够在第一处理器3603上运行的计算机程序;
第一处理器3603,用于在运行所述计算机程序时,执行:
确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
可以理解,本申请实施例中的第一存储器3602可以是易失性存储器或非易失性存储器,或可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synchlink DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请描述的系统和方法的第一存储器3602旨在包括但不限于这些和任意其它适合类型的存储器。
而第一处理器3603可能是一种集成电路芯片,具有信号的处理能力。在实现过程中,上述方法的各步骤可以通过第一处理器3603中的硬件的集成逻辑电路或者软件形式的指令完成。上述的第一处理 器3603可以是通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。可以实现或者执行本申请实施例中的公开的各方法、步骤及逻辑框图。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。结合本申请实施例所公开的方法的步骤可以直接体现为硬件译码处理器执行完成,或者用译码处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于第一存储器3602,第一处理器3603读取第一存储器3602中的信息,结合其硬件完成上述方法的步骤。
可以理解的是,本申请描述的这些实施例可以用硬件、软件、固件、中间件、微码或其组合来实现。对于硬件实现,处理单元可以实现在一个或多个专用集成电路(Application Specific Integrated Circuits,ASIC)、数字信号处理器(Digital Signal Processing,DSP)、数字信号处理设备(DSP Device,DSPD)、可编程逻辑设备(Programmable Logic Device,PLD)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)、通用处理器、控制器、微控制器、微处理器、用于执行本申请所述功能的其它电子单元或其组合中。对于软件实现,可通过执行本申请所述功能的模块(例如过程、函数等)来实现本申请所述的技术。软件代码可存储在存储器中并通过处理器执行。存储器可以在处理器中或在处理器外部实现。
可选地,作为另一个实施例,第一处理器3603还配置为在运行所述计算机程序时,执行前述实施例中任一项所述的方法。
本实施例提供了一种编码器,在该编码器中,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。
在本申请的再一实施例中,基于前述实施例相同的发明构思,图37为本申请实施例提供的一种解码器的组成结构示意图。如图37所示,该解码器370可以包括第二确定单元3701和第二变换单元3702,其中:
第二确定单元3701,配置为确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;以及还配置为根据变换核组,确定当前块的变换核;
第二变换单元3702,配置为确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
在一些实施例中,参见图37,该解码器370还可以包括解码单元3703,配置为解码码流,确定当前块的变换核组索引。
在一些实施例中,解码单元3703,还配置为解码码流,确定第一语法元素的取值;第二确定单元3701,还配置为根据第一语法元素的取值,确定当前块的变换核组索引。
在一些实施例中,第二确定单元3701,还配置为确定当前块的样本个数;在当前块的样本个数大于或等于最小样本阈值时,执行确定当前块的变换核组索引的步骤。
在一些实施例中,第二确定单元3701,还配置为确定当前块的尺寸,其中,所述当前块的尺寸包括高度和宽度;在当前块的高度和宽度均大于或等于最小尺寸阈值时,执行确定当前块的变换核组索引的步骤。
在一些实施例中,第二确定单元3701,还配置为确定当前块的预测模式;在当前块的预测模式为第一预测模式集合中的其中一项时,执行确定当前块的变换核组索引的步骤;其中,第一预测模式集合包括以下预测模式中至少之一:DC模式、PLANAR模式和角度预测模式之外的其他预测模式。
在一些实施例中,第一预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式;复制帧内块进行预测的模式;使用外插滤波器进行预测的模式;使用矩阵运算进行预测的模式。
在一些实施例中,第一预测模式集合包括以下预测模式中至少之一:DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式。
在一些实施例中,在当前块的预测模式为第二预测模式集合中的其中一项时,第二确定单元3701,还配置为若变换核组索引为第i值,则基于当前块的预测模式推导的第i个帧内预测模式确定当前块的变换核组,i为正整数;其中,第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
在一些实施例中,在当前块的预测模式为第三预测模式集合中的其中一项时,第二确定单元3701,还配置为若变换核组索引为第一值,则基于当前块的预测模式推导的第一个帧内预测模式确定当前块的 变换核组;若变换核组索引为第二值,则基于PLANAR模式确定当前块的变换核组;其中,第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
在一些实施例中,第二确定单元3701,还配置为确定当前块的第一候选列表,第一候选列表指示至少两个候选变换核组;以及根据第一候选列表和变换核组索引,确定当前块的变换核组。
在一些实施例中,第二确定单元3701,还配置为确定基于当前块的预测模式推导的一个或多个帧内预测模式;以及根据一个或多个帧内预测模式确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,第二确定单元3701,还配置为在第一候选列表未填满时,确定当前块的预设纹理特征索引;以及根据预设纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,第二确定单元3701,还配置为在当前块的预测模式为第二预测模式集合中的其中一项时,确定基于当前块的预测模式推导的多个帧内预测模式;以及根据多个帧内预测模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表;其中,第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
在一些实施例中,第二确定单元3701,还配置为根据多个帧内预测模式和PLANAR模式确定多个候选变换核组,并将多个候选变换核组添加到第一候选列表。
在一些实施例中,第二确定单元3701,还配置为在当前块的预测模式为第三预测模式集合中的其中一项时,确定基于当前块的预测模式推导的帧内预测模式;以及根据帧内预测模式和PLANAR模式确定两个候选变换核组,并将两个候选变换核组添加到第一候选列表;其中,第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
在一些实施例中,第二确定单元3701,还配置为在第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;根据候选样本,确定当前块的一个或多个候选纹理特征索引;以及根据一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将一个或多个候选变换核组添加到第一候选列表。
在一些实施例中,第二确定单元3701,还配置为确定候选样本的水平梯度值和竖直梯度值;根据候选样本的水平梯度值和竖直梯度值,确定候选样本对应的纹理特征索引以及梯度强度值;根据候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表;以及根据纹理特征统计表,确定当前块的一个或多个候选纹理特征索引。
在一些实施例中,第二确定单元3701,还配置为在候选样本的数量为至少一个时,确定至少一个纹理特征索引以及对应的至少一个梯度强度值;根据至少一个纹理特征索引确定具有互异特性的至少一种参考纹理特征索引,以及根据至少一个梯度强度值,对归属于同一参考纹理特征索引的梯度强度值进行累加计算,确定至少一种参考纹理特征索引对应的梯度强度累加值;以及根据至少一种参考纹理特征索引和至少一种参考纹理特征索引对应的梯度强度累加值,确定纹理特征统计表。
在一些实施例中,第二确定单元3701,还配置为将纹理特征统计表按照梯度强度累加值由高向低进行排序,确定排序靠前的N个梯度强度累加值对应的参考纹理特征索引;其中,N为正整数;以及将N个参考纹理特征索引确定为当前块的一个或多个候选纹理特征索引。
在一些实施例中,第二确定单元3701,还配置为对N个参考纹理特征索引进行剪枝,确定当前块的一个或多个候选纹理特征索引。
在一些实施例中,在第一候选列表指示至少两个变换核组包括的变换核时,第二确定单元3701,还配置为确定当前块的变换核索引;以及根据第一候选列表以及变换核索引,确定当前块的变换核。
在一些实施例中,第二确定单元3701,还配置为确定当前块的变换核索引;以及根据变换核组以及变换核索引,确定当前块的变换核。
在一些实施例中,解码单元3703,还配置为解码码流,确定当前块的变换核索引。
在一些实施例中,解码单元3703,还配置为解码码流,确定第二语法元素的取值;第二确定单元3701,还配置为在第二语法元素指示当前块使用第一变换模式时,根据第二语法元素的取值,确定当前块的变换核索引。
在一些实施例中,解码单元3703,还配置为解码码流,确定当前块的量化系数;第二确定单元3701,还配置为对当前块的量化系数进行反量化,确定当前块的变换系数。
在一些实施例中,第二变换单元3702,还配置为根据变换核对当前块的变换系数进行不可分离基础变换,确定当前块的残差块;或者,根据变换核对当前块的变换系数进行低频不可分离变换,确定当前块的变换块,并对当前块的变换块进行离散余弦变换,确定当前块的残差块。
在一些实施例中,第二确定单元3701,还配置为对当前块进行帧内预测,确定当前块的预测块;以及根据当前块的预测块和当前块的残差块,确定当前块的重建块。
在一些实施例中,解码单元3703,还配置为解码码流,确定第三语法元素的取值;在第三语法元素指示当前序列允许使用多变换核组选择技术时,执行确定当前块的变换核组索引的步骤;其中,当前序列包括当前块。
在一些实施例中,解码单元3703,还配置为解码码流,确定第三语法元素的取值;以及在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第四语法元素的取值;第二确定单元3701,还配置为在第四语法元素指示当前图像允许使用多变换核组选择技术时,执行确定当前块的变换核组索引的步骤;其中,当前序列包括当前图像,当前图像包括当前块。
在一些实施例中,解码单元3703,还配置为解码码流,确定第三语法元素的取值;以及在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第五语法元素的取值;第二确定单元3701,还配置为在第五语法元素指示当前片允许使用多变换核组选择技术时,执行确定当前块的变换核组索引的步骤;其中,当前序列包括当前片,当前片包括当前块。
在一些实施例中,解码单元3703,还配置为解码码流,确定最小样本阈值。
在一些实施例中,解码单元3703,还配置为解码码流,确定第三语法元素的取值;第二确定单元3701,还配置为在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第六语法元素的取值;以及根据第六语法元素的取值,确定最小样本阈值。
在一些实施例中,解码单元3703,还配置为解码码流,确定最小尺寸阈值。
在一些实施例中,解码单元3703,还配置为解码码流,确定第三语法元素的取值;第二确定单元3701,还配置为在第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第七语法元素的取值;以及根据第七语法元素的取值,确定最小尺寸阈值。
可以理解地,在本实施例中,“单元”可以是部分电路、部分处理器、部分程序或软件等等,当然也可以是模块,还可以是非模块化的。而且在本实施例中的各组成部分可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。
在本申请的再一实施例中,图38为本申请实施例提供的一种解码器的具体硬件结构示意图。如图38所示,解码器370可以包括:第二通信接口3801、第二存储器3802和第二处理器3803;各个组件通过第二总线系统3804耦合在一起。可理解,第二总线系统3804用于实现这些组件之间的连接通信。第二总线系统3804除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图38中将各种总线都标为第二总线系统3804。其中,
第二通信接口3801,用于在与其他外部网元之间进行收发信息过程中,信号的接收和发送;
第二存储器3802,用于存储能够在第二处理器3803上运行的计算机程序;
第二处理器3803,用于在运行所述计算机程序时,执行:
确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。
可选地,作为另一个实施例,第二处理器3803还配置为在运行所述计算机程序时,执行前述实施例中任一项所述的方法。
可以理解,第二存储器3802与第一存储器3602的硬件功能类似,第二处理器3803与第一处理器3603的硬件功能类似;这里不再详述。
本实施例提供了一种解码器,在该解码器中,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。
在本申请的再一实施例中,图39为本申请实施例提供的一种编解码系统的组成结构示意图。如图39所示,编解码系统390可以包括编码器3901和解码器3902。
在本申请实施例中,编码器3901可以为前述实施例中任一项所述的编码器,解码器3902可以为前述实施例中任一项所述的解码器。
在一些实施例中,本申请实施例还提供了一种计算机可读存储介质,其上存储有计算机程序。该计算机程序被处理器(例如第一处理器或第二处理器)执行时实现如前述实施例中任一项所述的方法。
在一些实施例中,本申请实施例还提供了一种计算机程序产品,包括计算机程序或指令。该计算机程序或指令被处理器(例如第一处理器或第二处理器)执行时实现如前述实施例中任一项所述的方法。
在一些实施例中,本申请实施例还提供了一种计算机程序,该计算机程序被处理器(例如第一处理器或第二处理器)执行时实现如前述实施例中任一项所述的方法。
本领域普通技术人员可以意识到,结合本申请所公开的实施例描述的各示例的单元及算法步骤,能够以电子硬件、或者计算机软件和电子硬件的结合来实现。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
在本申请所提供的几个实施例中,应该理解到,所揭露的装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
需要说明的是,在本申请中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
本申请所提供的几个方法实施例中所揭露的方法,在不冲突的情况下可以任意组合,得到新的方法实施例。
本申请所提供的几个产品实施例中所揭露的特征,在不冲突的情况下可以任意组合,得到新的产品实施例。
本申请所提供的几个方法或设备实施例中所揭露的特征,在不冲突的情况下可以任意组合,得到新的方法实施例或设备实施例。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。
工业实用性
本申请实施例中,在编码端,确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的残差块,并根据变换核对当前块的残差块进行变换,确定当前块的变换系数;对当前块的变换系数进行编码处理,将所得到的编码比特写入码流。在解码端,确定当前块的变换核组索引,并根据当前块的变换核组索引确定当前块的变换核组;根据变换核组,确定当前块的变换核;确定当前块的变换系数,并根据变换核对当前块的变换系数进行变换,确定当前块的残差块。这样,无论是编码端还是解码端,都是根据多变换核组选择技术确定当前块的变换核组,然后从中确定当前块的变换核。也就是说,针对使用某些帧内预测模式进行预测的当前块,在确定当前块的变换核时,可以使用多变换核组选择技术推导出的多个候选纹理特征指导变换,提高了变换预测的准确性,从而能够提高压缩效率,进而提升编解码性能。

Claims (78)

  1. 一种解码方法,应用于解码器,所述方法包括:
    确定当前块的变换核组索引,并根据所述当前块的变换核组索引确定所述当前块的变换核组;
    根据所述变换核组,确定所述当前块的变换核;
    确定所述当前块的变换系数,并根据所述变换核对所述当前块的变换系数进行变换,确定所述当前块的残差块。
  2. 根据权利要求1所述的方法,其中,所述确定当前块的变换核组索引,包括:
    解码码流,确定所述当前块的变换核组索引。
  3. 根据权利要求1所述的方法,其中,所述确定当前块的变换核组索引,包括:
    解码码流,确定第一语法元素的取值;
    根据所述第一语法元素的取值,确定所述当前块的变换核组索引。
  4. 根据权利要求1所述的方法,其中,所述方法还包括:
    确定所述当前块的样本个数;
    在所述当前块的样本个数大于或等于最小样本阈值时,执行所述确定当前块的变换核组索引的步骤。
  5. 根据权利要求1所述的方法,其中,所述方法还包括:
    确定所述当前块的尺寸,其中,所述当前块的尺寸包括高度和宽度;
    在所述当前块的高度和宽度均大于或等于最小尺寸阈值时,执行所述确定当前块的变换核组索引的步骤。
  6. 根据权利要求1所述的方法,其中,所述方法还包括:
    确定所述当前块的预测模式;
    在所述当前块的预测模式为第一预测模式集合中的其中一项时,执行所述确定当前块的变换核组索引的步骤;
    其中,所述第一预测模式集合包括以下预测模式中至少之一:DC模式、PLANAR模式和角度预测模式之外的其他预测模式。
  7. 根据权利要求6所述的方法,其中,所述第一预测模式集合包括以下预测模式中至少之一:
    使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式;
    复制帧内块进行预测的模式;
    使用外插滤波器进行预测的模式;
    使用矩阵运算进行预测的模式。
  8. 根据权利要求6所述的方法,其中,所述第一预测模式集合包括以下预测模式中至少之一:
    DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式。
  9. 根据权利要求6所述的方法,其中,在所述当前块的预测模式为第二预测模式集合中的其中一项时,所述根据所述当前块的变换核组索引确定所述当前块的变换核组,包括:
    若所述变换核组索引为第i值,则基于所述当前块的预测模式推导的第i个帧内预测模式确定所述当前块的变换核组,i为正整数;
    其中,所述第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
  10. 根据权利要求6所述的方法,其中,在所述当前块的预测模式为第三预测模式集合中的其中一项时,所述根据所述当前块的变换核组索引确定所述当前块的变换核组,包括:
    若所述变换核组索引为第一值,则基于所述当前块的预测模式推导的第一个帧内预测模式确定所述当前块的变换核组;
    若所述变换核组索引为第二值,则基于PLANAR模式确定所述当前块的变换核组;
    其中,所述第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
  11. 根据权利要求6所述的方法,其中,所述根据所述当前块的变换核组索引确定所述当前块的变换核组,包括:
    确定所述当前块的第一候选列表,所述第一候选列表指示至少两个候选变换核组;
    根据所述第一候选列表和所述变换核组索引,确定所述当前块的变换核组。
  12. 根据权利要求11所述的方法,其中,所述确定所述当前块的第一候选列表,包括:
    确定基于所述当前块的预测模式推导的一个或多个帧内预测模式;
    根据所述一个或多个帧内预测模式确定一个或多个候选变换核组,并将所述一个或多个候选变换核组添加到所述第一候选列表。
  13. 根据权利要求12所述的方法,其中,所述方法还包括:
    在所述第一候选列表未填满时,确定所述当前块的预设纹理特征索引;
    根据所述预设纹理特征索引确定一个或多个候选变换核组,并将所述一个或多个候选变换核组添加到所述第一候选列表。
  14. 根据权利要求11所述的方法,其中,所述确定所述当前块的第一候选列表,包括:
    在所述当前块的预测模式为第二预测模式集合中的其中一项时,确定基于所述当前块的预测模式推导的多个帧内预测模式;
    根据所述多个帧内预测模式确定多个候选变换核组,并将所述多个候选变换核组添加到所述第一候选列表;
    其中,所述第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
  15. 根据权利要求14所述的方法,其中,所述方法还包括:
    根据所述多个帧内预测模式和PLANAR模式确定多个候选变换核组,并将所述多个候选变换核组添加到所述第一候选列表。
  16. 根据权利要求11所述的方法,其中,所述确定所述当前块的第一候选列表,包括:
    在所述当前块的预测模式为第三预测模式集合中的其中一项时,确定基于所述当前块的预测模式推导的帧内预测模式;
    根据所述帧内预测模式和PLANAR模式确定两个候选变换核组,并将所述两个候选变换核组添加到所述第一候选列表;
    其中,所述第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
  17. 根据权利要求11所述的方法,其中,所述方法还包括:
    在所述第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;
    根据所述候选样本,确定所述当前块的一个或多个候选纹理特征索引;
    根据所述一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将所述一个或多个候选变换核组添加到所述第一候选列表。
  18. 根据权利要求17所述的方法,其中,所述根据所述候选样本,确定所述当前块的一个或多个候选纹理特征索引,包括:
    确定所述候选样本的水平梯度值和竖直梯度值;
    根据所述候选样本的水平梯度值和竖直梯度值,确定所述候选样本对应的纹理特征索引以及梯度强度值;
    根据所述候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表;
    根据所述纹理特征统计表,确定所述当前块的一个或多个候选纹理特征索引。
  19. 根据权利要求18所述的方法,其中,所述根据所述候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表,包括:
    在所述候选样本的数量为至少一个时,确定至少一个纹理特征索引以及对应的至少一个梯度强度值;
    根据所述至少一个纹理特征索引确定具有互异特性的至少一种参考纹理特征索引,以及根据所述至少一个梯度强度值,对归属于同一参考纹理特征索引的梯度强度值进行累加计算,确定所述至少一种参考纹理特征索引对应的梯度强度累加值;
    根据所述至少一种参考纹理特征索引和所述至少一种参考纹理特征索引对应的梯度强度累加值,确定所述纹理特征统计表。
  20. 根据权利要求19所述的方法,其中,所述根据所述纹理特征统计表,确定所述当前块的一个或多个候选纹理特征索引,包括:
    将所述纹理特征统计表按照所述梯度强度累加值由高向低进行排序,确定排序靠前的N个梯度强度累加值对应的参考纹理特征索引;其中,N为正整数;
    将所述N个参考纹理特征索引确定为所述当前块的一个或多个候选纹理特征索引。
  21. 根据权利要求20所述的方法,其中,所述方法还包括:
    对所述N个参考纹理特征索引进行剪枝,确定所述当前块的一个或多个候选纹理特征索引。
  22. 根据权利要求11所述的方法,其中,在所述第一候选列表指示所述至少两个变换核组包括的 变换核时,所述方法还包括:
    确定所述当前块的变换核索引;
    根据所述第一候选列表以及所述变换核索引,确定所述当前块的变换核。
  23. 根据权利要求1所述的方法,其中,所述根据所述变换核组,确定所述当前块的变换核,包括:
    确定所述当前块的变换核索引;
    根据所述变换核组以及所述变换核索引,确定所述当前块的变换核。
  24. 根据权利要求22或23所述的方法,其中,所述确定所述当前块的变换核索引,包括:
    解码码流,确定所述当前块的变换核索引。
  25. 根据权利要求22或23所述的方法,其中,所述确定所述当前块的变换核索引,包括:
    解码码流,确定第二语法元素的取值;
    在所述第二语法元素指示所述当前块使用第一变换模式时,根据所述第二语法元素的取值,确定所述当前块的变换核索引。
  26. 根据权利要求1至25中任一项所述的方法,其中,所述确定所述当前块的变换系数,包括:
    解码码流,确定所述当前块的量化系数;
    对所述当前块的量化系数进行反量化,确定所述当前块的变换系数。
  27. 根据权利要求1至25中任一项所述的方法,其中,所述根据所述变换核对所述当前块的变换系数进行变换,确定所述当前块的残差块,包括:
    根据所述变换核对所述当前块的变换系数进行不可分离基础变换,确定所述当前块的残差块;或者,
    根据所述变换核对所述当前块的变换系数进行低频不可分离变换,确定所述当前块的变换块,并对所述当前块的变换块进行离散余弦变换,确定所述当前块的残差块。
  28. 根据权利要求1至25中任一项所述的方法,其中,所述方法还包括:
    对所述当前块进行帧内预测,确定所述当前块的预测块;
    根据所述当前块的预测块和所述当前块的残差块,确定所述当前块的重建块。
  29. 根据权利要求1至25中任一项所述的方法,其中,所述方法还包括:
    解码码流,确定第三语法元素的取值;
    在所述第三语法元素指示当前序列允许使用多变换核组选择技术时,执行所述确定当前块的变换核组索引的步骤;其中,所述当前序列包括所述当前块。
  30. 根据权利要求1至25中任一项所述的方法,其中,所述方法还包括:
    解码码流,确定第三语法元素的取值;
    在所述第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第四语法元素的取值;
    在所述第四语法元素指示当前图像允许使用多变换核组选择技术时,执行所述确定当前块的变换核组索引的步骤;其中,所述当前序列包括所述当前图像,所述当前图像包括所述当前块。
  31. 根据权利要求1至25中任一项所述的方法,其中,所述方法还包括:
    解码码流,确定第三语法元素的取值;
    在所述第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第五语法元素的取值;
    在所述第五语法元素指示当前片允许使用多变换核组选择技术时,执行所述确定当前块的变换核组索引的步骤;其中,所述当前序列包括所述当前片,所述当前片包括所述当前块。
  32. 根据权利要求4所述的方法,其中,所述方法还包括:
    解码码流,确定所述最小样本阈值。
  33. 根据权利要求4所述的方法,其中,所述方法还包括:
    解码码流,确定第三语法元素的取值;
    在所述第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第六语法元素的取值;
    根据所述第六语法元素的取值,确定所述最小样本阈值。
  34. 根据权利要求5所述的方法,其中,所述方法还包括:
    解码码流,确定所述最小尺寸阈值。
  35. 根据权利要求5所述的方法,其中,所述方法还包括:
    解码码流,确定第三语法元素的取值;
    在所述第三语法元素指示当前序列允许使用多变换核组选择模式时,解码码流,确定第七语法元素的取值;
    根据所述第七语法元素的取值,确定所述最小尺寸阈值。
  36. 一种编码方法,应用于编码器,所述方法包括:
    确定当前块的变换核组;
    根据所述变换核组,确定所述当前块的变换核;
    确定所述当前块的残差块,并根据所述变换核对所述当前块的残差块进行变换,确定所述当前块的变换系数;
    对所述当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
  37. 根据权利要求36所述的方法,其中,所述确定所述当前块的变换核组,包括:
    根据至少两个候选变换核组对所述当前块进行编码代价计算,确定所述至少两个候选变换核组各自对应的代价结果;
    在所述至少两个候选变换核组各自对应的代价结果中确定最小代价结果,将所述最小代价结果对应的候选变换核组确定为所述当前块的变换核组。
  38. 根据权利要求37所述的方法,其中,所述方法还包括:
    确定所述当前块的变换核组索引;其中,所述变换核组索引用于指示所述当前块的变换核组在所述至少两个候选变换核组中的编号;
    对所述当前块的变换核组索引进行编码处理,将所得到的编码比特写入码流。
  39. 根据权利要求38所述的方法,其中,所述对所述当前块的变换核组索引进行编码处理,将所得到的编码比特写入码流,包括:
    根据所述当前块的变换核组索引,确定第一语法元素的取值;
    对所述第一语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  40. 根据权利要求36所述的方法,其中,所述方法还包括:
    确定所述当前块的样本个数;
    在所述当前块的样本个数大于或等于最小样本阈值时,执行所述确定当前块的变换核组的步骤。
  41. 根据权利要求36所述的方法,其中,所述方法还包括:
    确定所述当前块的尺寸,其中,所述当前块的尺寸包括高度和宽度;
    在所述当前块的高度和宽度均大于或等于最小尺寸阈值时,执行所述确定当前块的变换核组的步骤。
  42. 根据权利要求36所述的方法,其中,所述方法还包括:
    确定所述当前块的预测模式;
    在所述当前块的预测模式为第一预测模式集合中的其中一项时,执行所述确定当前块的变换核组的步骤;
    其中,所述第一预测模式集合包括以下预测模式中至少之一:DC模式、PLANAR模式和角度预测模式之外的其他预测模式。
  43. 根据权利要求42所述的方法,其中,所述第一预测模式集合包括以下预测模式中至少之一:
    使用至少两个帧内预测模式进行组合预测的模式,所述帧内预测模式包括下述至少之一:DC模式、PLANAR模式和角度预测模式;
    复制帧内块进行预测的模式;
    使用外插滤波器进行预测的模式;
    使用矩阵运算进行预测的模式。
  44. 根据权利要求42所述的方法,其中,所述第一预测模式集合包括以下预测模式中至少之一:DIMD模式、TIMD模式、SGPM模式、MIP模式、EIP模式、ITMP模式和IBC模式。
  45. 根据权利要求42所述的方法,其中,在所述当前块的预测模式为第二预测模式集合中的其中一项时,所述根据所述当前块的变换核组索引确定所述当前块的变换核组,包括:
    若所述变换核组索引为第i值,则基于所述当前块的预测模式推导的第i个帧内预测模式确定所述当前块的变换核组,i为正整数;
    其中,所述第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
  46. 根据权利要求42所述的方法,其中,在所述当前块的预测模式为第三预测模式集合中的其中一项时,所述根据所述当前块的变换核组索引确定所述当前块的变换核组,包括:
    若所述变换核组索引为第一值,则基于所述当前块的预测模式推导的第一个帧内预测模式确定所述当前块的变换核组;
    若所述变换核组索引为第二值,则基于PLANAR模式确定所述当前块的变换核组;
    其中,所述第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使 用矩阵运算进行预测的模式。
  47. 根据权利要求42所述的方法,其中,所述根据所述当前块的变换核组索引确定所述当前块的变换核组,包括:
    确定所述当前块的第一候选列表,所述第一候选列表指示至少两个候选变换核组;
    根据所述第一候选列表和所述变换核组索引,确定所述当前块的变换核组。
  48. 根据权利要求47所述的方法,其中,所述确定所述当前块的第一候选列表,包括:
    确定基于所述当前块的预测模式推导的一个或多个帧内预测模式;
    根据所述一个或多个帧内预测模式确定一个或多个候选变换核组,并将所述一个或多个候选变换核组添加到所述第一候选列表。
  49. 根据权利要求48所述的方法,其中,所述方法还包括:
    在所述第一候选列表未填满时,确定所述当前块的预设纹理特征索引;
    根据所述预设纹理特征索引确定一个或多个候选变换核组,并将所述一个或多个候选变换核组添加到所述第一候选列表。
  50. 根据权利要求47所述的方法,其中,所述确定所述当前块的第一候选列表,包括:
    在所述当前块的预测模式为第二预测模式集合中的其中一项时,确定基于所述当前块的预测模式推导的多个帧内预测模式;
    根据所述多个帧内预测模式确定多个候选变换核组,并将所述多个候选变换核组添加到所述第一候选列表;
    其中,所述第二预测模式集合包括以下预测模式中至少之一:使用至少两种帧内预测模式进行组合预测的模式,所述帧内预测模式至少包括下述其中之一:DC模式、PLANAR模式和角度预测模式。
  51. 根据权利要求50所述的方法,其中,所述方法还包括:
    根据所述多个帧内预测模式和PLANAR模式确定多个候选变换核组,并将所述多个候选变换核组添加到所述第一候选列表。
  52. 根据权利要求47所述的方法,其中,所述确定所述当前块的第一候选列表,包括:
    在所述当前块的预测模式为第三预测模式集合中的其中一项时,确定基于所述当前块的预测模式推导的帧内预测模式;
    根据所述帧内预测模式和PLANAR模式确定两个候选变换核组,并将所述两个候选变换核组添加到所述第一候选列表;
    其中,所述第三预测模式集合包括以下预测模式中至少之一:使用外插滤波器进行预测的模式和使用矩阵运算进行预测的模式。
  53. 根据权利要求47所述的方法,其中,所述方法还包括:
    在所述第一候选列表未填满时,确定用于推导纹理特征索引的候选样本;
    根据所述候选样本,确定所述当前块的一个或多个候选纹理特征索引;
    根据所述一个或多个候选纹理特征索引确定一个或多个候选变换核组,并将所述一个或多个候选变换核组添加到所述第一候选列表。
  54. 根据权利要求53所述的方法,其中,所述根据所述候选样本,确定所述当前块的一个或多个候选纹理特征索引,包括:
    确定所述候选样本的水平梯度值和竖直梯度值;
    根据所述候选样本的水平梯度值和竖直梯度值,确定所述候选样本对应的纹理特征索引以及梯度强度值;
    根据所述候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表;
    根据所述纹理特征统计表,确定所述当前块的一个或多个候选纹理特征索引。
  55. 根据权利要求54所述的方法,其中,所述根据所述候选样本对应的纹理特征索引以及梯度强度值,确定纹理特征统计表,包括:
    在所述候选样本的数量为至少一个时,确定至少一个纹理特征索引以及对应的至少一个梯度强度值;
    根据所述至少一个纹理特征索引确定具有互异特性的至少一种参考纹理特征索引,以及根据所述至少一个梯度强度值,对归属于同一参考纹理特征索引的梯度强度值进行累加计算,确定所述至少一种参考纹理特征索引对应的梯度强度累加值;
    根据所述至少一种参考纹理特征索引和所述至少一种参考纹理特征索引对应的梯度强度累加值,确定所述纹理特征统计表。
  56. 根据权利要求55所述的方法,其中,所述根据所述纹理特征统计表,确定所述当前块的一个或多个候选纹理特征索引,包括:
    将所述纹理特征统计表按照所述梯度强度累加值由高向低进行排序,确定排序靠前的N个梯度强度累加值对应的参考纹理特征索引;其中,N为正整数;
    将所述N个参考纹理特征索引确定为所述当前块的一个或多个候选纹理特征索引。
  57. 根据权利要求56所述的方法,其中,所述方法还包括:
    对所述N个参考纹理特征索引进行剪枝,确定所述当前块的一个或多个候选纹理特征索引。
  58. 根据权利要求47所述的方法,其中,在所述第一候选列表指示所述至少两个变换核组包括的变换核时,所述方法还包括:
    确定所述第一候选列表指示的至少两个候选变换核;
    根据所述至少两个候选变换核对所述当前块进行编码代价计算,确定所述至少两个候选变换核各自对应的代价结果;
    在所述至少两个候选变换核各自对应的代价结果中确定最小代价结果,将所述最小代价结果对应的候选变换核确定为所述当前块的变换核。
  59. 根据权利要求58所述的方法,其中,所述根据所述变换核组,确定所述当前块的变换核,包括:
    确定所述变换核组包括的至少两个候选变换核;
    根据所述至少两个候选变换核对所述当前块进行编码代价计算,确定所述至少两个候选变换核各自对应的代价结果;
    在所述至少两个候选变换核各自对应的代价结果中确定最小代价结果,将所述最小代价结果对应的候选变换核确定为所述当前块的变换核。
  60. 根据权利要求58或59所述的方法,其中,所述方法还包括:
    确定所述当前块的变换核索引;其中,所述变换核索引用于指示所述当前块的变换核在所述第一候选列表或者所述当前块的变换核组中的编号;
    对所述当前块的变换核索引进行编码处理,将所得到的编码比特写入码流。
  61. 根据权利要求58或59所述的方法,其中,所述方法还包括:
    确定第二语法元素的取值;其中,所述第二语法元素用于指示所述当前块是否使用第一变换模式以及对应使用的变换核索引,所述变换核索引用于指示所述当前块的变换核在所述第一候选列表或者所述当前块的变换核组中的编号;
    对所述第一语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  62. 根据权利要求36至61中任一项所述的方法,其中,所述确定所述当前块的残差块,包括:
    对所述当前块进行帧内预测,确定所述当前块的预测块;
    根据所述当前块的初始块和所述当前块的预测块,确定所述当前块的残差块。
  63. 根据权利要求36至61中任一项所述的方法,其中,所述根据所述变换核对所述当前块的残差块进行变换,确定所述当前块的变换系数,包括:
    根据所述变换核对所述当前块的残差块进行不可分离基础变换,确定所述当前块的变换系数;或者,
    对所述当前块的残差块进行离散余弦变换,确定所述当前块的变换块,并根据所述变换核对所述当前块的变换块进行低频不可分离变换,确定所述当前块的变换系数。
  64. 根据权利要求36至61中任一项所述的方法,其中,所述对所述当前块的变换系数进行编码处理,将所得到的编码比特写入码流,包括:
    对所述当前块的变换系数进行量化,确定所述当前块的量化系数;
    对所述当前块的量化系数进行编码处理,将所得到的编码比特写入码流。
  65. 根据权利要求36至61中任一项所述的方法,其中,所述方法还包括:
    确定第三语法元素的取值;其中,所述第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,所述当前序列包括所述当前块;
    对所述第三语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  66. 根据权利要求36至61中任一项所述的方法,其中,所述方法还包括:
    确定第三语法元素的取值和第四语法元素的取值;其中,所述第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,所述第四语法元素用于指示当前图像是否允许使用多变换核组选择技术,所述当前序列包括所述当前图像,且所述当前图像包括所述当前块;
    对所述第三语法元素的取值和所述第四语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  67. 根据权利要求36至61中任一项所述的方法,其中,所述方法还包括:
    确定第三语法元素的取值和第五语法元素的取值;其中,所述第三语法元素用于指示当前序列是否 允许使用多变换核组选择技术,所述第五语法元素用于指示当前片是否允许使用多变换核组选择技术,所述当前序列包括所述当前片,且所述当前片包括所述当前块;
    对所述第三语法元素的取值和所述第五语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  68. 根据权利要求40所述的方法,其中,所述方法还包括:
    对所述最小样本阈值进行编码处理,将所得到的编码比特写入码流。
  69. 根据权利要求40所述的方法,其中,所述方法还包括:
    在当前序列允许使用多变换核组选择技术时,确定第六语法元素的取值;其中,所述第六语法元素用于指示所述最小样本阈值;
    对所述第六语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  70. 根据权利要求41所述的方法,其中,所述方法还包括:
    对所述最小尺寸阈值进行编码处理,将所得到的编码比特写入码流。
  71. 根据权利要求41所述的方法,其中,所述方法还包括:
    在当前序列允许使用多变换核组选择技术时,确定第七语法元素的取值;其中,所述第七语法元素用于指示所述最小尺寸阈值;
    对所述第七语法元素的取值进行编码处理,将所得到的编码比特写入码流。
  72. 一种码流,其中,所述码流是根据待编码信息进行比特编码生成的;其中,待编码信息包括下述至少一项:当前块的量化系数、所述当前块的变换核索引、所述当前块的变换核组索引、最小样本阈值、最小尺寸阈值、第一语法元素的取值、第二语法元素的取值、第三语法元素的取值、第四语法元素的取值、第五语法元素的取值、第六语法元素的取值和第七语法元素的取值;
    其中,所述第一语法元素用于指示所述当前块的变换核组索引,所述第二语法元素用于指示所述当前块是否使用第一变换模式以及对应使用的变换核索引,所述第三语法元素用于指示当前序列是否允许使用多变换核组选择技术,所述第四语法元素用于指示当前图像是否允许使用多变换核组选择技术,所述第五语法元素用于指示当前片是否允许使用多变换核组选择技术,所述第六语法元素用于指示所述最小样本阈值,所述第七语法元素用于指示所述最小尺寸阈值。
  73. 一种编码器,所述编码器包括第一确定单元、第一变换单元和编码单元,其中:
    所述第一确定单元,配置为确定当前块的变换核组;以及还配置为根据所述变换核组,确定所述当前块的变换核;
    所述第一变换单元,配置为确定所述当前块的残差块,并根据所述变换核对所述当前块的残差块进行变换,确定所述当前块的变换系数;
    所述编码单元,配置为对所述当前块的变换系数进行编码处理,将所得到的编码比特写入码流。
  74. 一种编码器,所述编码器包括第一存储器和第一处理器,其中:
    所述第一存储器,用于存储能够在所述第一处理器上运行的计算机程序;
    所述第一处理器,用于在运行所述计算机程序时,执行如权利要求36至71中任一项所述的方法。
  75. 一种解码器,所述解码器包括第二确定单元和第二变换单元,其中:
    所述第二确定单元,配置为确定当前块的变换核组索引,并根据所述当前块的变换核组索引确定所述当前块的变换核组;以及还配置为根据所述变换核组,确定所述当前块的变换核;
    所述第二变换单元,配置为确定所述当前块的变换系数,并根据所述变换核对所述当前块的变换系数进行变换,确定所述当前块的残差块。
  76. 一种解码器,所述解码器包括第二存储器和第二处理器,其中:
    所述第二存储器,用于存储能够在所述第二处理器上运行的计算机程序;
    所述第二处理器,用于在运行所述计算机程序时,执行如权利要求1至35中任一项所述的方法。
  77. 一种计算机可读存储介质,其上存储有计算机程序,其中,所述计算机程序被处理器执行时实现如权利要求1至35中任一项所述的方法、或者实现如权利要求36至71中任一项所述的方法。
  78. 一种计算机程序产品,包括计算机程序或指令,其中,所述计算机程序或指令被处理器执行时实现如权利要求1至35中任一项所述的方法、或者实现如权利要求36至71中任一项所述的方法。
PCT/CN2024/084081 2024-03-27 2024-03-27 编解码方法、码流、编码器、解码器以及存储介质 Pending WO2025199802A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/084081 WO2025199802A1 (zh) 2024-03-27 2024-03-27 编解码方法、码流、编码器、解码器以及存储介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/084081 WO2025199802A1 (zh) 2024-03-27 2024-03-27 编解码方法、码流、编码器、解码器以及存储介质

Publications (1)

Publication Number Publication Date
WO2025199802A1 true WO2025199802A1 (zh) 2025-10-02

Family

ID=97218672

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/084081 Pending WO2025199802A1 (zh) 2024-03-27 2024-03-27 编解码方法、码流、编码器、解码器以及存储介质

Country Status (1)

Country Link
WO (1) WO2025199802A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109922340A (zh) * 2017-12-13 2019-06-21 华为技术有限公司 图像编解码方法、装置、系统及存储介质
CN112262579A (zh) * 2018-06-13 2021-01-22 华为技术有限公司 基于比特流标志位的用于视频编码的帧内锐化和/或去振铃滤波器
WO2023059056A1 (ko) * 2021-10-05 2023-04-13 엘지전자 주식회사 비분리 1차 변환에 기반한 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장하는 기록 매체
CN116547966A (zh) * 2021-10-01 2023-08-04 腾讯美国有限责任公司 用于帧间帧内联合预测模式的二次变换
CN116868567A (zh) * 2021-10-13 2023-10-10 腾讯美国有限责任公司 自适应多变换集合选择

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109922340A (zh) * 2017-12-13 2019-06-21 华为技术有限公司 图像编解码方法、装置、系统及存储介质
CN112262579A (zh) * 2018-06-13 2021-01-22 华为技术有限公司 基于比特流标志位的用于视频编码的帧内锐化和/或去振铃滤波器
CN116547966A (zh) * 2021-10-01 2023-08-04 腾讯美国有限责任公司 用于帧间帧内联合预测模式的二次变换
WO2023059056A1 (ko) * 2021-10-05 2023-04-13 엘지전자 주식회사 비분리 1차 변환에 기반한 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장하는 기록 매체
CN116868567A (zh) * 2021-10-13 2023-10-10 腾讯美国有限责任公司 自适应多变换集合选择

Similar Documents

Publication Publication Date Title
CN113491116B (zh) 基于帧内预测的视频信号处理方法和装置
CN115695787A (zh) 基于神经网络的视频编解码中的分割信息
JP7781987B2 (ja) 動画処理方法、およびストリーム生成方法
WO2022218385A1 (en) Unified neural network filter model
WO2023245194A1 (en) Partitioning information in neural network-based video coding
WO2023051653A1 (en) Method, apparatus, and medium for video processing
WO2024007116A1 (zh) 解码方法、编码方法、解码器以及编码器
WO2023051654A1 (en) Method, apparatus, and medium for video processing
CN118474373A (zh) 编解码方法和装置
US11202082B2 (en) Image processing apparatus and method
CN114175653B (zh) 用于视频编解码中的无损编解码模式的方法和装置
JP7768746B2 (ja) イントラ予測装置、復号装置、及びプログラム
WO2024081872A1 (en) Method, apparatus, and medium for video processing
WO2025199802A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
JP2025512076A (ja) デコーディング方法、エンコーディング方法、デコーダー及びエンコーダー
CN116800985A (zh) 编解码方法和装置
WO2025147924A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025067391A1 (en) On designing an improved neural network-based super-resolution for video coding
CN113615202A (zh) 用于屏幕内容编解码的帧内预测的方法和装置
WO2023198057A9 (en) Method, apparatus, and medium for video processing
WO2025145288A9 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025260253A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025065658A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025065670A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2024207136A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24932164

Country of ref document: EP

Kind code of ref document: A1