WO2024259568A1 - 编解码方法、码流、编码器、解码器以及存储介质 - Google Patents
编解码方法、码流、编码器、解码器以及存储介质 Download PDFInfo
- Publication number
- WO2024259568A1 WO2024259568A1 PCT/CN2023/101156 CN2023101156W WO2024259568A1 WO 2024259568 A1 WO2024259568 A1 WO 2024259568A1 CN 2023101156 W CN2023101156 W CN 2023101156W WO 2024259568 A1 WO2024259568 A1 WO 2024259568A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- current block
- value
- filter
- target
- determining
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/11—Selection of coding mode or of prediction mode among a plurality of spatial predictive coding modes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/157—Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
- H04N19/159—Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/189—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the adaptation method, adaptation tool or adaptation type used for the adaptive coding
- H04N19/196—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the adaptation method, adaptation tool or adaptation type used for the adaptive coding being specially adapted for the computation of encoding parameters, e.g. by averaging previously computed encoding parameters
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/593—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
- H04N19/82—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation involving filtering within a prediction loop
Definitions
- the embodiments of the present application relate to the field of video coding and decoding technology, and in particular, to a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium.
- high-resolution videos such as HD and UHD have emerged.
- high-resolution videos usually have more information and therefore require more bandwidth.
- video coding standards involving video compression have been introduced.
- the interpolation-based intra-frame prediction technology has been proposed in the video coding standard. Specifically, the interpolation filter coefficients are obtained through the reconstructed pixel values around the current block, and then used to perform intra-frame prediction on the current block.
- the existing technical solutions still have some defects, resulting in low cost-effectiveness in encoding and decoding performance and time complexity.
- the embodiments of the present application provide a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium, which can reduce the time complexity while ensuring the coding and decoding performance.
- an embodiment of the present application provides a decoding method, which is applied to a decoder, and the method includes:
- An intra-frame prediction is performed on the current block according to the filter coefficient to determine a prediction value of the current block.
- an embodiment of the present application provides an encoding method, which is applied to an encoder, and the method includes:
- An intra-frame prediction is performed on the current block according to the filter coefficient to determine a prediction value of the current block.
- an embodiment of the present application provides a code stream, which is generated by bit encoding according to information to be encoded; wherein the information to be encoded includes at least one of the following:
- a target filtering mode of a current block a residual value of the current block, and a transform kernel index value of the current block.
- an encoder comprising a first determination unit and a first prediction unit, wherein:
- a first determining unit is configured to determine a target filtering mode of the current block; and determine a reference area of the current block according to a size parameter of the current block and the target filtering mode;
- the first prediction unit is configured to determine a filter coefficient of the current block according to a reference area of the current block; and perform intra-frame prediction on the current block according to the filter coefficient to determine a prediction value of the current block.
- an embodiment of the present application provides an encoder, the encoder comprising a first memory and a first processor; wherein,
- a first memory for storing a computer program that can be run on the first processor
- the first processor is used to execute the method described in the second aspect when running a computer program.
- an embodiment of the present application provides a decoder, the decoder comprising a decoding unit, a second determining unit, and a second predicting unit, wherein:
- a decoding unit configured to decode the bitstream and determine a target filtering mode for a current block
- a second determining unit is configured to determine a reference area of the current block according to a size parameter of the current block and a target filtering mode
- the second prediction unit is configured to determine a filter coefficient of the current block according to a reference area of the current block; and perform intra-frame prediction on the current block according to the filter coefficient to determine a prediction value of the current block.
- an embodiment of the present application provides a decoder, the decoder comprising a second memory and a second processor; wherein:
- a second memory for storing a computer program that can be run on a second processor
- the second processor is used to execute the method described in the first aspect when running a computer program.
- an embodiment of the present application provides a computer-readable storage medium, which stores a computer program.
- the computer program When executed, it implements the method as described in the first aspect, or implements the method as described in the second aspect.
- the embodiment of the present application provides a coding and decoding method, a code stream, an encoder, a decoder and a storage medium.
- the reference area of the current block is determined according to the size parameter of the current block and the target filtering mode; then the filtering coefficient of the current block is determined according to the reference area of the current block; and then the current block is intra-predicted according to the filtering coefficient to determine the prediction value of the current block.
- the intra-frame prediction technology based on interpolation filtering is not only related to the target filtering mode when determining the reference area for calculating the filtering coefficient, but also to the size parameter of the current block.
- a large reference area can be used when the size of the current block is large, and a small reference area can be used when the size of the current block is small; in this way, the computational complexity can be reduced and the encoding time can be reduced; at the same time, the accuracy of intra-frame prediction can be improved, thereby improving the coding and decoding performance.
- FIG. 1A is a schematic diagram of a calculation method for obtaining a value m
- FIG1B is a second schematic diagram of a calculation method for obtaining the value m
- FIG1C is a third schematic diagram of a calculation method for obtaining the value m
- FIG2A is a schematic diagram of a positional relationship between a current block and a reconstruction area
- FIG2B is a second schematic diagram of the positional relationship between the current block and the reconstruction area
- FIG2C is a third schematic diagram of the positional relationship between the current block and the reconstruction area
- FIG3A is a schematic diagram of a shape of an interpolation filter
- FIG3B is a second schematic diagram of the shape of an interpolation filter
- FIG3C is a third schematic diagram of the shape of an interpolation filter
- FIG4 is a schematic diagram of a structure for obtaining input and output at possible positions of an interpolation filter
- FIG5 is a schematic diagram of a prediction direction based on interpolation filtering
- FIG6 is a schematic diagram of an angle mode of intra-frame prediction
- FIG7 is a schematic diagram of a 3 ⁇ 3 window sliding in a prediction block
- FIG8 is a schematic diagram of the accumulated gradient amplitude values at different angles
- FIG9A is a schematic block diagram of a composition of an encoder provided in an embodiment of the present application.
- FIG9B is a schematic block diagram of a decoder provided in an embodiment of the present application.
- FIG10 is a schematic diagram of a network architecture of a coding and decoding system provided in an embodiment of the present application.
- FIG11 is a flowchart diagram 1 of a decoding method provided in an embodiment of the present application.
- FIG12A is a schematic diagram 1 of a reference area of a current block provided in an embodiment of the present application.
- FIG12B is a second schematic diagram of a reference area of a current block provided in an embodiment of the present application.
- FIG12C is a third schematic diagram of a reference area of a current block provided in an embodiment of the present application.
- FIG13 is a second flow chart of a decoding method provided in an embodiment of the present application.
- FIG14A is a first schematic diagram of the distribution of linear terms and nonlinear terms of a target filter provided in an embodiment of the present application
- FIG14B is a second schematic diagram of the distribution of linear terms and nonlinear terms of an interpolation filter provided in an embodiment of the present application.
- FIG14C is a third schematic diagram of the distribution of linear terms and nonlinear terms of an interpolation filter provided in an embodiment of the present application.
- FIG15A is a first schematic diagram of distribution of linear terms and nonlinear terms of another target filter provided in an embodiment of the present application.
- FIG15B is a second schematic diagram of distribution of linear terms and nonlinear terms of another interpolation filter provided in an embodiment of the present application.
- FIG15C is a third schematic diagram of distribution of linear terms and nonlinear terms of another interpolation filter provided in an embodiment of the present application.
- FIG16A is a first schematic diagram of distribution of linear terms and nonlinear terms of another target filter provided in an embodiment of the present application.
- FIG16B is a second schematic diagram of distribution of linear terms and nonlinear terms of another interpolation filter provided in an embodiment of the present application.
- FIG16C is a third schematic diagram of distribution of linear terms and nonlinear terms of another interpolation filter provided in an embodiment of the present application.
- FIG17 is a flowchart diagram 1 of an encoding method provided in an embodiment of the present application.
- FIG18A is a first schematic diagram of dividing a reference area provided in an embodiment of the present application.
- FIG18B is a second schematic diagram of division of a reference area provided in an embodiment of the present application.
- FIG18C is a third schematic diagram of division of a reference area provided in an embodiment of the present application.
- FIG19 is a second flow chart of an encoding method provided in an embodiment of the present application.
- FIG20 is a schematic diagram of the composition structure of an encoder provided in an embodiment of the present application.
- FIG21 is a schematic diagram of a specific hardware structure of an encoder provided in an embodiment of the present application.
- FIG22 is a schematic diagram of the composition structure of a decoder provided in an embodiment of the present application.
- FIG23 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application.
- FIG. 24 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application.
- first ⁇ second ⁇ third involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that “first ⁇ second ⁇ third” can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
- JVET Joint Video Exploration Team
- VVC Versatile Video Coding
- VVC reference software test platform
- VTM VVC Test Model
- ECM Enhanced Compression Mode
- MTS Multiple Transform Selection
- DCT Discrete Cosine Transform
- DST Discrete Sine Transform
- NPT Non-Separable Primary Transform
- PLANAR Planar mode
- IBC Intra Block Copy
- WAIP Wide Angle Intra Prediction
- MSE Mean Squared Error
- the intra prediction technology based on interpolation refers to a technology that obtains interpolation filter coefficients through reconstructed pixel values around the current block to perform intra prediction on the current block.
- the intra prediction technology based on interpolation filtering may include one or more of the following features:
- the number of taps of the interpolation filter should be greater than or equal to 2.
- the interpolation filter can have a variety of shapes. The shape of the selected interpolation filter is controlled using syntax elements.
- the reconstructed pixels used to obtain the interpolation filter coefficients should be in one or several regions around the current block, and the region used to obtain the interpolation filter coefficients is selected using syntax elements.
- Interpolation filter prediction can be used for intra block prediction of luma or chroma.
- the input of the interpolation filter is the reconstructed pixel value and/or the predicted pixel value, or it can be the reconstructed value and the predicted value minus a certain value.
- the maximum value and the minimum value of the reconstructed pixels are found in a reconstruction area of 13 rows and 13 columns around the current block.
- the maximum value and the minimum value here can be used to limit the range of the prediction result.
- the value m to be subtracted from the input and added to the output of the interpolation filter is obtained according to the following method: m is the value used in DC mode prediction, and m is a positive integer.
- the calculation method for obtaining the value m can be divided into three cases:
- three 15-tap interpolation filters and three reconstruction areas are defined.
- FIG2A shows a schematic diagram of the positional relationship between a current block and a reconstruction area.
- the reconstruction area may include an upper adjacent area adjacent to the upper side of the current block and a left adjacent area adjacent to the left side of the current block; wherein the length of the upper adjacent area is 2 ⁇ Width+13 and the width is 13; the length of the left adjacent area is 2 ⁇ Height+13 and the width is 13.
- FIG2B shows another schematic diagram of the positional relationship between a current block and a reconstruction area.
- the reconstruction area may include an upper adjacent area adjacent to the upper side of the current block; wherein the length of the upper adjacent area is 2 ⁇ Width+13 and the width is 13.
- FIG2C shows another schematic diagram of the positional relationship between a current block and a reconstruction area.
- the reconstruction area may include a left adjacent area adjacent to the left side of the current block; wherein the length of the left adjacent area is 2 ⁇ Height+13 and the width is 13.
- Height and Width represent the current height and width, respectively. It should be noted that, for the reconstructed area in FIG. 2A , FIG. 2B and FIG. 2C , that is, the reconstructed pixels in 13 rows and/or 13 columns around the current block may be used for obtaining interpolation filter coefficients.
- FIG3A shows a schematic diagram of the shape of an interpolation filter.
- the shape of the interpolation filter is a 4 ⁇ 4 square.
- FIG3B shows a schematic diagram of the shape of another interpolation filter.
- the shape of the interpolation filter is a 2 ⁇ 8 rectangle.
- FIG3C shows a schematic diagram of the shape of yet another interpolation filter.
- the shape of the interpolation filter is an 8 ⁇ 2 rectangle.
- the grid-filled portion represents the input position of the interpolation filter
- the black-filled portion represents the output position of the interpolation filter.
- 3 ⁇ 3 different filtering modes can be derived by combining the three reconstruction areas with the three interpolation filter shapes in different ways (each filter shape combined with each reconstruction area can derive a filtering mode), and the encoder decides a combination of filter shape and reconstruction area through rate-distortion cost.
- the encoder and decoder first determine the coefficients of the interpolation filter according to the determined filter shape and reconstruction area.
- the input of the interpolation filter is the pixel value without the mean value (i.e., the reconstructed pixel value minus the mean value). Then, when obtaining the parameters, the selected interpolation filter is slid on the selected area with a horizontal and vertical sliding step of 1 pixel distance. Specifically, see FIG4 , which shows a schematic diagram of the structure of a 4 ⁇ 4 interpolation filter obtaining inputs and outputs at possible positions of the interpolation filter on the selected reconstruction area. The autocorrelation coefficient matrix and the cross-correlation coefficient vector are constructed by the obtained inputs and outputs. When the selected reconstruction area includes pixel values that have not been reconstructed, the pixel values will not be counted in the samples used to obtain the interpolation filter parameters.
- c 0 ...c N-1 are the coefficients of the interpolation filter to be solved (also called "filter coefficients"), and m is a value subtracted from the input of the interpolation filter (a value to be added to the output at this time).
- the prediction process starts from the upper left corner of the current block and predicts toward the lower left corner in a certain order.
- the prediction formula is as follows:
- a and b in Clip(a,b,c) represent the output range of the limited prediction results.
- pred r is the prediction result of position r in the current block
- min, max are the minimum and maximum values obtained above
- m is a certain value obtained above.
- the interpolation filter is predicted in the diagonal direction, wherein the grid-filled part represents the input position of the interpolation filter, and the black-filled part represents the output position of the interpolation filter.
- the points to be predicted on the same diagonal line can be predicted in parallel.
- a prediction block of the current block can be obtained, and the prediction block includes a prediction value of at least one pixel in the current block.
- different angle modes are suitable for using different transforms, including a primary transform MTS, NSPT and a secondary transform LFNST.
- MTS includes some traditional transforms, such as DCT transform and DST transform.
- NSPT and LFNST are a series of transform coefficients obtained through a universal training set based on the optimal transform. The difference between NSPT and LFNST is that NSPT is directly used to transform the residual coefficients, while LFNST further transforms the transform coefficients after DCT2 transform.
- NST non-separable primary transform
- LNNST non-separable secondary transform
- the conventional intra prediction modes may include:
- PLANAR mode the intra prediction mode index is 0;
- Angular mode The intra prediction mode index is 2 to 66.
- the intra prediction mode may include angle modes of 2 to 66, and wide angle modes of -1 to 14 and 67 to 80.
- the arrows in FIG6 point to the directions predicted for the angle modes existing in VVC, and the intra prediction mode indexes used in encoding and decoding are 2 to 66.
- the current block is a non-square block, some angle directions will be replaced with wide angle modes (such as -1 to -14 and 67 to 80 in FIG6 ).
- NSPT and LFNST divide the transformation kernels of the traditional prediction mode into 35 groups, each of which has 3 selectable transformation kernels.
- Table 2 shows the correspondence between the traditional prediction mode and the transformation kernel groups.
- a method for matching a prediction block based on interpolation filtering to a traditional prediction mode is proposed here, and then the prediction block based on interpolation filtering is matched to a different transformation kernel of a preset primary transformation (separable or inseparable) or secondary transformation (separable or inseparable) through the matched traditional prediction mode.
- the prediction block based on interpolation filtering is matched to the PLANAR mode or the mode of angle direction 2 to 66 through the prediction value in the prediction block.
- the following steps may be included:
- a sliding 3 ⁇ 3 window is used to calculate the horizontal and vertical gradient values G x and G y of each 3 ⁇ 3 window in the prediction block based on interpolation filtering.
- G x and G y are obtained by multiplying the 3 ⁇ 3 horizontal gradient operator M x and the vertical gradient operator My with the prediction value within the window position.
- Figure 7 is a schematic diagram of a 3 ⁇ 3 window sliding in a prediction block, which can slide in the horizontal and vertical directions. Assuming that the prediction block based on interpolation filtering is a block with a width and height of (w, h), the sliding 3 ⁇ 3 window can calculate G x and G y at (w-2) ⁇ (h-2) positions at the center of the prediction block.
- the calculation process of atan() can be simplified and completed by table lookup or some transformation.
- the gradient magnitude value G at each position is accumulated on the traditional angle mode derived from it, and a histogram of the gradient magnitude value is obtained as shown in Figure 8.
- the traditional angle mode with the largest accumulated gradient magnitude value is selected from the histogram as the prediction mode corresponding to the current block; in particular, when the gradient magnitude values derived from all traditional angle modes are zero, the current block will be matched to the traditional PLANAR mode as the corresponding prediction mode.
- the traditional prediction mode derived from the interpolation filter prediction will be used for the selection of the transformation kernel group of NSPT and LFNST.
- an embodiment of the present application proposes a coding method to determine a target filtering mode of a current block; determine a reference area of the current block according to a size parameter of the current block and the target filtering mode; determine a filtering coefficient of the current block according to the reference area of the current block; perform intra-frame prediction on the current block according to the filtering coefficient to determine a prediction value of the current block.
- the embodiment of the present application proposes a decoding method, which decodes a bit stream and determines a target filtering mode of a current block; determines a reference area of the current block according to a size parameter of the current block and the target filtering mode; determines a filtering coefficient of the current block according to the reference area of the current block; and performs intra-frame prediction on the current block according to the filtering coefficient to determine a prediction value of the current block.
- the intra-frame prediction technology based on interpolation filtering when determining the reference area for calculating the filter coefficient, is not only related to the target filtering mode, but also related to the size parameters of the current block. For example, when the size of the current block is large, a large reference area can be used, and when the size of the current block is small, a small reference area can be used. In this way, while ensuring the encoding and decoding performance, the calculation complexity can be reduced, the encoding time can be reduced, and the cost-effectiveness of the encoding and decoding performance and the encoding complexity can be improved. At the same time, the intra-frame prediction accuracy can be improved, thereby improving the encoding and decoding efficiency.
- the encoder 100 may include a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control analysis unit 107, a filtering unit 108, an encoding unit 109, and a decoded image cache unit 110, etc.
- the filtering unit 108 may implement deblocking filtering and sample adaptive offset (Sample Adaptive Offset, SAO) filtering
- the encoding unit 109 may implement header information encoding and context-based adaptive binary arithmetic coding (Context-based Adaptive Binary Arithmetic Coding, CABAC).
- a video coding block can be obtained by dividing the coding tree unit (CTU), and then the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transformation and quantization unit 101 to transform the video coding block, including transforming the residual information from the pixel domain to the transform domain, and quantizing the obtained transform coefficients to further reduce the bit rate;
- the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to perform intra-frame prediction on the video coding block; specifically, the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to determine the intra-frame prediction mode to be used to encode the video coding block;
- the motion compensation unit 104 and the motion estimation unit 105 are used to perform inter-frame prediction coding of the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information;
- the motion estimation performed by the motion estimation unit 105 is a process of generating a motion vector, and the motion vector can estimate the motion of the video coding block, and then
- the motion vector determined by the motion estimation unit 105 performs motion compensation; after determining the intra-frame prediction mode, the intra-frame prediction unit 103 is also used to provide the selected intra-frame prediction data to the encoding unit 109, and the motion estimation unit 105 also sends the calculated and determined motion vector data to the encoding unit 109; in addition, the inverse transform and inverse quantization unit 106 is used to reconstruct the video coding block, reconstruct the residual block in the pixel domain, and the reconstructed residual block is removed by the filter control analysis unit 107 and the filtering unit 108.
- the encoding unit 109 is used to encode various coding parameters and quantized transform coefficients.
- the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-frame prediction mode and output the code stream of the video signal; and the decoded image buffer unit 110 is used to store the reconstructed video coding block for prediction reference. As the video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoded image buffer unit 110 .
- the decoder 200 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205, and a decoded image cache unit 206, etc., wherein the decoding unit 201 can implement header information decoding and CABAC decoding, and the filtering unit 205 can implement deblocking filtering and SAO filtering.
- the decoding unit 201 can implement header information decoding and CABAC decoding
- the filtering unit 205 can implement deblocking filtering and SAO filtering.
- a code stream of the video signal is output; the code stream is input to the decoder 200, and first passes through the decoding unit 201 to obtain the decoded transform coefficients; the transform coefficients are processed by the inverse transform and inverse quantization unit 202 to generate residual blocks in the pixel domain; the intra-frame prediction unit 203 can be used to generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and the data from the previously decoded block of the current frame or picture; the motion compensation unit 204 is to determine the prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and use The prediction information is used to generate a predictive block of the video decoding block being decoded; a decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 and the corresponding predictive block generated by the intra-frame prediction unit 203 or the motion compensation unit 204; the decoded video signal passes through the filtering unit 205 to remove the block effect artifacts
- the embodiment of the present application also provides a network architecture of a codec system including an encoder and a decoder, wherein FIG. 10 shows a schematic diagram of the network architecture of a coding and decoding system provided by an embodiment of the present application.
- the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01.
- the electronic device can be various types of devices with video coding and decoding functions during implementation.
- the electronic device can include a smart phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not specifically limited in the embodiment of the present application.
- the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.
- the method of the embodiment of the present application is mainly applied to the intra-frame prediction unit 103 part shown in Figure 9A and the intra-frame prediction unit 203 part shown in Figure 9B. That is to say, the embodiment of the present application can be applied to both the encoder and the decoder, and can even be applied to both the encoder and the decoder at the same time, but the embodiment of the present application is not specifically limited.
- the "current block” specifically refers to the coding block currently to be intra-frame predicted; when applied to the intra-frame prediction unit 203, the "current block” specifically refers to the decoding block currently to be intra-frame predicted.
- FIG11 a schematic flow chart of a decoding method provided by an embodiment of the present application is shown. As shown in FIG11, the method may include:
- S1101 Decode the bitstream and determine the target filtering mode of the current block.
- the decoding method of the embodiment of the present application may be an intra-frame prediction method, specifically referring to an improvement of an intra-frame prediction mode based on interpolation filtering to enhance the cost-effectiveness of performance and complexity.
- the current block includes at least a first color component and a second color component.
- the block at this time can be simply referred to as a first color component block; and when the first color component is a brightness component, the first color component block can also be referred to as a brightness block.
- the block at this time can be simply referred to as a second color component block; and when the second color component is a chrominance component, the second color component block can also be referred to as a chrominance block.
- the target filter mode may refer to a mode of using a target filter to perform intra-frame prediction on the current block, wherein the target filter may refer to an interpolation filter.
- the target filtering mode can be implemented using the first syntax element identification information. That is, in some embodiments, the code stream is decoded to determine the value of the first syntax element identification information; when the value of the first syntax element identification information is a first value, the prediction mode of the current block is determined to be the target filtering mode; when the value of the first syntax element identification information is a second value, the prediction mode of the current block is determined to be a non-target filtering mode.
- the first value is different from the second value, and the first value and the second value can be in parameter form or in digital form.
- the first syntax element identification information can be a parameter written in the profile or a value of a flag, which is not specifically limited here.
- the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to 0 and the second value can be set to 1; or, the first value can be set to true and the second value can be set to false; or, the first value can be set to false and the second value can be set to true.
- the first value is set to 1 and the second value is set to 0, but this is not specifically limited.
- the target filtering mode may include the reference area category of the current block and the shape of the target filter.
- the reference area category of the current block may include a first category, a second category, and a third category.
- the method may further include:
- the reference area category of the current block is the first category, determining that the reference area of the current block includes an upper adjacent area and a left adjacent area;
- the reference area category of the current block is the second category, determining that the reference area of the current block includes the upper adjacent area
- the reference area category of the current block is the third category, it is determined that the reference area of the current block includes the left adjacent area.
- the reference area of the current block refers to the reconstructed area around the current block.
- the upper adjacent area may refer to the reconstructed area adjacent to the upper side of the current block, and the left adjacent area may refer to the reconstructed area adjacent to the left side of the current block.
- the reference area category shown in FIG. 2A is the first category
- the reference area category shown in FIG. 2B is the second category
- the reference area category shown in FIG. 2C is the third category.
- the shape of the target filter may include a first shape, a second shape, and a third shape, wherein the first shape may be a 4 ⁇ 4 square, the second shape may be a 2 ⁇ 8 rectangle, and the third shape may be an 8 ⁇ 2 rectangle; however, this is not specifically limited here.
- the target filter shown in FIG. 3A is of a first shape
- the target filter shown in FIG. 3B is of a second shape
- the target filter shown in FIG. 3C is of a third shape.
- the candidate filter mode for the current block can be obtained by combining three reference area categories and three target filter shapes.
- nine candidate filter modes can be combined here in total, and the target filter mode is one of the nine candidate filter modes.
- S1102 Determine a reference area of the current block according to a size parameter of the current block and a target filtering mode.
- the reference area of the current block may be determined in combination with the size parameter of the current block, wherein the size parameter of the current block may include the height and width of the current block.
- the size of the reference area of the current block is associated with the shape and minimum parameter of the target filter. In short, if the size of the current block is large, a large reference area can be used; if the size of the current block is small, a small reference area can be used. The number of rows and columns of the reference area can be derived according to the size of the current block.
- FIG. 12A is a schematic diagram of a reference area of a current block
- FIG. 12B is a schematic diagram of another reference area of a current block
- FIG. 12C is a schematic diagram of a reference area of yet another current block.
- the area within the dotted box is the reference area of the current block, which depends on the shape of the target filter used by the current block and the size of the variable tplSize.
- the size of the variable tplSize is equal to the smaller value of the width and height of the current block. For example, for a 4 ⁇ 8 current block, the value of the variable tplSize is 4; for a 16 ⁇ 16 current block, the value of the variable tplSize is 16.
- the enabling of the filtering mode can also be limited according to the size parameters of the current block.
- the method may further include:
- determining the reference region category in the target prediction mode to be any one item except the second category
- the reference region category in the target prediction mode is determined to be any one item except the third category.
- the value of the first factor may be a first preset constant.
- the value of the first factor may be set to 2, but this is not specifically limited.
- the reference area category of the current block can be disabled as the second category, that is, it is prohibited to use the upper adjacent area of the current block to calculate the filter coefficient.
- the reference area category in the target prediction mode may only be the first category or the third category; if the ratio of the width to the height of the current block is greater than the first factor, that is, the height of the current block and the multiple of the first factor are less than the width of the current block, then the reference area category of the current block can be disabled as the third category, that is, it is prohibited to use the left adjacent area of the current block to calculate the filter coefficient.
- the reference area category in the target prediction mode may only be the first category or the second category.
- the number of candidate filter modes will be reduced accordingly. For example, if the reference area category of the current block is disabled as the second category (that is, the use of the upper adjacent area of the current block for calculating the filter coefficient is prohibited), the number of candidate filter modes will be reduced to six. In other words, since some interpolation filter modes are restricted according to the ratio of the width to the height of the current block (referred to as "aspect ratio"), the number of candidate filter modes allowed to be used under different aspect ratios is different; therefore, decoding can be performed based on the context model when parsing the target filter mode.
- decoding the bitstream and determining the target filtering mode of the current block may include: determining a context model of the current block; and decoding the bitstream based on the context model to determine the target filtering mode of the current block.
- the determination of the context model is associated with at least one of the following parameters:
- the ratio of the width to the height of the current block is the ratio of the width to the height of the current block.
- the selection of a context model may be related to factors such as the shape and aspect ratio of the current block.
- factors such as the shape and aspect ratio of the current block.
- the selection of a context model may be related to factors such as the shape and aspect ratio of the current block.
- there are multiple context models in the decoding end and which context model to use for decoding may be determined based on factors such as the shape and aspect ratio of the current block.
- the reason is that for a current block of narrow shape, since there may be fewer interpolation filter modes to select, the length of the codeword required to indicate a selected interpolation filter mode is short; while current blocks of other shapes allow different numbers of interpolation filter modes to be selected, and the length of the codeword required to indicate a certain interpolation filter mode is also long, which makes the probability of selecting an interpolation filter mode under different shapes different.
- different probabilities require the selection of different context models, and here different context model indexes can be used to determine which specific context model is used.
- the value of the first syntax element identification information can be decoded according to the context model, and then the target filtering mode of the current block can be determined. It can also reduce the computational complexity.
- S1103 Determine a filter coefficient of the current block according to a reference area of the current block.
- the filter coefficient of the current block is mainly determined according to the reference area of the current block and the shape of the target filter.
- determining the filter coefficient of the current block according to the reference area of the current block may include:
- the target filter determines the input value of the target filter corresponding to at least one reference pixel in the reference area and the output value of the target filter; according to the input value of the target filter corresponding to at least one reference pixel, determine the autocorrelation coefficient matrix; according to the input value of the target filter corresponding to at least one reference pixel and the output value of the target filter, determine the cross-correlation coefficient vector; determine the coefficients of the target filter according to the autocorrelation coefficient matrix and the cross-correlation coefficient vector; determine the coefficients of the target filter as the filter coefficients of the current block.
- the intra-frame prediction technology based on interpolation filtering determines the filter shape and reference area category corresponding to the current block by parsing relevant syntax elements at the decoding end, then traverses each position in the reference area to construct an autocorrelation coefficient matrix and a cross-correlation coefficient vector, and then obtains the filter coefficient by solving the set of equations.
- the autocorrelation coefficient matrix can be represented by A
- the mutual correlation coefficient vector can be represented by Y, as follows:
- t represents the reconstructed pixel value
- r represents the coordinate position of the reference area
- p 0 ...p N-1 represents the coordinate relationship relative to position r
- the relative coordinates they refer to are the relative coordinate relationship between the input position and the output position of the target filter.
- c 0 ...c N-1 are the filter coefficients to be solved
- m is a value subtracted from the input of the target filter (a value added to the output at this time).
- S1104 Perform intra-frame prediction on the current block according to the filter coefficient to determine a prediction value of the current block.
- intra-frame prediction is performed on pixels in the current block according to the filter coefficient to determine the predicted value of the pixels in the current block, which may include: determining a reference sample value corresponding to the pixel to be predicted in the current block; determining the predicted value of the pixel to be predicted in the current block according to the reference sample value and the filter coefficient corresponding to the pixel to be predicted in the current block.
- determining the reference sample value corresponding to the pixel to be predicted in the current block may include: based on the shape of the target filter, if the reference sample value is located in a reference area of the current block, determining the reconstructed value at the corresponding position in the reference area as the reference sample value; if the reference sample value is located inside the current block, determining the predicted value at the corresponding position in the current block as the reference sample value.
- the target filter that is, the reference sample value corresponding to the pixel to be predicted in the current block
- the reconstructed value is used as the input of the target filter; or, if the corresponding position is in the current block, then the predicted value that has been predicted is used as the input of the target filter.
- the interpolation filtering is predicted in the diagonal direction; and the pixels to be predicted on the same diagonal line can be predicted in parallel, as shown in FIG. 5 for details.
- the method may include:
- S1301 Determine a first input value of a target filter based on reference sample values corresponding to pixels to be predicted in a current block.
- determining the first input value of the target filter based on the reference sample value corresponding to the pixel to be predicted in the current block may include: determining a second factor; performing a subtraction operation on the reference sample value and the second factor to obtain the first input value of the target filter.
- S1302 Determine a first output value of a target filter based on a first input value and a filter coefficient.
- determining the first output value of the target filter based on the first input value and the filter coefficient may include: determining the second output value of the target filter based on the first input value and the filter coefficient; performing a first processing on the second output value to determine the first output value of the target filter.
- determining the second output value of the target filter based on the first input value and the filter coefficient may include: calculating the product of the first input value and the corresponding filter coefficient; setting the second output value of the target filter to be equal to the sum of n products; wherein n represents the number of input items corresponding to the target filter, and n is a positive integer.
- the reference sample value corresponding to the pixel r to be predicted in the current block can be
- the second output value of the target filter is represented by P out1 , as shown in the following formula:
- performing the first processing on the second output value to determine the first output value of the target filter may include: performing an addition operation on the second output value and the second factor to obtain the first output value of the target filter.
- the first output value of the target filter can be represented by P out2 , where:
- the value of the second factor may be a second preset constant.
- the method may further include: determining a reconstruction value of at least one reference pixel in the reference area; performing mean calculation on the reconstruction value of at least one reference pixel to obtain a first mean; and setting the value of the second factor to be equal to the first mean.
- the second factor may be obtained by calculating the mean value of the reconstructed values in the reference area, or may be a preset constant, or may even be a specific value, such as the reconstructed value of the upper left corner of the current block, which is not specifically limited here.
- the second factor is the mean value of the reference area
- the input of the target filter needs to subtract the mean value
- the output of the target filter needs to add the mean value to serve as the final prediction result.
- performing a first processing on the second output value to determine the first output value of the target filter may include: determining the third output value of the target filter; determining the fourth output value of the target filter based on the second output value and the third output value; and adding the fourth output value and the second factor to obtain the first output value of the target filter.
- the number of input items includes not only the number of linear items, but also the number of nonlinear items and/or the number of bias items.
- the third output value can be calculated based on the number of nonlinear items and/or the number of bias items
- the second output value can be calculated based on the number of linear items.
- the second output value of the target filter it can be specifically: calculate the product of the first input value and the corresponding filter coefficient; set the second output value of the target filter to be equal to the sum of n products; wherein n represents the number of first type input items corresponding to the target filter, and n is a positive integer.
- the third output value is calculated based on the number of nonlinear terms.
- determining the third output value of the target filter may include: determining the number of first-type input terms corresponding to the target filter based on the shape of the target filter; if the number of first-type input terms corresponding to the target filter is p, then determining p+q filter coefficients of the target filter, where p and q are both positive integers; and determining the third output value of the target filter based on q filter coefficients and q second-type input terms among the p+q filter coefficients.
- the third output value is calculated based on the number of bias items.
- determining the third output value of the target filter may include: determining the number of first-type input items corresponding to the target filter based on the shape of the target filter; if the number of first-type input items corresponding to the target filter is p, then determining p+m filter coefficients of the target filter, where p and m are both positive integers; and determining the third output value of the target filter based on m filter coefficients and m third-type input items among the p+m filter coefficients.
- the third output value is calculated based on the number of nonlinear terms and the number of bias terms.
- the number of first-type input items is the number of linear items
- the number of second-type input items is the number of nonlinear items
- the number of third-type input items is the number of bias items.
- the calculation formula of the first output value of the current position is:
- the corresponding nonlinear term values should also be added when constructing the autocorrelation coefficient matrix and the mutual correlation coefficient vector; in addition, when there is a bias term, the bias term value should also be further increased; this is set according to actual conditions and is not specifically limited here.
- nonlinear terms in addition to using three nonlinear terms, the embodiments of the present application may also use more nonlinear terms.
- five nonlinear terms are used in FIG. 16A , FIG. 16B , and FIG. 16C , and the positions of the five nonlinear terms are specifically five positions filled with dots.
- the number of nonlinear terms should be a positive integer, the specific number is not limited, and different designs can be performed according to performance complexity requirements.
- S1303 Determine a prediction value of a pixel to be predicted in the current block according to the first output value.
- determining the predicted value of the pixel to be predicted in the current block according to the first output value may include: performing a second processing on the first output value to obtain the predicted value of the pixel to be predicted in the current block.
- the second processing may be to set the prediction value of the to-be-predicted pixel in the current block to be equal to the first output value.
- the second processing may be to limit the first output value within a preset value range, or it may also be referred to as a "clip operation" herein, wherein the lower limit value of the preset value range is the minimum reconstruction value (min) in the reference area, and the upper limit value of the preset value range is the maximum reconstruction value (max) in the reference area.
- the method may also include:
- the luminance component of the current block uses intra prediction based on the filter coefficient, determining a derived intra prediction mode of the luminance component of the current block;
- the direct mode is set as the derived intra prediction mode to determine the prediction value of the chrominance component of the current block.
- the derived intra-frame prediction mode may be a traditional PLANAR mode, a DC mode or an angle mode, etc., which may be specifically determined according to the aforementioned method of constructing a gradient histogram.
- an efficient intra-frame chrominance prediction mode is used in many standards when performing intra-frame prediction.
- the chrominance block selects to use the DM mode, the chrominance block will obtain the mode selected by the luminance block at the corresponding position for intra-frame prediction.
- the interpolation filtering technology described in the above embodiments only works on intra-frame block prediction of luminance.
- a direct approach is to extend the mode to chrominance, but this will result in the need to derive filter coefficients for chrominance, which will bring high computational complexity.
- the chrominance block selects the DM mode, the DM mode will be set to the PLANAR mode for prediction.
- a traditional prediction mode can be derived by constructing a gradient histogram. This traditional mode can be used when the chrominance mode selects the DM mode and the luminance block at the corresponding position selects the interpolation filtering mode.
- the method may also include:
- the reference block uses intra prediction based on filter coefficients, determining a derived intra prediction mode for the reference block;
- the derived intra prediction mode is added to the intra prediction mode candidate list for the current block.
- the current block satisfies a preset condition, including at least one of the following:
- the current block is an inter-frame prediction block
- the current block is an IBC block.
- the IBC block and the inter-frame block are not intra-coded blocks, so they do not have an intra-frame prediction mode
- the initial reference blocks of the IBC block and the inter-frame block are intra-frame prediction blocks.
- the intra-frame prediction mode of the reference block is also transferred to the current block at the same time.
- These intra-frame prediction modes are traditional intra-frame prediction modes (PLANAR, DC, angle mode). These transferred traditional intra-frame prediction modes will be used when the surrounding blocks are IBC blocks or inter-frame blocks when constructing the intra-frame prediction mode candidate list for the current block.
- the traditional intra-frame prediction mode corresponding to the interpolation filter mode is used for transmission.
- the method may also include:
- a reconstructed value of the current block is determined according to the predicted value of the current block and the residual value of the current block.
- decoding the code stream and determining the residual value of the current block may include: decoding the code stream and determining the quantization coefficient of the current block; performing inverse quantization processing on the quantization coefficient to obtain the transformation coefficient of the current block; performing inverse transformation processing on the transformation coefficient to obtain the residual value of the current block.
- the encoder will calculate the residual value based on the original value and the predicted value, and the residual value will be further transformed and quantized to obtain the quantization coefficient, and then transmitted to the decoder through the code stream.
- the decoder can obtain the quantization coefficient of the current block through decoding, and then obtain the residual value of the current block through inverse quantization and inverse transformation; then, the residual value of the current block and the predicted value of the current block are further added to obtain the reconstructed value of the current block.
- performing inverse transform processing on the transform coefficients to obtain the residual value of the current block may include: when the current block uses a multi-transform selection mode and the target filtering mode is an interpolation filtering mode, determining a target transform kernel of the current block; performing inverse transform processing on the transform coefficients according to the target transform kernel to obtain the residual value of the current block.
- the determination of the target transformation kernel may be associated with at least one of the following parameters:
- the prediction result of the interpolation filter prediction is derived into a gradient histogram and matched to the traditional prediction mode, and the method of further selecting an inseparable transformation kernel is used.
- the selection of the transformation kernel is the same as that of the PLANAR mode.
- the interpolation filter mode has different characteristics from the PLANAR mode, and the selection of the basic transformation kernel should be more optimized.
- the basic transformation can be divided into horizontal and vertical directions.
- the transformation modes allowed in each direction include the following 7 types: ⁇ 'DCT2', 'DCT8', 'DST7', 'DCT5', 'DST4', 'DST1', 'IDTR' ⁇ .
- DCT2, DCT8, and DCT5 are subclasses of discrete cosine transform
- DST7, DST4, and DST1 are subclasses of discrete sine transform
- IDTR is Identity transform, which means no transformation.
- the most commonly used basic transform mode is DCT2 in both horizontal and vertical directions, here written as DCT2-DCT2, which is used as a transform before the inseparable secondary transform LFNST, and is also used as a transform when the multi-transform selection MTS technology is turned off.
- DCT2-DCT2 is used as a transform before the inseparable secondary transform LFNST, and is also used as a transform when the multi-transform selection MTS technology is turned off.
- the transform process will be a combination of the basic transforms in the horizontal and vertical directions, rather than an inseparable transform.
- the method may further include: decoding a code stream to determine non-zero coefficient information of a current block; and determining at least one candidate transform kernel according to the non-zero coefficient information of the current block.
- the number of at least one candidate transform kernel is less than or equal to 6. That is, in the reference software ECM, according to the characteristics of the non-zero coefficients in the current block analyzed, the current block can have at most 6 non-DCT2-DCT2 transform kernels to choose from.
- the residual MTS basic transform kernel should be related to whether the current block selects the interpolation filtering mode. More specifically, it can be related to which interpolation filtering mode is selected and/or the size and shape of the current block.
- determining a target transform core of a current block may include: decoding a bitstream to determine a transform core index value of the current block; and determining a target transform core of the current block from at least one candidate transform core according to the transform core index value.
- the optional basic transform kernel of MTS is related to whether the interpolation filter prediction mode is selected for the current block. If the current block uses the interpolation filter prediction mode, the 6 optional MTS transform kernels are as follows (the transform kernel is: horizontal transform-vertical transform), as shown in Table 3.
- a corresponding target transform core is selected from six transform cores according to the parsed MTS transform core index value for inverse transformation.
- determining the target transform core of the current block may include: decoding a bitstream to determine a transform core index value of the current block; and determining the target transform core of the current block from at least one candidate transform core according to the transform core index value and a size parameter of the current block.
- the optional basic transformation of MTS is related to whether the interpolation filter mode is selected for the current block and the size and shape of the current block.
- the shape size of the current block is: height ⁇ width. In one embodiment, it can be as shown in Table 4.
- the prediction mode of the current block is the interpolation prediction mode
- the corresponding target transform kernel is selected for inverse transformation according to the parsed MTS transform kernel index value and the shape and size of the current block.
- the interpolation filter prediction mode can be applied to 4 ⁇ 4 to 32 ⁇ 32 luminance blocks.
- the method for obtaining the candidate MTS transformation core may include:
- Step 1 encoding an image set or a video set using an encoder including an interpolation filtering prediction mode
- Step 2 The residual value of the block with the selected interpolation filter mode is classified into possible horizontal-vertical transform kernels one by one according to the classification (e.g., the shape and size of the block, the interpolation filter mode, etc.).
- the transform kernel selection criteria can be the size of SAD, the size of SSE, or other criteria, such as transform coding gain, which are not specifically limited here.
- the transform coding gain is defined as the arithmetic mean transform coefficient variance divided by the geometric mean transform coefficient variance.
- This embodiment provides a decoding method, which decodes a bit stream, determines a target filtering mode of a current block; determines a reference area of the current block according to the size parameter of the current block and the target filtering mode; then determines a filtering coefficient of the current block according to the reference area of the current block; and then performs intra-frame prediction on the current block according to the filtering coefficient to determine a prediction value of the current block.
- the intra-frame prediction technology based on interpolation filtering is not only related to the target filtering mode when determining the reference area for calculating the filtering coefficient, but also to the size parameter of the current block.
- a large reference area can be used when the size of the current block is large, and a small reference area can be used when the size of the current block is small.
- the computational complexity can be reduced, the encoding time can be reduced, and the accuracy of intra-frame prediction can be improved, thereby improving the encoding and decoding performance.
- FIG17 a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application is shown. As shown in FIG17, the method may include:
- S1801 Determine the target filtering mode of the current block.
- the encoding method of the embodiment of the present application may be an intra-frame prediction method, specifically referring to an improvement of an intra-frame prediction mode based on interpolation filtering to enhance the cost-effectiveness of performance and complexity.
- the current block includes at least a first color component and a second color component.
- the block at this time can be simply referred to as a first color component block; and when the first color component is a brightness component, the first color component block can also be referred to as a brightness block.
- the block at this time can be simply referred to as a second color component block; and when the second color component is a chrominance component, the second color component block can also be referred to as a chrominance block.
- the target filter mode may refer to a mode of using a target filter to perform intra-frame prediction on the current block, wherein the target filter may refer to an interpolation filter.
- determining a target filtering mode for a current block may include:
- the wave mode is determined as the target filtering mode for the current block.
- the number of at least one candidate filtering mode may be determined based on the number of reference region categories of the current block and the number of shapes of the target filter.
- the distortion value method can be used to determine the cost result, specifically, the rate-distortion cost method can be used to determine the cost result; however, it can also be the size of SAD, the size of MSE, the size of SSE or other standards for judging the cost, which are not specifically limited here.
- the reference area category of the current block may include a first category, a second category, and a third category.
- the method may further include:
- the reference area category of the current block is the first category, determining that the reference area of the current block includes an upper adjacent area and a left adjacent area;
- the reference area category of the current block is the second category, determining that the reference area of the current block includes the upper adjacent area
- the reference area category of the current block is the third category, it is determined that the reference area of the current block includes the left adjacent area.
- the reference area of the current block refers to the reconstructed area around the current block.
- the upper adjacent area may refer to the reconstructed area adjacent to the upper side of the current block
- the left adjacent area may refer to the reconstructed area adjacent to the left side of the current block.
- the reference area category shown in FIG. 2A is the first category
- the reference area category shown in FIG. 2B is the second category
- the reference area category shown in FIG. 2C is the third category.
- the shape of the target filter may include a first shape, a second shape, and a third shape, wherein the first shape may be a 4 ⁇ 4 square, the second shape may be a 2 ⁇ 8 rectangle, and the third shape may be an 8 ⁇ 2 rectangle; however, this is not specifically limited here.
- the target filter shown in FIG. 3A is of a first shape
- the target filter shown in FIG. 3B is of a second shape
- the target filter shown in FIG. 3C is of a third shape.
- the candidate filter mode for the current block can be obtained by combining three reference area categories and three target filter shapes.
- nine candidate filter modes can be combined here in total, and the target filter mode is one of the nine candidate filter modes.
- S1802 Determine a reference area of the current block according to a size parameter of the current block and a target filtering mode.
- the reference area of the current block can be determined in combination with the size parameter of the current block, where the size parameter of the current block can include the height and width of the current block.
- the size of the reference area of the current block is associated with the shape and minimum parameters of the target filter. Simply put, if the size of the current block is large, a large reference area can be used; if the size of the current block is small, a small reference area can be used. Among them, the number of rows and columns of the reference area can be derived according to the size of the current block. Exemplarily, as shown in Figures 12A, 12B, and 12C, the area in the dotted box is the reference area of the current block, which depends on the shape of the target filter used by the current block and the size of the variable tplSize.
- the size of the variable tplSize is equal to the smaller value of the width and height of the current block. For example, for a 4 ⁇ 8 current block, the value of the variable tplSize is 4; for a 16 ⁇ 16 current block, the value of the variable tplSize is 16.
- the enabling of the filtering mode can also be limited according to the size parameters of the current block.
- the method may also include:
- the reference area category of the current block is prohibited from being the third category, and the number of reference area categories of the current block is determined based on other reference area categories except the third category.
- the value of the first factor may be a first preset constant.
- the value of the first factor may be set to 2, but this is not specifically limited.
- the reference area category of the current block can be disabled as the second category, that is, it is prohibited to use the upper adjacent area of the current block to calculate the filter coefficient.
- the reference area category in the target prediction mode can only be the first category or the third category; if the ratio of the width to the height of the current block is greater than the first factor, that is, the height of the current block and the multiple of the first factor are less than the width of the current block, then the current block can be disabled.
- the reference area category of the block is the third category, that is, it is forbidden to use the left adjacent area of the current block to calculate the filter coefficient.
- the reference area category in the target prediction mode can only be the first category or the second category.
- the number of candidate filter modes will be reduced accordingly. For example, if the reference area category of the current block is disabled as the second category (i.e., the use of the upper adjacent area of the current block for calculating the filter coefficient is prohibited), the number of candidate filter modes will be reduced to six. In other words, since some interpolation filter modes are restricted according to the ratio of the width to the height of the current block (referred to as "aspect ratio"), the number of candidate filter modes allowed to be used under different aspect ratios is different; therefore, encoding of the target filter mode can be performed based on the context model.
- the method may further include: encoding the target filtering mode of the current block, and writing the obtained encoding bits into a bitstream.
- encoding the target filtering mode of the current block and writing the obtained coded bits into the bitstream may include: determining a context model of the current block; encoding the target filtering mode of the current block based on the context model, and writing the obtained coded bits into the bitstream.
- the determination of the context model is associated with at least one of the following parameters:
- the ratio of the width to the height of the current block is the ratio of the width to the height of the current block.
- the selection of a context model may be related to factors such as the shape and aspect ratio of the current block.
- factors such as the shape and aspect ratio of the current block.
- the selection of a context model may be related to factors such as the shape and aspect ratio of the current block.
- there are multiple context models in the decoding end and which context model to use for decoding may be determined based on factors such as the shape and aspect ratio of the current block.
- the reason is that for a current block of narrow shape, since there may be fewer interpolation filter modes to select, the length of the codeword required to indicate a selected interpolation filter mode is short; while current blocks of other shapes allow different numbers of interpolation filter modes to be selected, and the length of the codeword required to indicate a certain interpolation filter mode is also long, which makes the probability of selecting an interpolation filter mode under different shapes different.
- different probabilities require the selection of different context models, and here different context model indexes can be used to determine which specific context model is used.
- the target prediction mode can be encoded according to the context model, so that the target prediction mode of the current block can be obtained by decoding the bit stream at the decoding end according to the context model selected according to the shape of the current block.
- the target filtering mode may include the reference area category of the current block and the shape of the target filter.
- the target filtering mode may be written into the bitstream via the first syntax element identification information. That is, in some embodiments, the value of the first syntax element identification information is determined; the value of the first syntax element identification information is encoded based on the context model, and the obtained encoded bits are written into the bitstream.
- determining the value of the first syntax element identification information may include: if the current block uses a target filtering mode for prediction encoding, determining the value of the first syntax element identification information to be a first value; if the current block uses a non-target filtering mode for prediction encoding, determining the value of the first syntax element identification information to be a second value.
- the first value is different from the second value, and the first value and the second value can be in parameter form or in digital form.
- the first syntax element identification information can be a parameter written in the profile or a value of a flag, which is not specifically limited here.
- the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to 0 and the second value can be set to 1; or, the first value can be set to true and the second value can be set to false; or, the first value can be set to false and the second value can be set to true.
- the first value is set to 1 and the second value is set to 0, but this is not specifically limited.
- the decoding end can subsequently determine whether the prediction mode of the current block is the target prediction mode by parsing the value of the first syntax element identification information. Exemplarily, if the value of the first syntax element identification information obtained by parsing is 1, then it can be determined that the prediction mode of the current block is the target prediction mode. In this way, not only the prediction accuracy can be improved, but also the computational complexity can be reduced.
- S1803 Determine a filter coefficient of the current block according to a reference area of the current block.
- the reference area of the current block can be classified and divided to construct an autocorrelation coefficient matrix and a cross-correlation coefficient vector that do not contain repeated areas, so that the encoding end can derive the filter coefficients under each combination mode.
- the method may also include:
- the first candidate sub-reference region is any one of the multiple candidate sub-reference regions.
- determining the autocorrelation coefficient matrix and the cross-correlation coefficient vector of the first candidate sub-reference region may specifically include: According to the shapes of the first candidate sub-reference region and the target filter, the input value of the target filter and the output value of the target filter corresponding to at least one reference pixel in the first candidate sub-reference region are determined; according to the input value of the target filter corresponding to at least one reference pixel, the autocorrelation coefficient matrix of the first candidate sub-reference region is determined; according to the input value of the target filter and the output value of the target filter corresponding to at least one reference pixel, the cross-correlation coefficient vector of the first candidate sub-reference region is determined. In this way, the autocorrelation coefficient matrix and the cross-correlation coefficient vector of each of the multiple candidate sub-reference regions can be determined.
- the encoder needs to select from 9 combinations of 3 reference areas and 3 filter shapes. When a certain combination is selected, the corresponding syntax element identification information will be written into the bitstream; then, if the decoder parses and finds that a certain combination is selected, only one filter coefficient needs to be derived. This makes the complexity of the technology at the encoder much higher than that at the decoder.
- the reference region of the current block can be divided into R0, R1 and R2.
- the reference region of the current block can be divided into R0 and R1.
- the reference region of the current block can be divided into R0 and R2.
- f 0 , f 1 , and f 2 are used to represent three filter shapes, and R all , R top , and R left are used to represent three reference regions.
- the autocorrelation coefficient matrix and the cross-correlation coefficient vector under the nine combinations can be written as follows:
- A represents the autocorrelation coefficient matrix
- Y represents the cross-correlation coefficient vector
- R all , R top , and R left can all be composed of R 0 , R 1 , and R 2 , the above 9 combinations can be further decomposed into:
- determining the filter coefficient of the current block according to the reference area of the current block may include: dividing the reference area of the current block to determine at least one sub-reference area; obtaining the autocorrelation coefficient matrix and cross-correlation coefficient vector of each of the at least one sub-reference area from a preset buffer; determining the coefficient of the target filter according to the autocorrelation coefficient matrix and cross-correlation coefficient vector of each of the at least one sub-reference area; and determining the coefficient of the target filter as the filter coefficient of the current block.
- the encoding end when the encoding end derives the filter coefficients based on the current combination and performs rate-distortion optimization, each time an autocorrelation coefficient matrix and a cross-correlation coefficient vector that have not been constructed are encountered, they need to be cached for use in subsequent other combinations; thereby reducing the computational complexity of the encoding end.
- S1804 Perform intra-frame prediction on the current block according to the filter coefficients to determine a prediction value of the current block.
- intra-frame prediction is performed on pixels in the current block according to the filter coefficient to determine the predicted value of the pixels in the current block, which may include: determining a reference sample value corresponding to the pixel to be predicted in the current block; determining the predicted value of the pixel to be predicted in the current block according to the reference sample value and the filter coefficient corresponding to the pixel to be predicted in the current block.
- determining the reference sample value corresponding to the pixel to be predicted in the current block may include: based on the shape of the target filter, if the reference sample value is located in a reference area of the current block, determining the reconstructed value at the corresponding position in the reference area as the reference sample value; if the reference sample value is located inside the current block, determining the predicted value at the corresponding position in the current block as the reference sample value.
- the target filter that is, the reference sample value corresponding to the pixel to be predicted in the current block
- the reconstructed value is used as the input of the target filter; or, if the corresponding position is in the current block, then the predicted value that has been predicted is used as the input of the target filter.
- the interpolation filtering is predicted in the diagonal direction; and the pixels to be predicted on the same diagonal line can be predicted in parallel, as shown in FIG. 5 for details.
- determining the predicted value of the pixel to be predicted in the current block according to the reference sample value and the filter coefficient corresponding to the pixel to be predicted in the current block may include:
- a prediction value of a pixel to be predicted in the current block is determined.
- determining the first input value of the target filter based on the reference sample value corresponding to the pixel to be predicted in the current block may include: determining a second factor; performing a subtraction operation on the reference sample value and the second factor to obtain the first input value of the target filter.
- determining the first output value of the target filter based on the first input value and the filter coefficient may include: determining the second output value of the target filter based on the first input value and the filter coefficient; performing a first processing on the second output value to determine the first output value of the target filter.
- determining the second output value of the target filter based on the first input value and the filter coefficient may include: calculating the product of the first input value and the corresponding filter coefficient; setting the second output value of the target filter to be equal to the sum of n products; wherein n represents the number of input items corresponding to the target filter, and n is a positive integer.
- the reference sample value corresponding to the pixel r to be predicted in the current block can be
- the second output value of the target filter is represented by P out1 , as shown in the following formula:
- performing the first processing on the second output value to determine the first output value of the target filter may include: performing an addition operation on the second output value and the second factor to obtain the first output value of the target filter.
- the first output value of the target filter can be represented by P out2 , where:
- the value of the second factor may be a second preset constant.
- the method may further include: determining a reconstruction value of at least one reference pixel in the reference area; performing mean calculation on the reconstruction value of at least one reference pixel to obtain a first mean; and setting the value of the second factor to be equal to the first mean.
- the second factor may be obtained by calculating the mean value of the reconstructed values in the reference area, or may be a preset constant, or may even be a specific value, such as the reconstructed value of the upper left corner of the current block, which is not specifically limited here.
- the second factor is the mean value of the reference area
- the input of the target filter needs to subtract the mean value
- the output of the target filter needs to add the mean value to serve as the final prediction result.
- performing a first processing on the second output value to determine the first output value of the target filter may include: determining the third output value of the target filter; determining the fourth output value of the target filter based on the second output value and the third output value; and adding the fourth output value and the second factor to obtain the first output value of the target filter.
- the number of input items includes not only the number of linear items, but also the number of nonlinear items and/or the number of bias items.
- the third output value can be calculated based on the number of nonlinear items and/or the number of bias items
- the second output value can be calculated based on the number of linear items.
- the second output value of the target filter it can be specifically: calculate the product of the first input value and the corresponding filter coefficient; set the second output value of the target filter to be equal to the sum of n products; wherein n represents the number of first type input items corresponding to the target filter, and n is a positive integer.
- the third output value is calculated based on the number of nonlinear terms.
- determining the third output value of the target filter may include: determining the number of first-type input terms corresponding to the target filter based on the shape of the target filter; if the number of first-type input terms corresponding to the target filter is p, determining p+q filter coefficients of the target filter, where p and q are both positive integers; A third output value of the target filter is determined according to q filter coefficients among the p+q filter coefficients and q second type input items.
- the third output value is calculated based on the number of bias items.
- determining the third output value of the target filter may include: determining the number of first-type input items corresponding to the target filter based on the shape of the target filter; if the number of first-type input items corresponding to the target filter is p, then determining p+m filter coefficients of the target filter, where p and m are both positive integers; and determining the third output value of the target filter based on m filter coefficients and m third-type input items among the p+m filter coefficients.
- the third output value is calculated based on the number of nonlinear terms and the number of bias terms.
- the number of first-type input items is a linear number of items
- the number of second-type input items is a nonlinear number of items
- the number of third-type input items is a biased number of items.
- the linear items of 15 taps are grid filling positions
- the nonlinear items of 3 taps are dot filling positions
- the black filling positions represent the current positions to be predicted.
- the calculation formula of the first output value of the current position is:
- the corresponding nonlinear term values should also be added when constructing the autocorrelation coefficient matrix and the mutual correlation coefficient vector; in addition, when there is a bias term, the bias term value should also be further increased; this is set according to actual conditions and is not specifically limited here.
- Figures 15A, 15B, and 15C can also be shown in Figures 15A, 15B, and 15C. Compared with Figures 14A, 14B, and 14C, Figures 15A, 15B, and 15C all add three nonlinear terms, but since different filter shapes use the same nonlinear terms, the calculation is simpler, further reducing the complexity.
- the embodiment of the present application can also use more nonlinear terms, for example, 5 nonlinear terms are used in FIG. 16A, FIG. 16B, and FIG. 16C, and the positions of the 5 nonlinear terms are specifically five positions filled with dots.
- the number of nonlinear terms should be a positive integer, the specific number is not limited, and different designs can be performed according to the performance complexity requirements.
- the predicted value of the pixel to be predicted in the current block can be determined based on the first output value, which can specifically include: performing a second processing on the first output value to obtain the predicted value of the pixel to be predicted in the current block.
- the second processing may be to set the prediction value of the to-be-predicted pixel in the current block to be equal to the first output value.
- the second processing may be to limit the first output value within a preset value range, or it may also be referred to as a "clip operation" herein, wherein the lower limit value of the preset value range is the minimum reconstruction value (min) in the reference area, and the upper limit value of the preset value range is the maximum reconstruction value (max) in the reference area.
- the method may also include:
- the luminance component of the current block uses intra prediction based on the filter coefficient, determining a derived intra prediction mode of the luminance component of the current block;
- the direct mode is set as the derived intra prediction mode to determine the prediction value of the chrominance component of the current block.
- the derived intra-frame prediction mode may be the traditional PLANAR mode, DC mode or The angle mode, etc. can be specifically determined according to the aforementioned method of constructing the gradient histogram.
- an efficient intra-frame chrominance prediction mode is used in many standards when performing intra-frame prediction.
- the chrominance block selects to use the DM mode, the chrominance block will obtain the mode selected by the luminance block at the corresponding position for intra-frame prediction.
- the interpolation filtering technology described in the above embodiments only works on intra-frame block prediction of luminance.
- a direct approach is to extend the mode to chrominance, but this will result in the need to derive filter coefficients for chrominance, which will bring high computational complexity.
- the chrominance block selects the DM mode, the DM mode will be set to the PLANAR mode for prediction.
- a traditional prediction mode can be derived by constructing a gradient histogram. This traditional mode can be used when the chrominance mode selects the DM mode and the luminance block at the corresponding position selects the interpolation filtering mode.
- the method may also include:
- the reference block uses intra prediction based on filter coefficients, determining a derived intra prediction mode for the reference block;
- the derived intra prediction mode is added to the intra prediction mode candidate list for the current block.
- the current block satisfies a preset condition, including at least one of the following:
- the current block is an inter-frame prediction block
- the current block is an IBC block.
- IBC blocks and inter-frame blocks they are not intra-coded blocks, so they do not have intra-frame prediction modes, and the initial reference blocks of IBC blocks and inter-frame blocks are intra-frame prediction blocks.
- the intra-frame prediction mode of the reference block is also transferred to the current block at the same time.
- These intra-frame prediction modes are traditional intra-frame prediction modes (PLANAR, DC, angle mode).
- PLANAR, DC, angle mode traditional intra-frame prediction modes
- These transferred traditional intra-frame prediction modes will be used when the surrounding blocks are IBC blocks or inter-frame blocks when constructing the intra-frame prediction mode candidate list for the current block. In this way, when the position of the IBC block or inter-frame block reference is the interpolation filter mode, the traditional intra-frame prediction mode corresponding to the interpolation filter mode is used for transmission.
- the method may further include:
- S2004 Encode the quantization coefficients of the current block, and write the obtained coded bits into the bitstream.
- determining the residual value of the current block may include: determining the original value of the current block; determining the residual value of the current block according to the original value of the current block and the predicted value of the current block. Then, encoding the residual value of the current block, and writing the obtained coded bits into the bitstream.
- the residual value of the current block can be determined by performing a subtraction operation based on the original value of the current block and the predicted value of the current block.
- the residual value needs to be transformed and quantized first, and the obtained quantization coefficient is written into the bit stream, and then transmitted to the decoding end through the bit stream.
- transforming the residual value to obtain the transform coefficient of the current block may include: when the current block uses a multi-transform selection mode and the target filtering mode is an interpolation filtering mode, determining the target transform kernel of the current block; transforming the residual value according to the target transform kernel to obtain the transform coefficient of the current block.
- the determination of the target transformation kernel may be associated with at least one of the following parameters:
- the prediction result of the interpolation filter prediction is derived into a gradient histogram and matched to the traditional prediction mode, and the method of further selecting an inseparable transformation kernel is used.
- the selection of the transformation kernel is the same as that of the PLANAR mode.
- the interpolation filter mode has different characteristics from the PLANAR mode, and the selection of the basic transformation kernel should be more optimized.
- the basic transformation can be divided into horizontal and vertical directions.
- the transformation modes allowed in each direction include the following 7 types: ⁇ 'DCT2', 'DCT8', 'DST7', 'DCT5', 'DST4', 'DST1', 'IDTR' ⁇ .
- DCT2, DCT8, and DCT5 are subclasses of discrete cosine transform
- DST7, DST4, and DST1 are subclasses of discrete sine transform
- IDTR is Identity transform, which means no transformation.
- DCT2 in both horizontal and vertical directions, written here as DCT2-DCT2, which is used as a transform before the inseparable secondary transform LFNST and is also used in the multi-transform selection MTS technology.
- DCT2-DCT2 is used as a transform before the inseparable secondary transform LFNST and is also used in the multi-transform selection MTS technology.
- MTS mode the transformation process will be a combination of the basic transformations in the horizontal and vertical directions, rather than an inseparable transformation.
- the method may further include: determining non-zero coefficient information of the current block; and determining at least one candidate transform kernel according to the non-zero coefficient information of the current block.
- the number of at least one candidate transform kernel is less than or equal to 6. That is, in the reference software ECM, according to the characteristics of the non-zero coefficients in the current block determined after transformation and quantization, the current block can have at most 6 non-DCT2-DCT2 transform kernels to choose from.
- the residual MTS basic transform kernel should be related to whether the current block selects the interpolation filtering mode. More specifically, it can be related to which interpolation filtering mode is selected and/or the size and shape of the current block.
- determining the target transformation core of the current block may include: determining at least one candidate transformation core; performing cost calculation on the at least one candidate transformation core to determine the cost result of the at least one candidate transformation core; determining the minimum cost result from the cost results of the at least one candidate transformation core, and determining the candidate transformation core corresponding to the minimum cost result as the target transformation core of the current block.
- the cost result can be determined by using a distortion value, specifically, the cost result can be determined by using a rate-distortion cost; however, it can also be the size of the SAD, the size of the MSE, the size of the SSE or other judgment criteria, such as the transform coding gain, which is not specifically limited here.
- the method may further include: determining a transform core index value of a current block, wherein the transform core index value is used to indicate an index number of a target transform core in at least one candidate transform core; encoding the transform core index value of the current block, and writing the resulting encoded bits into a bitstream.
- the optional basic transform kernel of MTS is related to whether the interpolation filter prediction mode is selected for the current block. If the current block uses the interpolation filter prediction mode, the 6 optional MTS transform kernels are as follows (the transform kernel is: horizontal transform-vertical transform), as shown in Table 3.
- the MTS transform core index value can be determined and written into the bitstream; so that the subsequent decoding end can select the corresponding target transform core from the six transform cores for inverse transformation according to the parsed MTS transform core index value.
- the method may further include: determining a transform core index value of the current block, wherein the transform core index value is used to indicate an index number of a target transform core in at least one candidate transform core, and at least one candidate transform core has an associated relationship with a size parameter of the current block; encoding the transform core index value of the current block, and writing the resulting encoded bits into a bitstream.
- the optional basic transform kernel of the MTS is related to whether the interpolation filter mode is selected for the current block and the size and shape of the current block; as shown in Table 4.
- the shape size of the current block is: height ⁇ width.
- the MTS transform kernel index value can be determined in combination with the size parameter of the current block and written into the bitstream; so that the subsequent decoding end can select the corresponding target transform kernel for inverse transformation according to the parsed MTS transform kernel index value and the shape and size of the current block.
- the interpolation filter prediction mode can be applied to luminance blocks of 4 ⁇ 4 to 32 ⁇ 32.
- the method for obtaining the candidate MTS transformation core may include:
- Step 1 encoding an image set or a video set using an encoder including an interpolation filtering prediction mode
- Step 2 The residual value of the block with the selected interpolation filter mode is classified into possible horizontal-vertical transform kernels one by one according to the classification (e.g., the shape and size of the block, the interpolation filter mode, etc.).
- the transform kernel selection criteria can be the size of SAD, the size of SSE, or other criteria, such as transform coding gain, which are not specifically limited here.
- the transform coding gain is defined as the arithmetic mean transform coefficient variance divided by the geometric mean transform coefficient variance.
- an embodiment of the present application further provides a code stream, which is generated by bit encoding according to the information to be encoded; wherein the information to be encoded includes at least one of the following:
- the target filtering mode of the current block the residual value of the current block, and the transform kernel index value of the current block.
- the target filter mode of the current block when writing the bitstream, can also be written into the bitstream by taking the value of the first syntax element identification information.
- the residual value of the current block can also be the quantization coefficient obtained after the residual value is transformed and quantized, and written into the bitstream.
- the encoding end In order to facilitate the decoding end to quickly determine the target transform kernel used, the encoding end also needs to write the transform kernel index value of the current block into the bitstream; thereby improving the encoding and decoding efficiency.
- This embodiment provides a coding method, which determines the target filtering mode of the current block; determines the reference area of the current block according to the size parameter of the current block and the target filtering mode; then determines the filter coefficient of the current block according to the reference area of the current block; and then performs intra-frame prediction on the current block according to the filter coefficient to determine the prediction value of the current block.
- the intra-frame prediction technology based on interpolation filtering is not only related to the target filtering mode, but also to the size parameters of the current block, such as the size of the current block.
- a large reference area can be used, and when the size of the current block is small, a small reference area can be used; in this way, the computational complexity can be reduced and the encoding time can be reduced; at the same time, the intra-frame prediction accuracy can be improved, thereby improving the encoding and decoding performance.
- the intra-frame prediction mode based on interpolation filtering is improved, and the improvements are described in detail from several aspects below.
- the reference area always uses a reconstructed area composed of 13 rows and/or 13 columns of reconstructed pixel values, which results in much higher computational complexity on small blocks than on large blocks.
- the encoder needs to decide the division of the blocks, and increasing the amount of calculation for small blocks is more likely to lead to an increase in encoding time.
- an embodiment of the present application proposes that large blocks use a large reference area and small blocks use a small reference area, and the number of rows and columns of the reference area can be derived according to the size of the block. See Figures 12A, 12B, and 12C above for details.
- interpolation filter mode when encoding and decoding the interpolation filter mode, some interpolation filter sub-modes are restricted according to the aspect ratio, so that the number of interpolation filter sub-modes allowed to be used under different aspect ratios is different. Therefore, when parsing the syntax element identifier of the interpolation filter, the selection of its context model should be related to the shape of the block and the aspect ratio factor. It should be noted that, assuming that three reference area categories and three filter shapes can form nine interpolation filter modes, each of which can be regarded as an interpolation filter sub-mode; in other words, the interpolation filter mode can include nine interpolation filter sub-modes.
- the reference areas can be classified and divided to construct an autocorrelation coefficient matrix and a cross-correlation coefficient vector that do not contain repeated areas, so as to be used by the encoding end to derive the filter coefficients of each combination.
- the interpolation filter prediction technology determines the shape of the interpolation filter selected by the current block and the reference area category by parsing relevant syntax elements at the decoding end, traverses each position in the reference area to construct an autocorrelation coefficient matrix and a cross-correlation coefficient vector, and solves the equation group to obtain the filter coefficient.
- t represents the reconstructed pixel value
- r represents the coordinate position of the reference area
- p 0 ...p N-1 represents the coordinate relationship relative to position r
- the relative coordinates they refer to are the relative coordinate relationship between the input position and the output position of the interpolation filter.
- c 0 ...c N-1 are the filter coefficients to be solved
- m is a value subtracted from the input of the interpolation filter (a value added to the output at this time).
- the encoding end needs to screen from a total of 9 combinations of 3 reference areas and 3 filter shapes.
- the corresponding syntax element is encoded into the bit stream. If the decoding end parses out that a certain combination is selected, the filter coefficient only needs to be derived once; this makes the complexity of the technology at the encoding end much higher than that at the decoding end.
- FIG. 2A, FIG. 2B and FIG. 2C they are all composed of three parts, R0, R1 and R2, as shown in FIG. 18A, FIG. 18B and FIG. 18C.
- f 0 , f 1 , f 2 are used to represent three filter shapes
- R all , R top , R left are used to represent three reference regions.
- the autocorrelation coefficient matrix and the cross-correlation coefficient vector under the 9 combinations can be written as follows:
- A represents the autocorrelation coefficient matrix and Y represents the cross-correlation coefficient vector.
- R all , R top , and R left can all be composed of R 0 , R 1 , and R 2 , the above 9 combinations can be further decomposed into:
- the encoder when the encoder derives filter coefficients based on the current combination and performs rate-distortion optimization, each time it encounters an unconstructed autocorrelation coefficient matrix and cross-correlation coefficient vector, it needs to cache them for use in subsequent other combinations; thereby reducing the computational complexity of the encoding end.
- the DM mode is an efficient intra-frame chrominance prediction mode used in many standards for prediction.
- the chrominance block selects the DM mode, the chrominance block will obtain the mode selected by the luminance block at the corresponding position for intra-frame prediction.
- the interpolation filtering technology described in the aforementioned embodiment only acts on the prediction of intra-frame blocks of luminance. A direct approach is to expand the mode to chrominance, but this will cause the chrominance to also need to derive filtering parameters, which will bring high computational complexity. In the related art, there is no interpolation filtering prediction mode for chrominance.
- the DM mode When the DM mode is selected for the chrominance intra-frame block, the DM mode will be set to the PLANAR mode.
- a traditional prediction mode can be derived by constructing a gradient histogram. This traditional mode can be used when the chrominance mode selects the DM mode and the luminance block at the corresponding position selects the interpolation filter mode.
- IBC blocks and inter-frame blocks they are not intra-coded blocks, so they do not have intra-frame prediction modes, and the initial reference blocks of IBC blocks and inter-frame blocks are intra-frame prediction blocks.
- the intra-frame prediction mode of the reference block is also transferred to the current block at the same time.
- These intra-frame prediction modes are traditional intra-frame prediction modes (PLANAR, DC, angle mode). These transferred traditional intra-frame prediction modes are used when the surrounding blocks are IBC blocks or inter-frame blocks when constructing the intra-frame prediction mode candidate list for the current block.
- the technology for constructing an intra prediction candidate list may include the following:
- Intra-coded blocks These blocks can use a range of intra prediction techniques, such as spatial geometric partitioning mode (SGPM), template-based multiple reference line intra prediction (TMRL), most probable mode (MPM), and template-based intra mode derivation (TIMD);
- SGPM spatial geometric partitioning mode
- TMRL template-based multiple reference line intra prediction
- MPM most probable mode
- TMD template-based intra mode derivation
- IBC Intra block copy
- the traditional mode corresponding to the interpolation filter mode is used for transmission.
- the encoder After the current block is predicted, the encoder will calculate the residual value by dividing the predicted value with the original value. The residual value will be further transformed and quantized. At the decoder, the quantization coefficient parsed from the bitstream will be inverse quantized and inverse transformed to obtain the reconstructed residual value. The reconstructed residual value is added to the predicted value to obtain the reconstructed value.
- the above-mentioned embodiment introduces a method of deriving a gradient histogram from the interpolation filter prediction result and matching it to the traditional prediction mode, and further selecting an inseparable transformation kernel.
- the selection of the transformation kernel is the same as that of the PLANAR mode.
- the interpolation filter mode has different characteristics from the PLANAR mode, and the selection of the basic transformation kernel should be more optimized.
- the basic transform is divided into horizontal and vertical directions, and the allowed transform modes in each direction include the following 7 types: ⁇ 'DCT2', 'DCT8', 'DST7', 'DCT5', 'DST4', 'DST1', 'IDTR' ⁇ .
- DCT2, DCT8, DCT5 are several subclasses of discrete cosine transform
- DST7, DST4, DST1 are several subclasses of discrete sine transform
- IDTR is Identity transform, which means no transform.
- the most commonly used basic transform mode is DCT2 in both horizontal and vertical directions, written as DCT2-DCT2. It is used as a transform before the inseparable secondary transform LFNST, and is also used as a transform when the multi-transform selection (MTS) technology is turned off.
- MTS multi-transform selection
- the transform process will be a combination of basic transforms in the horizontal and vertical directions, rather than an inseparable transform.
- the current block can have up to 6 non-DCT2-DCT2 transform cores to select.
- the residual MTS basic transform kernel should be related to whether the interpolation filtering mode is selected for the current block. More specifically, it may be related to which sub-mode of the interpolation filtering mode is selected and/or the size and shape of the current block.
- the embodiments of the present application provide two implementations of basic change core candidates that can be used under the current MTS design of ECM.
- the MTS optional basic transform kernel is related to whether the interpolation filter prediction mode is selected for the current block. If the current block uses the interpolation filter prediction mode, the six MTS transform kernels are as follows (the transform kernel is: horizontal transform-vertical transform). See Table 3 for details. When MTS is selected and the prediction mode of the current block is the interpolation prediction mode, the corresponding target transform kernel is selected from the six transform kernels according to the parsed MTS transform kernel index value for inverse transformation.
- the optional basic transform kernel of MTS is related to whether the interpolation filter mode is selected for the current block and the size and shape of the current block.
- the shape and size of the block are: height ⁇ width.
- the corresponding target transform kernel is selected for inverse transformation according to the parsed MTS transform kernel index value and the shape and size of the block.
- the interpolation filter prediction mode can be applied to luminance blocks of 4 ⁇ 4 to 32 ⁇ 32.
- the method for obtaining the candidate MTS transformation core may include:
- Step 1 Encode a set of images or a set of videos using an encoder that includes an interpolation filter prediction mode
- Step 2 The residuals of the blocks with the selected interpolation filter mode are classified into possible horizontal-vertical transform kernels one by one according to the categories (e.g., block shape and size, interpolation filter mode).
- the transform kernel selection criteria can be the size of SAD, the size of SSE, or other metrics, such as transform coding gain.
- Transform coding gain is defined as the arithmetic mean of the transform coefficient variance divided by the geometric mean of the transform coefficient variance.
- the prediction of the interpolation filtering does not contain nonlinear terms or bias terms.
- nonlinear terms or bias terms can also be added to the interpolation filtering.
- the 15 linear terms used in this implementation process are three cases as shown in Figures 3A, 3B and 3C.
- the linear terms of the 15 taps of the interpolation filter are grid filling positions, and the black filling position is the current position to be predicted.
- 3 tap nonlinear terms can be added, and the reconstructed pixel positions used by the nonlinear terms are shown in Figures 14A, 14B and 14C, specifically three dot filling positions.
- the calculation formula of the prediction value is as follows:
- the corresponding nonlinear term value should also be increased when constructing the autocorrelation coefficient matrix and the mutual correlation coefficient vector; and/or, when there is a bias term, the bias term value should also be further increased when constructing the autocorrelation coefficient matrix and the mutual correlation coefficient vector.
- 3 tap nonlinear terms are added to the 15 tap linear terms of the target filter, as shown in FIG.
- the nonlinear term is specifically three positions filled with dots, and the black filled position indicates the current position to be predicted.
- FIG15A, FIG15B, and FIG15C all add three nonlinear terms compared to FIG14A, FIG14B, and FIG14C, but because different filter shapes use the same nonlinear term, the calculation is simpler, further reducing the complexity.
- the embodiment of the present application can also use more nonlinear terms, for example, 5 nonlinear terms are used in FIG. 16A, FIG. 16B, and FIG. 16C.
- 5 nonlinear terms are used in FIG. 16A, FIG. 16B, and FIG. 16C.
- the number of nonlinear terms should be a positive integer, the specific number is not limited, and different designs can be performed according to the performance complexity requirements.
- the encoder 220 may include a first determination unit 2201 and a first prediction unit 2202, wherein:
- the first determining unit 2201 is configured to determine a target filtering mode of the current block; and determine a reference area of the current block according to a size parameter of the current block and the target filtering mode;
- the first prediction unit 2202 is configured to determine a filter coefficient of the current block according to a reference area of the current block; and perform intra-frame prediction on the current block according to the filter coefficient to determine a prediction value of the current block.
- the first determination unit 2201 is further configured to determine at least one candidate filtering mode; perform cost calculation on at least one candidate filtering mode to determine the cost result of at least one candidate filtering mode; and determine the minimum cost result from the cost results of at least one candidate filtering mode, and determine the candidate filtering mode corresponding to the minimum cost result as the target filtering mode of the current block.
- the number of at least one candidate filtering mode is determined based on the number of reference region categories of the current block and the number of shapes of the target filter.
- the target filtering mode includes a reference region category of the current block and a shape of a target filter.
- the first determination unit 2201 is further configured to, if the reference area category of the current block is the first category, determine that the reference area of the current block includes an upper adjacent area and a left adjacent area; if the reference area category of the current block is the second category, determine that the reference area of the current block includes an upper adjacent area; if the reference area category of the current block is the third category, determine that the reference area of the current block includes a left adjacent area; wherein the upper adjacent area refers to a reconstructed area adjacent to the upper side of the current block, and the left adjacent area refers to a reconstructed area adjacent to the left side of the current block.
- the size parameters of the current block include the height and width of the current block; the first determination unit 2201 is further configured to determine the minimum parameter from the height and width of the current block; and determine the reference area of the current block based on the minimum parameter and the target filtering mode.
- the size of the reference area of the current block has an associated relationship with the shape and minimum parameter of the target filter.
- the first determination unit 2201 is further configured to, if the width of the current block and the multiple of the first factor are smaller than the height of the current block, prohibit the reference area category of the current block from being the second category, and determine that the number of reference area categories of the current block is determined based on reference area categories other than the second category; if the height of the current block and the multiple of the first factor are smaller than the width of the current block, prohibit the reference area category of the current block from being the third category, and determine that the number of reference area categories of the current block is determined based on reference area categories other than the third category.
- the value of the first factor is a first preset constant.
- the encoder 220 may further include an encoding unit 2203 configured to encode a target filtering mode of a current block and write the obtained encoded bits into a bitstream.
- the first determining unit 2201 is further configured to determine a context model of the current block
- the encoding unit 2203 is further configured to encode the target filtering mode of the current block based on the context model, and write the obtained encoding bits into the bitstream.
- the determination of the context model is associated with at least one of the following parameters:
- the ratio of the width to the height of the current block is the ratio of the width to the height of the current block.
- the first determination unit 2201 is further configured to determine multiple candidate sub-reference areas of the current block, and the multiple candidate sub-reference areas do not overlap with each other; determine the autocorrelation coefficient matrices and mutual correlation coefficient vectors of the multiple candidate sub-reference areas according to the shapes of the multiple candidate sub-reference areas and the target filter; and store the autocorrelation coefficient matrices and mutual correlation coefficient vectors of the multiple candidate sub-reference areas in a preset cache area.
- the first determining unit 2201 is further configured to, based on the shape of the first candidate sub-reference region and the target filter, Determine an input value of a target filter and an output value of the target filter corresponding to at least one reference pixel in a first candidate sub-reference area; determine an autocorrelation coefficient matrix of the first candidate sub-reference area based on the input value of the target filter corresponding to at least one reference pixel; and determine a mutual correlation coefficient vector of the first candidate sub-reference area based on the input value of the target filter and the output value of the target filter corresponding to at least one reference pixel; wherein the first candidate sub-reference area is any one of a plurality of candidate sub-reference areas.
- the first determination unit 2201 is further configured to divide the reference area of the current block to determine at least one sub-reference area; obtain the autocorrelation coefficient matrix and mutual correlation coefficient vector of at least one sub-reference area from a preset cache area; determine the coefficients of the target filter based on the autocorrelation coefficient matrix and mutual correlation coefficient vector of at least one sub-reference area; and determine the coefficients of the target filter as the filter coefficients of the current block.
- the first determining unit 2201 is further configured to determine a reference sample value corresponding to a pixel to be predicted in the current block;
- the first prediction unit 2202 is further configured to determine a prediction value of a pixel to be predicted in the current block according to a reference sample value and a filter coefficient corresponding to the pixel to be predicted in the current block.
- the first determination unit 2201 is further configured to, based on the shape of the target filter, determine the reconstructed value at the corresponding position in the reference area as the reference sample value if the reference sample value is located in the reference area of the current block; and determine the predicted value at the corresponding position in the current block as the reference sample value if the reference sample value is located inside the current block.
- the first prediction unit 2202 is further configured to determine a first input value of the target filter based on a reference sample value corresponding to the pixel to be predicted in the current block; determine a first output value of the target filter based on the first input value and the filter coefficient; and determine a predicted value of the pixel to be predicted in the current block based on the first output value.
- the first determination unit 2201 is further configured to determine a second factor; and perform a subtraction operation on the reference sample value and the second factor to obtain a first input value of the target filter.
- the first determination unit 2201 is further configured to determine a second output value of the target filter based on the first input value and the filter coefficient; and perform a first process on the second output value to determine a first output value of the target filter.
- the first determination unit 2201 is also configured to calculate the product of the first input value and the corresponding filter coefficient; and set the second output value of the target filter to be equal to the sum of n products; wherein n represents the number of input items corresponding to the target filter, and n is a positive integer.
- the first determination unit 2201 is further configured to perform an addition operation on the second output value and the second factor to obtain a first output value of the target filter.
- the first determination unit 2201 is further configured to determine the third output value of the target filter; determine the fourth output value of the target filter based on the second output value and the third output value; and add the fourth output value and the second factor to obtain the first output value of the target filter.
- the first determination unit 2201 is also configured to determine the number of first type input items corresponding to the target filter based on the shape of the target filter; if the number of first type input items corresponding to the target filter is p, then determine p+q filter coefficients of the target filter, where p and q are both positive integers; and determine the third output value of the target filter based on q filter coefficients and q second type input items among the p+q filter coefficients.
- the first determination unit 2201 is also configured to determine the number of first type input items corresponding to the target filter based on the shape of the target filter; if the number of first type input items corresponding to the target filter is p, then determine p+m filter coefficients of the target filter, where p and m are both positive integers; and determine the third output value of the target filter based on m filter coefficients among the p+m filter coefficients and m third type input items.
- the number of third type input items is preset bias information.
- the value of the second factor is a second preset constant.
- the first determination unit 2201 is further configured to determine a reconstruction value of at least one reference pixel in the reference area; perform mean calculation on the reconstruction value of at least one reference pixel to obtain a first mean; and set the value of the second factor to be equal to the first mean.
- the first prediction unit 2202 is further configured to perform a second process on the first output value to obtain a predicted value of a pixel to be predicted in the current block.
- the first prediction unit 2202 is further configured such that the second processing is to set the prediction value of the pixel to be predicted in the current block to be equal to the first output value.
- the first prediction unit 2202 is further configured to limit the first output value to a preset value range. wherein the lower limit of the preset numerical range is the minimum reconstruction value in the reference area, and the upper limit of the preset numerical range is the maximum reconstruction value in the reference area.
- the first determination unit 2201 is further configured to determine a derived intra-frame prediction mode for the luminance component of the current block if the luminance component of the current block uses intra-frame prediction based on a filter coefficient; if the chrominance component of the current block uses intra-frame prediction in a direct mode, the direct mode is set to the derived intra-frame prediction mode to determine the predicted value of the chrominance component of the current block.
- the first determination unit 2201 is further configured to determine a reference block of the current block when the current block meets a preset condition; if the reference block uses intra-frame prediction based on a filter coefficient, determine a derived intra-frame prediction mode of the reference block; and add the derived intra-frame prediction mode to the intra-frame prediction mode candidate list of the current block.
- the current block satisfies a preset condition, including at least one of the following:
- the current block is an inter-frame prediction block
- the first determining unit 2201 is further configured to determine an original value of the current block; and determine a residual value of the current block according to the original value of the current block and the predicted value of the current block;
- the encoding unit 2203 is further configured to encode the residual value of the current block and write the obtained encoding bits into the bit stream.
- the encoding unit 2203 is further configured to transform the residual value to obtain the transform coefficient of the current block; quantize the transform coefficient to obtain the quantization coefficient of the current block; and encode the quantization coefficient of the current block and write the obtained coded bits into the bit stream.
- the encoding unit 2203 is also configured to determine the target transform kernel of the current block when the current block uses a multi-transform selection mode and the target filtering mode is an interpolation filtering mode; and perform transform processing on the residual value according to the target transform kernel to obtain the transform coefficient of the current block.
- the determination of the target transformation kernel is associated with at least one of the following parameters:
- the first determination unit 2201 is further configured to determine at least one candidate transformation core; perform cost calculation on at least one candidate transformation core to determine a cost result of at least one candidate transformation core; and determine a minimum cost result from the cost results of at least one candidate transformation core, and determine the candidate transformation core corresponding to the minimum cost result as the target transformation core of the current block.
- the first determining unit 2201 is further configured to determine non-zero coefficient information of the current block; and determine at least one candidate transform kernel according to the non-zero coefficient information of the current block.
- the number of the at least one candidate transform kernel is less than or equal to six.
- the first determining unit 2201 is further configured to determine a transform core index value of the current block, wherein the transform core index value is used to indicate an index number of a target transform core in at least one candidate transform core;
- the encoding unit 2203 is further configured to encode the transform core index value of the current block and write the obtained encoding bits into the bitstream.
- the first determining unit 2201 is further configured to determine a transform core index value of the current block, wherein the transform core index value is used to indicate an index number of a target transform core in at least one candidate transform core, and at least one candidate transform core is associated with a size parameter of the current block;
- the encoding unit 2203 is further configured to encode the transform core index value of the current block and write the obtained encoding bits into the bitstream.
- a "unit” may be a part of a circuit, a part of a processor, a part of a program or software, etc., and of course, it may be a module, or it may be non-modular.
- the components in the present embodiment may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional module.
- the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium.
- the technical solution of this embodiment is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product.
- the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in this embodiment.
- the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
- an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 220.
- the computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, the method described in any one of the aforementioned embodiments is implemented.
- the encoder 220 may include: a first communication interface 2301, a first memory 2302 and the first processor 2303; each component is coupled together through a first bus system 2304. It can be understood that the first bus system 2304 is used to realize the connection and communication between these components. In addition to the data bus, the first bus system 2304 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are marked as the first bus system 2304 in FIG. 21.
- the first communication interface 2301 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
- a first memory 2302 used to store a computer program that can be run on the first processor 2303;
- the first processor 2303 is configured to, when running the computer program, execute:
- An intra-frame prediction is performed on the current block according to the filter coefficient to determine a prediction value of the current block.
- the first memory 2302 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
- the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
- the volatile memory can be a random access memory (RAM), which is used as an external cache.
- RAM static RAM
- DRAM dynamic RAM
- SDRAM synchronous DRAM
- DDRSDRAM double data rate synchronous DRAM
- ESDRAM enhanced SDRAM
- SLDRAM synchronous link DRAM
- DRRAM direct RAM bus RAM
- the first processor 2303 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the first processor 2303.
- the above-mentioned first processor 2303 can be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
- DSP Digital Signal Processor
- ASIC Application Specific Integrated Circuit
- FPGA Field Programmable Gate Array
- the methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed.
- the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
- the steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be executed.
- the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc.
- the storage medium is located in the first memory 2302, and the first processor 2303 reads the information in the first memory 2302 and completes the steps of the above method in combination with its hardware.
- the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processors (Digital Signal Processing, DSP), digital signal processing devices (DSP Device, DSPD), programmable logic devices (Programmable Logic Device, PLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA), general processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application or a combination thereof.
- ASIC Application Specific Integrated Circuits
- DSP Digital Signal Processing
- DSP Device digital signal processing devices
- PLD programmable logic devices
- FPGA field programmable gate array
- general processors controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application or a combination thereof.
- the technology described in this application can be implemented by a module (such as a process, function, etc.) that performs the functions described in this application.
- the software code can be stored in a memory and executed by a processor.
- the memory can be implemented in the processor or outside the processor.
- the first processor 2303 is further configured to execute the method described in any one of the aforementioned embodiments when running the computer program.
- the present embodiment provides an encoder, for which the intra-frame prediction technology based on interpolation filtering, when determining the reference area for calculating the filter coefficient, is not only related to the target filtering mode, but also related to the size parameters of the current block. For example, when the size of the current block is large, a large reference area can be used, and when the size of the current block is small, a small reference area can be used; in this way, the calculation complexity can be reduced and the encoding time can be reduced; at the same time, the accuracy of intra-frame prediction can be improved, thereby improving the encoding and decoding performance.
- the decoder 240 may include a decoding unit 2401, a second determination unit 2402, and a second prediction unit 2403, wherein:
- the decoding unit 2401 is configured to decode the bitstream and determine a target filtering mode for a current block
- the second determining unit 2402 is configured to determine a reference area of the current block according to a size parameter of the current block and a target filtering mode
- the second prediction unit 2403 is configured to determine a filter coefficient of the current block according to a reference area of the current block; and perform intra-frame prediction on the current block according to the filter coefficient to determine a prediction value of the current block.
- the target filtering mode includes a reference region category of the current block and a shape of a target filter.
- the second determination unit 2402 is further configured to, if the reference area category of the current block is the first category, determine that the reference area of the current block includes an upper adjacent area and a left adjacent area; if the reference area category of the current block is the second category, determine that the reference area of the current block includes an upper adjacent area; if the reference area category of the current block is the third category, determine that the reference area of the current block includes a left adjacent area; wherein the upper adjacent area refers to a reconstructed area adjacent to the upper side of the current block, and the left adjacent area refers to a reconstructed area adjacent to the left side of the current block.
- the size parameters of the current block include the height and width of the current block; the second determination unit 2402 is further configured to determine the minimum parameter from the height and width of the current block; and determine the reference area of the current block according to the minimum parameter and the target filtering mode.
- the size of the reference area of the current block has an associated relationship with the shape and minimum parameter of the target filter.
- the second determination unit 2402 is further configured to determine that the reference area category in the target prediction mode is any one item except the second category if the width of the current block and the multiple of the first factor are smaller than the height of the current block; and to determine that the reference area category in the target prediction mode is any one item except the third category if the height of the current block and the multiple of the first factor are smaller than the width of the current block.
- the value of the first factor is a first preset constant.
- the second determining unit 2402 is further configured to determine a context model of the current block
- the decoding unit 2401 is further configured to decode the code stream based on the context model and determine a target filtering mode for the current block.
- the determination of the context model is associated with at least one of the following parameters:
- the ratio of the width to the height of the current block is the ratio of the width to the height of the current block.
- the second determination unit 2402 is further configured to determine, based on the reference area of the current block and the shape of the target filter, an input value of the target filter corresponding to at least one reference pixel in the reference area and an output value of the target filter; determine an autocorrelation coefficient matrix based on the input value of the target filter corresponding to at least one reference pixel; determine a mutual correlation coefficient vector based on the input value of the target filter corresponding to at least one reference pixel and the output value of the target filter; determine the coefficients of the target filter based on the autocorrelation coefficient matrix and the mutual correlation coefficient vector; and determine the coefficients of the target filter as the filter coefficients of the current block.
- the second prediction unit 2403 is further configured to determine a reference sample value corresponding to the pixel to be predicted in the current block; and determine a predicted value of the pixel to be predicted in the current block based on the reference sample value and filter coefficient corresponding to the pixel to be predicted in the current block.
- the second determination unit 2402 is further configured to, based on the shape of the target filter, determine the reconstructed value at the corresponding position in the reference area as the reference sample value if the reference sample value is located in the reference area of the current block; and determine the predicted value at the corresponding position in the current block as the reference sample value if the reference sample value is located inside the current block.
- the second prediction unit 2403 is further configured to determine a first input value of the target filter based on a reference sample value corresponding to the pixel to be predicted in the current block; determine a first output value of the target filter based on the first input value and the filter coefficient; and determine a predicted value of the pixel to be predicted in the current block based on the first output value.
- the second determination unit 2402 is further configured to determine a second factor; and perform a subtraction operation on the reference sample value and the second factor to obtain a first input value of the target filter.
- the second determination unit 2402 is further configured to determine a second output value of the target filter based on the first input value and the filter coefficient; and perform a first process on the second output value to determine a first output value of the target filter.
- the second determination unit 2402 is further configured to calculate the product of the first input value and the corresponding filter coefficient; and set the second output value of the target filter to be equal to the sum of n products; wherein n represents the number of input items corresponding to the target filter, and n is a positive integer.
- the second determining unit 2402 is further configured to perform an addition operation on the second output value and the second factor to obtain a first output value of the target filter.
- the second determination unit 2402 is further configured to determine the third output value of the target filter; determine the fourth output value of the target filter based on the second output value and the third output value; and add the fourth output value and the second factor to obtain the first output value of the target filter.
- the second determination unit 2402 is further configured to determine the number of first type input items corresponding to the target filter based on the shape of the target filter; if the number of first type input items corresponding to the target filter is p, then determine p+q filter coefficients of the target filter, where p and q are both positive integers; and determine the third output value of the target filter based on q filter coefficients and q second type input items among the p+q filter coefficients.
- the second determination unit 2402 is further configured to determine the number of first type input items corresponding to the target filter based on the shape of the target filter; if the number of first type input items corresponding to the target filter is p, then determine p+m filter coefficients of the target filter, where p and m are both positive integers; and determine the third output value of the target filter based on m filter coefficients among the p+m filter coefficients and m third type input items.
- the number of third type input items is preset bias information.
- the value of the second factor is a second preset constant.
- the second determination unit 2402 is further configured to determine a reconstruction value of at least one reference pixel in the reference area; perform mean calculation on the reconstruction value of at least one reference pixel to obtain a first mean; and set the value of the second factor to be equal to the first mean.
- the second prediction unit 2403 is further configured to perform a second process on the first output value to obtain a predicted value of a pixel to be predicted in the current block.
- the second prediction unit 2403 is further configured such that the second processing is to set the prediction value of the pixel to be predicted in the current block to be equal to the first output value.
- the second prediction unit 2403 is further configured so that the second processing is to limit the first output value within a preset numerical range; wherein the lower limit value of the preset numerical range is the minimum reconstructed value in the reference area, and the upper limit value of the preset numerical range is the maximum reconstructed value in the reference area.
- the second determination unit 2402 is further configured to determine a derived intra-frame prediction mode for the luminance component of the current block if the luminance component of the current block uses intra-frame prediction based on a filter coefficient; if the chrominance component of the current block uses intra-frame prediction in a direct mode, the direct mode is set to the derived intra-frame prediction mode to determine the predicted value of the chrominance component of the current block.
- the second determination unit 2402 is further configured to determine a reference block of the current block when the current block meets a preset condition; if the reference block uses intra-frame prediction based on a filter coefficient, determine a derived intra-frame prediction mode of the reference block; and add the derived intra-frame prediction mode to the intra-frame prediction mode candidate list of the current block.
- the current block satisfies a preset condition, including at least one of the following:
- the current block is an inter-frame prediction block
- the current block is an intra block copy (IBC) block.
- IBC intra block copy
- the decoding unit 2401 is further configured to decode the bitstream and determine a residual value of the current block
- the second determining unit 2402 is further configured to determine a reconstructed value of the current block according to the predicted value of the current block and the residual value of the current block.
- the decoding unit 2401 is further configured to decode the code stream to determine the quantization coefficient of the current block; perform inverse quantization processing on the quantization coefficient to obtain the transformation coefficient of the current block; and perform inverse transformation processing on the transformation coefficient to obtain the residual value of the current block.
- the decoding unit 2401 is also configured to determine the target transform kernel of the current block when the current block uses a multi-transform selection mode and the target filtering mode is an interpolation filtering mode; and perform inverse transform processing on the transform coefficients according to the target transform kernel to obtain a residual value of the current block.
- the determination of the target transformation kernel is associated with at least one of the following parameters:
- the decoding unit 2401 is further configured to decode the bitstream and determine a transform kernel index value of the current block;
- the second determining unit 2402 is further configured to determine a target transform core of the current block from at least one candidate transform core according to the transform core index value.
- the decoding unit 2401 is further configured to decode the bitstream and determine a transform kernel index value of the current block;
- the second determining unit 2402 is further configured to determine a target transform core of the current block from at least one candidate transform core according to the transform core index value and a size parameter of the current block.
- the second determining unit 2402 is further configured to decode the bitstream and determine the non-zero coefficient information of the current block;
- the second determining unit 2402 is further configured to determine at least one candidate transform kernel according to the non-zero coefficient information of the current block.
- the number of the at least one candidate transform kernel is less than or equal to six.
- a "unit" can be a part of a circuit, a part of a processor, a part of a program or software, etc., and of course it can also be a module, or it can be non-modular.
- the components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
- the above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional module.
- the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium.
- this embodiment provides a computer-readable storage medium, which is applied to the decoder 240, and the computer-readable storage medium stores a computer program. When the computer program is executed by the second processor, the method described in any one of the above embodiments is implemented.
- the decoder 240 may include: a second communication interface 2501, a second memory 2502 and a second processor 2503; each component is coupled together through a second bus system 2504. It can be understood that the second bus system 2504 is used to achieve connection and communication between these components.
- the second bus system 2504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are marked as the second bus system 2504 in Figure 23. Among them,
- the second communication interface 2501 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
- the second memory 2502 is used to store a computer program that can be run on the second processor 2503;
- the second processor 2503 is configured to, when running the computer program, execute:
- An intra-frame prediction is performed on the current block according to the filter coefficient to determine a prediction value of the current block.
- the second processor 2503 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.
- the present embodiment provides a decoder.
- the intra-frame prediction technology based on interpolation filtering when determining the reference area for calculating the filter coefficient, is not only related to the target filtering mode, but also related to the size parameters of the current block. For example, when the size of the current block is large, a large reference area can be used, and when the size of the current block is small, a small reference area can be used. In this way, the calculation complexity can be reduced, and the accuracy of intra-frame prediction can be improved, thereby improving the encoding and decoding performance.
- a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application is shown.
- a coding and decoding system 260 may include an encoder 2601 and a decoder 2602 .
- the encoder 2601 may be the encoder described in any one of the aforementioned embodiments
- the decoder 2602 may be the decoder described in any one of the aforementioned embodiments.
- the reference area of the current block is determined according to the size parameters of the current block and the target filtering mode; then the filter coefficient of the current block is determined according to the reference area of the current block; and then the current block is intra-predicted according to the filter coefficient to determine the prediction value of the current block.
- the intra-frame prediction technology based on interpolation filtering is not only related to the target filtering mode when determining the reference area for calculating the filter coefficient, but also to the size parameters of the current block.
- a large reference area can be used when the size of the current block is large, and a small reference area can be used when the size of the current block is small; in this way, while ensuring the encoding and decoding performance, the computational complexity can be reduced, the encoding time can be reduced, so that the cost performance of the encoding and decoding performance and the encoding complexity can be improved, and at the same time, the accuracy of the intra-frame prediction can be improved, thereby improving the encoding and decoding efficiency.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computing Systems (AREA)
- Theoretical Computer Science (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
G=|Gx|+|Gy| (5)
pred=Clip(min,max,Pout2) (12)
pred=Clip(min,max,Pout2) (21)
Claims (88)
- 一种解码方法,应用于解码器,所述方法包括:解码码流,确定当前块的目标滤波模式;根据所述当前块的尺寸参数和所述目标滤波模式,确定所述当前块的参考区域;根据所述当前块的参考区域,确定所述当前块的滤波系数;根据所述滤波系数对所述当前块进行帧内预测,确定所述当前块的预测值。
- 根据权利要求1所述的方法,其中,所述目标滤波模式包括所述当前块的参考区域类别和目标滤波器的形状。
- 根据权利要求2所述的方法,其中,所述方法还包括:若所述当前块的参考区域类别为第一类别时,则确定所述当前块的参考区域包括上相邻区域和左相邻区域;若所述当前块的参考区域类别为第二类别时,则确定所述当前块的参考区域包括上相邻区域;若所述当前块的参考区域类别为第三类别时,则确定所述当前块的参考区域包括左相邻区域;其中,所述上相邻区域是指与所述当前块的上侧相邻的已重建区域,所述左相邻区域是指与所述当前块的左侧相邻的已重建区域。
- 根据权利要求3所述的方法,其中,所述当前块的尺寸参数包括所述当前块的高度和宽度;所述基于所述当前块的尺寸参数和所述目标滤波模式,确定所述当前块的参考区域,包括:从所述当前块的高度与宽度中确定最小参数;根据所述最小参数和所述目标滤波模式,确定所述当前块的参考区域。
- 根据权利要求4所述的方法,其中,所述当前块的参考区域的大小与所述目标滤波器的形状和所述最小参数具有关联关系。
- 根据权利要求3所述的方法,其中,所述方法还包括:若所述当前块的宽度与第一因子的倍数小于所述当前块的高度,则确定所述目标预测模式中的参考区域类别为除所述第二类别之外的任意一项;若所述当前块的高度与第一因子的倍数小于所述当前块的宽度,则确定所述目标预测模式中的参考区域类别为除所述第三类别之外的任意一项。
- 根据权利要求6所述的方法,其中,所述第一因子的取值为第一预设常数。
- 根据权利要求1所述的方法,其中,所述解码码流,确定当前块的目标滤波模式,包括:确定所述当前块的上下文模型;基于所述上下文模型解码码流,确定所述当前块的目标滤波模式。
- 根据权利要求8所述的方法,其中,所述上下文模型的确定与下述参数中的至少一项具有关联关系:所述当前块的形状;所述当前块的宽度与高度的比值。
- 根据权利要求2所述的方法,其中,所述根据所述当前块的参考区域,确定所述当前块的滤波系数,包括:根据所述当前块的参考区域和所述目标滤波器的形状,确定所述参考区域中至少一个参考像素对应的所述目标滤波器的输入值和所述目标滤波器的输出值;根据所述至少一个参考像素对应的所述目标滤波器的输入值,确定自相关系数矩阵;根据所述至少一个参考像素对应的所述目标滤波器的输入值和所述目标滤波器的输出值,确定互相关系数向量;根据所述自相关系数矩阵和所述互相关系数向量,确定所述目标滤波器的系数;将所述目标滤波器的系数确定为所述当前块的滤波系数。
- 根据权利要求2至10中任一项所述的方法,其中,所述根据所述滤波系数对所述当前块中的像素点进行帧内预测,确定所述当前块中的像素点的预测值,包括:确定所述当前块中待预测像素对应的参考样值;根据所述当前块中待预测像素对应的参考样值和所述滤波系数,确定所述当前块中待预测像素的预测值。
- 根据权利要求11所述的方法,其中,所述确定所述当前块中待预测像素对应的参考样值,包括:基于所述目标滤波器的形状,若所述参考样值位于所述当前块的参考区域,则将所述参考区域中对应位置处的重建值确定为所述参考样值;若所述参考样值位于所述当前块的内部,则将所述当前块中对应位置处的预测值确定为所述参考样值。
- 根据权利要求11所述的方法,其中,所述根据所述当前块中待预测像素对应的参考样值和所述滤波系数,确定所述当前块中待预测像素的预测值,包括:基于所述当前块中待预测像素对应的参考样值,确定所述目标滤波器的第一输入值;基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第一输出值;根据所述第一输出值,确定所述当前块中待预测像素的预测值。
- 根据权利要求13所述的方法,其中,所述基于所述当前块中待预测像素对应的参考样值,确定所述目标滤波器的第一输入值,包括:确定第二因子;对所述参考样值与所述第二因子进行减法运算,得到所述目标滤波器的第一输入值。
- 根据权利要求14所述的方法,其中,所述基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第一输出值,包括:基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第二输出值;对所述第二输出值进行第一处理,确定所述目标滤波器的第一输出值。
- 根据权利要求15所述的方法,其中,所述基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第二输出值,包括:计算所述第一输入值与对应的所述滤波系数的乘积;将所述目标滤波器的第二输出值设置为等于n个所述乘积之和;其中,n表示所述目标滤波器对应的输入项数,且n为正整数。
- 根据权利要求15所述的方法,其中,所述对所述第二输出值进行第一处理,确定所述目标滤波器的第一输出值,包括:对所述第二输出值与所述第二因子进行加法运算,得到所述目标滤波器的第一输出值。
- 根据权利要求15所述的方法,其中,所述对所述第二输出值进行第一处理,确定所述目标滤波器的第一输出值,包括:确定所述目标滤波器的第三输出值;根据所述第二输出值和所述第三输出值,确定所述目标滤波器的第四输出值;对所述第四输出值与所述第二因子进行加法运算,得到所述目标滤波器的第一输出值。
- 根据权利要求18所述的方法,其中,所述确定所述目标滤波器的第三输出值,包括:基于所述目标滤波器的形状,确定所述目标滤波器对应的第一类型输入项数;若所述目标滤波器对应的第一类型输入项数为p,则确定所述目标滤波器的p+q个滤波系数,p、q均为正整数;根据所述p+q个滤波系数中的q个滤波系数和q个第二类型输入项数,确定所述目标滤波器的第三输出值。
- 根据权利要求18所述的方法,其中,所述确定所述目标滤波器的第三输出值,包括:基于所述目标滤波器的形状,确定所述目标滤波器对应的第一类型输入项数;若所述目标滤波器对应的第一类型输入项数为p,则确定所述目标滤波器的p+m个滤波系数,p、m均为正整数;根据所述p+m个滤波系数中的m个滤波系数和m个第三类型输入项数,确定所述目标滤波器的第三输出值。
- 根据权利要求18所述的方法,其中,所述确定所述目标滤波器的第三输出值,包括:基于所述目标滤波器的形状,确定所述目标滤波器对应的第一类型输入项数;若所述目标滤波器对应的第一类型输入项数为p,则确定所述目标滤波器的p+k个滤波系数,p、k均为正整数;根据所述p+k个滤波系数中的i个滤波系数和i个第二类型输入项数以及所述p+k个滤波系数中的j个滤波系数和j个第三类型输入项数,确定所述目标滤波器的第三输出值;其中,i、j均为正整数,且k=i+j。
- 根据权利要求21所述的方法,其中,所述第一类型输入项数与所述参考样值之间具有线性关系,所述第二类型输入项数与所述参考样值之间具有非线性关系,所述第三类型输入项数为预设的偏置信息。
- 根据权利要求14、17或18所述的方法,其中,所述第二因子的取值为第二预设常数。
- 根据权利要求14、17或18所述的方法,其中,所述确定第二因子,包括:确定所述参考区域中至少一个参考像素的重建值;对所述至少一个参考像素的重建值进行均值计算,得到第一均值;将所述第二因子的取值设置为等于所述第一均值。
- 根据权利要求13所述的方法,其中,所述根据所述第一输出值,确定所述当前块中待预测像素的预测值,包括:对所述第一输出值进行第二处理,得到所述当前块中待预测像素的预测值。
- 根据权利要求25所述的方法,其中,所述方法还包括:所述第二处理是将所述当前块中待预测像素的预测值设置为等于所述第一输出值。
- 根据权利要求25所述的方法,其中,所述方法还包括:所述第二处理是将所述第一输出值限制在预设数值范围之内;其中,所述预设数值范围的下限值为所述参考区域中的最小重建值,所述预设数值范围的上限值为所述参考区域中的最大重建值。
- 根据权利要求1至27中任一项所述的方法,其中,所述方法还包括:若所述当前块的亮度分量使用基于所述滤波系数的帧内预测,则确定所述当前块的亮度分量的推导帧内预测模式;若所述当前块的色度分量使用直接模式的帧内预测,则将所述直接模式设置为所述推导帧内预测模式,以确定所述当前块的色度分量的预测值。
- 根据权利要求1至27中任一项所述的方法,其中,所述方法还包括:在所述当前块满足预设条件时,确定所述当前块的参考块;若所述参考块使用基于所述滤波系数的帧内预测,确定所述参考块的推导帧内预测模式;将所述推导帧内预测模式添加至所述当前块的帧内预测模式候选列表中。
- 根据权利要求29所述的方法,其中,所述当前块满足预设条件,至少包括下述其中一项:所述当前块为帧间预测块;所述当前块为帧内块拷贝IBC块。
- 根据权利要求1至30中任一项所述的方法,其中,所述方法还包括:解码码流,确定所述当前块的残差值;根据所述当前块的预测值和所述当前块的残差值,确定所述当前块的重建值。
- 根据权利要求31所述的方法,其中,所述解码码流,确定所述当前块的残差值,包括:解码码流,确定所述当前块的量化系数;对所述量化系数进行反量化处理,得到所述当前块的变换系数;对所述变换系数进行反变换处理,得到所述当前块的残差值。
- 根据权利要求32所述的方法,其中,所述对所述变换系数进行反变换处理,得到所述当前块的残差值,包括:在所述当前块使用多变换选择模式且所述目标滤波模式为插值滤波模式时,确定所述当前块的目标变换核;根据所述目标变换核对所述变换系数进行反变换处理,得到所述当前块的残差值。
- 根据权利要求33所述的方法,其中,所述目标变换核的确定与下述参数中的至少一项具有关联关系:所述当前块的目标滤波模式;所述当前块的尺寸参数;所述当前块的形状。
- 根据权利要求33所述的方法,其中,所述确定所述当前块的目标变换核,包括:解码码流,确定所述当前块的变换核索引值;根据所述变换核索引值,从至少一个候选变换核中确定所述当前块的目标变换核。
- 根据权利要求33所述的方法,其中,所述确定所述当前块的目标变换核,包括:解码码流,确定所述当前块的变换核索引值;根据所述变换核索引值和所述当前块的尺寸参数,从至少一个候选变换核中确定所述当前块的目标变换核。
- 根据权利要求35或36所述的方法,其中,所述方法还包括:解码码流,确定所述当前块的非零系数信息;根据所述当前块的非零系数信息,确定所述至少一个候选变换核。
- 根据权利要求37所述的方法,其中,所述至少一个候选变换核的个数小于或等于6个。
- 一种编码方法,应用于编码器,所述方法包括:确定当前块的目标滤波模式;根据所述当前块的尺寸参数和所述目标滤波模式,确定所述当前块的参考区域;根据所述当前块的参考区域,确定所述当前块的滤波系数;根据所述滤波系数对所述当前块进行帧内预测,确定所述当前块的预测值。
- 根据权利要求39所述的方法,其中,所述确定当前块的目标滤波模式,包括:确定至少一种候选滤波模式;对所述至少一种候选滤波模式进行代价计算,确定所述至少一种候选滤波模式的代价结果;从所述至少一种候选滤波模式的代价结果中确定最小代价结果,将所述最小代价结果对应的候选滤波模式确定为所述当前块的目标滤波模式。
- 根据权利要求40所述的方法,其中,所述至少一种候选滤波模式的个数是基于所述当前块的参考区域类别数量和目标滤波器的形状数量确定的。
- 根据权利要求40所述的方法,其中,所述目标滤波模式包括所述当前块的参考区域类别和目标滤波器的形状。
- 根据权利要求42所述的方法,其中,所述方法还包括:若所述当前块的参考区域类别为第一类别时,则确定所述当前块的参考区域包括上相邻区域和左相邻区域;若所述当前块的参考区域类别为第二类别时,则确定所述当前块的参考区域包括上相邻区域;若所述当前块的参考区域类别为第三类别时,则确定所述当前块的参考区域包括左相邻区域;其中,所述上相邻区域是指与所述当前块的上侧相邻的已重建区域,所述左相邻区域是指与所述当前块的左侧相邻的已重建区域。
- 根据权利要求43所述的方法,其中,所述当前块的尺寸参数包括所述当前块的高度和宽度;所述基于所述当前块的尺寸参数和所述目标滤波模式,确定所述当前块的参考区域,包括:从所述当前块的高度与宽度中确定最小参数;根据所述最小参数和所述目标滤波模式,确定所述当前块的参考区域。
- 根据权利要求44所述的方法,其中,所述当前块的参考区域的大小与所述目标滤波器的形状和所述最小参数具有关联关系。
- 根据权利要求43所述的方法,其中,所述方法还包括:若所述当前块的宽度与第一因子的倍数小于所述当前块的高度,则禁止所述当前块的参考区域类别为所述第二类别,以及确定所述当前块的参考区域类别数量是基于除所述第二类别之外的其他参考区域类别确定的;若所述当前块的高度与第一因子的倍数小于所述当前块的宽度,则禁止所述当前块的参考区域类别为所述第三类别,以及确定所述当前块的参考区域类别数量是基于除所述第三类别之外的其他参考区域类别确定的。
- 根据权利要求46所述的方法,其中,所述第一因子的取值为第一预设常数。
- 根据权利要求39所述的方法,其中,所述方法还包括:对所述当前块的目标滤波模式进行编码,将所得到的编码比特写入码流。
- 根据权利要求48所述的方法,其中,所述对所述当前块的目标滤波模式进行编码,将所得到的编码比特写入码流,包括:确定所述当前块的上下文模型;基于所述上下文模型对所述当前块的目标滤波模式进行编码,将所得到的编码比特写入码流。
- 根据权利要求49所述的方法,其中,所述上下文模型的确定与下述参数中的至少一项具有关联关系:所述当前块的形状;所述当前块的宽度与高度的比值。
- 根据权利要求42所述的方法,其中,所述方法还包括:确定所述当前块的多个候选子参考区域,且所述多个候选子参考区域互不重叠;根据所述多个候选子参考区域和所述目标滤波器的形状,确定所述多个候选子参考区域各自的自相关系数矩阵和互相关系数向量;将所述多个候选子参考区域各自的自相关系数矩阵和互相关系数向量存储至预设缓存区中。
- 根据权利要求51所述的方法,其中,所述根据所述多个候选子参考区域和所述目标滤波器的形状,确定所述多个候选子参考区域各自的自相关系数矩阵和互相关系数向量,包括:根据第一候选子参考区域和所述目标滤波器的形状,确定所述第一候选子参考区域中至少一个参考像素对应的所述目标滤波器的输入值和所述目标滤波器的输出值;根据所述至少一个参考像素对应的所述目标滤波器的输入值,确定所述第一候选子参考区域的自相关系数矩阵;根据所述至少一个参考像素对应的所述目标滤波器的输入值和所述目标滤波器的输出值,确定所述第一候选子参考区域的互相关系数向量;其中,所述第一候选子参考区域为所述多个候选子参考区域中的任意一个。
- 根据权利要求51所述的方法,其中,所述根据所述当前块的参考区域,确定所述当前块的滤波系数,包括:对所述当前块的参考区域进行划分,确定至少一个子参考区域;从所述预设缓存区中获取所述至少一个子参考区域各自的自相关系数矩阵和互相关系数向量;根据所述至少一个子参考区域各自的自相关系数矩阵和互相关系数向量,确定所述目标滤波器的系数;将所述目标滤波器的系数确定为所述当前块的滤波系数。
- 根据权利要求42至53中任一项所述的方法,其中,所述根据所述滤波系数对所述当前块中的像素点进行帧内预测,确定所述当前块中的像素点的预测值,包括:确定所述当前块中待预测像素对应的参考样值;根据所述当前块中待预测像素对应的参考样值和所述滤波系数,确定所述当前块中待预测像素的预测值。
- 根据权利要求54所述的方法,其中,所述确定所述当前块中待预测像素对应的参考样值,包括:基于所述目标滤波器的形状,若所述参考样值位于所述当前块的参考区域,则将所述参考区域中对应位置处的重建值确定为所述参考样值;若所述参考样值位于所述当前块的内部,则将所述当前块中对应位置处的预测值确定为所述参考样值。
- 根据权利要求54所述的方法,其中,所述根据所述当前块中待预测像素对应的参考样值和所述滤波系数,确定所述当前块中待预测像素的预测值,包括:基于所述当前块中待预测像素对应的参考样值,确定所述目标滤波器的第一输入值;基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第一输出值;根据所述第一输出值,确定所述当前块中待预测像素的预测值。
- 根据权利要求56所述的方法,其中,所述基于所述当前块中待预测像素对应的参考样值,确定所述目标滤波器的第一输入值,包括:确定第二因子;对所述参考样值与所述第二因子进行减法运算,得到所述目标滤波器的第一输入值。
- 根据权利要求57所述的方法,其中,所述基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第一输出值,包括:基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第二输出值;对所述第二输出值进行第一处理,确定所述目标滤波器的第一输出值。
- 根据权利要求58所述的方法,其中,所述基于所述第一输入值和所述滤波系数,确定所述目标滤波器的第二输出值,包括:计算所述第一输入值与对应的所述滤波系数的乘积;将所述目标滤波器的第二输出值设置为等于n个所述乘积之和;其中,n表示所述目标滤波器对应的输入项数,且n为正整数。
- 根据权利要求58所述的方法,其中,所述对所述第二输出值进行第一处理,确定所述目标滤波器的第一输出值,包括:对所述第二输出值与所述第二因子进行加法运算,得到所述目标滤波器的第一输出值。
- 根据权利要求58所述的方法,其中,所述对所述第二输出值进行第一处理,确定所述目标滤波器的第一输出值,包括:确定所述目标滤波器的第三输出值;根据所述第二输出值和所述第三输出值,确定所述目标滤波器的第四输出值;对所述第四输出值与所述第二因子进行加法运算,得到所述目标滤波器的第一输出值。
- 根据权利要求61所述的方法,其中,所述确定所述目标滤波器的第三输出值,包括:基于所述目标滤波器的形状,确定所述目标滤波器对应的第一类型输入项数;若所述目标滤波器对应的第一类型输入项数为p,则确定所述目标滤波器的p+q个滤波系数,p、q均为正整数;根据所述p+q个滤波系数中的q个滤波系数和q个第二类型输入项数,确定所述目标滤波器的第三输出值。
- 根据权利要求61所述的方法,其中,所述确定所述目标滤波器的第三输出值,包括:基于所述目标滤波器的形状,确定所述目标滤波器对应的第一类型输入项数;若所述目标滤波器对应的第一类型输入项数为p,则确定所述目标滤波器的p+m个滤波系数,p、m均为正整数;根据所述p+m个滤波系数中的m个滤波系数和m个第三类型输入项数,确定所述目标滤波器的第三输出值。
- 根据权利要求61所述的方法,其中,所述确定所述目标滤波器的第三输出值,包括:基于所述目标滤波器的形状,确定所述目标滤波器对应的第一类型输入项数;若所述目标滤波器对应的第一类型输入项数为p,则确定所述目标滤波器的p+k个滤波系数,p、k均为正整数;根据所述p+k个滤波系数中的i个滤波系数和i个第二类型输入项数以及所述p+k个滤波系数中的j个滤波系数和j个第三类型输入项数,确定所述目标滤波器的第三输出值;其中,i、j均为正整数,且k=i+j。
- 根据权利要求64所述的方法,其中,所述第一类型输入项数与所述参考样值之间具有线性关系,所述第二类型输入项数与所述参考样值之间具有非线性关系,所述第三类型输入项数为预设的偏置信息。
- 根据权利要求57、60或61所述的方法,其中,所述第二因子的取值为第二预设常数。
- 根据权利要求57、60或61所述的方法,其中,所述确定第二因子,包括:确定所述参考区域中至少一个参考像素的重建值;对所述至少一个参考像素的重建值进行均值计算,得到第一均值;将所述第二因子的取值设置为等于所述第一均值。
- 根据权利要求56所述的方法,其中,所述根据所述第一输出值,确定所述当前块中待预测像素的预测值,包括:对所述第一输出值进行第二处理,得到所述当前块中待预测像素的预测值。
- 根据权利要求68所述的方法,其中,所述方法还包括:所述第二处理是将所述当前块中待预测像素的预测值设置为等于所述第一输出值。
- 根据权利要求68所述的方法,其中,所述方法还包括:所述第二处理是将所述第一输出值限制在预设数值范围之内;其中,所述预设数值范围的下限值为所述参考区域中的最小重建值,所述预设数值范围的上限值为所述参考区域中的最大重建值。
- 根据权利要求39至70中任一项所述的方法,其中,所述方法还包括:若所述当前块的亮度分量使用基于所述滤波系数的帧内预测,则确定所述当前块的亮度分量的推导帧内预测模式;若所述当前块的色度分量使用直接模式的帧内预测,则将所述直接模式设置为所述推导帧内预测模式,以确定所述当前块的色度分量的预测值。
- 根据权利要求39至70中任一项所述的方法,其中,所述方法还包括:在所述当前块满足预设条件时,确定所述当前块的参考块;若所述参考块使用基于所述滤波系数的帧内预测,确定所述参考块的推导帧内预测模式;将所述推导帧内预测模式添加至所述当前块的帧内预测模式候选列表中。
- 根据权利要求72所述的方法,其中,所述当前块满足预设条件,至少包括下述其中一项:所述当前块为帧间预测块;所述当前块为帧内块拷贝IBC块。
- 根据权利要求39至73中任一项所述的方法,其中,所述方法还包括:确定所述当前块的原始值;根据所述当前块的原始值和所述当前块的预测值,确定所述当前块的残差值;对所述当前块的残差值进行编码,将所得到的编码比特写入码流。
- 根据权利要求74所述的方法,其中,所述对所述当前块的残差值进行编码,将所得到的编码比特写入码流,包括:对所述残差值进行变换处理,得到所述当前块的变换系数;对所述变换系数进行量化处理,得到所述当前块的量化系数;对所述当前块的量化系数进行编码,将所得到的编码比特写入码流。
- 根据权利要求75所述的方法,其中,所述对所述残差值进行变换处理,得到所述当前块的变换系数,包括:在所述当前块使用多变换选择模式且所述目标滤波模式为插值滤波模式时,确定所述当前块的目标变换核;根据所述目标变换核对所述残差值进行变换处理,得到所述当前块的变换系数。
- 根据权利要求76所述的方法,其中,所述目标变换核的确定与下述参数中的至少一项具有关联关系:所述当前块的目标滤波模式;所述当前块的尺寸参数;所述当前块的形状。
- 根据权利要求76所述的方法,其中,所述确定所述当前块的目标变换核,包括:确定至少一个候选变换核;对所述至少一个候选变换核进行代价计算,确定所述至少一个候选变换核的代价结果;从所述至少一个候选变换核的代价结果中确定最小代价结果,将所述最小代价结果对应的候选变换核确定为所述当前块的目标变换核。
- 根据权利要求78所述的方法,其中,所述方法还包括:确定所述当前块的非零系数信息;根据所述当前块的非零系数信息,确定所述至少一个候选变换核。
- 根据权利要求79所述的方法,其中,所述至少一个候选变换核的个数小于或等于6个。
- 根据权利要求78所述的方法,其中,所述方法还包括:确定所述当前块的变换核索引值,其中,所述变换核索引值用于指示所述目标变换核在所述至少一个候选变换核中的索引序号;对所述当前块的变换核索引值进行编码,将所得到的编码比特写入码流。
- 根据权利要求78所述的方法,其中,所述方法还包括:确定所述当前块的变换核索引值,其中,所述变换核索引值用于指示所述目标变换核在所述至少一个候选变换核中的索引序号,且所述至少一个候选变换核与所述当前块的尺寸参数具有关联关系;对所述当前块的变换核索引值进行编码,将所得到的编码比特写入码流。
- 一种码流,所述码流是根据待编码信息进行比特编码生成的;其中,所述待编码信息包括下述至少一项:当前块的目标滤波模式、所述当前块的残差值和所述当前块的变换核索引值。
- 一种编码器,所述编码器包括第一确定单元和第一预测单元,其中:所述第一确定单元,配置为确定当前块的目标滤波模式;以及根据所述当前块的尺寸参数和所述目标滤波模式,确定所述当前块的参考区域;所述第一预测单元,配置为根据所述当前块的参考区域,确定所述当前块的滤波系数;以及根据所述滤波系数对所述当前块进行帧内预测,确定所述当前块的预测值。
- 一种编码器,所述编码器包括第一存储器和第一处理器,其中:所述第一存储器,用于存储能够在所述第一处理器上运行的计算机程序;所述第一处理器,用于在运行所述计算机程序时,执行如权利要求39至82中任一项所述的方法。
- 一种解码器,所述解码器包括解码单元、第二确定单元和第二预测单元,其中:所述解码单元,配置为解码码流,确定当前块的目标滤波模式;所述第二确定单元,配置为根据所述当前块的尺寸参数和所述目标滤波模式,确定所述当前块的参考区域;所述第二预测单元,配置为根据所述当前块的参考区域,确定所述当前块的滤波系数;以及根据所述滤波系数对所述当前块进行帧内预测,确定所述当前块的预测值。
- 一种解码器,所述解码器包括第二存储器和第二处理器,其中:所述第二存储器,用于存储能够在所述第二处理器上运行的计算机程序;所述第二处理器,用于在运行所述计算机程序时,执行如权利要求1至38中任一项所述的方法。
- 一种计算机可读存储介质,其中,所述计算机可读存储介质存储有计算机程序,所述计算机程序被执行时实现如权利要求1至38中任一项所述的方法、或者实现如权利要求39至82中任一项所述的方法。
Priority Applications (6)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202380099393.3A CN121511593A (zh) | 2023-06-19 | 2023-06-19 | 编解码方法、码流、编码器、解码器以及存储介质 |
| AU2023459194A AU2023459194A1 (en) | 2023-06-19 | 2023-06-19 | Encoding method, decoding method, code stream, encoder, decoder, and storage medium |
| PCT/CN2023/101156 WO2024259568A1 (zh) | 2023-06-19 | 2023-06-19 | 编解码方法、码流、编码器、解码器以及存储介质 |
| KR1020257042002A KR20260025916A (ko) | 2023-06-19 | 2023-06-19 | 부호화 방법, 복호화 방법, 비트스트림, 인코더, 디코더 및 저장 매체 |
| US19/420,029 US20260106981A1 (en) | 2023-06-19 | 2025-12-15 | Encoding method, decoding method, bitstream, encoder, decoder, and storage medium |
| MX2025015301A MX2025015301A (es) | 2023-06-19 | 2025-12-16 | Metodo de codificacion, metodo de decodificacion, flujo de codigo, codificador, decodificador y medio de almacenamiento |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/101156 WO2024259568A1 (zh) | 2023-06-19 | 2023-06-19 | 编解码方法、码流、编码器、解码器以及存储介质 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US19/420,029 Continuation US20260106981A1 (en) | 2023-06-19 | 2025-12-15 | Encoding method, decoding method, bitstream, encoder, decoder, and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024259568A1 true WO2024259568A1 (zh) | 2024-12-26 |
Family
ID=93934627
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/101156 Ceased WO2024259568A1 (zh) | 2023-06-19 | 2023-06-19 | 编解码方法、码流、编码器、解码器以及存储介质 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US20260106981A1 (zh) |
| KR (1) | KR20260025916A (zh) |
| CN (1) | CN121511593A (zh) |
| AU (1) | AU2023459194A1 (zh) |
| MX (1) | MX2025015301A (zh) |
| WO (1) | WO2024259568A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20130105114A (ko) * | 2012-03-16 | 2013-09-25 | 주식회사 아이벡스피티홀딩스 | 인트라 예측 모드에서의 영상 복호화 방법 |
| CN108293111A (zh) * | 2015-10-16 | 2018-07-17 | Lg电子株式会社 | 用于改善在图像编码系统中进行预测的滤波方法和装置 |
| CN111247796A (zh) * | 2017-10-20 | 2020-06-05 | 韩国电子通信研究院 | 图像编码/解码方法和装置以及存储比特流的记录介质 |
| CN111837388A (zh) * | 2018-03-09 | 2020-10-27 | 韩国电子通信研究院 | 使用样点滤波的图像编码/解码方法和设备 |
| CN112425161A (zh) * | 2018-07-11 | 2021-02-26 | 英迪股份有限公司 | 基于帧内预测的视频编码方法和装置 |
-
2023
- 2023-06-19 CN CN202380099393.3A patent/CN121511593A/zh active Pending
- 2023-06-19 KR KR1020257042002A patent/KR20260025916A/ko active Pending
- 2023-06-19 WO PCT/CN2023/101156 patent/WO2024259568A1/zh not_active Ceased
- 2023-06-19 AU AU2023459194A patent/AU2023459194A1/en active Pending
-
2025
- 2025-12-15 US US19/420,029 patent/US20260106981A1/en active Pending
- 2025-12-16 MX MX2025015301A patent/MX2025015301A/es unknown
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20130105114A (ko) * | 2012-03-16 | 2013-09-25 | 주식회사 아이벡스피티홀딩스 | 인트라 예측 모드에서의 영상 복호화 방법 |
| CN108293111A (zh) * | 2015-10-16 | 2018-07-17 | Lg电子株式会社 | 用于改善在图像编码系统中进行预测的滤波方法和装置 |
| CN111247796A (zh) * | 2017-10-20 | 2020-06-05 | 韩国电子通信研究院 | 图像编码/解码方法和装置以及存储比特流的记录介质 |
| CN111837388A (zh) * | 2018-03-09 | 2020-10-27 | 韩国电子通信研究院 | 使用样点滤波的图像编码/解码方法和设备 |
| CN112425161A (zh) * | 2018-07-11 | 2021-02-26 | 英迪股份有限公司 | 基于帧内预测的视频编码方法和装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| KR20260025916A (ko) | 2026-02-24 |
| AU2023459194A1 (en) | 2026-01-15 |
| US20260106981A1 (en) | 2026-04-16 |
| CN121511593A (zh) | 2026-02-10 |
| MX2025015301A (es) | 2026-02-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7540841B2 (ja) | ループ内フィルタリングの方法、コンピュータ可読記憶媒体及びプログラム | |
| AU2024287236A1 (en) | Video image component prediction method and apparatus, and computer storage medium | |
| WO2024152385A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| CN114866783B (zh) | 帧间预测方法、编码器、解码器以及计算机存储介质 | |
| JP2025513048A (ja) | 符号化・復号化方法およびその装置、符号化機器、復号化機器、並びに記憶媒体 | |
| CN113840142B (zh) | 图像分量预测方法、装置及计算机存储介质 | |
| WO2024259568A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| CN120548703A (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| CN120476592A (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| WO2024216632A1 (zh) | 视频编解码方法、装置、设备、系统、及存储介质 | |
| CN119032561A (zh) | 编解码方法、装置、编码设备、解码设备以及存储介质 | |
| CN113411588B (zh) | 预测方向的确定方法、解码器以及计算机存储介质 | |
| WO2025043408A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| CN113196770A (zh) | 图像分量预测方法、装置及计算机存储介质 | |
| WO2025065696A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| CN120500846A (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| WO2024153241A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| WO2024207136A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| CA3280451A1 (en) | Encoding method, decoding method, code stream, encoder, decoder, and storage medium | |
| WO2024212251A9 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| WO2025147830A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| WO2024192733A9 (zh) | 视频编解码方法、装置、设备、系统、及存储介质 | |
| WO2026085742A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| WO2024098263A1 (zh) | 编解码方法、码流、编码器、解码器以及存储介质 | |
| JP2026513796A (ja) | 符号化・復号化方法、ビットストリーム、エンコーダ、デコーダ及び記憶媒体 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23941875 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202517126747 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: AU2023459194 Country of ref document: AU Ref document number: MX/A/2025/015301 Country of ref document: MX |
|
| WWP | Wipo information: published in national office |
Ref document number: 202517126747 Country of ref document: IN |
|
| ENP | Entry into the national phase |
Ref document number: 2023459194 Country of ref document: AU Date of ref document: 20230619 Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2026100704 Country of ref document: RU |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: MX/A/2025/015301 Country of ref document: MX |
|
| WWP | Wipo information: published in national office |
Ref document number: 2026100704 Country of ref document: RU |