WO2024007366A1 - 一种帧内预测融合方法、视频编解码方法、装置和系统 - Google Patents

一种帧内预测融合方法、视频编解码方法、装置和系统 Download PDF

Info

Publication number
WO2024007366A1
WO2024007366A1 PCT/CN2022/106337 CN2022106337W WO2024007366A1 WO 2024007366 A1 WO2024007366 A1 WO 2024007366A1 CN 2022106337 W CN2022106337 W CN 2022106337W WO 2024007366 A1 WO2024007366 A1 WO 2024007366A1
Authority
WO
WIPO (PCT)
Prior art keywords
current block
mode
fusion
intra prediction
prediction
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/106337
Other languages
English (en)
French (fr)
Inventor
徐陆航
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Guangdong Oppo Mobile Telecommunications Corp Ltd
Original Assignee
Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Guangdong Oppo Mobile Telecommunications Corp Ltd filed Critical Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority to KR1020257003534A priority Critical patent/KR20250035555A/ko
Priority to JP2024577022A priority patent/JP2025521765A/ja
Priority to CN202280097794.0A priority patent/CN119487831A/zh
Priority to TW112125248A priority patent/TW202406347A/zh
Publication of WO2024007366A1 publication Critical patent/WO2024007366A1/zh
Priority to MX2024016089A priority patent/MX2024016089A/es
Priority to US18/990,493 priority patent/US20250126291A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/184Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being bits, e.g. of the compressed video stream
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • H04N19/11Selection of coding mode or of prediction mode among a plurality of spatial predictive coding modes
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/157Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
    • H04N19/159Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/593Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques

Definitions

  • Embodiments of the present disclosure relate to, but are not limited to, video technology, and more specifically, relate to an intra prediction fusion method, video encoding and decoding method, device and system.
  • Digital video compression technology mainly compresses huge digital image and video data to facilitate transmission and storage.
  • Current common video encoding and decoding standards such as H.266/Versatile Video Coding (VVC), all use block-based hybrid coding frameworks.
  • Each frame in the video is divided into square largest coding units (LCU: largest coding unit) of the same size (such as 128x128, 64x64, etc.).
  • Each maximum coding unit can be divided into rectangular coding units (CU: coding unit) according to rules.
  • Coding units may also be divided into prediction units (PU: prediction unit), transformation units (TU: transform unit), etc.
  • the hybrid coding framework includes prediction, transform, quantization, entropy coding, in loop filter and other modules.
  • the prediction module includes intra prediction and inter prediction, which are used to reduce or remove the inherent redundancy of the video.
  • Intra-frame blocks are predicted using the surrounding pixels of the block as a reference, while inter-frame blocks refer to spatially adjacent block information and reference information in other frames.
  • the residual information is encoded into a code stream through block-based transformation, quantization and entropy encoding.
  • An embodiment of the present disclosure provides an intra prediction fusion method, including:
  • the selected intra prediction mode of the current block includes the angle mode, determine whether the restriction condition for the current block to use intra prediction to fuse IPF is established;
  • An embodiment of the present disclosure also provides a video decoding method, including:
  • the reconstructed value of the current block is determined based on the predicted value of the current block.
  • An embodiment of the present disclosure also provides a video encoding method, including:
  • An embodiment of the present disclosure also provides a video encoding method, including:
  • the decoding determines that the current block uses the template-based multi-reference line intra prediction TMRL mode, continue to decode the TMRL mode index and TMRL fusion flag of the current block;
  • the weighted sum of the first prediction result and the second prediction result is used as the final prediction result of the current block
  • the first prediction result is obtained by predicting the current block according to the extended reference line and the intra prediction mode
  • the second prediction result is obtained by predicting the current block according to another reference line and the intra prediction mode.
  • the intra prediction mode is angle mode.
  • An embodiment of the present disclosure also provides a video encoding method, including:
  • the candidate list is filled with a combination of the extended reference line and the intra prediction mode of the current block candidate; through rate distortion optimization, select the current block A combination of reference line and intra prediction modes;
  • the TMRL mode flag of the current block is encoded to indicate that the current block uses the TMRL mode
  • the TMRL mode index of the current block is encoded to indicate the position of the selected combination in the candidate list
  • the encoding condition at least includes: the selected combination is in the candidate list.
  • An embodiment of the present disclosure also provides a method for constructing a multi-reference line intra prediction mode candidate list, including:
  • N extended reference lines and M intra prediction modes of the current block N ⁇ M original combinations of extended reference lines and angle modes are obtained;
  • K combinations corresponding to the errors are filled in the candidate list of the template-based multi-reference line intra prediction TMRL mode of the current block, where K, N, M are the set positive values.
  • An embodiment of the present disclosure also provides a method for constructing a multi-reference line intra prediction mode candidate list, including:
  • N extended reference lines and M intra prediction modes of the current block N ⁇ M original combinations of extended reference lines and intra prediction modes are obtained;
  • fusion For each combination of the K original combinations with the smallest error that includes a predetermined angle pattern, determine whether fusion is required. If fusion is required, fill in the fusion combination corresponding to the original combination into the template-based multi-reference line intra prediction of the current block. For the candidate list of TMRL mode, if fusion is not required, fill in the original combination into the candidate list, where K, N, M are set positive integers, 1 ⁇ K ⁇ N ⁇ M.
  • An embodiment of the present disclosure also provides a method for constructing a multi-reference line intra prediction mode candidate list, including:
  • N extended reference lines and M intra prediction modes of the current block N ⁇ M original combinations of extended reference lines and intra prediction modes are obtained;
  • Perform fusion processing on the N ⁇ M original combinations includes: for each original combination including a predetermined angle pattern, when the set conditions are met, replace the original combination with the corresponding fusion combination;
  • K combinations corresponding to the errors are filled in the candidate list of the template-based multi-reference line intra prediction TMRL mode of the current block;
  • K, N, M are set positive integers, 1 ⁇ K ⁇ N ⁇ M;
  • the weighted sum of the first prediction result and the second prediction result is used as the prediction value of the current block; wherein the first prediction result is based on the original combination.
  • the second prediction result is the prediction result of the current block based on the second reference row and the angle mode in the original combination.
  • the second reference row is the prediction result of the current block. It is the adjacent row of the first reference row or the reference row with index 0.
  • An embodiment of the present disclosure also provides a code stream, wherein the code stream is generated by the video encoding method described in any embodiment of the present disclosure.
  • An embodiment of the present disclosure also provides an intra prediction fusion device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the method described in any embodiment of the present disclosure. Intra-frame prediction fusion method.
  • An embodiment of the present disclosure also provides a device for constructing a multi-reference line intra prediction mode candidate list, including a processor and a memory storing a computer program, wherein the processor can implement any of the tasks herein when executing the computer program.
  • a device for constructing a multi-reference line intra prediction mode candidate list including a processor and a memory storing a computer program, wherein the processor can implement any of the tasks herein when executing the computer program.
  • An embodiment of the present disclosure also provides a video decoding device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the video decoding described in any embodiment of the present disclosure. method.
  • An embodiment of the present disclosure also provides a video encoding device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the video encoding described in any embodiment of the present disclosure. method.
  • An embodiment of the present disclosure also provides a video encoding and decoding system, which includes the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.
  • An embodiment of the present disclosure also provides a non-transitory computer-readable storage medium.
  • the computer-readable storage medium stores a computer program, wherein the computer program implements any embodiment of the present disclosure when executed by a processor.
  • the intra prediction fusion method, the construction method of the multi-reference line intra prediction mode candidate list according to any embodiment of the present disclosure, or the video decoding method according to any embodiment of the present disclosure, or the implementation of the present disclosure The video encoding method according to any embodiment.
  • Figure 1A is a schematic diagram of a coding and decoding system according to an embodiment of the present disclosure
  • Figure 1B is a frame diagram of the encoding end according to an embodiment of the present disclosure.
  • Figure 1C is a frame diagram of the decoding end according to an embodiment of the present disclosure.
  • Figure 2 is a schematic diagram of an intra prediction mode according to an embodiment of the present disclosure
  • Figure 3 is a schematic diagram of adjacent intra prediction blocks of the current block according to an embodiment of the present disclosure
  • Figure 4 is a schematic diagram of the template and template reference area of the current block according to an embodiment of the present disclosure
  • Figure 5 is a schematic diagram of multiple reference lines around the current block according to an embodiment of the present disclosure.
  • Figure 6 is a flow chart of a video encoding method according to an embodiment of the present disclosure.
  • Figure 7 is a flow chart of a method for constructing a TMRL pattern candidate list according to an embodiment of the present disclosure.
  • Figure 8A is a schematic diagram of the template area and extended reference lines around the current block according to an embodiment of the present disclosure
  • Figure 8B is a schematic diagram of the template area and extended reference lines around the current block according to another embodiment of the present disclosure.
  • Figure 9 is a flow chart of a video decoding method according to an embodiment of the present disclosure.
  • Figure 10 is a flow chart of an intra prediction fusion method according to an embodiment of the present disclosure.
  • Figure 11 is a flow chart of a video encoding method according to another embodiment of the present disclosure.
  • Figure 12 is a flow chart of a video decoding method according to another embodiment of the present disclosure.
  • Figure 13 is a flow chart of a video decoding method according to another embodiment of the present disclosure.
  • Figures 14, 15 and 16 are respectively TMRL candidate list construction methods and flow charts in three embodiments of the present disclosure.
  • Figure 17 is a schematic diagram of a device for constructing a TMRL pattern candidate list according to an embodiment of the present disclosure
  • Figure 18 is a flow chart of a video encoding method according to another embodiment of the present disclosure.
  • the words “exemplary” or “such as” are used to mean an example, illustration, or explanation. Any embodiment described in this disclosure as “exemplary” or “such as” is not intended to be construed as preferred or advantageous over other embodiments.
  • "And/or” in this article is a description of the relationship between associated objects, indicating that there can be three relationships, for example, A and/or B, which can mean: A exists alone, A and B exist simultaneously, and they exist alone. B these three situations.
  • "Plural” means two or more than two.
  • words such as “first” and “second” are used to distinguish the same or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as “first” and “second” do not limit the number and execution order, and words such as “first” and “second” do not limit the number and execution order.
  • the local illumination compensation method and video encoding and decoding method proposed by the embodiments of the present disclosure can be applied to various video encoding and decoding standards, such as: H.264/Advanced Video Coding (AVC), H.265/High Efficiency Video Coding (High Efficiency Video Coding, HEVC), H.266/Versatile Video Coding (VVC), AVS (Audio Video coding Standard, audio and video coding standard), and MPEG (Moving Picture Experts Group, Moving Picture Experts Group) , other standards formulated by AOM (Alliance for Open Media), JVET (Joint Video Experts Team) and extensions of these standards, or any other customized standards, etc.
  • FIG. 1A is a block diagram of a video encoding and decoding system that can be used in embodiments of the present disclosure. As shown in the figure, the system is divided into an encoding end device 1 and a decoding end device 2.
  • the encoding end device 1 generates a code stream.
  • the decoding end device 2 can decode the code stream.
  • the decoding end device 2 can receive the code stream from the encoding end device 1 via the link 3 .
  • Link 3 includes one or more media or devices capable of moving the code stream from the encoding end device 1 to the decoding end device 2 .
  • the link 3 includes one or more communication media that enable the encoding end device 1 to directly send the code stream to the decoding end device 2 .
  • the encoding end device 1 modulates the code stream according to the communication standard (such as a wireless communication protocol), and sends the modulated code stream to the decoding end device 2 .
  • the one or more communication media may include wireless and/or wired communication media and may form part of a packet network.
  • the code stream can also be output from the output interface 15 to a storage device, and the decoding end device 2 can read the stored data from the storage device via streaming or downloading.
  • the code end device 1 includes a data source 11, a video encoding device 13 and an output interface 15.
  • Data sources 11 include a video capture device (eg, a video camera), an archive containing previously captured data, a feed interface to receive data from a content provider, a computer graphics system to generate the data, or a combination of these sources.
  • the video encoding device 13 encodes the data from the data source 11 and outputs the data to the output interface 15.
  • the output interface 15 may include at least one of a regulator, a modem and a transmitter.
  • the decoding end device 2 includes an input interface 21 , a video decoding device 23 and a display device 25 .
  • the input interface 21 includes at least one of a receiver and a modem.
  • the input interface 21 may receive the code stream via link 3 or from a storage device.
  • the video decoding device 23 decodes the received code stream.
  • the display device 25 is used to display the decoded data.
  • the display device 25 can be integrated with other devices of the decoding end device 2 or set up separately.
  • the display device 25 is optional for the decoding end. In other examples, the decoding end may include other devices or devices that apply decoded data.
  • FIG. 1B is a block diagram of an exemplary video encoding device that can be used in embodiments of the present disclosure.
  • the video encoding device 1000 includes a prediction unit 1100, a division unit 1101, a residual generation unit 1102 (indicated by a circle with a plus sign after the division unit 1101 in the figure), a transformation processing unit 1104, a quantization unit 1106, Inverse quantization unit 1108, inverse transform processing unit 1110, reconstruction unit 1112 (indicated by a circle with a plus sign after the inverse transform processing unit 1110 in the figure), filter unit 1113, decoded image buffer 1114, and entropy encoding unit 1115.
  • the prediction unit 1100 includes an inter prediction unit 1121 and an intra prediction unit 1126, and the decoded image buffer 1114 may also be called a decoded image buffer, a decoded picture buffer, a decoded picture buffer, etc.
  • Video encoder 20 may also include more, fewer, or different functional components than this example, such that transform processing unit 1104, inverse transform processing unit 1110, etc. may be eliminated in some cases.
  • the dividing unit 1101 cooperates with the prediction unit 1100 to divide the received video data into slices, coding tree units (CTU: Coding Tree Unit) or other larger units.
  • the video data received by the dividing unit 1101 may be a video sequence including video frames such as I frames, P frames, or B frames.
  • the prediction unit 1100 can divide the CTU into coding units (CU: Coding Unit), and perform intra prediction encoding or inter prediction encoding on the CU.
  • CU Coding Unit
  • the CU can be divided into one or more prediction units (PU: prediction unit).
  • the inter prediction unit 1121 may perform inter prediction on the PU to generate prediction data for the PU, including prediction blocks of the PU, motion information of the PU, and various syntax elements.
  • the inter prediction unit 1121 may include a motion estimation (ME: motion estimation) unit and a motion compensation (MC: motion compensation) unit.
  • the motion estimation unit may be used for motion estimation to generate motion vectors, and the motion compensation unit may be used to obtain or generate prediction blocks based on the motion vectors.
  • Intra prediction unit 1126 may perform intra prediction on the PU to generate prediction data for the PU.
  • the prediction data of the PU may include the prediction block of the PU and various syntax elements.
  • Residual generation unit 1102 may generate a residual block of the CU based on the original block of the CU minus the prediction blocks of the PU into which the CU is divided.
  • the transformation processing unit 1104 may divide the CU into one or more transformation units (TU: Transform Unit), and the divisions of prediction units and transformation units may be different.
  • the residual block associated with the TU is the sub-block obtained by dividing the residual block of the CU.
  • a TU-associated coefficient block is generated by applying one or more transforms to the TU-associated residual block.
  • the quantization unit 1106 can quantize the coefficients in the coefficient block based on the selected quantization parameter, and the degree of quantization of the coefficient block can be adjusted by adjusting the quantization parameter (QP: Quantizer Parameter).
  • QP Quantizer Parameter
  • Inverse quantization unit 1108 and inverse transform unit 1110 may apply inverse quantization and inverse transform to the coefficient block, respectively, to obtain a TU-associated reconstructed residual block.
  • the reconstruction unit 1112 may add the reconstruction residual block and the prediction block generated by the prediction unit 1100 to generate a reconstructed image.
  • the filter unit 1113 performs loop filtering on the reconstructed image, and stores the filtered reconstructed image in the decoded image buffer 1114 as a reference image.
  • Intra prediction unit 1126 may extract reference images of blocks adjacent to the PU from decoded image buffer 1114 to perform intra prediction.
  • the inter prediction unit 1121 may perform inter prediction on the PU of the current frame image using the reference image of the previous frame buffered by the decoded image buffer 1114 .
  • the entropy encoding unit 1115 may perform an entropy encoding operation on received data (such as syntax elements, quantized coefficient blocks, motion information, etc.).
  • the video decoding device 101 includes an entropy decoding unit 150, a prediction unit 152, an inverse quantization unit 154, an inverse transform processing unit 156, and a reconstruction unit 158 (indicated by a circle with a plus sign after the inverse transform processing unit 155 in the figure). ), filter unit 159, and decoded image buffer 160.
  • the video decoder 30 may include more, fewer, or different functional components, such as the inverse transform processing unit 155 may be eliminated in some cases.
  • the entropy decoding unit 150 may perform entropy decoding on the received code stream, and extract syntax elements, quantized coefficient blocks, motion information of the PU, etc.
  • the prediction unit 152, the inverse quantization unit 154, the inverse transform processing unit 156, the reconstruction unit 158 and the filter unit 159 may all perform corresponding operations based on syntax elements extracted from the code stream.
  • Inverse quantization unit 154 may inversely quantize the quantized TU-associated coefficient block.
  • Inverse transform processing unit 156 may apply one or more inverse transforms to the inverse quantized coefficient block to produce a reconstructed residual block of the TU.
  • Prediction unit 152 includes inter prediction unit 162 and intra prediction unit 164 .
  • intra prediction unit 164 may determine the intra prediction mode of the PU based on the syntax elements decoded from the codestream, based on the determined intra prediction mode and the PU's neighbors obtained from decoded image buffer 160 Intra prediction is performed on the reconstructed reference information to generate the prediction block of the PU.
  • inter prediction unit 162 may determine one or more reference blocks for the PU based on the motion information of the PU and corresponding syntax elements, generated based on the reference blocks obtained from decoded image buffer 160 Prediction block of PU.
  • Reconstruction unit 158 may obtain a reconstructed image based on the reconstruction residual block associated with the TU and the prediction block of the PU generated by prediction unit 152 .
  • the filter unit 159 may perform loop filtering on the reconstructed image, and the filtered reconstructed image is stored in the decoded image buffer 160 .
  • the decoded image buffer 160 can provide a reference image for subsequent motion compensation, intra-frame prediction, inter-frame prediction, etc., and can also output the filtered reconstructed image as decoded video data for presentation on the display device.
  • a frame of image is divided into blocks, and intra-frame prediction or inter-frame prediction or other algorithms are performed on the current block to generate the prediction of the current block.
  • Block use the original block of the current block to subtract the prediction block to obtain the residual block, transform and quantize the residual block to obtain the quantization coefficient, and perform entropy encoding on the quantization coefficient to generate a code stream.
  • intra-frame prediction or inter-frame prediction is performed on the current block to generate the prediction block of the current block.
  • the quantized coefficients obtained from the decoded code stream are inversely quantized and inversely transformed to obtain the residual block.
  • the prediction block and residual The blocks are added to obtain the reconstructed block, the reconstructed block constitutes the reconstructed image, and the reconstructed image is loop filtered based on the image or block to obtain the decoded image.
  • the encoding end also obtains the decoded image through similar operations as the decoding end.
  • the decoded image obtained by the encoding end is also usually called a reconstructed image.
  • the decoded image can be used as a reference frame for inter-frame prediction of subsequent frames.
  • the block division information determined by the encoding end, mode information and parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc. can be written into the code stream if necessary.
  • the decoding end determines the same block division information as the encoding end by decoding the code stream or analyzing the existing information, and determines the mode information and parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc., thereby ensuring the decoding obtained by the encoding end.
  • the image is the same as the decoded image obtained at the decoding end.
  • block-based hybrid coding framework is used as an example above, the embodiments of the present disclosure are not limited thereto. With the development of technology, one or more modules in the framework, and one or more steps in the process Can be replaced or optimized.
  • the current block can be a block-level coding and decoding unit such as the current coding unit (current CU) or the current prediction unit (current PU) in the current image.
  • the encoding end When the encoding end performs intra-frame prediction, it usually uses various angle modes and non-angle modes to predict the current block to obtain the predicted block; based on the rate distortion information calculated between the predicted block and the original block, the optimal intra-frame prediction is selected for the current block. mode, and the intra prediction mode encoding is transmitted to the decoder through the code stream.
  • the decoder obtains the currently selected intra prediction mode through decoding, and performs intra prediction on the current block according to the intra prediction mode.
  • the reference line and intra prediction mode selected for the current block are also expressed as the reference line and intra prediction mode selected for the current block.
  • the angle directions of some angle modes in the angle modes with indexes 2 to 66 can be replaced with wider angle directions, as shown in the figure, the index is -
  • the angle modes of 14 ⁇ -1, 67 ⁇ 80 are the angle modes obtained by wide-angle replacement.
  • the selection of these angle modes does not need to be represented by the flag bit, but by the shape of the current block and the prediction mode index selected by the current block. (2 ⁇ 66) are obtained through the corresponding relationship.
  • the variable whRatio is Abs(Log2(blockWidth)-Log2(blockHeight)).
  • predModeIntra (2 ⁇ 66) is replaced by a wide angle according to whether the following conditions are met.
  • predModeIntra is represented by an index:
  • predModeIntra will be equal to (predModeIntra+65).
  • predModeIntra will be equal to (predModeIntra-67).
  • the height of the current block is greater than the width
  • each angle prediction mode predModeIntra (-14 ⁇ 80) will have an angle value intraPredAngle.
  • the angle value of the angle mode will be used for subsequent angle prediction.
  • each angle mode predModeIntra corresponds to an angle.
  • the angle of each angle prediction mode is the angle in the rectangular coordinate system of the line segment corresponding to the angle prediction mode in Figure 1.
  • angle mode with index number 34, angle The value intraPredAngle is -32, and the angle is -45° or 45° or 135°, which is related to the 0° direction defined by this Cartesian coordinate system.
  • the intra prediction mode mentioned in this article refers to the traditional intra prediction mode including Planar mode, DC mode and angle mode, unless there are other limitations.
  • ECM Enhanced Compression Model
  • MPM first builds an MPM list, and fills the MPM list with the six intra prediction modes most likely to be selected by the current block. If the intra prediction mode selected in the current block is in the MPM list, you only need to encode its index number (only 3 bits are needed). If the intra prediction mode selected in the current block is not in the MPM list but in the 61 non-MPM (non -MPM) mode, the intra prediction mode is encoded using the truncated binary code (TBC) in the entropy coding stage.
  • TBC truncated binary code
  • the MPM list has 6 prediction modes.
  • the MPM in ECM is divided into MPM and Secondary MPM (Secondary MPM).
  • MPM and Secondary MPM use lists of length 6 and length 16 respectively.
  • Planar mode is always filled in the first position in the MPM. The remaining 5 positions are filled in the following three steps in sequence until 5 positions are filled. The extra modes will automatically enter Secondary. MPM.
  • the first step is to fill in the intra prediction modes used by the prediction blocks in the five adjacent positions around the current block in turn; as shown in Figure 3, the five positions include the upper left (AL), upper (A), and upper right of the current block. (AR), left (L) and lower left (BL) positions.
  • a pattern derived using a gradient histogram is used based on the reconstructed pixels around the current block
  • the third step is an angle mode that is similar to the angle mode selected in the first step.
  • the Secondary MPM list can be composed of some main angle modes except the intra prediction mode in MPM.
  • the encoding and decoding order of the MPM flag (mpm_flag) is after the MRL mode, the encoding and decoding of MPM in ECM needs to depend on the MRL flag bit.
  • the MPM flag needs to be decoded to determine whether the current block uses MPM.
  • the current block uses MRL mode, there is no need to decode the MPM flag, and the current block uses MPM by default.
  • Template based intra mode derivation (TIMD: Template based intra mode derivation) and decoder-side intra mode derivation (DIMD: Decoder-side intra mode derivation) are two frames that are not in the VVC standard but are adopted into the ECM reference software Intra-prediction technology, these two technologies can derive the intra-prediction mode of the current block based on the reconstructed pixel values around the current block at the decoder, thereby eliminating the need to encode the index of the intra-prediction mode and thereby saving bits.
  • TDD Template based intra mode derivation
  • DIMD Decoder-side intra mode derivation
  • TIMD is an intra prediction mode for luminance frames.
  • the TIMD mode is generated by the candidate intra prediction mode and the template (Template) area (referred to as the template) in the MPM list.
  • the left adjacent area and the upper adjacent area of the current block (such as the current CU) 11 constitute the template area 12 of the current block.
  • the adjacent area on the left is called the left template area (referred to as the left template), and the adjacent area above is called the upper template area or the upper template area (referred to as the upper template).
  • a template reference (reference of the template) area 13 is provided outside the template area 12 (referring to the left and upper sides).
  • the exemplary size and position of each area are as shown in the figure.
  • the width L1 of the left template and the height L2 of the upper template are both 4.
  • the template reference area 13 may be an adjacent row above the template area or an adjacent column to the left.
  • TIMD assumes that the distribution characteristics of the current block and the template area of the current block are consistent, uses the reconstruction value of the template reference area as the reconstruction value of the reference row, traverses all intra prediction modes in MPM and Secondary MPM to predict the template area, and obtains the prediction result. . Then calculate the error between the reconstruction value on the template area and the prediction result of each mode, expressed by the sum of absolute transformed differences (SATD: Sum of absolute transformed differences), and select the intra prediction mode with the smallest or optimal SATD. This intra prediction mode is used as the TIMD mode of the current block.
  • the decoder can derive the TIMD mode through the same derivation method. If the sequence allows the use of TIMD, each current block requires a flag bit to indicate whether to use TIMD.
  • the current block uses the TIMD mode for prediction, and the decoding process of the remaining syntax elements related to intra prediction such as ISP, MPM, etc. can be skipped, thereby greatly reducing the mode Encoding bits.
  • the final TIMD mode used can be determined according to the following method:
  • mode1 and mode2 are the two angle modes used for intra prediction in MPM.
  • mode1 is the angle mode with the smallest SATD, and its SATD is cost1;
  • mode2 is the angle mode with the second smallest SATD, and its SATD is cost2:
  • the prediction mode that weights the prediction results of mode1 and mode2 is used as the TMID mode of the current block, also called TIMD fusion mode.
  • the weighting method and weight are as follows:
  • Pred Pred mode1 ⁇ w1+Pred mode2 ⁇ w2
  • Pred is the prediction result of the current block using TIMD fusion mode
  • Pred mode1 is the prediction result of the current block using mode1
  • Pred mode2 is the prediction result of the current block using mode2
  • w1 and w2 are the weights calculated based on cost1 and cost2.
  • the prediction mode at this time combines three intra-frame prediction modes: planar mode and the two angle modes with the highest and second highest amplitude values. It is called DIMD fusion mode in the article. In the absence of the highest and second highest amplitude angle modes, prediction using DIMD is equivalent to planar mode prediction.
  • VVC adopts multiple reference line (MRL: Multiple reference line) intra prediction technology.
  • MRL Multiple reference line
  • VVC can also use the reference with index 1.
  • the reference line (Reference line1) and the reference line with index 2 (Reference line2) are used as extended reference lines for intra prediction.
  • MRL is only used in non-planar mode in MPM. When the encoding end predicts each angle mode, all three reference lines must be tried.
  • the current block selects the reference line with the smallest rate distortion cost (RD Cost), and the index of the selected reference line is The encoding is sent to the decoding end.
  • the decoding end decodes to obtain the index of the reference line, and then determines the reference line selected by the current block based on the index of the reference line, which is used for prediction of the current block.
  • reference line 0 reference line0
  • reference line 2 reference line 2
  • reference line 3 is separated from the current block by 3 lines.
  • reference line 3 is the reference line with index 3.
  • reference row is called a "row", this is for convenience of expression.
  • a reference row actually includes one row and one column.
  • the reconstructed value of the reference row used in prediction also includes the reconstruction of one row and one column. value, which is the same as the usual description method in the industry.
  • MRL mode can use more reference lines.
  • the indexes of multiple candidate reference lines are filled in a list.
  • This list is called a multi-reference row index list, abbreviated as MRL index list, and can also be called a multi-reference row list, a candidate reference row list, a reference row index list, etc.
  • MRL index list When the current block does not use TIMD, the length of the MRL index list is 6, that is, there are 6 positions in total, and the indexes of 6 reference rows can be filled in.
  • MRL index list the index filled in the first position is 0, which is the index of the reference row closest to the current block.
  • the indexes filled in the second to sixth positions are 1, 3, 5, and 7 respectively.
  • ,12 is the index of the 5 extended reference lines arranged in order from nearest to farthest from the current block.
  • the encoded MRL index can represent the selected reference row. Taking the MRL index list ⁇ 0,1,3,5,7,12 ⁇ as an example, the MRL indexes corresponding to the 1st position to the 6th position are 0 to 5 respectively. Assuming that the current block selects the reference row with index 0, the MRL index is 0. Assuming that the current block selects the reference row with index 7, the MRL index is 4, and so on in other cases.
  • the MRL index can be encoded by a unary truncation code based on the context model.
  • binary identifiers After encoding, multiple binary identifiers based on the context model are obtained.
  • the binary identifiers can also be called binary identifiers, binary symbols, binary bits, etc. The smaller the value of the MRL index, the shorter the code length and the faster the decoding.
  • MRL mode can also be used at the same time as TIMD mode.
  • One embodiment provides a template-based multiple reference line & intra_intra prediction mode, abbreviated as TMRL mode.
  • the TMRL mode is a candidate list based on a combination of extended reference lines and intra prediction modes.
  • a prediction mode that is encoded and decoded by a combination of an extended reference line and an intra prediction mode.
  • the video coding method in this embodiment is applied to the encoder, as shown in Figure 6, including:
  • Step 110 Construct a candidate list of the TMRL mode of the current block.
  • the candidate list is filled with a combination of the extended reference line of the current block candidate and the intra prediction mode;
  • Step 120 Through rate-distortion optimization, the current block selects a combination of reference line and intra prediction mode for intra prediction;
  • the reference line includes the reference line with index 0 and the extended reference line.
  • the combination of a reference line and intra prediction mode selected in the current block may be a combination of the reference line with index 0 and an intra prediction mode. It is also possible a combination of an extended reference line and an intra prediction mode.
  • Step 130 When the encoding condition of the TMRL mode of the current block is met, encode the TMRL mode flag of the current block to indicate that the current block uses the TMRL mode, and encode the TMRL mode index of the current block to indicate that the selected combination is in the candidate list. s position;
  • the encoding condition at least includes: the selected combination is in the candidate list.
  • the candidate list filled with the combination of the extended reference line of the current block candidate and the intra prediction mode may also be called a candidate list of the TMRL mode.
  • the candidate list is filled with the combination of the extended reference line of the current block candidate and the intra prediction mode, which means that the combination in the candidate list needs to participate in the rate-distortion optimization of the current block, that is, to participate in the selection through the rate-distortion cost Mode selection process for the current block prediction mode. This makes it possible for combinations in the candidate list to be selected.
  • the TMRL mode candidate list constructed in this embodiment is filled with a combination of the extended reference line of the current block candidate and the intra prediction mode, and is no longer a single list of candidate extended reference lines or candidate intra prediction modes. list.
  • the combination of the reference line and intra prediction mode selected in the current block is in the candidate list (at this time, the extended reference line is selected in the current block), that is, the selected combination is a combination in the candidate list, and other encoding conditions are met.
  • the decoding end uses the TMRL mode flag and the TMRL mode index
  • the extended reference line and intra prediction mode selected for the current block can be determined.
  • the combined encoding and decoding method of this embodiment can reduce encoding costs and improve encoding performance.
  • the TMRL mode is based on N ⁇ M combinations of N extended reference lines and M intra prediction modes of the current block. It predicts the template area of the current block respectively, and calculates the reconstruction value of the template area and the prediction value obtained by prediction. The error between them; fill in the corresponding K combinations into the candidate list of the TMRL mode of the current block in order of error from small to large, 1 ⁇ K ⁇ N ⁇ M.
  • the template area is predicted based on 25 combinations of extended reference lines and intra prediction modes in the ECM, and the 25 combinations are sorted according to the ascending order of the errors, and the 12 combinations with the smallest error, that is, the most likely to be selected, are
  • This combination is filled in the candidate example table of the TMRL pattern, so that the TMRL pattern index can use fewer encoding bits to complete the encoding.
  • the extended reference rows ⁇ 1,3,5,7,12 ⁇ such as the extended reference rows with indexes 7 and 12
  • the most likely selected ones can be sorted based on 15 combinations.
  • the encoding bits of the TMRL pattern index can still be fully and effectively utilized.
  • Sorting in ascending order of error allows the reference rows and prediction modes that are more likely to be selected for prediction to be left in the candidate list, and the combinations that are more likely to be selected are ranked higher in the list, which reduces the encoding cost.
  • the creation of candidate lists for TMRL patterns is described in more detail below.
  • the encoding condition further includes: the current block does not use TIMD; the method further includes:
  • the encoding of the TMRL mode flag and TMRL mode index of the current block is skipped;
  • the TMRL mode flag of the current block is encoded to indicate that the current block does not use TMRL mode, and the encoding of the TMRL mode index of the current block is skipped.
  • This embodiment is based on the situation where the TIMD mode is encoded and decoded before the TMRL mode. If the current block uses TIMD mode, there is no need to use TMRL mode, so the encoding of TMRL mode flag and TMRL mode index is skipped. And if the current block does not use TIMD mode, there are two situations: the combination selected by the current block is in the candidate list of TMRL mode, or it is not in the candidate list of TMRL mode. If the selected combination is not in the candidate list. The need to encode TMRL mode flag indicates that the current block does not use TMRL mode and skips encoding of the TMRL mode index. If the selected combination is in the candidate list, both the TMRL pattern flag and the TMRL pattern index need to be encoded.
  • the TMRL mode flag and TMRL mode index provided in this embodiment can replace the original multi-reference row index multiRefIdx.
  • a multi-reference row index can still be used to represent the selected reference row, and the multi-reference row index can be encoded.
  • constructing a candidate list of the TMRL mode of the current block includes: only constructing the candidate list to allow the use of the TMRL mode when the set conditions for allowing the use of the TMRL mode in the current block are met.
  • the conditions include any one or more of the following:
  • the current block is a block in the brightness frame; that is, the TMRL mode is only used for brightness frames (ie, brightness images);
  • the current block is not located at the upper boundary of the coding tree unit CTU; if it is located at the upper boundary of the CTU, no reference line can be used above the current block, so this embodiment uses not being located at the upper boundary of the CTU as a condition to allow the use of TMRL mode;.
  • Condition 3 The current block allows the use of multi-reference line MRL, that is, the TMRL mode is allowed only when the multi-reference line mode is allowed.
  • the size of the current block is not larger than the maximum size of the current block that can use the TMRL mode; the maximum size can be preset. Larger blocks are generally smoother and less likely to have angular details. For such large blocks, the use of TMRL mode can be limited.
  • Condition 5 The aspect ratio of the current block meets the requirements for the aspect ratio of the current block using TMRL mode.
  • TMRL mode is only allowed to be used when the aspect ratio of the current block is not greater than the preset value.
  • the TMRL pattern index is encoded using the Golomb-Rice encoding method.
  • the use of Columbus Rice coding can more reasonably classify candidate combinations into categories with different codeword lengths for encoding and decoding, thereby improving coding efficiency.
  • the method further includes: skipping the encoding of syntax elements in any one or more of the following modes: MPM mode, intra subframe Block division ISP mode, multi-transform selection MTS mode, low-frequency indivisible transform LFNST mode, TIMD mode.
  • the encoding conditions of the TMRL mode no longer include that the current block does not use TIMD.
  • the current block uses the TMRL mode, the current block is not allowed to use TIMD, and the TIMD mode can be skipped. Encoding of syntax elements.
  • the TMRL mode flag and TMRL mode index can simultaneously represent the reference line and intra prediction mode selected in the current block. At this time, there is no need to encode and decode MPM related syntax elements.
  • the TMRL mode can be restricted from being used simultaneously with the multiple transform selection (MTS: Multiple transform selection) mode and/or the low frequency non-separable transform (LFNST: Low frequency non-separable transform) mode.
  • MTS Multiple transform selection
  • LNNST Low frequency non-separable transform
  • An embodiment provides a method for constructing a TMRL pattern candidate list, which can be applied to an encoder or a decoder. As shown in Figure 7, the method includes:
  • Step 210 Obtain N ⁇ M combinations of extended reference lines and intra prediction modes based on the N extended reference lines and M intra prediction modes of the current block, N ⁇ 1, M ⁇ 1, N ⁇ M ⁇ 2;
  • Step 220 Predict the template area of the current block according to the N ⁇ M combinations, and calculate the error between the reconstructed value of the template area and the predicted value;
  • the error in this step can be expressed by the sum of absolute errors (SAD: the sum of absolute difference), or the sum of absolute errors transformed (SATD: Sum of Absolute Transformed Difference), but is not limited to this, and can also be expressed by the difference. Represented by sum of squares (SSD: Sum of Squared Difference), mean absolute difference (MAD: Mean Absolute Difference), mean squared error (MSE: Mean Squared Error), etc.
  • Step 230 Fill in the K combinations corresponding to the errors into the candidate list of the TMRL mode of the current block in order of the errors from small to large, 1 ⁇ K ⁇ N ⁇ M.
  • the candidate list created in this embodiment can implement combined coding of extended reference lines and intra prediction modes, thereby improving coding efficiency.
  • K that is more likely to be selected can be selected from N ⁇ M combinations based on the similarity in distribution characteristics between the current block and the current block template area. combinations, and the combinations with a high probability of being selected are ranked at the front of the candidate list, so that the TMRL pattern index of the selected combination is smaller during encoding, reducing the actual encoding cost.
  • the template area of the current block is set on a reference line closest to the current block; or, the template area of the current block is set on multiple reference lines closest to the current block,
  • the N extended reference lines participating in the combination are extended reference lines located outside the template area.
  • the template area of the current block is set on the reference line 30 with index 0, in the candidate list for building the TMRL pattern, from the predefined extended reference with index ⁇ 1,3,5,7,12 ⁇ Select N extended reference rows that can be used in the row. If there are more than 13 reference lines between the top of the current block and the CTU boundary, 5 extended reference lines with indexes ⁇ 1,3,5,7,12 ⁇ are selected to participate in the combination. If there are 6 or 7 reference lines from above the current block to the CTU boundary, then 3 extended reference lines with indexes ⁇ 1,3,5 ⁇ are selected to participate in the combination, and so on.
  • the template area of the current block is set on the reference line with index 0.
  • the reference line with index 0 is called the reference line where the template area is located, and the reference lines with indexes 1 to 3 are called outside the template area. reference line. If the template area of the current block is set on the reference line with indexes 0 and 1, and the reference line where the template area is located includes the extended reference line, then the reference lines with indexes 0 and 1 are the reference lines where the template area is located, and the index
  • the reference lines 2 and 3 are the reference lines outside the template area (upper and left sides).
  • Figure 8A shows five extended reference rows participating in the combination: reference row 31 with index 1, reference row 33 with index 3, reference row 35 with index 5, reference row 37 with index 7, and reference row 37 with index 12 Reference line 39.
  • the template area 40 of the current block is set on two reference lines with indexes 0 and 1, and the extended reference lines participating in the combination are 5 extended reference lines, respectively. These are the reference row 42 with index 2, the reference row 43 with index 3, the reference row 45 with index 5, the reference row 47 with index 7, and the reference row 49 with index 12. That is, this example selects N extended reference rows that can be used from the predefined extended reference rows with indexes ⁇ 2,3,5,7,12 ⁇ . There are many choices for the template area and extended reference lines.
  • the N extended reference lines of the current block are extended reference lines located outside the template area of the current block and not exceeding the CTU boundary among the predefined N max extended reference lines; where, N max It is the maximum number of extended reference lines that can be used in TMRL mode.
  • N max It is the maximum number of extended reference lines that can be used in TMRL mode.
  • This embodiment limits the N extended reference lines used for combination to an area outside the template area of the current block and not exceeding the CTU boundary. However, if the hardware can provide support, you can also choose extended reference lines beyond the CTU boundary to participate in the combination.
  • N max 5
  • the five predefined extended reference rows are references with indexes ⁇ 1,3,5,7,12 ⁇ or ⁇ 2,3,5,7,12 ⁇ OK.
  • the predefined N max extension reference rows are the N max extension reference rows closest to the current block with indexes starting from 1; or, the N max extension reference rows with indexes starting from 1 and having an odd number are the closest to the current block.
  • N max extended reference rows alternatively, the N max extended reference rows closest to the current block starting from 2 and having an even index. Selecting odd or even reference rows simplifies the operation.
  • the M intra prediction modes are only allowed to be selected from the angle mode, or are only allowed to be selected from the angle mode and the DC mode, or are allowed to be selected from the angle mode, the DC mode and the Planar mode. out.
  • the M intra prediction modes are selected in the following manner, M ⁇ 5:
  • the first step Determine the intra prediction modes used by the prediction blocks in the five adjacent positions around the current block, select the intra prediction modes that are allowed to be selected in sequence, and remove duplicate modes;
  • the five adjacent positions are upper left, upper, upper right, left and lower left in order; as shown in Figure 3.
  • the angle mode expansion operation is performed in sequence to obtain the expanded angle mode, and the expanded angle mode that is different from all the angle modes that have been selected is selected until the angle mode has been selected.
  • the total number of selected intra prediction modes is equal to M.
  • the third step determines the set of predefined intra prediction modes allowed to be selected.
  • the selected intra prediction mode is selected from the determined intra prediction modes that are allowed to be selected, and intra prediction modes that are different from all the selected intra prediction modes are selected in sequence until the selected intra prediction mode is The total number is equal to M.
  • only the angle mode is allowed to be selected.
  • only the angle mode and the DC mode are allowed to be selected; in yet another example, in the process of selecting M intra prediction modes, only the angle mode and the DC mode are allowed to be selected.
  • Angle mode, DC mode and Planar mode are available. The combination of Planar mode and extended reference row has limited effect and does not need to participate in the combination. The situation is similar in DC mode. However, if you can accept the increased computational complexity, you can also add Planar mode and DC mode to the candidate list to participate in the combination.
  • the extended operation of the angle mode includes any one or more of the following operations: adding 1 and subtracting 1 to the angle mode; adding 2 and subtracting 2 to the angle mode; and, adding 1 and subtracting 1 to the angle mode.
  • Mode's add 3 and subtract 3 operations.
  • the M intra prediction modes use part or all of the intra prediction modes in the MPM except the Planar mode; or use part or all of the MPM and the second MPM except the Planar mode.
  • Intra prediction mode or, use some or all intra prediction modes in the most probable mode MPM except Planar mode and DC mode; or use some or all frames in MPM and the second MPM except Planar mode and DC mode
  • Intra prediction mode or, use some or all intra prediction modes in the most probable mode MPM except Planar mode, DC mode and DIMD mode; or, use the M intra prediction modes in MPM and the second MPM except Planar mode. , some or all intra prediction modes except DC mode and DIMD mode.
  • This example is to use all intra prediction modes in MPM except Planar mode as predefined intra prediction modes, or to use all intra prediction modes in MPM and the second MPM except Planar mode as predefined intra prediction modes.
  • mode when the reference line of the current block includes all predefined extended reference lines, all predefined intra prediction modes are used. When the reference line of the current block only includes a predefined partial extended reference line, the partially predefined intra prediction mode is used.
  • M intra prediction modes are selected in the following manner:
  • the template of the current block is predicted respectively and the error between the reconstructed value of the template and the predicted predicted value is calculated to obtain M' error;
  • M intra prediction modes with the smallest corresponding errors are selected from the M' intra prediction modes as the M intra prediction modes participating in the combination, M ⁇ M'.
  • the template of the current block is used to select M intra prediction modes from M' intra prediction modes, and the template area of the aforementioned current block is used to select K combinations from N ⁇ M combinations.
  • the two can be Different, but can also occupy the same area.
  • various methods for selecting M intra prediction modes in the above embodiments can be used, for example, directly selecting from the list of MPM and second MPM, or using the aforementioned implementation.
  • the method in the example is selected through the first step, or through the first and second steps, or through the first, second and third steps, and so on.
  • K max is the maximum number of candidate combinations allowed in TMRL mode.
  • N, M and K have at least two sets of values.
  • the first set of values are N 1 , M 1 , K 1
  • the second set of values are N 2 , M 2 . K 2 , where N 1 ⁇ N 2 , M 1 ⁇ M 2 , K 1 ⁇ K 2 , and N 1 ⁇ M 1 ⁇ N 2 ⁇ M 2 ;
  • the first set of values is the current value of the first size block is used when building the candidate list, and the second set of values is used when building the candidate list for the current block of a second size, the first size being smaller than the second size.
  • the first size and the second size here may respectively represent multiple sizes.
  • the first size may include 4 ⁇ 4, 4 ⁇ 8, 8 ⁇ 8, etc.
  • the second size may include 16 ⁇ 8, 16 ⁇ 16, 8 ⁇ 16 etc.
  • N, M and K for current blocks of different sizes.
  • the current block size is small, smaller values are used to build the candidate list of TMRL mode.
  • the current block size is large, a larger value is used to build the candidate list of the TMRL mode. A better balance can be achieved between computational complexity and performance.
  • predicting the template area of the current block according to the N ⁇ M combinations includes: when the current block is located at the left boundary of the image (picture), predicting the template area according to the N ⁇ M combinations. M combinations respectively predict the upper template area of the current block, but do not predict the left template area of the current block. This embodiment can simplify the operation and reduce the time required for the operation without affecting the performance.
  • predicting the template area of the current block according to the N ⁇ M combinations includes: for each of the N ⁇ M combinations, proceed in the following manner predict:
  • the initial prediction value of the template region is calculated according to the reconstruction value of the extended reference line in the combination and the intra prediction mode in the combination, wherein the reconstruction value of the extended reference line is the original reconstruction of the extended reference line. value or filtered reconstructed value;
  • the reconstructed value of the extended reference line may not be filtered, and the original reconstructed value may be used for calculation, or a shorter-tap filter (such as a 4-tap filter) may be used.
  • a shorter-tap filter such as a 4-tap filter
  • the template region of the current block is predicted respectively according to the N ⁇ M combinations, and the error between the reconstructed value of the template region and the predicted predicted value is calculated, including :
  • the entire template area of the current block is predicted according to K combinations, and the corresponding K errors are obtained to form an error set, and the maximum error in the error set is recorded as D max ;
  • the K combinations corresponding to the K errors in the error set are used as the K combinations with the smallest corresponding errors.
  • the errors in the error set can be arranged in order from small to large.
  • D 2 should be inserted into a position such that the errors in the error set are still arranged in order from small to large.
  • the K errors in the error set can be sorted.
  • the combinations can be sorted, which can reduce the computational complexity and speed up the calculation.
  • the K combinations corresponding to the errors are filled into the candidate list of the TMRL mode of the current block in the order of the errors from small to large, including: starting from the first of the candidate list Starting from position 1, K combinations corresponding to the errors are filled in the candidate list in order of the errors from small to large.
  • the candidate list of the TMRL mode in this example is only filled with the combination of the extended reference line and the intra prediction mode.
  • the combination of the reference line with index 0 and the intra prediction mode is indicated by other legacy modes, such as MPM.
  • the candidate list of the TMRL mode in this example is not only filled with the combination of the extended reference line and the intra prediction mode, but can also be filled with the combination of the reference line with index 0 and the intra prediction mode.
  • the selected combination of the current block is the combination of the reference line with index 0 and the intra prediction mode, it can also be represented by the TMRL mode index.
  • the TMRL mode flag at this time can still be used.
  • One embodiment provides a video decoding method related to TMRL mode, which is applied to the decoder. As shown in Figure 9, the method includes:
  • Step 310 decode the multi-reference line intra prediction TMRL mode flag of the current block and determine whether the current block uses the TMRL mode;
  • Step 320 If it is determined that the current block uses the TMRL mode, continue to decode the TMRL mode index of the current block, and construct a candidate list of the TMRL mode of the current block.
  • the candidate list is filled with the extended reference lines and frames of the current block candidates.
  • Step 330 Determine the combination of the extended reference line and intra prediction mode selected in the current block according to the candidate list and the TMRL mode index, and predict the current frame according to the selected combination;
  • the TMRL mode index is used to represent the position of the combination of the selected extended reference line and the intra prediction mode in the candidate list.
  • the combination of the extended reference line and the intra prediction mode is filled in the candidate list of the TMRL mode, and the selected TMRL mode of the current block is determined by the decoded TMRL mode index and candidate list. Combine and make predictions. That is, the TMRL mode index can simultaneously indicate the extended reference line and intra prediction mode selected in the current block, and there is no need to use two indexes to complete. Can reduce encoding cost.
  • the method before decoding the TMRL mode flag of the current block, the method further includes: decoding the TMRL mode flag of the current block when all the conditions for allowing the use of the TMRL mode in the current block are met.
  • the TMRL mode flag is allowed to be used in the current block. Conditions include any one or more of the following:
  • the current block is the block in the luma frame
  • the current block is not located at the upper boundary of the coding tree unit CTU;
  • the current block does not derive TIMD using template-based intra prediction mode.
  • the use of TMRL mode is not allowed, and the decoding of the TMRL mode flag and TMRL mode index can be skipped.
  • this is not necessarily the case in other embodiments.
  • the use of TIMD in the current block cannot be used as a condition for not allowing the use of TMRL mode.
  • the hardware can support obtaining reference lines outside the CTU boundary in the future, the current block being located at the upper boundary of the CTU will no longer be a condition for not allowing the use of TMRL mode, and so on.
  • the method further includes: decoding the multi-reference line index of the current block when the decoding determines that the current block is allowed to use MRL, the current block is not located at the upper boundary of the CTU, and the current block uses TIMD,
  • the multi-reference row index is used to represent the position of the reference row selected in the current block in the multi-reference row index list.
  • the TMRL mode is not allowed, but MRL is still allowed. Therefore, the reference line selected by the current block can still be determined by decoding the multi-reference line index of the current block, and then combined with the current block selection
  • the TIMD mode can predict the current block.
  • the method when it is determined that the current block uses the TMRL mode according to the TMRL mode flag, the method further includes: skipping the decoding of syntax elements in any one or more of the following modes: MPM mode, Sub-partition mode, multi-transform selection MTS mode, low-frequency indivisible transform LFNST mode, TIMD mode.
  • MPM mode MPM mode
  • Sub-partition mode multi-transform selection MTS mode
  • low-frequency indivisible transform LFNST mode low-frequency indivisible transform LFNST mode
  • TIMD mode TIMD mode.
  • the decoding end determines through decoding that the current block uses the TMRL mode flag. Skip decoding these patterns.
  • An embodiment also provides a video decoding method, which mainly involves the decoding process of intra-frame prediction, and also provides corresponding explanations on the encoding end.
  • the encoding end constructs a candidate list of the TMRL mode. When the combination in the candidate list is selected through mode selection, the syntax elements of the TMRL mode are encoded and decoded, and the extended reference line and intra prediction mode are combined to encode and decode. .
  • a template is constructed at the position of the reference line with index 0, that is, reference line 0, based on the predefined N extended reference lines and M intra prediction modes. See template area 30 shown in Figure 8A.
  • (x,-1), (-1,y) are the coordinates relative to the upper left corner (0,0) of the current block respectively.
  • the figure also adds five predefined extended reference lines, with indexes ⁇ 1,3,5,7,12 ⁇ .
  • the construction of the candidate list of the TMRL mode is an operation that both the encoder and the decoder need to perform.
  • the encoding end encodes the TMRL mode flag to indicate the use of the TMRL mode when the combination selected in the current block satisfies the encoding conditions such as the candidate list.
  • the TMRL pattern index is also determined based on the position of the selected combination in the candidate list. For example, at the first position, the TMRL pattern index is 0, at the second position, the TMRL pattern index is 1, and so on.
  • the TMRL pattern index can be encoded using Golomb-Rice encoding, but is not limited to this.
  • the indexes of the 5 predefined extended reference lines are ⁇ 1, 3, 5, 7, 12 ⁇ , and the 6 intra prediction modes are selected step by step.
  • Step 1 Decode the grammatical elements related to the TMRL pattern
  • the decoder parses the relevant syntax elements of the intra prediction mode, including relevant syntax elements of TIMD, MRL and other modes.
  • the TMRL mode proposed in this embodiment can be regarded as an evolution of the MRL mode, and the syntax elements of the TMRL mode can also be regarded as the MRL mode syntax. part of the element. Of course, the two can also be regarded as two different modes.
  • the decoding method of MRL mode syntax elements remains unchanged. If the current block does not use the TIMD mode, you need to decode the syntax elements of the TMRL mode.
  • the syntax related to the current block decoding is as shown in the following table:
  • “cu_tmrl_flag” in the table is the TMRL mode flag. When equal to 1, it means that the current block uses TMRL mode, which means that the intra-frame prediction type of the current brightness sample is defined as a template-based multi-reference line frame prediction mode; "cu_tmrl_flag" when equal to 0 means that the current block uses TMRL mode.
  • the block does not use the TMRL mode, that is, the intra prediction mode type defined for the current brightness sample is not a template-based multi-reference line frame prediction mode.
  • tmrl_idx in the table is the TMRL mode index.
  • the position of the combination of the extended reference line and intra prediction mode selected in the current block in the TMRL mode candidate list can also be said to define the selected combination in the TMRL mode.
  • the index in the candidate list (the index indicating the position of the combination).
  • "tmrl_idx” can be encoded and decoded using Columbus Rice method, which will not be described here.
  • the current block is allowed to use MRL (that is, whether sps_mrl_enabled_flag is 1 is true), the current block is not located at the upper boundary of the CTU (that is, whether (y0%CtbSizeY)>0 holds), and the current block does not use TIMD.
  • the current block is allowed to use MRL (that is, whether sps_mrl_enabled_flag is 1 is true)
  • the current block is not located at the upper boundary of the CTU (that is, whether (y0%CtbSizeY)>0 holds)
  • the current block does not use TIMD.
  • decode cu_tmrl_flag If the other two conditions are true and the current block uses TIMD, decode the multi-reference row index intra_luma_ref_idx of the current block.
  • the ISP mode flag (intra_subpartitions_mode_flag) in the table is decoded after the TMRL mode related syntax elements. If the current block does not use TMRL mode (!cu_tmrl_flag is established), intra_subpartitions_mode_flag is decoded again. Similarly, if the current block does not use TMRL mode (!cu_tmrl_flag is established), then decode the MPM related syntax elements.
  • Step 2 Construct a candidate list of the TMRL mode, and determine the extended reference line and intra prediction mode selected in the current block according to the TMRL mode index and the candidate list;
  • a candidate list of the TMRL mode needs to be constructed, and the extended reference line and intra prediction mode selected for the current block are determined based on the TMRL mode index and the candidate list.
  • Candidate extended reference lines are selected from predefined extended reference lines. Determine which of the predefined extended reference lines can be used based on the position of the current block in the image. In principle, the upper reference lines that can be used by the current block should not exceed the upper CTU boundary. In one example, among the extended reference rows with indexes ⁇ 1, 3, 5, 7, 12 ⁇ , all extended reference rows that do not exceed the CTU boundary are added to the candidate extended reference rows. In order to obtain better encoding and decoding performance, or to reduce complexity, more or fewer extended reference lines can also be used.
  • the TMRL mode is not bound to the MPM (it can also be bound in other embodiments), but a candidate list of intra prediction modes is constructed, and the intra prediction mode used for combination will be selected from this candidate list.
  • the candidate list is exported as follows:
  • the Planar mode and the DC mode are removed, or only the Planar mode is removed and the DC mode is retained.
  • the removed mode is not added to the candidate list, that is, it is not used as an intra prediction mode participating in the combination in the TMRL mode.
  • the length of the candidate prediction mode list to be constructed in this embodiment is 6.
  • non-overlapping intra prediction modes are selected sequentially from the intra prediction modes used by prediction blocks at five adjacent positions around the current block to fill the candidate prediction mode list.
  • perform an angle mode expansion operation on the modes that have been filled in the list. Specifically, you can add one and subtract one to the angle mode, select the non-duplicate extended angle modes, and fill in the candidate prediction mode list in turn. If the candidate list When the number of filled patterns reaches 6, the filling stops.
  • angle mode-1 refers to the angle mode obtained by subtracting 1 from the index of the filled angle mode.
  • the filled angle mode is mode 3
  • angle mode-1 is angle mode 2.
  • angle mode + 1 refers to the angle mode obtained by adding 1 to the index of the filled angle mode.
  • the filled angle mode is mode 3, and angle mode + 1 is angle mode 4.
  • angle mode -1 is smaller than angle mode 2, for example, angle mode -1 results in angle mode 1 (the index of angle mode is numbered from 2, and angle mode 1 does not exist), then select the angle mode in the opposite direction of -1 angle, Assume that there are 65 angle modes in total, and the angle mode in the opposite direction is angle mode 66. If the angle mode after angle mode +1 is larger than angle mode 66, then selecting the angle mode in the opposite direction of +1 angle is similar. If the filled angle mode is angle mode 66, then the angle mode +1 does not exist. The angle mode in the opposite direction of the +1 angle selected at this time is angle mode 2.
  • the operation of adding 1 and subtracting 1 is performed on the angle mode.
  • This pattern set includes some angle patterns filtered out according to statistical rules, as follows:
  • mpm_default[] ⁇ DC_IDX,VER_IDX,HOR_IDX,VER_IDX-4,VER_IDX+4,14,22,42,58,10,26,38,62,6,30,34,66,2,48,52,16 ⁇ ;
  • DC_IDX represents DC mode
  • VER_IDX represents vertical mode
  • HOR_IDX represents horizontal mode
  • the remaining numbers represent the angle mode corresponding to the number.
  • the length of the candidate prediction mode list is 6. You can also try more angle modes for performance and set the length to a value greater than 6. You can also try fewer modes to reduce complexity and set the length to A value less than 6.
  • this embodiment excludes Planar mode and DC mode, or only Planar mode. However, if complexity is not considered, these two modes may not be excluded, that is, Planar mode and DC mode. And all angle modes can participate as candidate intra prediction modes in combination with the extended reference line.
  • this embodiment only limits the use of the TMRL mode when the current block occupies the first row of the CTU.
  • the TMRL mode can still be used. In this case, due to the reference on the left Line 0 is already outside the image boundary, so the left template is not used during prediction, that is, only the upper template area is predicted.
  • the prediction process in the template area can be completely consistent with other normal intra-frame angle prediction processes, that is, the reconstruction value of the reference line is first filtered and then used as the initial prediction value of the template area, and based on the filtered reconstruction value of the reference line and the combination After predicting the template area in the intra prediction mode, the initial prediction result is filtered with 4 or 6 taps and then used as the predicted value. Considering the complexity of the operation, the filtering step of the reconstructed value of the reference row can be omitted, or a filter with shorter taps can be used. In this embodiment, when predicting the predicted value of the template area, the reconstructed value of the reference row pixel is not filtered, and the initial prediction result is filtered by 4-tap interpolation with 1/32 accuracy at a non-integer angle.
  • the template region is predicted by the filter based on the angle in the current combination, the reference row, and the filter. Calculate the SAD between the predicted value of the predicted template area and the reconstructed value of the template area, sort according to SAD in ascending order, and select the 12 combinations with the smallest SAD to fill in the candidate list of the TMRL mode.
  • a fast algorithm can be used in the sorting process.
  • the first 12 combinations are predicted After obtaining the corresponding SAD, starting from the 13th combination, you only need to keep the 12 combinations with the smallest error (also called the cost) and update them. Starting from the 13th combination, only the upper template area is predicted and the corresponding SAD is calculated. When the SAD calculated based on the upper template is already greater than the one with the largest error among the 12 combinations with the smallest error, you can skip the prediction on the left side.
  • the prediction and error calculation of the template area please refer to the foregoing embodiments for details.
  • Step 3 Determine the combination of the extended reference line and intra prediction mode selected in the current block based on the constructed TMRL mode candidate list and the decoded TMRL mode index, and perform intra prediction on the current block based on the selection.
  • the index refIdx of the reference row and the variable predModeIntra define the mode used for intra prediction, which are determined based on the TMRL mode index "tmrl_idx" and the candidate list of the TMRL mode.
  • Encoding Time, encoding time, 10X% means that when the reference row sorting technology is integrated, the encoding time is 10X% compared to before it is not integrated, which means that there is an X% increase in encoding time.
  • DecT Decoding Time, decoding time, 10X% means that when the reference row sorting technology is integrated, the decoding time is 10X% compared to before it is not integrated, which means that there is an X% increase in decoding time.
  • ClassA1 and Class A2 are test video sequences with a resolution of 3840x2160
  • ClassB is a test sequence with a resolution of 1920x1080
  • ClassC is 832x480
  • ClassD is 416x240
  • ClassE is 1280x720
  • ClassF is a screen content sequence of several different resolutions (Screen content) .
  • Y, U, and V are the three color components.
  • the columns of Y, U, and V represent the BD-rate of the test results on Y, U, and V. Indicator, the smaller the value, the better the encoding performance.
  • All intra represents the test configuration of the full intra frame configuration.
  • a template area of 1 row and 1 column is used, and SAD ascending order is used for sorting and filtering.
  • extended reference lines if all extended reference lines (including reference line 1) are sorted, only the template with 1 row and 1 column can be used.
  • more reference rows can be used like the TIMD mode to obtain more accurate results.
  • the manner in which the TMRL mode determines the candidate prediction mode list may also be changed.
  • TMRL mode candidate list when you need to build a candidate prediction mode list with a length of 6, you can first build a list with a length greater than 6 according to the same construction and filling method as in this embodiment, and then use the 4 rows and 4 columns closest to the current block as a template , use the fifth reference row and the intra prediction mode in the candidate prediction mode list to predict the template, calculate the error (SAD or SATD) between the predicted value and the reconstructed value of the template, sort in ascending order of error, select Six of the intra prediction modes with small errors are selected as the intra prediction modes in the TMRL mode candidate list of length 6 to be constructed.
  • the length of the candidate list of TMRL mode is 6 is just an example, and the value can be adjusted according to the situation.
  • angles can also be expanded to 129 or more to obtain better performance.
  • intra-frame prediction The number of filters should also be increased accordingly, for example, 1/64 precision filtering is used when 129 angles are used.
  • IPF Intra prediction fusion
  • IPF allows the angle mode to be weighted using the prediction results of two adjacent reference lines to obtain the final prediction result of the current block.
  • the weighting method is as follows:
  • p a is the prediction result of the current block using the reference line with index a (reference line a) and the angle mode
  • p b is the result of using the reference line with index a+1 (reference line a+1) and the angle mode.
  • the angle mode predicts the current block
  • p fusion is the fusion prediction result
  • w a is the weight of p a when l is weighted
  • w b is the weight of p b when weighting
  • w a is 3/4
  • w b is 1 /4.
  • the above fused prediction result is used as the final prediction result of the current block.
  • the angle mode selected in the current block is not an angle mode with an integer slope
  • the width multiplied by the height of the current block is greater than 16
  • the intra block division ISP mode is not selected for the current block.
  • the angle mode selected in the current block is the angle mode with integer slope
  • the width multiplied by the height of the current block is less than or equal to 16;
  • the current block selects ISP mode.
  • the angle mode is an angle mode with an integer slope.
  • the intra prediction mode selected for the current block may be the DIMD fusion mode selected using DIMD (fusion of planar mode and two angle modes), or it may be the TIMD fusion mode selected using TIMD. From the perspective of hardware implementation, the less fusion, the better. When DIMD fusion mode is selected, there are already three intra-frame prediction results fused on the current block. If IPF is used again, too much fusion will make the prediction phase The complexity increases.
  • An embodiment of the present disclosure provides an intra prediction fusion method, which can be applied to an encoder or a decoder. As shown in Figure 10, the method includes:
  • Step 410 If the selected intra prediction mode of the current block includes the angle mode, determine whether the restriction condition of the current block using intra prediction to fuse IPF is established;
  • Step 430 When at least one of the restriction conditions is true, restrict the use of IPF when performing intra prediction on the current block.
  • the use of IPF is restricted when performing intra prediction on the current block, either by not allowing the use of IPF, or by only allowing the use of IPF when the selected intra prediction mode includes multiple angle modes that meet the IPF usage conditions.
  • Some of the angle modes are IPF fused.
  • the angle pattern that satisfies the IPF usage conditions refers to an angle pattern that is not an integer slope, or an angle pattern other than -45°, 0°, 45°, 90°, or 135°.
  • IPF is not allowed to be used when a certain restriction condition is met, which is a sufficient condition for not using IPF when predicting the current block.
  • IPF is allowed to be used, which is a necessary condition for using IPF when predicting the current block. If other constraints are true, IPF may still not be allowed to be used.
  • the restriction includes the following restriction on the number of modes: using IPF will cause the prediction of the current block to incorporate more than N intra prediction modes, where N is an integer greater than or equal to 3. .
  • the intra prediction mode selected in the current block includes the angle mode
  • TIMD fusion mode When the TIMD fusion mode is selected for the current block and TIMD fuses two angle modes that meet the IPF usage conditions, it is determined that the restriction on the number of modes is established. When performing intra prediction on the current block, only the two angle modes are allowed. One performs IPF fusion;
  • TIMD fusion mode When the TIMD fusion mode is selected for the current block and only one of the two intra prediction modes of TIMD fusion is an angle mode that meets the IPF usage conditions, it is determined that the restriction on the number of modes does not hold, and intra prediction is allowed for the current block. Perform IPF fusion on this angle pattern.
  • the restriction includes the following mode number restriction: using IPF will cause the prediction of the current block to fuse more than M angle modes, where M is an integer greater than or equal to 2.
  • the restriction condition of using IPF for the current block is established; when the restriction When at least one of the conditions is true, the use of IPF is restricted when performing intra prediction on the current block, including:
  • TIMD fusion mode When the TIMD fusion mode is selected for the current block and TIMD fuses two angle modes that meet the IPF usage conditions, it is determined that the restriction on the number of modes is established. When performing intra prediction on the current block, only the two angle modes are allowed. One performs IPF fusion. For example, only the angle mode with the smallest or second-lowest cost among the two angle modes is allowed to be IPF fused;
  • TIMD fusion mode When the TIMD fusion mode is selected for the current block and only one of the two intra prediction modes of TIMD fusion is an angle mode that meets the IPF usage conditions, it is determined that the restriction on the number of modes does not hold, and intra prediction is allowed for the current block. Perform IPF fusion on this angle pattern.
  • the restriction condition of using IPF for the current block is established; when the restriction When at least one of the conditions is true, the use of IPF is restricted when performing intra prediction on the current block, including:
  • the DIMD fusion mode is selected for the current block and the two angle modes of DIMD fusion both meet the IPF usage conditions, it is determined that the restriction on the number of modes is established.
  • the two angle modes are allowed.
  • the current block uses the DIMD fusion mode and only one of the two angle modes of DIMD fusion satisfies the IPF usage conditions, it is determined that the mode number restriction condition is not established, and the IPF usage conditions are allowed when performing intra prediction on the current block.
  • This angle mode performs IPF fusion.
  • the intra prediction mode selected in the current block includes the angle mode
  • TIMD fusion mode When the TIMD fusion mode is selected for the current block and TIMD fuses two angle modes, it is determined that the restriction on the number of modes is established, and IPF is not allowed to be used when performing intra prediction on the current block;
  • the TIMD fusion mode is selected for the current block and only one of the two intra prediction modes of TIMD fusion is the angle mode, and the angle mode satisfies the IPF usage conditions, it is determined that the restriction on the number of modes does not hold, and the current block is framed.
  • IPF fusion is allowed for the angle pattern that meets the IPF usage conditions.
  • the intra prediction of the current block allows IPF fusion of an angle mode, including: using the weighted sum of two or more prediction results as the final prediction result of the current block;
  • the prediction results include the prediction results obtained by predicting the current block based on the first reference line selected by the current block and the angle mode, and the prediction results obtained by predicting the current block based on the second reference line that is different from the first reference line and the angle mode. Predict the predicted results.
  • the second reference line may be a reference line adjacent to the first reference line or a reference line with an index of 0.
  • the restrictions include any one or more of the following restrictions:
  • Restriction 1 The TIMD fusion mode is selected for the current block
  • Restriction 3 The current block uses multi-reference line MRL;
  • the index of the reference row selected in the current block is greater than or equal to K, and K is an integer greater than or equal to 3;
  • Restriction 5 The width of the current block is less than or equal to the set value
  • Restriction 5 The height of the current block is less than or equal to the set value
  • Restriction 7 The width multiplied by the height of the current block is less than or equal to the set value
  • Restriction 8 The current block uses intra-frame sub-block division ISP mode
  • the angle mode selected in the current block is any one of -45°, 0°, 45°, 90°, and 135°, or the angle mode selected in the current block is an angle mode with an integer slope;
  • the current frame to which the current block belongs is an inter-frame
  • the current frame to which the current block belongs is a chroma frame, that is, only IPF is used for luminance frames and IPF is not used for chroma frames;
  • IPF is not allowed to be used when performing intra prediction on the current block.
  • this embodiment can limit the number of fused intra prediction modes, limit the number of fused angle modes, and limit IPF and TIMD, DIMD, and other prediction modes that may fuse multiple modes. Using etc. at the same time can avoid excessive fusion during prediction, resulting in an undue increase in the complexity of the prediction stage.
  • This embodiment can limit the size of the current block, and only use IPF when the size of the current block is larger than the set value. This is because when the current block is smaller than a certain size, it usually has more textures, and using fusion prediction has limited improvement in performance.
  • IPF is not allowed to be used when the index of the reference row selected in the current block is greater than or equal to K, that is, IPF is not allowed to be used when the extended reference row selected in the current block is far away from the current block. In this case, the performance improvement of using IPF is limited. .
  • This embodiment can reduce calculation complexity by limiting the simultaneous use of IPF and MRL. This embodiment does not allow the use of IPF when the current frame is an inter-frame frame (such as a B frame or a P frame), which can reduce the cost of inter-frame encoding and decoding.
  • an inter-frame frame such as a B frame or a P frame
  • the method further includes: when none of the restrictive conditions are established, using IPF when performing intra prediction on the current block; using IPF when performing intra prediction on the current block.
  • IPF including:
  • the weighted sum of the first prediction result and the second prediction result is used as the final prediction result of the current block; wherein, the first prediction result is The result of predicting the current block according to the first reference row selected by the current block and the first angle mode selected by the current block.
  • the second prediction result is the prediction of the current block according to the second reference row and the first angle mode.
  • the second reference row is an adjacent row of the first reference row or a reference row with an index of 0.
  • IPF when none of the above-mentioned restriction conditions are met, IPF is used when performing intra prediction on the current block. It does not mean that IPF is allowed to be used only when intra prediction is performed on the current block if all the constraints are not met. For example, when the aforementioned restriction on the number of modes holds true, IPF fusion can still be performed on some angle modes that meet the IPF usage conditions.
  • the weighted sum of the first prediction result and the second prediction result is calculated according to the following formula:
  • p a is the first prediction result
  • p b is the second prediction result
  • p fusion is the final prediction result of the current block
  • w a is the weight of p a
  • w b is the weight of p b ;
  • the above algorithm can avoid decimals during the operation and improve algorithm efficiency.
  • the method further includes: when none of the restrictive conditions are established, using IPF when performing intra prediction on the current block; using IPF when performing intra prediction on the current block.
  • IPF including:
  • the weighted sum of more than three prediction results is used as the final prediction result of the current block; wherein the three or more prediction results include: based on the selected prediction result of the current block.
  • the weight given to the first prediction result is greater than the weight given to the second prediction result.
  • the first prediction result is given a weight of 3/4
  • the second prediction result is given a weight of 1/4.
  • the second reference row when the index of the first reference row is greater than or equal to K, the second reference row is adjacent to the first reference row and is greater than the first reference row. Close to the current block, K is an integer greater than or equal to 1.
  • the second reference row may be determined according to at least one of the following methods:
  • IPF can be used simultaneously with TMRL mode. That is, when the current block selects an extended reference line and an angle mode in the candidate list of the TMRL mode, and the above constraints are not established, calculate the first prediction method for the current block based on the extended reference line and the angle mode. The prediction result, and the second prediction result of predicting the current block according to another reference line and the angle mode, the weighted sum of the first prediction result and the second prediction result is used as the final prediction result of the current block, where, the The other reference row is the reference row with index 0 or an adjacent row of the extended reference row.
  • An embodiment of the present disclosure also provides a video encoding method, applied to an encoder, as shown in Figure 11, including:
  • Step 520 Predict the current block according to the intra prediction fusion method described in any embodiment of the present disclosure to obtain the prediction value of the current block;
  • the prediction value of the current block is obtained based on the final prediction result of the current block.
  • the video coding method of this embodiment uses the intra prediction fusion method of any embodiment of the present disclosure to predict the current block, and can achieve various effects of the intra prediction fusion method.
  • the prediction value of the current block is obtained based on the final prediction result of the current block.
  • the video coding method of this embodiment uses the intra prediction fusion method of any embodiment of the present disclosure to predict the current block, and can achieve various effects of the intra prediction fusion method.
  • the intra prediction method of the template-based multi-reference line intra prediction (TMRL) mode used in the above embodiment is to construct a candidate list based on the combination of the extended reference line and the intra prediction mode.
  • a single reference row is used for prediction.
  • the single reference row usually contains noise, which will affect the accuracy of prediction.
  • the TMRL mode of this embodiment can use the method of the previous embodiment when constructing the candidate list and encoding and decoding.
  • the difference is that the stage of generating the prediction value of the current block is to combine the prediction value generated by the selected angle mode and the selected reference line with The selected angle pattern is weighted with another reference row to produce a predicted value to obtain the final predicted value for the current block.
  • An embodiment of the present disclosure provides a video decoding method, which is applied to a decoder. As shown in Figure 13, the method includes:
  • Step 710 If the decoding determines that the current block uses the template-based multi-reference line intra prediction TMRL mode, continue to decode the TMRL mode index and TMRL fusion flag of the current block;
  • Step 720 Construct a candidate list of the TMRL mode of the current block, and determine the selected extended reference line and intra prediction mode of the current block according to the candidate list and the TMRL mode index;
  • the candidate list may be constructed in the same manner as in the previous embodiment, and the extended reference line and intra prediction mode selected in the current block may be determined.
  • Step 730 When the TMRL fusion flag indicates that intra-frame prediction is used to fuse IPF, the weighted sum of the first prediction result and the second prediction result is used as the final prediction result of the current block;
  • the weighted sum of the first prediction result and the second prediction result can be calculated according to the aforementioned formula 1 and formula 2.
  • the first prediction result is obtained by predicting the current block according to the extended reference line and the intra prediction mode
  • the second prediction result is obtained by predicting the current block according to another reference line and the intra prediction mode.
  • the intra prediction mode is angle mode. The weighting of the two prediction results can be calculated using the formula in the previous embodiment.
  • the current block is predicted according to the extended reference line and the intra prediction mode, and the final prediction result of the current block is obtained.
  • this embodiment sets the TMRL fusion flag to indicate whether to use IPF.
  • the encoding end can encode the TMRL fusion flag according to whether the same or similar restrictions in the previous embodiment are true. If it is determined to use IPF, the TMRL fusion flag will be used. The TMRL fusion flag is set to 1. When it is determined not to use IPF, the TMRL fusion flag is set to 0. On the decoding side, there is no need to combine these constraints for judgment. According to the TMRL fusion flag, it can be determined whether to perform IPF fusion on the angle mode selected when using the TMRL mode. It can simplify the processing on the decoding end and improve the flexibility of setting IPF usage conditions. Using IPF on the basis of TMRL mode can improve the accuracy of prediction and improve the performance of video encoding.
  • the other reference row is a reference row adjacent to the extended reference row and located inside the extended reference row, or is adjacent to the extended reference row and completely outside the extended reference row.
  • the inner side refers to the side close to the current block, and the outer side refers to the side far away from the current block.
  • the set conditions include any one or more of the following conditions:
  • the height of the current block is greater than the set value
  • the width multiplied by the height of the current block is greater than the set value
  • the current frame to which the current block belongs is not an inter frame; that is, intra blocks using fusion mode are prohibited from appearing in B frames and P frames.
  • this embodiment uses the size of the current block to be larger than the set value as a condition for decoding the TMRL fusion flag.
  • the TMRL fusion flag may not be encoded or decoded to improve encoding efficiency.
  • the candidate list for constructing the TMRL mode of the current block includes:
  • N extended reference lines and M intra prediction modes of the current block N ⁇ M combinations of extended reference lines and intra prediction modes are obtained, N ⁇ 1, M ⁇ 1, N ⁇ M ⁇ 2;
  • K combinations corresponding to the errors are filled in the candidate list of the template-based multi-reference line intra prediction TMRL mode of the current block, 1 ⁇ K ⁇ N ⁇ M;
  • the M intra prediction modes participating in the combination are limited to other angle modes except some specific angles, which is beneficial to the combination with IPF to improve the accuracy of prediction.
  • the sps_mrl_enabled_flag in the above table is a sequence-level identifier. A value of 1 indicates that the current sequence can use MRL, and a value of 1 indicates that MRL is not used. (y0%CtbSizeY)>0 means that the current CU position is not the first row of the CTU. If tmrl_fusion_flag is 1, it means that the subsequent TMRL mode will use the fusion mode, otherwise the fusion mode will not be used.
  • M, N, L in the table are the setting values of the size, all are positive integers, M can be equal to N.
  • one or more of the following intra prediction modes can be excluded: PLANAR, DC, Angle mode with an angle of horizontal angle, Angle mode with an angle of vertical angle, Angle mode with an angle of -45 degrees , the angle mode with an angle of 45 degrees, and the angle mode with an angle of 135 degrees.
  • Fusion in the prediction stage needs to be based on the value of tmrl_fusion_flag.
  • tmrl_fusion_flag 1
  • the selected intra prediction mode and the selected reference line generate the prediction signal p a
  • the selected intra prediction mode and the reference line reference line 0 generate the prediction signal p b
  • p a and p b are fused to generate the final Predictive signals. Otherwise, p a is directly used as the final prediction signal.
  • embodiments of the present disclosure also provide a video encoding method, applied to the encoder, as shown in Figure 18, the method includes:
  • Step 1110 Construct a candidate list of the template-based multi-reference line intra prediction TMRL mode of the current block.
  • the candidate list is filled with a combination of the extended reference line and intra prediction mode of the current block candidate;
  • Step 1130 When the encoding conditions of the TMRL mode of the current block are met, encode the TMRL mode flag of the current block to indicate that the current block uses the TMRL mode, and encode the TMRL mode index of the current block to indicate that the selected combination is in the candidate list. s position;
  • Step 1140 encode the TMRL fusion flag of the current block to indicate that the current block uses intra prediction fusion IPF or does not use IPF;
  • encoding the TMRL fusion flag of the current block to indicate that the current block uses intra prediction fusion IPF or does not use IPF includes: selecting a combination of angle modes in the candidate list for the current block. In this case, if at least one of the set constraints is true, encode the TMRL fusion flag of the current block to indicate that IPF is not used; if none of the set constraints are true, encode the TMRL fusion flag of the current block to indicate that IPF is not used. Indicates the use of IPF, where:
  • the TIMD fusion mode is selected for the current block
  • the current block selects DIMD mode
  • the index of the reference row selected in the current block is greater than or equal to K, and K is an integer greater than or equal to 3;
  • the height of the current block is less than or equal to the set value
  • the width multiplied by the height of the current block is less than or equal to the set value
  • the current block uses intra-frame sub-block division ISP mode
  • the current frame to which the current block belongs is an inter-frame
  • the current frame to which the current block belongs is the chroma frame.
  • the intra prediction modes participating in the combination exclude angle modes with angles of -45°, 0°, 45°, 90°, and 135°. .
  • the encoding end When the encoding end selects a combination in the candidate list of TMRL modes for the current block and uses IPF, it uses IPF to predict the current block based on the combination. The result of the combination and the angle mode in the combination are compared with another reference. The prediction results of the row are weighted to obtain the prediction value of the current block. In turn, the reconstruction value of the current block can be determined.
  • the fusion prediction mode based on the combination of the extended reference line and the angle mode can also be used as a combination in the candidate list when constructing the candidate list of TMRL.
  • An embodiment of the present disclosure provides a method for constructing a multi-reference line intra prediction mode candidate list, which can be applied to an encoder or a decoder. As shown in Figure 14, the method includes:
  • Step 810 Obtain N ⁇ M original combinations of extended reference lines and intra prediction modes based on the N extended reference lines and M intra prediction modes of the current block;
  • Step 820 Predict the template area of the current block according to the N ⁇ M original combinations, and calculate the error between the reconstructed value of the template area and the predicted value;
  • Step 830 For each combination of the K original combinations with the smallest error that includes a predetermined angle pattern, determine whether fusion is required. If fusion is required, fill in the fusion combination corresponding to the original combination into the template-based multi-reference row of the current block. For the candidate list of the intra-frame prediction TMRL mode, if fusion is not required, the original combination is filled in the candidate list, where K, N, and M are set positive integers, 1 ⁇ K ⁇ N ⁇ M.
  • each of the K original combinations with the smallest error including a predetermined angle pattern is determined whether fusion is required. If fusion is required, the fusion combination corresponding to the original combination is filled in the candidate list of the TMRL pattern of the current block.
  • the weighted sum of the first prediction result and the second prediction result is used as the prediction value of the current block; wherein the first prediction result is based on the extended reference in the original combination.
  • the second prediction result is the result of predicting the current block according to the second reference row and the angle mode in the original combination, and the second reference row is the result of the prediction of the current block.
  • the predetermined angle pattern includes all angle patterns, or includes other angle patterns except angle patterns with integer slopes.
  • the weighted sum of the first prediction result and the second prediction result in this embodiment can be calculated according to the above-mentioned formula 1 or formula 2.
  • K combinations are selected as candidate lists by sorting the non-fusion modes. Then, it is further determined whether the reference rows of each of the K combinations need to be fused. The determination method is based on the error size between the predicted signal and the reconstructed signal generated in fusion or non-fusion mode based on the reference row selected by the current combination. If it is small in fusion mode, then fusion is used for this combination, otherwise no fusion is performed.
  • This embodiment only requires prediction and error value calculation of 5xM+K combinations on the template area.
  • Another embodiment of the present disclosure provides a method for constructing a multi-reference line intra prediction mode candidate list, which can be applied to an encoder or a decoder. As shown in Figure 15, the method includes:
  • Step 910 Obtain N ⁇ M original combinations of extended reference lines and intra prediction modes based on the N extended reference lines and M intra prediction modes of the current block;
  • Step 920 Perform fusion processing on the N ⁇ M original combinations.
  • the fusion processing includes: for each original combination including a predetermined angle pattern, when the set conditions are met, replace the original combination with the corresponding fusion combination;
  • Step 930 Predict the template area of the current block according to the N ⁇ M combinations after the fusion process, and calculate the error between the reconstructed value of the template area and the predicted value;
  • Step 940 Fill in the candidate list of the template-based multi-reference line intra prediction TMRL mode of the current block with the K combinations corresponding to the errors in the order of the errors from small to large;
  • K, N, M are set positive integers, 1 ⁇ K ⁇ N ⁇ M;
  • the weighted sum of the first prediction result and the second prediction result is used as the prediction value of the current block; wherein the first prediction result is based on the original combination.
  • the second prediction result is the prediction result of the current block based on the second reference row and the angle mode in the original combination.
  • the second reference row is the prediction result of the current block. It is the adjacent row of the first reference row or the reference row with index 0.
  • the setting conditions include any one or more of the following conditions: the size of the current block is greater than NxM, and N and M are set positive integers; and, within the frame selected by the current block
  • the prediction mode is not an angle mode with an integer slope; the predetermined angle mode includes all angle modes, or includes other angle modes except the angle mode with an integer slope.
  • the weighted sum of the first prediction result and the second prediction result in this embodiment can be calculated according to the above-mentioned formula 1 or formula 2.
  • the fusion mode of adjacent reference rows is used to participate in template sorting and generate a combined ⁇ reference row, prediction mode ⁇ list; otherwise, the fusion mode is not used:
  • the selected reference row and prediction mode are determined based on the tmrl_idx obtained by decoding and the candidate list generated based on template sorting, and based on the same conditions, it is determined whether to use adjacent reference rows to generate fused prediction results. That is, when one or more of the following conditions are met, IPF fusion is performed on the selected angle mode:
  • Yet another embodiment of the present disclosure provides a method for constructing a multi-reference line intra prediction mode candidate list, which can be applied to an encoder or a decoder. As shown in Figure 16, the method includes:
  • Step 1010 Obtain N ⁇ M original combinations of extended reference lines and angle modes based on the N extended reference lines and M intra prediction modes of the current block;
  • Step 1020 Obtain the corresponding fusion combination according to the combinations including the predetermined angle pattern among the N ⁇ M original combinations;
  • Step 1030 Predict the template area of the current block according to the N ⁇ M original combinations and the obtained fusion combinations, and calculate the error between the reconstructed value of the template area and the predicted predicted value;
  • Step 1040 Fill in the K combinations corresponding to the errors into the candidate list of the template-based multi-reference line intra prediction TMRL mode of the current block in order of the errors from small to large, where K, N, M are A definite positive integer, 1 ⁇ K ⁇ N ⁇ M.
  • This embodiment obtains corresponding fusion combinations based on the combinations including predetermined angle patterns among the N ⁇ M original combinations, including: obtaining a fusion combination according to each original combination; when predicting the current block based on the fusion combination, The weighted sum of the first prediction result and the second prediction result is used as the prediction value of the current block; wherein the first prediction result is the result of predicting the current block according to the extended reference line and angle mode in the original combination, so The second prediction result is the result of predicting the current block based on the second reference line and the angle pattern in the original combination.
  • the second reference line is an adjacent line of the first reference line or is an index of 0. reference line;
  • the predetermined angle pattern includes all angle patterns, or includes other angle patterns except angle patterns with integer slopes.
  • the weighted sum of the first prediction result and the second prediction result in this embodiment can be calculated according to the above-mentioned formula 1 or formula 2.
  • the fusion combination corresponding to the original combination may be, for example, the original combination is a reference row with an index of 1
  • the corresponding fusion combination can be a combination of the two reference rows with indexes 1 and 2 and the angle mode.
  • the reference row with index 1 will be used.
  • the prediction result of the row and the angle mode, and the prediction result of the reference row with index 2 and the angle mode are weighted to obtain the final prediction result of the current block.
  • the corresponding fusion combination may be a combination of two reference rows with indexes 0 and 1 and the angle pattern.
  • the current block will be fusion predicted according to the prediction method of the fusion combination. If If the original combination in the candidate list is selected, the current block will be predicted according to the prediction method of the original combination without using IPF.
  • An embodiment of the present disclosure also provides a code stream, wherein the code stream is generated by the video encoding method described in any embodiment of the present disclosure.
  • An embodiment of the present disclosure also provides a device for constructing a multi-reference line intra prediction mode candidate list, as shown in Figure 17, including a processor 71 and a memory 73 storing a computer program, wherein the processor 71 executes
  • the computer program can implement the method for constructing a multi-reference line intra prediction mode candidate list described in any embodiment of this article.
  • An embodiment of the present disclosure also provides an intra prediction fusion device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the method described in any embodiment of the present disclosure. Intra-frame prediction fusion method.
  • An embodiment of the present disclosure also provides a video decoding device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the video decoding described in any embodiment of the present disclosure. method.
  • An embodiment of the present disclosure also provides a video encoding device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the video encoding described in any embodiment of the present disclosure. method.
  • An embodiment of the present disclosure also provides a video encoding and decoding system, which includes the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.
  • An embodiment of the present disclosure also provides a non-transitory computer-readable storage medium.
  • the computer-readable storage medium stores a computer program, wherein the computer program implements any embodiment of the present disclosure when executed by a processor.
  • the intra prediction fusion method, the construction method of the multi-reference line intra prediction mode candidate list according to any embodiment of the present disclosure, or the video decoding method according to any embodiment of the present disclosure, or the implementation of the present disclosure The video encoding method according to any embodiment.
  • the processor in the above embodiments of the present disclosure may be a general-purpose processor, including a central processing unit (CPU), a network processor (Network Processor, NP for short), a microprocessor, etc., or it may be other conventional processors, etc.;
  • the processor may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), a discrete logic or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, or other equivalent integrated or discrete logic circuits, or a combination of the above devices.
  • DSP digital signal processor
  • ASIC application specific integrated circuit
  • FPGA off-the-shelf programmable gate array
  • the processor in the above embodiments can be any processing device or device combination that implements the methods, steps and logical block diagrams disclosed in the embodiments of the present invention. If embodiments of the present disclosure are implemented in part in software, instructions for the software may be stored in a suitable non-volatile computer-readable storage medium and may be executed in hardware using one or more processors. Instructions are provided to perform the methods of embodiments of the present disclosure.
  • the term "processor” as used herein may refer to the structure described above or any other structure suitable for implementing the techniques described herein.
  • Computer-readable media may include computer-readable storage media that corresponds to tangible media, such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another according to a communications protocol.
  • Computer-readable media generally may correspond to non-transitory, tangible computer-readable storage media or communication media such as a signal or carrier wave.
  • Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and/or data structures for implementing the techniques described in this disclosure.
  • a computer program product may include computer-readable media.
  • Such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory or may be used to store instructions or data. Any other medium that stores the desired program code in the form of a structure and that can be accessed by a computer.
  • any connection is also termed a computer-readable medium if, for example, a connection is sent from a website, server, or using any of the following: coaxial cable, fiber-optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio, and microwave or other remote source transmits instructions, then coaxial cable, fiber optic cable, twin-wire, DSL or wireless technologies such as infrared, radio and microwave are included in the definition of medium.
  • coaxial cable, fiber optic cable, twin-wire, DSL or wireless technologies such as infrared, radio and microwave are included in the definition of medium.
  • disks and optical discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, or Blu-ray discs. Disks usually reproduce data magnetically, while optical discs use lasers to reproduce data. Regenerate data optically. Combinations of the above should also be included within the scope of computer-readable media.
  • the functionality described herein may be provided within dedicated hardware and/or software modules configured for encoding and decoding, or incorporated in a combined codec. Furthermore, the techniques may be implemented entirely in one or more circuits or logic elements.
  • inventions of the present disclosure may be implemented in a wide variety of devices or equipment, including wireless handsets, integrated circuits (ICs), or a set of ICs (eg, chipsets).
  • ICs integrated circuits
  • a set of ICs eg, chipsets.
  • Various components, modules or units are depicted in embodiments of the present disclosure to emphasize functional aspects of devices configured to perform the described techniques, but do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperating hardware units (including one or more processors as described above) in conjunction with suitable software and/or firmware.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

一种帧内预测融合方法、视频编解码方法、装置和系统,在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用帧内预测融合IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF。通过限制IPF的使用,可以避免预测时融合过多的模式,以及提高编码的性能。本实施例还提供了使用该帧内预测融合方法的视频编解码方法、相应的装置和系统。

Description

一种帧内预测融合方法、视频编解码方法、装置和系统
交叉引用
本申请要求在2022年7月8日提交中国专利局、申请号为202210806435.X、名称为“一种候选列表构建方法、视频编解码方法、装置和系统”的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本公开实施例涉及但不限于视频技术,更具体地,涉及一种帧内预测融合方法、视频编解码方法、装置和系统。
背景技术
数字视频压缩技术主要是将庞大的数字影像视频数据进行压缩,以便于传输以及存储等。目前通用的视频编解码标准,如H.266/Versatile Video Coding(多功能视频编码,VVC),都采用基于块的混合编码框架。视频中的每一帧被分割成相同大小(如128x128,64x64等)的正方形的最大编码单元(LCU:largest coding unit)。每个最大编码单元可根据规则划分成矩形的编码单元(CU:coding unit)。编码单元可能还会划分预测单元(PU:prediction unit),变换单元(TU:transform unit)等。混合编码框架包括预测(prediction)、变换(transform)、量化(quantization)、熵编码(entropy coding)、环路滤波(in loop filter)等模块。预测模块包括用于减少或取出视频内在冗余的帧内预测(intra prediction)和帧间预测(inter prediction)。帧内块通过块周边像素作为参考进行预测,帧间块则参考空间上的邻近块信息和其他帧里的参考信息。与预测信号相对,残差信息通过以块为单位的变换、量化和熵编码成码流。这些技术被描述在标准里并实施在的各种与视频压缩相关的领域。
随着互联网视频的激增以及人们对视频清晰度的要求越来越高,尽管已有的数字视频压缩标准能够节省不少视频数据,但目前仍然需要追求更好的数字视频压缩技术,以减少数字视频传输的带宽和流量压力。
发明概述
以下是对本文详细描述的主题的概述。本概述并非是为了限制权利要求的保护范围。
本公开一实施例提供了一种帧内预测融合方法,包括:
在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用帧内预测融合IPF的限制条件是否成立;
在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF。
本公开一实施例还提供了一种视频解码方法,包括:
解码码流,确定当前块选中的参考行和帧内预测模式;
按照如本公开任一实施例述的帧内预测融合方法对当前块进行预测,得到当前块的预测值;
根据当前块的预测值确定当前块的重建值。
本公开一实施例还提供了一种视频编码方法,包括:
通过模式选择确定当前块选中的参考行和帧内预测模式;
按照如本公开任一实施例所述的帧内预测融合方法对当前块进行预测,得到当前块的预测值;
根据当前块的原始值和预测值确定当前块的残差。
本公开一实施例还提供了一种视频编码方法,包括:
解码确定当前块使用基于模板的多参考行帧内预测TMRL模式的情况下,继续解码当前块的TMRL模式索引和TMRL融合标志;
构建当前块的TMRL模式的候选列表,根据所述候选列表和TMRL模式索引确定当前块选中的扩展参考行和帧内预测模式;
在所述TMRL融合标志表示使用帧内预测融合IPF的情况下,将第一预测结果和第二预测结果的 加权和作为当前块最终的预测结果;
其中,所述第一预测结果是根据该扩展参考行和该帧内预测模式对当前块进行预测得到的,所述第二预测结果是根据另一参考行和该帧内预测模式对当前块进行预测得到的,该帧内预测模式为角度模式。
本公开一实施例还提供了一种视频编码方法,包括:
构建当前块基于模板的多参考行帧内预测TMRL模式的候选列表,所述候选列表中填入有当前块候选的扩展参考行与帧内预测模式的组合;通过率失真优化,为当前块选中一种参考行和帧内预测模式的组合;
在当前块的TMRL模式的编码条件满足时,编码当前块的TMRL模式标志以表示当前块使用TMRL模式,编码当前块的TMRL模式索引以表示所述选中的组合在所述候选列表中的位置;
编码当前块的TMRL融合标志以表示当前块使用帧内预测融合IPF或者不使用IPF;
其中,所述编码条件至少包括:所述选中的组合在所述候选列表中。
本公开一实施例还提供了一种多参考行帧内预测模式候选列表的构建方法,包括:
根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与角度模式的N×M种原始组合;
根据所述N×M种原始组合中包括预定角度模式的组合得到对应的融合组合;
根据所述N×M种原始组合和得到的融合组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,其中,K,N,M为设定的正整数,1≤K≤N×M。
本公开一实施例还提供了一种多参考行帧内预测模式候选列表的构建方法,包括:
根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种原始组合;
根据所述N×M种原始组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
对误差最小的K种原始组合中每一种包括预定角度模式的组合,分别确定是否需要融合,如需要融合,将该原始组合对应的融合组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,如不需要融合,将该原始组合填入所述候选列表,其中,K,N,M为设定的正整数,1≤K≤N×M。
本公开一实施例还提供了一种多参考行帧内预测模式候选列表的构建方法,包括:
根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种原始组合;
对所述N×M种原始组合进行融合处理,所述融合处理包括:对每一种包括预定角度模式的原始组合,在满足设定条件时,将该原始组合替换为对应的融合组合;
根据融合处理后的N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表;
其中,K,N,M为设定的正整数,1≤K≤N×M;
其中,基于该原始组合对应的融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
本公开一实施例还提供了一种码流,其中,所述码流通过本公开任一实施例所述的视频编码方法生成。
本公开一实施例还提供了一种帧内预测融合装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现如本公开任一实施例所述的帧内预测融合方法。
本公开一实施例还提供了一种多参考行帧内预测模式候选列表的构建装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现本文任一实施例所述的多参考行帧内预测模式候选列表的构建方法。
本公开一实施例还提供了一种视频解码装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现本公开任一实施例所述的视频解码方法。
本公开一实施例还提供了一种视频编码装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现本公开任一实施例所述的视频编码方法。
本公开一实施例还提供了一种视频编解码系统,其中,包括本公开任一实施例所述的视频编码装置和本公开任一实施例所述的视频解码装置。
本公开一实施例还提供了一种非瞬态计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,其中,所述计算机程序时被处理器执行时实现本公开任一实施例所述的帧内预测融合方法,本公开任一实施例所述的多参考行帧内预测模式候选列表的构建方法,或实现本公开任一实施例所述的视频解码方法,或实现本公开任一实施例所述的视频编码方法。
在阅读并理解了附图和详细描述后,可以明白其他方面。
附图概述
附图用来提供对本公开实施例的理解,并且构成说明书的一部分,与本公开实施例一起用于解释本公开的技术方案,并不构成对本公开技术方案的限制。
图1A是本公开一实施例编解码系统的示意图;
图1B是本公开一实施例编码端的框架图;
图1C是本公开一实施例解码端的框架图;
图2是本公开一实施例帧内预测模式的示意图;
图3是本公开一实施例当前块的相邻帧内预测块的示意图;
图4是本公开一实施例当前块的模板及模板参考区域的示意图;
图5是本公开一实施例当前块周围的多个参考行的示意图;
图6是本公开一实施例视频编码方法的流程图;
图7是本公开一实施例TMRL模式候选列表构建方法的流程图
图8A是本公开一实施例当前块周围模板区域及扩展参考行的示意图;
图8B是本公开另一实施例当前块周围模板区域及扩展参考行的示意图;
图9是本公开一实施例视频解码方法的流程图
图10是本公开一实施例帧内预测融合方法的流程图;
图11是本公开另一实施例视频编码方法的流程图;
图12是本公开另一实施例视频解码方法的流程图;
图13是本公开另一实施例视频解码方法的流程图;
图14、图15和图16分别是本公开三个实施例中TMRL候选列表构建方法和流程图;
图17是本公开一实施例TMRL模式候选列表的构建装置的示意图;
图18是本公开另一实施例视频编码方法的流程图。
详述
本公开描述了多个实施例,但是该描述是示例性的,而不是限制性的,并且对于本邻域的普通技术人员来说显而易见的是,在本公开所描述的实施例包含的范围内可以有更多的实施例和实现方案。
本公开的描述中,“示例性的”或者“例如”等词用于表示作例子、例证或说明。本公开中被描述为“示例性的”或者“例如”的任何实施例不应被解释为比其他实施例更优选或更具优势。本文中的“和/或”是对关联对象的关联关系的一种描述,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。“多个”是指两个或多于两个。另外,为了便于清楚描述本公开实施例的技术方案,采用了“第一”、“第二”等字样对功能和作用基本相同的相同项或相似项进行区分。本邻域技术人员可以理解“第一”、“第二”等字样并不对数量和执行次序进行限定,并且“第一”、“第二”等字样也并不限定一定不同。
在描述具有代表性的示例性实施例时,说明书可能已经将方法和/或过程呈现为特定的步骤序列。然而,在该方法或过程不依赖于本文所述步骤的特定顺序的程度上,该方法或过程不应限于所述的特定顺序的步骤。如本邻域普通技术人员将理解的,其它的步骤顺序也是可能的。因此,说明书中阐述的步骤的特定顺序不应被解释为对权利要求的限制。此外,针对该方法和/或过程的权利要求不应限于按照所写顺序执行它们的步骤,本邻域技术人员可以容易地理解,这些顺序可以变化,并且仍然保持在本公开实施例的精神和范围内。
本公开实施例提出的局部光照补偿方法、视频编解码方法可以应用于各种视频编解码标准,例如:H.264/Advanced Video Coding(高级视频编码,AVC),H.265/High Efficiency Video Coding(高效视频编码,HEVC),H.266/Versatile Video Coding(多功能视频编码,VVC),AVS(Audio Video coding Standard,音视频编码标准),以及MPEG(Moving Picture Experts Group,动态图像专家组)、AOM(开放媒体联盟,Alliance for Open Media)、JVET(联合视频专家组,Joint Video Experts Team)制订的其他标准以及这些标准的拓展,或任何自定义的其他标准等。
图1A是可用于本公开实施例的一种视频编解码系统的框图。如图所示,该系统分为编码端装置1和解码端装置2,编码端装置1产生码流。解码端装置2可对码流进行解码。解码端装置2可经由链路3从编码端装置1接收码流。链路3包括能够将码流从编码端装置1移动到解码端装置2的一个或多个媒体或装置。在一个示例中,链路3包括使得编码端装置1能够将码流直接发送到解码端装置2的一个或多个通信媒体。编码端装置1根据通信标准(例如无线通信协议)来调制码流,将经调制的码流发送到解码端装置2。所述一个或多个通信媒体可包含无线和/或有线通信媒体,可形成分组网络的一部分。在另一示例中,也可将码流从输出接口15输出到一个存储装置,解码端装置2可经由流式传输或下载从该存储装置读取所存储的数据。
如图所示,码端装置1包含数据源11、视频编码装置13和输出接口15。数据源11包括视频捕获装置(例如摄像机)、含有先前捕获的数据的存档、用以从内容提供者接收数据的馈入接口,用于产生数据的计算机图形系统,或这些来源的组合。视频编码装置13对来自数据源11的数据进行编码后输出到输出接口15,输出接口15可包含调节器、调制解调器和发射器中的至少之一。解码端装置2包含输入接口21、视频解码装置23和显示装置25。输入接口21包含接收器和调制解调器中的至少之一。输入接口21可经由链路3或从存储装置接收码流。视频解码装置23对接收的码流进行解码。显示装置25用于显示解码后的数据,显示装置25可以与解码端装置2的其他装置集成在一起或者单独设置,显示装置25对于解码端来说是可选的。在其他示例中,解码端可以包含应用解码后数据的其他装置或设备。
基于图1A所示的视频编解码系统,可以使用各种视频编解码方法来实现视频压缩和解压缩。
图1B为可用于本公开实施例的一示例性的视频编码装置的框图。如图所示,该视频编码装置1000包含预测单元1100、划分单元1101、残差产生单元1102(图中用划分单元1101后的带加号的圆圈表示)、变换处理单元1104、量化单元1106、反量化单元1108、反变换处理单元1110、重建单元1112(图中用反变换处理单元1110后的带加号的圆圈表示)、滤波器单元1113、解码图像缓冲器1114以及熵编码单元1115。其中,预测单元1100包含帧间预测单元1121和帧内预测单元1126,解码图像缓冲器1114也可以称为已解码图像缓冲器、解码图片缓冲器、已解码图片缓冲器等。视频编码器20也可以包含比该示例更多、更少或不同功能组件,如在某些情况下可以取消变换处理单元1104、反变换处理单元1110等。
划分单元1101与预测单元1100配合将接收的视频数据划分为切片(Slice)、编码树单元(CTU:Coding Tree Unit)或其它较大的单元。划分单元1101接收的视频数据可以是包括I帧、P帧或B帧等视频帧的视频序列。
预测单元1100可以将CTU划分为编码单元(CU:Coding Unit),对CU执行帧内预测编码或帧间预测编码。对CU做帧内预测和帧间预测时,可以将CU划分为一个或多个预测单元(PU:prediction unit)。
帧间预测单元1121可对PU执行帧间预测,产生PU的预测数据,所述预测数据包括PU的预测块、PU的运动信息和各种语法元素。帧间预测单元1121可以包括运动估计(ME:motion estimation)单元和运动补偿(MC:motion compensation)单元。运动估计单元可以用于运动估计以产生运动矢量,运动补偿单元可以用于根据运动矢量获得或生成预测块。
帧内预测单元1126可对PU执行帧内预测,产生PU的预测数据。PU的预测数据可包含PU的预测块和各种语法元素。
残差产生单元1102可基于CU的原始块减去CU划分成的PU的预测块,产生CU的残差块。
变换处理单元1104可将CU划分为一个或多个变换单元(TU:Transform Unit),预测单元和变换单元的划分可以不同。TU关联的残差块是CU的残差块划分得到的子块。通过将一种或多种变换应用于TU关联的残差块来产生TU关联的系数块。
量化单元1106可基于选定的量化参数对系数块中的系数进行量化,通过调整量化参数(QP:Quantizer Parameter)可以调整对系数块的量化程度。
反量化单元1108和反变换单元1110可分别将反量化和反变换应用于系数块,得到TU关联的重建残差块。
重建单元1112可将所述重建残差块和预测单元1100产生的预测块相加,产生重建图像。
滤波器单元1113对重建图像执行环路滤波,将滤波后的重建图像存储在解码图像缓冲器1114中作为参考图像。帧内预测单元1126可以从解码图像缓冲器1114中提取PU邻近的块的参考图像以执行帧内预测。帧间预测单元1121可使用解码图像缓冲器1114缓存的上一帧的参考图像对当前帧图像的PU执行帧间预测。
熵编码单元1115可以对接收的数据(如语法元素、量化后的系数块、运动信息等)执行熵编码操作。
图1C为可用于本公开实施例的一示例性的视频解码装置的框图。如图所示,视频解码装置101包含熵解码单元150、预测单元152、反量化单元154、反变换处理单元156、重建单元158(图中用反变换处理单元155后的带加号的圆圈表示)、滤波器单元159,以及解码图像缓冲器160。在其它实施例中,视频解码器30可以包含更多、更少或不同的功能组件,如在某些情况下可以取消反变换处理单元155等。
熵解码单元150可对接收的码流进行熵解码,提取语法元素、量化后的系数块和PU的运动信息等。预测单元152、反量化单元154、反变换处理单元156、重建单元158以及滤波器单元159均可基于从码流提取的语法元素来执行相应的操作。
反量化单元154可对量化后的TU关联的系数块进行反量化。
反变换处理单元156可将一种或多种反变换应用于反量化后的系数块以便产生TU的重建残差块。
预测单元152包含帧间预测单元162和帧内预测单元164。如果PU使用帧内预测编码,帧内预测单元164可基于从码流解码出的语法元素确定PU的帧内预测模式,根据确定的帧内预测模式和从解码图像缓冲器160获取的PU邻近的已重建参考信息执行帧内预测,产生PU的预测块。如果PU使用帧间预测编码,帧间预测单元162可基于PU的运动信息和相应的语法元素来确定PU的一个或多个参考块,基于从解码图像缓冲器160获取的所述参考块来产生PU的预测块。
重建单元158可基于TU关联的重建残差块和预测单元152产生的PU的预测块,得到重建图像。
滤波器单元159可对重建图像执行环路滤波,滤波后的重建图像存储在解码图像缓冲器160中。解码图像缓冲器160可提供参考图像以用于后续运动补偿、帧内预测、帧间预测等,也可将滤波后的重建图像作为已解码视频数据输出,在显示装置上的呈现。
基于上述视频编码装置和视频解码装置,可以执行以下基本的编解码流程,在编码端,将一帧图像划分成块,对当前块进行帧内预测或帧间预测或其他算法产生当前块的预测块,使用当前块的原始 块减去预测块得到残差块,对残差块进行变换和量化得到量化系数,对量化系数进行熵编码生成码流。在解码端,对当前块进行帧内预测或帧间预测产生当前块的预测块,另一方面对解码码流得到的量化系数进行反量化、反变换得到残差块,将预测块和残差块相加得到重建块,重建块组成重建图像,基于图像或基于块对重建图像进行环路滤波得到解码图像。编码端同样通过和解码端类似的操作以获得解码图像,编码端获得的解码图像通常也叫做重建图像。解码图像可以作为对后续帧进行帧间预测的参考帧。编码端确定的块划分信息,预测、变换、量化、熵编码、环路滤波等模式信息和参数信息如果需要可以写入码流。解码端通过解码码流或根据已有信息进行分析,确定与编码端相同的块划分信息,预测、变换、量化、熵编码、环路滤波等模式信息和参数信息,从而保证编码端获得的解码图像和解码端获得的解码图像相同。
以上虽然是以基于块的混合编码框架为示例,但本公开实施例并不局限于此,随着技术的发展,该框架中的一个或多个模块,及该流程中的一个或多个步骤可以被替换或优化。
本文中,当前块(current block)可以是当前图像中的当前编码单元(current CU)、当前预测单元(current PU)等块级的编解码单位。
编码端在帧内预测时,通常借助各种角度模式与非角度模式对当前块进行预测得到预测块;根据预测块与原始块计算得到的率失真信息,为当前块选中最优的帧内预测模式,将该帧内预测模式编码经码流传输到解码端。解码端通过解码得到为当前选中的该帧内预测模式,按照该帧内预测模式对当前块进行帧内预测。本文中,为当前块选中的参考行和帧内预测模式也表述为当前块选中的参考行和帧内预测模式。
在VVC和ECM中,多种的传统帧内预测模式使用当前块周围已经重建的信息预测当前块。包括:平面模式(即Planar模式,模式索引为0)、均值模式(即DC模式、模式索引为1)和65个角度预测模式(模式索引2~66)。图2示出了模式索引为2~66的角度模式的角度方向。文中,也将角度预测模式简称为角度模式。
由于在VVC中引入了长方形的预测块,对于长方形的块,索引为2~66的角度模式中一些角度模式的角度方向可以被替换成了更宽的角度方向,如图所示,索引为-14~-1,67~80的角度模式是因宽角度替换而得到的角度模式,这些角度模式的选中不需要通过标识位去表示,而是通过当前块的形状和当前块选中的预测模式索引(2~66)通过对应关系得到。
在VVC的标准文本中,宽角度的替换方式如下:
变量whRatio为Abs(Log2(blockWidth)-Log2(blockHeight)).
对于非正方形的预测块,角度预测模式predModeIntra(2~66)按照以下条件是否满足进行宽角度替换,predModeIntra用索引表示:
如果以下三个条件都满足,predModeIntra将等于(predModeIntra+65).
当前块宽度大于高度
角度模式索引大于等于2
角度模式索引小于(whRatio>1)?(8+2*whRatio):8
否则,如果以下三个条件都满足,predModeIntra将等于(predModeIntra-67).
当前块的高度大于宽度
角度模式索引小于等于66
角度模式索引大于(whRatio>1)?(60-2*whRatio):60
完成了宽角度模式的匹配后,每个角度预测模式predModeIntra(-14~80)都会有一个角度值intraPredAngle。
表1.角度预测模式与角度值的对应关系
predModeIntra -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 2 3 4
intraPredAngle 512 341 256 171 128 102 86 73 64 57 51 45 39 35 32 29 26
predModeIntra 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
intraPredAngle 23 20 18 16 14 12 10 8 6 4 3 2 1 0 -1 -2 -3
predModeIntra 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38
intraPredAngle -4 -6 -8 -10 -12 -14 -16 -18 -20 -23 -26 -29 -32 -29 -26 -23 -20
predModeIntra 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55
intraPredAngle -18 -16 -14 -12 -10 -8 -6 -4 -3 -2 -1 0 1 2 3 4 6
predModeIntra 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72
intraPredAngle 8 10 12 14 16 18 20 23 26 29 32 35 39 45 51 57 64
predModeIntra 73 74 75 76 77 78 79 80                  
intraPredAngle 73 86 102 128 171 256 341 512                  
角度模式的角度值将用于后续的角度预测。
本文中,每个角度模式predModeIntra均对应一个角度,每个角度预测模式的角度是图1中该角度预测模式对应的线段在直角坐标系中的角度,例如,索引号为34的角度模式,角度值intraPredAngle为-32,而角度为-45°或45°或135°,这与该直角坐标系统定义的0°方向有关。
本文中所说的帧内预测模式在没有其他限定时,是指包括Planar模式、DC模式和角度模式的传统帧内预测模式。
如果直接对当前块的帧内预测模式进行编码,67种模式需要7bit来编码,数据量很大。根据统计特性,距离当前块越近的像素区域往往越容易与当前块选中同样的帧内预测模式,根据这一特性,在HEVC、VVC以及增强的压缩模型(ECM:Enhanced Compression Model)中都采纳了最可能模式(MPM:most probable mode)技术。ECM是基于VTM-10.0参考软件,集成各类新工具,从而进一步挖掘编解码性能的参考软件。
MPM先构建MPM列表,在MPM列表中填充最有可能被当前块选中的6种帧内预测模式。如果当前块选中的帧内预测模式在MPM列表中,只需要编码其索引号(只需要3bit)即可,如果当前块选中的帧内预测模式不在MPM列表中而是在61个非MPM(non-MPM)模式中,则在熵编码阶段使用截断二元码(Truncated Binary Code,TBC)编码该帧内预测模式。
在VVC内,无论是否应用多参考行(MRL:Multiple reference line)和帧内子块划分(ISP:Intra Sub-Partitions),MPM列表均有6种预测模式。而ECM中的MPM分为MPM和第二MPM(Secondary MPM),MPM和Secondary MPM分别使用长度为6与长度为16的列表。MPM列表的6个模式中,Planar模式始终填充在MPM中的第一个位置,其余5个位置的填充由下面三步依次进行,直到填充满5个位置为止,多出来的模式将自动进入Secondary MPM。
第一步,依次填充当前块周围邻近的5个位置上的预测块使用的帧内预测模式;如图3所示,该5个位置依次包括当前块左上(AL),上(A),右上(AR),左(L)及左下(BL)位置。
第二步,基于当前块周围重建像素使用梯度直方图导出的模式;
第三步,与第一步选中的角度模式的角度相近的角度模式。
Secondary MPM列表可以由除MPM中的帧内预测模式以外的一些主要角度模式构成。
由于MPM标志(mpm_flag)的编解码顺序在MRL模式之后,ECM中MPM的编解码需依赖于MRL标识位,在当前块不使用MRL模式时,需解码MPM标志以确定当前块是否使用MPM,而在当前块使用MRL模式时,无需解码MPM标志,默认当前块使用MPM。
基于模板的帧内模式推导(TIMD:Template based intra mode derivation)和解码端帧内预测模式导出(DIMD:Decoder-side intra mode derivation)是VVC标准中没有,但采纳进ECM参考软件的两项帧内预测技术,这两项技术可在解码端根据当前块周边已重建的像素值导出当前块的帧内预测模式,从而省去编码帧内预测模式的索引从而达到节省bit的作用。
TIMD是一种针对亮度帧的帧内预测模式,TIMD模式由MPM列表中候选的帧内预测模式和模板(Template)区域(简称模板)生成。在ECM中,如图4所示,当前块(如当前CU)11的左侧相邻区域和上方相邻区域构成当前块的模板区域12。其中的左侧相邻区域称为左侧模板区域(简称左模板),上方相邻区域称为上方模板区域或上侧模板区域(简称上模板)。
如图所示,模板区域12外侧(指左侧和上侧)设置有模板参考(reference of the template)区域13,各区域示例性的尺寸和位置如图所示。在一示例中,左模板的宽度L1和上模板的高度L2均为4。模板参考区域13可以是模板区域上方相邻的一行或左侧相邻的一列。
TIMD假定当前块和当前块的模板区域的分布特性一致,以模板参考区域的重建值作为参考行的重建值,遍历MPM和Secondary MPM中的所有帧内预测模式对模板区域进行预测,得到预测结果。再计算模板区域上的重建值与每种模式的预测结果之间的误差,用误差变换绝对值和(SATD:Sum of absolute transformed differences)表示,选出SATD最小即最优的帧内预测模式,将该帧内预测模式作为当前块的TIMD模式。解码端可以通过相同的推导方式推导TIMD模式。如果序列允许使用TIMD,每个当前块需要一个标志位表示是否使用TIMD。如果当前块选中的帧内预测模式为TIMD模式,则当前块使用TIMD模式进行预测,且剩余的与帧内预测相关的语法元素如ISP、MPM等的解码过程可以跳过,从而大幅降低模式的编码比特。
在求得模板区域的重建值与每种模式预测结果(模板区域的预测值)之间的SATD之后,可根据以下方式确定最终使用的TIMD模式:
假设mode1和mode2为MPM中用于帧内预测的两种角度模式,mode1为SATD最小的角度模式,其SATD为cost1;mode2为SATD次小的角度模式,其SATD为cost2:
当cost1×2≤cost2时,将mode1作为当前块的TIMD模式;
当cost1×2>cost2时,将对mode1和mode2的预测结果进行加权的预测模式作为当前块的TMID模式,也称TIMD融合(fusion)模式。
加权的方法和权重如下式所示:
Pred=Pred mode1×w1+Pred mode2×w2
Figure PCTCN2022106337-appb-000001
其中,Pred为当前块使用TIMD融合模式的预测结果,Pred mode1为当前块使用mode1的预测结果,Pred mode2为当前块使用mode2的预测结果,w1和w2为根据cost1和cost2计算得到的权重。
DIMD是以当前块周边已重建的像素值为模板,通过索贝尔(sobel)算子在模板上的每个3x3区域上扫描并计算水平方向和竖直方向的梯度,根据水平和竖直方向上求得梯度Dx和Dy,根据Dx和Dy求得每个位置上的幅度值Amp=abs(Dx)+abs(Dy),和角度值angular=arctan(Dy/Dx)。根据模板上每个位置的角度值对应到传统的角度模式,累加相同角度模式的幅度值得到幅度值与角度模式的直方图。在存在幅度值最高和次高的两种角度模式的情况下,对幅度值最高和次高的两种角度模式以及planar模式的预测值进行加权,可得到当前块使用DIMD时最终的预测结果。此时的预测模式融合了三种帧内预测模式:planar模式和幅度值最高和次高的两种角度模式。文中称为DIMD融合模式。在 不存在幅度最高和次高的角度模式的情况下,使用DIMD的预测与planar模式预测等同。
在HEVC中,帧内预测使用距离当前块最近的上一行和左一列作为参考进行预测,如果这一行和一列的重建值与原始像素值有着较大的误差,那么当前块的预测质量也会有很大的影响。为了解决这一问题,VVC中采纳了多参考行(MRL:Multiple reference line)帧内预测技术,除了可以使用索引为0的参考行(Reference line0)之外,VVC还可以使用索引为1的参考行(Reference line1)和索引为2的参考行(Reference line2)作为扩展参考行进行帧内预测。为了减小编码复杂度,MRL仅在MPM中的非planar模式上使用。编码端对每个角度模式进行预测时,要将这三个参考行都尝试过,通过率失真优化,当前块选中率失真代价(RD Cost)最小的一个参考行,选中的参考行的索引被编码发送到解码端。解码端解码得到参考行的索引,再根据参考行的索引确定当前块选中的参考行,用于对当前块的预测。
图5所示的示例中,示出了当前块的4个参考行,包括与当前块相邻的参考行0(reference line0)221即索引为0的参考行;与当前块间隔1行的参考行1(reference line1)222即索引为1的参考行;与当前块间隔2行的参考行2(reference line2)223即索引为2的参考行;以及,与当前块间隔3行的参考行3(reference line3)224即索引为3的参考行。当前块的参考行可以有更多,即可以有索引为4以上的参考行。在预测时可以只使用参考行部分的重建值。本文中,参考行的索引均按照图5所示的方式编号。
本文中,参考行虽然称之为“行”,但这是为了表述方便,一个参考行实际上包括一行和一列,一般情况下,预测时使用的参考行的重建值也包括一行和一列的重建值,这与业界通常的描述方法是一样的。
ECM中,MRL模式可以使用更多的参考行,为了编码当前块选中的参考行,将多个候选的参考行的索引填入一个列表。该列表称为多参考行索引列表,简写为MRL索引列表,也可以称为多参考行列表,候选参考行列表,参考行索引列表等。在当前块不使用TIMD的情况下,MRL索引列表的长度为6,即共有6个位置,可填入6个参考行的索引,这6个参考行的索引及其顺序是固定的,依次为0,1,3,5,7,12,可以用下式表示:MULTI_REF_LINE_IDX[6]={0,1,3,5,7,12}。该MRL索引列表中,第1个位置填入的索引为0,是距离当前块最近的参考行的索引,第2个位置至第6个位置填入的索引分别为1,3,5,7,12,是按照到当前块的距离从近到远的顺序排列的5个扩展参考行的索引。
当前块选中的参考行在MRL索引列表中的情况下,使用多参考行索引(multiRefIdx)表示当前块选中的参考行在MRL索引列表中的位置,编码MRL索引即可表示选中的参考行。以MRL索引列表{0,1,3,5,7,12}为例,其中第1个位置至第6个位置对应的MRL索引分别为0至5。假定当前块选中的是索引为0的参考行,则MRL索引为0。假定当前块选中的是索引为7的参考行,则MRL索引为4,其他情况依此类推。MRL索引可以采用基于上下文模型的一元截断码编码,编码后得到多个基于上下文模型的二元标识,二元标识也可以称为二元标识位、二进制符号、二进制位等。MRL索引的值越小,码长越少,解码越快。
MRL模式也可以与TIMD模式同时使用,在使用TIMD的情况下,MRL索引列表长度为3,可填入3个参考行的索引且索引之间的顺序固定,表示为MULTI_REF_LINE_IDX[3]={0,1,3}。
需要说明的是,在不同的标准中,同样的技术可能有不同的名称。例如象MPM那样利用当前块周围的块导出一个最可能模式的列表的技术,在AV2(AVM)中称为自适应帧内模式编码(AIMC:Adaptive Intra Mode Coding),在AVS3中,在屏幕内容编码情况下,称为基于频数信息的帧内编码(FIMC:Frequency-based Intra Mode Coding)。在非屏幕内容编码情况下,类似MPM的技术常开。而如MRL那样,利用多参考行进行帧内预测技术,在AV2(AVM)中称为用于帧内预测的多参考行选择(MRLS:Multiple reference line selection for intra prediction)。但这仅仅是名称的不同,本实施例使用MPM、MRL等术语时,同样应覆盖到其他标准中的这些实质相同的技术。
一实施例提供了一种基于模板的多参考行帧内预测(Multiple reference line&intra_intra prediction)模式,简写为TMRL模式,TMRL模式是一种基于扩展参考行与帧内预测模式的组合构建候选列表,对扩展参考行与帧内预测模式的组合进行编解码的预测模式。
本实施例视频编码方法应用于编码器,如图6所示,包括:
步骤110,构建当前块的TMRL模式的候选列表,所述候选列表中填入有当前块候选的扩展参考行与帧内预测模式的组合;
步骤120,通过率失真优化,当前块选中一种参考行和帧内预测模式的组合用于帧内预测;
本文中,参考行包括索引为0的参考行和扩展参考行,当前块选中的一种参考行和帧内预测模式的组合可能是索引为0的参考行与一种帧内预测模式的组合,也可能是一个扩展参考行与一种帧内预测模式的组合。
步骤130,在当前块的TMRL模式的编码条件满足时,编码当前块的TMRL模式标志以表示当前块使用TMRL模式,编码当前块的TMRL模式索引以表示所述选中的组合在所述候选列表中的位置;
其中,所述编码条件至少包括:所述选中的组合在所述候选列表中。
本文中,填入有当前块候选的扩展参考行与帧内预测模式的组合的候选列表也可以称为TMRL模式的候选列表。
本文中,候选列表中填入的是当前块候选的扩展参考行与帧内预测模式的组合,这意味着候选列表中的组合需要参与当前块的率失真优化,即参与通过率失真代价来选择当前块预测模式的模式选择过程。使得候选列表中的组合均存在被选中的可能。
本文中,N,M,K这些表示个数的参数均为正整数,无需另行说明。
该实施例所构建的TMRL模式的候选列表,填入的是当前块候选的扩展参考行与帧内预测模式的组合,不再是单一的候选扩展参考行的列表或者候选的帧内预测模式的列表。在当前块选中的参考行和帧内预测模式的组合在该候选列表中(此时当前块选中的是扩展参考行),也即选中的组合是该候选列表中的一种组合等编码条件满足时,编码当前块的TMRL模式标志以表示当前块使用TMRL模式,编码当前块的TMRL模式索引以表示所述选中的组合在所述候选列表中的位置;解码端根据TMRL模式标志和TMRL模式索引可确定当前块选中的扩展参考行和帧内预测模式。本实施例的组合编解码方式可以减少编码代价,提升编码性能。
TMRL模式是根据当前块的N个扩展参考行和M个帧内预测模式得到的N×M种组合,分别对当前块的模板区域进行预测,并计算模板区域的重建值和预测得到的预测值之间的误差;按照误差从小到大的顺序将对应的K种组合填入当前块的TMRL模式的候选列表,1≤K≤N×M。
在一个示例中,基于ECM中扩展参考行与帧内预测模式的25种组合分别对模板区域进行预测,根据所述误差的升序对该25种组合排序,将误差最小也即最可能选中的12种组合填入TMRL模式的候选例表,使得TMRL模式索引可以使用更少的编码比特完成编码。即使扩展参考行{1,3,5,7,12}中有一部分(如索引为7和12的扩展参考行)在CTU边界之外,也可以基于15种组合排序,将最有可能选中的12种组合加入候选列表,TMRL模式索引的编码比特仍然可以得到充分有效地利用。而按照误差升序排序,使得更有可能被选中用来预测的参考行排和预测模式被留在候选列表中,越可能被选中的组合排在列表越靠前的位置使得编码代价降低。TMRL模式的候选列表的创建将在下文中加以更详细的描述。
在本实施例的一示例性中,所述编码条件还包括:当前块不使用TIMD;所述方法还包括:
在当前块使用TIMD的情况下,跳过当前块的TMRL模式标志和TMRL模式索引的编码;
在当前块不使用TIMD但所述选中的组合不在所述候选列表中的情况下,编码当前块的TMRL模式标志以表示当前块不使用TMRL模式,跳过当前块的TMRL模式索引的编码。
本实施例是基于TIMD模式在TMRL模式之前编解码的情况。如果当前块使用TIMD模式,就无需再使用TMRL模式,故跳过TMRL模式标志和TMRL模式索引的编码。而如果当前块不使用TIMD模式,就存在当前块选中的组合在TMRL模式的候选列表中,或者不在TMRL模式的候选列表中这两种情况,如果选中的组合不在候选列表中。需要编码TMRL模式标志表示当前块不使用TMRL模式,并跳过TMRL模式索引的编码。如果选中的组合在候选列表中,则TMRL模式标志和TMRL模式索引均需要进行编码。
在当前块不使用TIMD的情况下,本实施例提供的TMRL模式标志和TMRL模式索引可以替代原有的多参考行索引multiRefIdx。而在当前块使用TIMD的情况下,仍然可使用多参考行索引来表示选中的参考行,并对该多参考行索引进行编码。
在本实施例的一示例中,构建当前块的TMRL模式的候选列表,包括:在设定的当前块允许使用TMRL模式的条件均成立的情况下,才构建所述候选列表,允许使用TMRL模式的条件包括以下任 意一种或多种:
条件1,当前块为亮度帧中的块;即TMRL模式只用于亮度帧(即亮度图像);
条件2,当前块不位于编码树单元CTU的上边界;如位于CTU上边界,当前块上方无参考行可使用,故本实施例将不位于CTU上边界作为允许使用TMRL模式的条件;.
条件3,当前块允许使用多参考行MRL,即允许使用多参考行模式时才允许使用TMRL模式。
条件4,当前块的尺寸不大于可使用TMRL模式的当前块的最大尺寸;该最大尺寸可以是预设的。较大的块一般比较平缓,不太容易有角度细节,对这样的大块,可以限制TMRL模式的使用。
条件5,当前块的长宽比满足使用TMRL模式对当前块长宽比的要求。例如当前块长宽比不大于预设值时,才允许使用TMRL模式。
在本实施例的一示例中,所TMRL模式索引采用哥伦布-莱斯编码方法进行编码。采用哥伦布莱斯编码可以将候选的组合更合理的归类成码字长度不同的类别进行编码和解码,提升编码效率。
在本实施例的一示例中,在当前块的TMRL模式的编码条件满足的情况下,所述方法还包括:跳过以下任意一种或多种模式的语法元素的编码:MPM模式,帧内子块划分ISP模式、多变换选择MTS模式、低频不可分割变换LFNST模式,TIMD模式。
例如,在TMRL模式放在TIMD模式之前编解码时,此时TMRL模式的编码条件不再包括当前块不使用TIMD,在当前块使用TMRL模式时,当前块不允许使用TIMD,可跳过TIMD模式语法元素的编码。
例如,在当前块使用TMRL模式的情况下,通过TMRL模式标志和TMRL模式索引可以同时表示当前块选中的参考行和帧内预测模式,此时无需要再对MPM相关语法元素进行编解码。
例如,在特定变换模式下,可以限制TMRL模式不与多变换选择(MTS:Multiple transform selection)模式和/或低频不可分割变换(LFNST:Low frequency non-separable transform)模式同时使用。
一实施例提供了一种TMRL模式候选列表的构建方法,可以应用于编码器,也可以应用于解码器。如图7所示,所述方法包括:
步骤210,根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种组合,N≥1,M≥1,N×M≥2;
步骤220,根据所述N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
本步骤的误差可以用绝对误差和(SAD:the sum of absolute difference)表示,或用误差变换绝对值和(SATD:Sum of Absolute Transformed Difference)表示,但不局限于此,也可以用差值的平方和(SSD:Sum of Squared Difference)、平均绝对差值(MAD:Mean Absolute Difference)、平均平方误差(MSE:Mean Squared Error)等表示。
步骤230,按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块的TMRL模式的候选列表,1≤K≤N×M。
本实施例创建的候选列表可以实现扩展参考行与帧内预测模式的组合编码,提升编码效率。而且通过使用不同组合对模板区域的预测和对误差的排序,能够基于当前块和当前块模板区域在分布特性上的相似性,从N×M种组合中选出被选中可能性较大的K种组合,且将被选中可能性大的组合排在候选列表的前面,使得编码时被选中的组合的TMRL模式索引较小,减少实际编码代价。
在本实施例的一示例中,所述当前块的模板区域设置在距离当前块最近的一个参考行上;或者,所述当前块的模板区域设置在距离当前块最近的多个参考行上,参与组合的N个扩展参考行是位于模板区域外侧的扩展参考行。图8A中,当前块的模板区域设置在索引为0的参考行30上,在构建TMRL模式的候选列表中,是从预定义的索引为{1,3,5,7,12}的扩展参考行中选择可以使用的N个扩展参考行。如果当前块上方到CTU边界之间有13个以上的参考行,则选择索引为{1,3,5,7,12}的5个扩展参考行参与组合。如果当前块上方到CTU边界之间有6个或7个参考行,则选择索引为{1,3,5}的3个扩展参考行参与组合,依此类推。
本示例中当前块的模板区域设置在索引为0的参考行上,则索引为0的参考行称之为模板区域所 在的参考行,索引为1至3的参考行称之为位于模板区域外侧的参考行。如果当前块的模板区域设置在索引为0和1的参考行上,此时模板区域所在的参考行包括扩展参考行,则索引为0和1的参考行是模板区域所在的参考行,而索引为2和3的参考行是模板区域外侧(上侧和左侧)的参考行。
图8A示出了参与组合的5个扩展参考行:索引为1的参考行31、索引为3的参考行33、索引为5的参考行35、索引为7的参考行37以及索引为12的参考行39。而与图8A不同的是,图8B所示的示例中,当前块的模板区域40设置在索引为0和1的两个参考行上,参与组合的扩展参考行是5个扩展参考行,分别是索引为2的参考行42、索引为3的参考行43、索引为5的参考行45、索引为7的参考行47、及索引为12的参考行49。即该示例是从预定义的索引为{2,3,5,7,12}的扩展参考行中选择可以使用的N个扩展参考行。对于模板区域和扩展参考行的选择还有很多种,例如将当前块的模板区域设置在索引为0,1,2的3个参考行上,将当前块的模板区域设置在索引为0~3的4个参考行上,等等。模板区域较宽时,预测相对更为准确。
在本实施例的一示例中,当前块的N个扩展参考行是预定义的N max个扩展参考行中,位于当前块的模板区域外侧且不超过CTU边界的扩展参考行;其中,N max为TMRL模式可使用的扩展参考行的最大个数。本实施例将用于组合的N个扩展参考行限制在当前块的模板区域外侧且不超过CTU边界的区域。但如果硬件可以提供支持,也可以选择超过CTU边界的扩展参考行参与组合。
在本实施例的一示例中,N max=5,预定义的5个扩展参考行是索引为{1,3,5,7,12}或者{2,3,5,7,12}的参考行。在另一示例中,预定义的N max个扩展参考行是索引从1开始的距离当前块最近的N max个扩展参考行;或者,是索引从1开始且索引为奇数的距离当前块最近的N max个扩展参考行;或者,是索引从2开始且索引为偶数的距离当前块最近的N max个扩展参考行。选择奇数参考行或偶数参考行可以简化运算。
在本实施例的一示例中,M个帧内预测模式只允许从角度模式中选出,或者只允许从角度模式和DC模式中选出,或者允许从角度模式、DC模式和Planar模式中选出。
在一示例中,所述M个帧内预测模式通过以下方式选出,M≥5:
第一步:确定当前块周围5个邻近位置上的预测块使用的帧内预测模式,将其中允许选出的帧内预测模式依次选出并去掉重复的模式;
该5个邻近位置依次为左上、上、右上、左和左下;如图3所示。
在第一步选出的帧内预测模式的个数等于M时,结束;在第一步选出的帧内预测模式的个数小于M且包括角度模式时,执行第二步;
第二步,从已选出的第一个角度模式开始依次进行角度模式的扩展操作,得到扩展的角度模式,将与已选出的所有角度模式均不同的扩展的角度模式选出,直到已选出的帧内预测模式的总数等于M。
本实施例中,如果第一步没有选出角度模式,或者经第二步后选出的帧内预测模式的总数仍小于M,进行第三步:确定预定义的帧内预测模式集中允许选出的帧内预测模式,从确定的允许选出的帧内预测模式中依次选出与已选出的所有帧内预测模式均不同的帧内预测模式,直到已选出的帧内预测模式的总数等于M。
在一示例中,在选出M个帧内预测模式的过程中,只允许选出角度模式。在另一示例中,在选出M个帧内预测模式的过程中,只允许选出角度模式和DC模式;在又一示例中,在选出M个帧内预测模式的过程中,允许选出角度模式、DC模式和Planar模式。Planar模式与扩展参考行结合效果有限,可以不参与组合。DC模式的情况类似。但是如果能够接受所增加的运算复杂度,也可以将Planar模式和DC模式加入到候选列表参与组合,
在一示例中,所述角度模式的扩展操作包括以下操作中的任意一种或多种:对角度模式的加1和减1操作;对角度模式的加2和减2操作;及,对角度模式的加3和减3操作。
在本实施例的一示例中,所述M个帧内预测模式使用MPM中除Planar模式外的部分或全部帧内预测模式;或者,使用MPM和第二MPM中除Planar模式外的部分或全部帧内预测模式;或者,使用最可能模式MPM中除Planar模式和DC模式外的部分或全部帧内预测模式;或者,使用MPM和第二MPM中除Planar模式和DC模式外的部分或全部帧内预测模式;或者,使用最可能模式MPM中除Planar模式、DC模式和DIMD模式外的部分或全部帧内预测模式;或者,所述M个帧内预测模式使用MPM和第二MPM中除Planar、DC模式和DIMD模式外的部分或全部帧内预测模式。
本示例是将MPM中除Planar模式外的全部帧内预测模式作为预定义的帧内预测模式,或者将MPM和第二MPM中除Planar模式外的全部帧内预测模式作为预定义的帧内预测模式,在当前块的参考行包括预定义的所有扩展参考行时,就使用全部预定义的帧内预测模式。在当前块的参考行只包括预定义的部分扩展参考行时,就使用部分预定义的帧内预测模式。
在本实施例的一示例中,M个帧内预测模式按以下方式选出:
选出M’个帧内预测模式;
根据当前块位于模板外侧的参考行和该M’个帧内预测模式,对当前块的模板分别进行预测并计算所述模板的重建值和预测得的预测值之间的误差,得到M’个误差;
从M’个帧内预测模式选出对应误差最小的M个帧内预测模式,作为参与组合的所述M个帧内预测模式,M<M’。
本示例当前块的模板用于从M’个帧内预测模式中选择M个帧内预测模式,而前述当前块的模板区域用于从N×M种组合中选择K种组合,两者可以是不同的,但也可以占据相同的区域。
本示例选出M’个帧内预测模式时,可以使用上述实施例选出M个帧内预测模式的各种方法,例如,从MPM和第二MPM的列表中直接选出,或者采用前述实施例的方法经第一步选出,或经第一步和第二步选出,或经第一步、第二步和第三步选出,等等。
在本实施例的一示例中,N≤N max,2≤N max≤12;2≤M≤18;K≤K max,6≤K max≤36;其中,N max为TMRL模式允许使用的扩展参考行的最大个数,K max为TMRL模式允许使用的候选组合的最大个数。虽然这里给出了N、M和K的相关参数的取值范围,但这仅仅是示例性地。
在本实施例的一示例中,所述N,M和K至少有两组取值,第一组取值为N 1,M 1,K 1,第二组取值为N 2,M 2,K 2,其中,N 1≤N 2,M 1≤M 2,K 1≤K 2,且N 1×M 1<N 2×M 2;所述第一组取值在为第一尺寸的当前块构建所述候选列表时使用,第二组取值在为第二尺寸的当前块构建所述候选列表时使用,所述第一尺寸小于所述第二尺寸。这里的第一尺寸和第二尺寸可以分别代表多种尺寸,例如第一尺寸可以包括4×4、4×8、8×8等等,第二尺寸可以包括16×8、16×16、8×16等等。本示例对于不同尺寸的当前块,采用不同的N,M和K,在当前块尺寸较小时,使用较小的取值来构建TMRL模式的候选列表。在当前块尺寸较大时,使用较大的取值来构建TMRL模式的候选列表。可以在运算的复杂和性能上取得更好的平衡。
在本实施例的一示例中,所述根据所述N×M种组合分别对当前块的模板区域进行预测,包括:在当前块位于图像(picture)左边界的情况下,根据所述N×M种组合分别对当前块的上方模板区域进行预测,不对当前块的左侧模板区域进行预测。本实施例在不对性能造成影响的基础上,可以简化运算,减小运算所需时间。
在本实施例的一示例中,所述根据所述N×M种组合分别对当前块的模板区域进行预测,包括:对所述N×M种组合中的每一种组合,按以下方式进行预测:
根据该组合中的扩展参考行的重建值和该组合中的帧内预测模式计算出所述模板区域初始的预测值,其中,所述扩展参考行的重建值是所述扩展参考行原始的重建值或经滤波后的重建值;
对所述模板区域初始的预测值进行4抽头滤波或6抽头滤波,滤波结果作为根据该组合预测得到的所述模板区域的预测值。
本实施例在对当前块的模板区域进行预测时,可以不对扩展参考行的重建值进行滤波,使用原始的重建值进行运算,还可以使用较短抽头的滤波器(如4抽头滤波器),来降低运算复杂度,加快运算速度。
在本实施例的一示例中,所述根据所述N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差,包括:
根据K种组合分别对当前块的整个模板区域进行预测,得到对应的K个误差组成一误差集,记录所述误差集中的最大误差为D max
对于余下的每一种组合,先根据该组合对当前块的单侧模板区域进行预测,计算当前块单侧模板区域的重建值和预测值之间的误差D 1,如D 1≥D max,完成该组合的预测,如D 1<D max,再根据该组合对当前块的另一侧模板区域进行预测,计算当前块整个模板区域的重建值和预测值之间的误差D 2,如 D 2<D max,将D 2加入所述误差集,将D max从所述误差集中删除并更新所述误差集中的最大误差D max,如D 2≥D max,则完成该组合的预测;
完成所述N×M种组合的预测后,将所述误差集中K个误差对应的K种组合作为对应误差最小的K种组合。
在一个示例中,可以将误差集中的误差按从小到大的顺序排列,将D 2加入误差集时、将D 2插入到的位置应使得误差集中的误差仍按照从小到大的顺序排列。但在其他示例中,也可以在完成所述N×M种组合的预测后,再将所述误差集中K个误差排序。
本示例不必对所有组合都进行整个模板区域的预测和误差计算,可以完成对组合的排序,可以降低运算复杂度,加快运算速度。
在本实施例的一示例中,所述按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块的TMRL模式的候选列表,包括:从所述候选列表的第1个位置开始,按照所述误差从小到大的顺序,将所述误差对应的K种组合填入所述候选列表。本示例TMRL模式的候选列表只填入扩展参考行与帧内预测模式的组合。索引为0的参考行与帧内预测模式的组合通过其他传统模式来指示,如MPM。
在本实施例的一示例中,从候选列表的第i个位置开始,按照所述误差从小到大的顺序,将误差对应的K种组合填入所述候选列表;候选列表的第i个位置之前填入有索引为0的参考行与一种或多种帧内预测模式的组合,i≥2。本示例TMRL模式的候选列表不仅填入扩展参考行与帧内预测模式的组合,还可以填入索引为0的参考行与帧内预测模式的组合。在此情况下,当前块选中的组合是索引为0的参考行与帧内预测模式的组合时,也可以通过TMRL模式索引来表示。此时的TMRL模式标志仍可以使用。
一实施例提供了TMRL模式相关的视频解码方法,应用于解码器,如图9所示,所述方法包括:
步骤310,解码当前块的多参考行帧内预测TMRL模式标志,确定当前块是否使用TMRL模式;
步骤320,确定当前块使用TMRL模式的情况下,继续解码当前块的TMRL模式索引,并构建当前块的TMRL模式的候选列表,所述候选列表中填入有当前块候选的扩展参考行与帧内预测模式的组合;
步骤330,根据所述候选列表和TMRL模式索引,确定当前块选中的扩展参考行和帧内预测模式的组合,根据选中的组合对当前帧进行预测;
其中,所述TMRL模式索引用于表示所述选中的扩展参考行与帧内预测模式的组合在所述候选列表中的位置。
本实施例通过解码TMRL模式标志确定当前块使用TMRL模式后,将扩展参考行与帧内预测模式的组合填入TMRL模式的候选列表,通过解码得到的TMRL模式索引和候选列表确定当前块选中的组合并进行预测。即通过TMRL模式索引可以同时指示当前块选中的扩展参考行和帧内预测模式,不需要使用两个索引来完成。可以减少编码代价。
在本实施例的一示例中,解码当前块的TMRL模式标志之前,还包括:在当前块允许使用TMRL模式的条件均成立时,再解码当前块的TMRL模式标志,所述允许使用TMRL模式的条件包括以下任意一种或多种:
当前块为亮度帧中的块;
当前块允许使用MRL;
当前块不位于编码树单元CTU的上边界;
当前块不使用基于模板的帧内预测模式导出TIMD。
本示例在以上条件中的一个成立时,不允许使用TMRL模式,可以跳过TMRL模式标志和TMRL模式索引的解码。但在其他实施例并不一定如此,例如,将TMRL模式放在TIMD之前编解码时,当前块使用TIMD就不能作为不允许使用TMRL模式的条件。又如,在将来硬件可以支持获取CTU边界之外的参考行时,当前块位于CTU的上边界也就不再作为不允许使用TMRL模式的条件,等等。
在本实施例的一示例中,所述方法还包括:在解码确定当前块允许使用MRL、当前块不位于CTU 的上边界且当前块使用TIMD的情况下,解码当前块的多参考行索引,所述多参考行索引用于表示当前块选中的参考行在多参考行索引列表中的位置。本示例在当前块使用TIMD的情况下,不允许使用TMRL模式,但仍然可以允许使用MRL,因而仍然可以通过解码当前块的多参考行索引来确定当前块选中的参考行,再结合当前块选中的TIMD模式,即可对当前块进行预测。
在本实施例的一示例中,根据所述TMRL模式标志确定当前块使用TMRL模式的情况下,所述方法还包括:跳过以下任意一种或多种模式的语法元素的解码:MPM模式、子分区模式、多变换选择MTS模式、低频不可分割变换LFNST模式、TIMD模式。与编码端相应的,如果编码时,在当前块使用TMRL模式标志的情况下,跳过上述模式中的一种或多种,则解码端通过解码确定当前块使用TMRL模式标志的情况下,也跳过对这些模式的解码。
一实施例还提供了一种视频解码方法,主要涉及帧内预测的解码过程,同时也对编码端做相应说明。本实施例编码端构建TMRL模式的候选列表,通过模式选择当前块选中该候选列表中的组合时,对TMRL模式的语法元素进行编码和解码,通过将扩展参考行和帧内预测模式组合编解码。
本实施例根据预定义的N个扩展参考行和M个帧内预测模式,在索引为0的参考行即reference line0的位置上构建模板,可参见图8A所示的模板区域30。(x,-1),(-1,y)分别是相对当前块左上角(0,0)位置的坐标。图中还增出了预定义的5个扩展参考行,索引为{1,3,5,7,12}。
构建TMRL模式的候选列表时,计算N×M种组合在模板上的预测值与重建值之间的SAD,根据SAD的升序对相应的组合排序,将NxM种组合中SAD较小的K种组合按照SAD从小到大的顺序填入TMRL模式的候选列表,N×M≥2。
TMRL模式的候选列表的构建是一个编码器和解码器都需要进行的操作,编码端在当前块选中的组合在该候选列表中等编码条件满足的情况下,编码TMRL模式标志以表示使用TMRL模式,还根据选中的组合在候选列表中位置确定TMRL模式索引,如在第1个位置时,TMRL模式索引为0,在第2个位置时,TMRL模式索引为1,依此类推。TMRL模式索引可使用哥伦布莱斯编码(golomb-rice)方式编码,但不局限于此。
以下以N=5,M=6和K=12为例子进行说明。预定义的5个扩展参考行的索引为{1,3,5,7,12},6个帧内预测模式分步选出。
本实施例的视频解码方法包括:
步骤一,解码TMRL模式相关的语法元素;
解码器解析帧内预测模式相关语法元素,包括TIMD、MRL等模式的相关语法元素,本实施例提出的TMRL模式可以视为对MRL模式的演进,TMRL模式的语法元素也可视为MRL模式语法元素的一部分。当然,也可以将两者视为两种不同的模式。
当前块使用TIMD模式的情况下,对MRL模式语法元素的解码方式不变。当前块不使用TIMD模式的情况下,则需要解码TMRL模式的语法元素,当前块解码相关的语法如下表所示:
Figure PCTCN2022106337-appb-000002
Figure PCTCN2022106337-appb-000003
表中的“cu_tmrl_flag”即TMRL模式标志,等于1表示当前块使用TMRL模式,也即定义了当前亮度样本的帧内预测类型为基于模板的多参考行帧预测模式;“cu_tmrl_flag”等于0表示当前块不使用TMRL模式,也即定义了当前亮度样本的帧内预测模式类型不为基于模板的多参考行帧人预测模式。
表中的“tmrl_idx”即TMRL模式索引,当前块选中的扩展参考行与帧内预测模式的组合在TMRL模式的候选列表中的位置,也可以说是定义了选中的组合在TMRL模式所排序的候选列表中的索引(表示组合所在位置的索引)。“tmrl_idx”可以使用哥伦布莱斯方式进行编解码,这里不再赘述。
从上表可以看出,在解码cu_tmrl_flag之前,先判断以下条件是否成立:当前块允许使用MRL(即sps_mrl_enabled_flag为1是否成立),当前块不位于CTU上边界(即(y0%CtbSizeY)>0是否成立),以及当前块不使用TIMD。在这些条件成立时,再解码cu_tmrl_flag。如果其他两个条件成立而当前块使用TIMD,则解码当前块的多参考行索引intra_luma_ref_idx。
表中的ISP模式标志(intra_subpartitions_mode_flag)在TMRL模式相关语法元素之后解码,在当前块不使用TMRL模式(!cu_tmrl_flag成立)的情况下,再解码intra_subpartitions_mode_flag。同样,在当前块不使用TMRL模式(!cu_tmrl_flag成立)的情况下,再解码MPM相关语法元素。
步骤二,构建TMRL模式的候选列表,根据TMRL模式索引和所述候选列表确定当前块选中的扩展参考行和帧内预测模式;
解析阶段结束后,对当前块进行预测前,如果当前块使用TMRL模式,则需要构建TMRL模式的候选列表,根据TMRL模式索引和所述候选列表确定当前块选中的扩展参考行和帧内预测模式。
构建TMRL模式的候选列表,需要先确定候选的扩展参考行、确定候选的帧内预测模式。
■确定候选的扩展参考行
候选的扩展参考行从预定义的扩展参考行中选出。根据当前块在图像中的位置来确定预定义的扩展参考行中哪些是可以使用的,原则上,对于当前块可以使用的上参考行,不应超出上CTU边界。在一个示例中,将索引为{1,3,5,7,12}的扩展参考行中,没有超出CTU边界的扩展参考行都加入候选的扩展参考行。为了获得更好的编解码性能,或者为了降低复杂度,也可以采用更多或更少的扩展参考行。
■确定候选的帧内预测模式
本实施例TMRL模式不与MPM绑定(在其他实施例也可以绑定),而是构建一个帧内预测模式的候选列表,用于组合的帧内预测模式将从此候选列表中选出。候选列表的导出方法如下:
首先,在67种传统的预测模式中,去除Planar模式和DC模式,或者仅去除Planar模式而保留DC模式,去除的模式不加入候选列表即不作为TMRL模式中参与组合的帧内预测模式。
本实施例待构建的候选预测模式列表长度为6,先从当前块周围5个邻近位置的预测块使用的帧内预测模式,依序选出不重复的帧内预测模式填充候选预测模式列表。然后,对已经填充进列表的模式进行角度模式的扩展操作,具体可以对角度模式进行加一减一操作,将不重复的扩展的角度模式选出,依次填入候选预测模式列表,若候选列表填充的模式数到达6,则停止填充。
加一减一的具体操作过程如下表所示:
Figure PCTCN2022106337-appb-000004
Figure PCTCN2022106337-appb-000005
上述“角度模式-1”是指将已填充的角度模式的索引减1后得到的角度模式,如已填充的角度模式为模式3,角度模式-1即角度模式2。上述“角度模式+1是指将已填充的角度模式的索引加1后得到的角度模式,例如已填充的角度模式为模式3,角度模式+1即角度模式4。
而若角度模式-1比角度模式2小,例如角度模式-1得到的是角度模式1(角度模式的索引从2编号,角度模式1不存在),则选取-1角度相反方向的角度模式,假定共有65种角度模式,此时相反方向的角度模式即角度模式66。而若角度模式+1后的角度模式比角度模式66大,则选取+1角度相反方向的角度模式是类似的,如已填充的角度模式是角度模式66,则该角度模式+1不存在,此时选取的+1角度相反方向的角度模式即角度模式2。
上述对角度模式进行扩展时,是对角度模式进行加1减1操作,在其他实施例中,也可以将加1减1操作扩展为从加1减1至加X减X,假定X=3,则可以对角度模式进行加1减1、加2减2、加3减3操作,直到候选预测模式列表填充满。
如果对已经填充进列表的模式进行角度模式的扩展操作之后,候选预测模式列表还没有填充满,则使用预定义的模式集中不重复的模式进行填充,直到候选预测模式列表填满。该模式集包括按照统计规律筛选出的一些角度模式,如下:
mpm_default[]={DC_IDX,VER_IDX,HOR_IDX,VER_IDX-4,VER_IDX+4,14,22,42,58,10,26,38,62,6,30,34,66,2,48,52,16};
其中,DC_IDX表示DC模式,VER_IDX表示垂直模式,HOR_IDX表示水平模式,其余数字表示该数字对应的角度模式。
本实施例中,候选预测模式列表的长度为6,也可以为了性能尝试更多的角度模式,将长度设置为大于6的值,也可以为了降低复杂度尝试较少的模式,将长度设置为小于6的值。
本实施例确定候选帧内预测模式时,将Planar模式和DC模式,或者仅将Planar模式排除,但在如果不考虑复杂度的情况,这两个模式也可以不排除,即Planar模式、DC模式和所有角度模式均可作为候选帧内预测模式而参与与扩展参考行的组合。
■构建TMRL模式的候选列表
确定候选的扩展参考行及帧内预测模式后,可逐一尝试扩展参考行列表和候选预测模式列表中的所有组合,对这些组合在reference line 0所在行的模板(Template如下图)区域上分别进行预测,参见图8A,计算模板区域的重建值和每种组合预测得到的预测值之间的误差,将误差最小的K种组合按误差升序填入TMRL模式的候选列表。
本实施例在预测过程中,只限制了TMRL模式在当前块占据CTU首行时的使用,当当前块位于图像左边界时,TMRL模式仍然可以使用,在这种情况下时由于左侧的reference line 0已经处于图像边界外,所以在预测时不使用左侧模板,即只进行上方模板区域的预测。
在模板区域上预测的过程可与其它正常的帧内角度预测过程完全一致,即先对参考行的重建值滤波再作为模板区域初始的预测值,而基于参考行滤波后的重建值和组合中的帧内预测模式对模板区域进行预测之后,将初始的预测结果进行4或6抽头滤波,再作为预测得到的预测值。考虑运算的复杂度,可以省去对参考行的重建值的滤波步骤,也可以使用较短抽头的滤波器。本实施例中,在预测得到模板区域的预测值时,参考行像素的重建值不经过滤波,而初始的预测结果是在非整数角度下经过1/32精度的4抽头插值滤波。
最后,根据当前组合中的角度,参考行,滤波器对模板区域预测。计算预测得到的模板区域的预测值与模板区域的重建值之间的SAD,按照SAD升序排序,选取SAD最小的12个组合填入TMRL模式的候选列表。
在排序过程中可以使用快速算法,在上述过程中需要尝试5个参考行,6种预测模式共30种组合,但只需要选择其中SAD最小的12个组合,本实施例在预测完前12种组合,得到相应的SAD后,从第13种组合开始,只需要保持12个误差(也可以叫代价)最小的组合并进行更新。从第13种组合开始,只对上方模板区域进行预测并计算相应的SAD,在根据上模板计算的SAD已经大于误差最小的12种组合中误差最大的一种组合时,可以跳过对左侧模板区域的预测和误差计算,具体可参见前述实施例。
步骤三,根据构建的TMRL模式的候选列表和解码得到的TMRL模式索引确定当前块选中的扩展参考行和帧内预测模式的组合,根据选中对当前块进行帧内预测。
当使用TMRL模式时,参考行的索引refIdx和变量predModeIntra定义了帧内预测所使用的模式,均根据TMRL模式索引“tmrl_idx”与TMRL模式的候选列表来确定。
在ECM-4.0参考软件上,使用本实施例所述的方法,以N=5(5个扩展参考行分别为1,3,5,7,12),M=6(6个预定义的预测模式,不包含Planar和DC模式)和K=12(只选取所有组合种的前12个SAD小的排序组合)的设定,在AI配置下测得结果如下:
Figure PCTCN2022106337-appb-000006
以N=5(5个扩展参考行分别为1,3,5,7,12),M=8(8个预定义的预测模式)和K=16的设定,在AI配置下测得结果如下,
Figure PCTCN2022106337-appb-000007
以N=5(5条扩展的参考行分别为1,3,5,7,12),M=12(12个预定义的预测模式)和K=24的设定,在AI配置下测得结果如下,
Figure PCTCN2022106337-appb-000008
以N=5(5条扩展的参考行分别为1,3,5,7,12),M=8(8个预定义的预测模式,不包含Planar和但包含DC模式)和K=16的设定,在AI配置下测得结果如下,
Figure PCTCN2022106337-appb-000009
表中的参数含义如下:
EncT:Encoding Time,编码时间,10X%代表当集成了参考行排序技术后,与没集成前相比,编码时间为10X%,这意味有X%的编码时间增加。
DecT:Decoding Time,解码时间,10X%代表当集成了参考行排序技术后,与没集成前相比,解码时间为10X%,这意味有X%的解码时间增加。
ClassA1和Class A2是分辨率为3840x2160的测试视频序列,ClassB为1920x1080分辨率的测试序列,ClassC为832x480,ClassD为416x240,ClassE为1280x720;ClassF为若干个不同分辨率的屏幕内容序列(Screen content)。
Y,U,V是颜色三分量,Y,U,V所在列表示测试结果在Y,U,V上的BD-rate
Figure PCTCN2022106337-appb-000010
指标,值越小表示编码性能越好。
All intra表示全帧内帧配置的测试配置。
由可见使用本实施例TMRL模式进行帧内预测编解码,可以取得编码性能的明显提升。
本实施例无论是扩展参考行或是预测模式,都是用1行1列的模板区域,并使用SAD升序进行了排序和筛选。对于扩展参考行而言,若排序所有的扩展参考行(包括reference line 1),则只能使用1行1列的模板。然而对于预测模式的筛选,为了筛选出更合适的预测模式,可以像TIMD模式那样,使用更多的参考行可以获得更准确的结果。在其他实施例中,也可以将TMRL模式确定候选的预测模式列表的方式加以变化。
例如,需要构建一个长度为6的候选预测模式列表时,可以根据与本实施例相同的构建和填充方法方法先构建一个长度大于6的列表,然后使用距离当前块最近的4行4列作为模板,使用第5个参考行和候选预测模式列表中的帧内预测模式对模板进行预测,计算预测得到的预测值与模板的重建值之间的误差(SAD或SATD),按照误差升序排序,选出其中6个误差小的帧内预测模式作为要构建的长度为6的TMRL模式候选列表中的帧内预测模式。此外,TMRL模式的候选列表的长度为6仅仅是一个示例,可根据情况做数值上的调整。
本实施例的角度模式以65种为例,但在其他实施例中,也可以将角度扩展到129种或更多以获得更好的性能,当扩展到更多的角度时,帧内预测的滤波器数量也应做相应提升,例如129种角度时使用1/64精度的滤波。
一实施例还提供了一种帧内预测融合(IPF:Intra prediction fusion)技术,IPF允许角度模式使用两条相邻的参考行的预测结果进行加权,以得到当前块最终的预测结果。加权的方式如下式所示,
p fusion=w a×p a+w b×p b
其中,p a为使用索引为a的参考行(reference line a)和该角度模式对当前块进行预测的结果,p b为使用索引为a+1的参考行(reference line a+1)和该角度模式对当前块进行预测的结果,p fusion为融合的预测结果;w a为l加权时p a的权重,w b为加权时p b的权重,w a为3/4,w b为1/4。
当前块使用IPF时,将上述融合的预测结果作为当前块最终的预测结果。
本实施例中,没有设置表示IPF是否使用(即是否开启)的标识符号,而是默认在以下条件满足 时,对当前块选中的角度模式时都开启:
当前块选中的角度模式不为整数斜率的角度模式;
当前块的宽乘高大于16
当前块没有选中帧内字块划分ISP模式。
换言之,在以下限制条件中的至少一种成立时,不允许使用IPF:
当前块选中的角度模式为整数斜率的角度模式;
当前块的宽乘高小于或等于16;
当前块选中ISP模式。
其中,当角度模式的角度值(intraPredAngle)除32余数为0时,该角度模式为整数斜率(integer slope)的角度模式,intraPredAngle与角度模式的对应可见上文的表1。
当当前模式满足IPF条件时,当前块选中的帧内预测模式有可能是使用DIMD选中的DIMD融合模式(融合planar模式和两种角度模式),也可能是使用TIMD选中的TIMD融合模式。从硬件实现的角度来说,融合越少越好,在选中DIMD融合模式时已经存在三个帧内预测结果在当前块上进行融合的情况,如果再使用IPF,过多的融合会使得预测阶段的复杂度增加。
本公开一实施例提供了一种帧内预测融合方法,可以应用于编码器,也可以应用于解码器,如图10所示,所述方法包括:
步骤410,在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用帧内预测融合IPF的限制条件是否成立;
步骤430,在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF。
本步骤中,对当前块进行帧内预测时限制使用IPF,可以是不允许使用IPF,也可以是在选中的帧内预测模式包括多种满足IPF使用条件的角度模式的情况下,只允许对其中的部分角度模式进行IPF融合。在一个示例中,满足IPF使用条件的角度模式是指不为整数斜率的角度模式,或是除角度为-45°、0°、45°、90°、135°之外的角度模式。
本文中,某一限制条件成立时不允许使用IPF,是对当前块进行预测时不使用IPF的充分条件。而某一限制条件不成立时,允许使用IPF,是对当前块进行预测时使用IPF的必要条件,如果其他限制条件成立,仍可能不允许使用IPF。
在本公开一示例性的实施例中,所述限制条件包括以下的模式数量限制条件:使用IPF会导致当前块的预测需要融合N种以上的帧内预测模式,N为大于或等于3的整数。
在本实施例的一个示例中,在N=3的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
当前块选中TIMD融合模式且TIMD融合了满足IPF使用条件的两种角度模式的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时只允许对该两种角度模式中的一种进行IPF融合;
当前块选中TIMD融合模式且TIMD融合的两种帧内预测模式中只有一种是满足IPF使用条件的角度模式的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对该角度模式进行IPF融合。
在本公开一示例性的实施例中,所述限制条件包括以下的模式数量限制条件:使用IPF会导致当前块的预测需要融合M种以上的角度模式,M为大于等于2的整数。
在本实施例的一个示例中,在M=3的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
当前块选中TIMD融合模式且TIMD融合了满足IPF使用条件的两种角度模式的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时只允许对该两种角度模式中的一种进行IPF融合,例如,只允许对该两种角度模式中代价最小或代价次小的角度模式进行IPF融合;
当前块选中TIMD融合模式且TIMD融合的两种帧内预测模式中只有一种是满足IPF使用条件的角度模式的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对该角度模式进行IPF融合。
在本实施例的一个示例中,在M=3的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
当前块选中DIMD融合模式且DIMD融合的两种角度模式均满足IPF使用条件的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时只允许对该两种角度模式中的一种进行IPF融合,例如,只允许对该两种角度模式中幅度值最高或幅度值次高的角度模式进行IPF融合。
在当前块使用DIMD融合模式且DIMD融合的两种角度模式中只有一种满足IPF使用条件的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对满足IPF使用条件的该角度模式进行IPF融合。
在本实施例的一个示例中,在M=2的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
当前块选中TIMD融合模式且TIMD融合了两种角度模式的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时不允许使用IPF;
当前块选中TIMD融合模式、TIMD融合的两种帧内预测模式中只有一种是角度模式,且该角度模式满足IPF使用条件的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对满足IPF使用条件的该角度模式进行IPF融合。
在上述实施例中,所述对当前块进行帧内预测时允许对一角度模式进行IPF融合,包括:将二种以上的预测结果的加权和作为当前块最终的预测结果;所述二种以上的预测结果包括根据当前块选中的第一参考行和该角度模式对当前块进行预测得到的预测结果,及根据不同于所述第一参考行的第二参考行和该角度模式对当前块进行预测得到的预测结果。例如,所述第二参考行可以是与所述第一参考行相邻的参考行或者是索引为0的参考行。
在本公开一示例性的实施例中,所述限制条件包括以下限制条件中的任意一种或多种:
限制条件一:当前块选中TIMD融合模式;
限制条件二:当前块选中DIMD融合模式;
限制条件三:当前块使用多参考行MRL;
限制条件四:当前块选中的参考行的索引大于等于K,K为大于等于3的整数;
限制条件五:当前块的宽小于或等于设定值;
限制条件五:当前块的高小于或等于设定值;
限制条件七:当前块的宽乘高小于或等于设定值;
限制条件八:当前块使用帧内子块划分ISP模式;
限制条件九:当前块选中的角度模式的角度为-45°、0°、45°、90°、135°中的任意一种,或当前块选中的角度模式为整数斜率的角度模式;
限制条件十:当前块所属的当前帧为帧间帧;
限制条件十一:当前块所属的当前帧为色度帧,即只对亮度帧使用IPF,对色度帧不使用IPF;
在所述限制条件一至限制条件十一中的至少一种成立的情况下,在当前块进行帧内预测时不允许使用IPF。
在本实施例的一示例中,所述限制条件包括限制条件四,其中,K=3或5或7或12。
本公开上述实施例通过限制IPF的使用,例如,本实施例可以限制融合的帧内预测模式的数量,限制融合的角度模式的数量,限制IPF与TIMD、DIMD等可能融合多种模式的预测模式同时使用等,可以避免在预测时出现过多的融合,导致预测阶段的复杂度不适当地增加。
本实施例可以限制当前块的尺寸,在当前块的尺寸大于设定值时才使用IPF,是因为当前块小于某个尺寸时,通常有着更多纹理,使用融合预测对性能的提升有限。本实施例在当前块选中的参考行的索引大于等于K时不允许使用IPF,即在当前块选中的扩展参考行距离当前块比较远时不允许使用IPF,此时使用IPF对性能的提升有限。
本实施例通过限制IPF和MRL同时使用,可以减少计算的复杂度。本实施例在当前帧为帧间帧(如B帧、P帧)时不允许使用IPF,可以降低帧间编解码的代价。
在本公开一示例性的实施例中,所述方法还包括:在所述限制条件均不成立的情况下,对当前块进行帧内预测时使用IPF;所述对当前块进行帧内预测时使用IPF,包括:
在当前块选中的帧内预测模式只有一种且为角度模式的情况下,将第一预测结果和第二预测结果的加权和作为当前块最终的预测结果;其中,所述第一预测结果是根据当前块选中的第一参考行和当前块选中的第一角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和所述第一角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
应当说明的是,上述在所述限制条件均不成立的情况下,对当前块进行帧内预测时使用IPF。并不意味着只有限制条件均不成立,对当前块进行帧内预测时才允许使用IPF。例如在前述模式数量限制条件成立时,仍可以对部分满足IPF使用条件的角度模式进行IPF融合。
在本实施例的一示例中:所述第一预测结果和第二预测结果的加权和按照以下公式计算:
p fusion=(w ap a+w bp b)>>shift      公式一
或者
p fusion=(w ap a+w bp b+offset)>>shift    公式二
其中,p a为所述第一预测结果,p b为所述第二预测结果,p fusion当前块最终的预测结果,w a为p a的权重,w b为p b的权重;
其中,w a+w b=(1<<shift),offset=1<<(shift-1),shift>=1,shift和offset为设定的参数。
以上算法可以避免在运算过程中出现小数,可以提高算法效率。
在本公开一示例性的实施例中,所述方法还包括:在所述限制条件均不成立的情况下,对当前块进行帧内预测时使用IPF;所述对当前块进行帧内预测时使用IPF,包括:
在当前块选中的帧内预测模式为融合模式的情况下,将三种以上预测结果的加权和作为当前块最终的预测结果;其中,所述三种以上预测结果包括:根据当前块选中的第一参考行和所述融合模式中的每一种帧内预测模式对当前块进行预测得到的多种预测结果,以及根据第二参考行和所述融合模式中允许进行IPF融合的一种或多种角度模式对当前块进行预测得到的一种或多种预测结果;所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
在本公开上述实施例中,计算所述第一预测结果和第二预测结果的加权和时,为所述第一预测结果赋予的权重大于为第二预测结果赋予的权重。例如,在第一参考行为索引为1的参考行,第二参考行为索引为0的参考行时,为第一预测结果赋予权重3/4,为第二预测结果赋予权重1/4。
在本公开上述实施例中,在所述第一参考行的索引大于或等于K的情况下,所述第二参考行与所述第一参考行相邻且比所述第一参考行更为靠近当前块,K为大于或等于1的整数。
例如,所述第二参考行可以根据以下方式中的至少一种确定:
在所述第一参考行的索引为1时,确定所述第二参考行的索引为0;
在所述第一参考行的索引为3时,确定所述第二参考行的索引为2;
在所述第一参考行的索引为5时,确定所述第二参考行的索引为4;
在所述第一参考行的索引为7时,确定所述第二参考行的索引为6;
在所述第一参考行的索引为12时,确定所述第二参考行的索引为11。
在第一参考行的索引大于或等于K的情况下,将第二参考行设置为与第一参考行相邻且比第一参考行更为靠近当前块的参考行,可以提高第二参考行对当前块预测的准确度,进而提高融合后的预测结果的准确度。
在本公开一示例性的实施例中,IPF可以与TMRL模式同时使用。即在当前块选中TMRL模式的候选列表中的一扩展参考行和一角度模式,且所述限制条件均不成立的情况下,计算根据该扩展参考行和该角度模式对当前块进行预测的第一预测结果,及根据另一参考行和该角度模式对当前块进行预测的第二预测结果,将所述第一预测结果和第二预测结果的加权和作为当前块最终的预测结果,其中,该另一参考行为索引为0的参考行或该扩展参考行的相邻行。
本公开一实施例还提供了一种视频编码方法,应用于编码器,如图11所示,包括:
步骤510,通过模式选择确定当前块选中的参考行和帧内预测模式;
步骤520,按照如本公开任一实施例所述的帧内预测融合方法对当前块进行预测,得到当前块的预测值;
本步骤中,在对当前块进行帧内预测时使用IPF的情况下,是根据当前块最终的预测结果得到当前块的预测值。
步骤530,根据当前块的原始值和预测值确定当前块的残差。
本实施例视频编码方法使用本公开任一实施例的帧内预测融合方法对当前块进行预测,可以取得该帧内预测融合方法的各种效果。
本公开一实施例还提供了一种视频解码方法,应用于解码器,如图12所示,包括:
步骤610,解码码流,确定当前块选中的参考行和帧内预测模式;
步骤620,按照如本公开任一实施例所述的帧内预测融合方法对当前块进行预测,得到当前块的预测值;
本步骤中,在对当前块进行帧内预测时使用IPF的情况下,是根据当前块最终的预测结果得到当前块的预测值。
步骤630,根据当前块的预测值确定当前块的重建值。
本实施例视频编码方法使用本公开任一实施例的帧内预测融合方法对当前块进行预测,可以取得该帧内预测融合方法的各种效果。
上述实施例中使用的基于模板的多参考行帧内预测(TMRL)模式的帧内预测方法,是基于扩展参考行与帧内预测模式的组合构建候选列表,对扩展参考行与帧内预测模式的组合进行编解码的预测模式。但无论是否基于模板或是否排序,都是使用单行的参考行进行预测,单行的参考行通常包含噪声,会影响预测的准确度。
故本公开一实施例提出一种在TMRL模式的基础上使用IPF的方法。在TMRL模式基础上使用IPF,即对选中的角度模式进行IPF融合,也就是根据选中的参考行和该角度模式对当前块进行预测的结果和根据另一参考行和该角度模式对当前块进行预测的结果加权,得到当前块最终的预测结果。通过将不同参考行的预测结果进行融合,从而提高预测准确度。
在TMRL模式的基础上使用IPF,可以不设置标识位表示当前块是否使用IPF,而是由编码器和解码器根据约定的条件进行判断。
在本公开一实施例中,IPF可以与TMRL模式同时使用时,在当前块选中TMRL模式的候选列表中的一扩展参考行和一角度模式,且限制条件均不成立的情况下,计算根据该扩展参考行和该角度模式对当前块进行预测的第一预测结果,及根据另一参考行和该角度模式对当前块进行预测的第二预测结果,将所述第一预测结果和第二预测结果的加权和作为当前块最终的预测结果,该另一参考行为索引为0的参考行或该扩展参考行的相邻行。
本实施例的TMRL模式在构建候选列表和编解码时均可以使用前述实施例的方法,区别在于生成当前块的预测值的阶段,是将选中的角度模式与选中的参考行产生的预测值与选中的角度模式与另一参考行产生预测值加权以得到当前块最终的预测值。
本实施例可以TMRL模式下的IPF融合加以限制,例如,仅在当选中的参考行为靠近当前块的某些参考行时,才进行IPF融合。具体地,仅当参考行为reference line 1,或参考行为reference line 1,3,或参考行为reference line 1,3,5,或参考行为reference line 1,3,5,7,或参考行为reference line 1,3,5,7,12时允许融合。此外,也可以限定当前块使用某些角度模式时,不使用IPF,例如角度模式的角度为水平角度,垂直角度,-45度角度,45度角度,135度角度时不允许使用IPF。此外,在当前块的尺寸小于或等于某个尺寸时,由于小块通常有着更多纹理,也可以不允许使用IPF即融合预测。
在TMRL模式的基础上使用IPF,也可以在TMRL模式的语法元素中增加标识位表示是否使用IPF。
本公开一实施例提供一种视频解码方法,应用于解码器,如图13所示,所述方法包括:
步骤710,解码确定当前块使用基于模板的多参考行帧内预测TMRL模式的情况下,继续解码当前块的TMRL模式索引和TMRL融合标志;
步骤720,构建当前块的TMRL模式的候选列表,根据所述候选列表和TMRL模式索引确定当前块选中的扩展参考行和帧内预测模式;
本步骤可以采用与前述实施例相同的方式构建所述候选列表,及确定当前块选中的扩展参考行和帧内预测模式。
步骤730,在所述TMRL融合标志表示使用帧内预测融合IPF的情况下,将第一预测结果和第二预测结果的加权和作为当前块最终的预测结果;
本步骤中,第一预测结果和第二预测结果的加权和可以根据前述的公式一和公式二计算得到。
其中,所述第一预测结果是根据该扩展参考行和该帧内预测模式对当前块进行预测得到的,所述第二预测结果是根据另一参考行和该帧内预测模式对当前块进行预测得到的,该帧内预测模式为角度模式。两个预测结果加权时可以使用前述实施例的公式计算。
本实施例中,在所述TMRL融合标志表示不使用IPF时,根据该扩展参考行和该帧内预测模式对当前块进行预测,得到当前块最终的预测结果。
本实施例在TMRL模式的基础上,通过设置TMRL融合标志来表示是否使用IPF,在编码端可以按照前述实施例相同或相似的限制条件是否成立来编码TMRL融合标志,如确定使用IPF时,将TMRL融合标志置1,在确定不使用IPF时,将TMRL融合标志置0。而在解码端,无须再结合这些限制条件进行判断,根据该TMRL融合标志即可确定是否对使用TMRL模式时选中的角度模式进行IPF融合。可简化解码端的处理,提高IPF使用条件设置的灵活性。而在TMRL模式的基础上使用IPF,可以提高预测的准确性,提升视频编码的性能。
在本实施例的一示例中,所述另一参考行为与该扩展参考行相邻且位于该扩展参考行内侧的参考行,或者为与该扩展参考行相邻且全于该扩展参考行外侧的参考行,或者为索引为0的参考行。所述内侧是指靠近当前块的一侧,外侧是指远离当前块的一侧。
在本实施例的一示例中,可以增加解码TMRL融合标志的条件,即所述继续解码当前块的TMRL模式索引和TMRL融合标志,包括:在解码当前块的TMRL模式索引之后,判断当前块的尺寸是否满足设定条件,在满足设定条件时,再解码当前块的TMRL融合标志,在不满足设定条件时,跳过对当前块的TMRL融合标志的解码;
其中,所述设定条件包括以下条件中的任意一种或多种:
当前块的宽大于设定值;
当前块的高大于设定值;
当前块的宽乘高大于设定值;
当前块所属的当前帧非帧间帧;即禁止在B帧和P帧中出现使用融合模式的帧内块。
因为小尺寸的块通常有着更多纹理,使用IPF的提升有限,故本实施例将当前块的尺寸大于设定值作为解码TMRL融合标志的条件。在当前块的尺寸不大于设定值时,可以不对TMRL融合标志编解码以提高编码效率。
在本实施例的一示例中,所述构建当前块的TMRL模式的候选列表,包括:
根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种 组合,N≥1,M≥1,N×M≥2;
根据所述N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,1≤K≤N×M;
其中,所述M个帧内预测模式从角度为-45°、0°、45°、90°、135°的角度模式之外的其他角度模式中选出。
本实施例在构建TMRL模式的候选列表时,将参与组合的M个帧内预测模式限制在除一些特定角度外的其他角度模式,有利于与IPF的结合,以提升预测的准确度。
本公开一实施例提供一种视频解码方法,在解析CU级的语法元素阶段,需额外解码标识符tmrl_fusion_flag,即TMRL融合标志。
例如下表的示例,
Figure PCTCN2022106337-appb-000011
上表中的sps_mrl_enabled_flag为序列级标识,为1标识当前序列可以使用MRL,为1表示不使用MRL。(y0%CtbSizeY)>0表示当前CU位置不是CTU的第一行。tmrl_fusion_flag为1则表示后续的TMRL模式将使用融合模式,否则不使用融合模式。
在另一实施例中,是否解码tmrl_fusion_flag还可以基于当前编码块Coding block(Cb)大小。例如下表的示例中,当编码块的宽cbWidth>N并且编码块的高cbHeight>M时,才解码tmrl_fusion_flag
Figure PCTCN2022106337-appb-000012
Figure PCTCN2022106337-appb-000013
在又一实施例中,如下表的示例中,当编码块cbWidth*cbHeight>L时,才解码tmrl_fusion_flag。
Figure PCTCN2022106337-appb-000014
表中的M,N,L是尺寸的设定值,均为正整数,M可以等于N。
本实施例在基于模板排序构建TMRL模式的候选列表时,对可以使用的N个参考行和M个帧内预测模式进行NxM个组合的遍历,在模板区域生成预测值,并与模板区域的重建值求得误差,按照误差的大小按照升序选出前K个,构成{参考行,预测模式}组合的候选列表。误差值可以为Sum of absolute difference(SAD)等。本实施例可以限制只部分模式下使用IPF。在构建候选列表时,可以将以下帧内预测模式的一个或多个排除在外:PLANAR,DC,角度为水平角度的角度模式、角度为垂直角度的角度模式、角度为-45度角度的角度模式、角度为45度角度的角度模式、角度为135度角度的角度模式。
在预测阶段的融合需要基于tmrl_fusion_flag的取值。当tmrl_fusion_flag为1时,则选中的帧内预测模式与选中的参考行产生预测信号p a,选中的帧内预测模式与参考行reference line 0产生预测信号p b,p a与p b融合产生最终预测信号。否则直接将p a作为最终的预测信号。
相应地,本公开实施例还提供了一种视频编码方法,应用于编码器,如图18所示,所述方法包括:
步骤1110,构建当前块基于模板的多参考行帧内预测TMRL模式的候选列表,所述候选列表中填入有当前块候选的扩展参考行与帧内预测模式的组合;
步骤1120,通过率失真优化,为当前块选中一种参考行和帧内预测模式的组合;
步骤1130,在当前块的TMRL模式的编码条件满足时,编码当前块的TMRL模式标志以表示当前块使用TMRL模式,编码当前块的TMRL模式索引以表示所述选中的组合在所述候选列表中的位置;
步骤1140,编码当前块的TMRL融合标志以表示当前块使用帧内预测融合IPF或者不使用IPF;
其中,所述编码条件至少包括:所述选中的组合在所述候选列表中。
本实施例的视频编码方法与融合无关的部分可以与前述使用TMRL但不进行融合的视频编码方法的实施例相同,如步骤1130中编码条件的设置可以相同,等等。
在本实施例的一示例中,所述编码当前块的TMRL融合标志以表示当前块使用帧内预测融合IPF或者不使用IPF,包括:在当前块选中所述候选列表中包括角度模式的组合的情况下,如设定的限制条件中的至少一种成立时,编码当前块的TMRL融合标志以表示不使用IPF;如所述设定的限制条件均不成立时,编码当前块的TMRL融合标志以表示使用IPF,其中:
所述设定的限制条件还包括以下限制条件中的任意一种或多种:
当前块选中TIMD融合模式;
当前块选中DIMD模式;
当前块选中的参考行的索引大于等于K,K为大于等于3的整数;
当前块的宽小于或等于设定值;
当前块的高小于或等于设定值;
当前块的宽乘高小于或等于设定值;
当前块使用帧内子块划分ISP模式;
当前块选中的角度模式的角度为-45°、0°、45°、90°、135°中的任意一种,或当前块选中的角度模式为整数斜率的角度模式;
当前块所属的当前帧为帧间帧;
当前块所属的当前帧为色度帧。
在本实施例的一示例中,所述构建当前块的TMRL模式的候选列表时,参与组合的帧内预测模式排除角度为-45°、0°、45°、90°、135°的角度模式。
编码端在当前块选中TMRL模式的候选列表中的组合且使用IPF的情况下,基于该组合使用IPF对当前块进行预测,将根据该组合的结果和根据该组合中的角度模式与另一参考行的预测结果加权,得到当前块的预测值。进而可以确定当前块的重建值。
为了在TMRL的基础上使用IPF,也可以在构建TMRL的候选列表时即将基于扩展参考行和角度模式的组合得到的融合预测模式也作为候选列表中的一种组合。
本公开一实施例提供了一种多参考行帧内预测模式候选列表的构建方法,可应用于编码器,也可应用于解码器,如图14所示,所述方法包括:
步骤810,根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种原始组合;
步骤820,根据所述N×M种原始组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
步骤830,对误差最小的K种原始组合中每一种包括预定角度模式的组合,分别确定是否需要融合,如需要融合,将该原始组合对应的融合组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,如不需要融合,将该原始组合填入所述候选列表,其中,K,N,M为设定的正整数,1≤K≤N×M。
本实施例对误差最小的K种原始组合中每一种包括预定角度模式的组合,分别确定是否需要融合,如需要融合,将该原始组合对应的融合组合填入当前块的TMRL模式的候选列表,包括:
根据该原始组合对应的融合组合对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差,在该原始组合对应的误差大于所述融合组合对应的误差时,确定需要融合;在该原始组合对应的误差小于等于所述融合组合对应的误差时,确定不需要融合;
其中,基于该融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行;
其中,所述预定角度模式包括全部角度模式,或包括除整数斜率的角度模式之外的其他角度模式。
本实施例的第一预测结果和第二预测结果的加权和可以按照上述公式一或公式二计算。
本实施例排序不融合的模式选出K个组合作为候选列表。接着逐一对K个组合中的每一个组合的参考行是否需要融合做进一步确定。确定方法为根据当前组合选中的参考行在融合与不融合模式下生成预测信号与重建信号之间的误差大小,若融合模式下小,则该组合使用融合,否则则不融合。本实施例只需要进行5xM+K种组合在模板区域上的预测和误差值计算。
本公开另一实施例提供了一种多参考行帧内预测模式候选列表的构建方法,可应用于编码器,也可应用于解码器,如图15所示,所述方法包括:
步骤910,根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种原始组合;
步骤920,对所述N×M种原始组合进行融合处理,所述融合处理包括:对每一种包括预定角度模式的原始组合,在满足设定条件时,将该原始组合替换为对应的融合组合;
步骤930,根据融合处理后的N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
步骤940,按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表;
其中,K,N,M为设定的正整数,1≤K≤N×M;
其中,基于该原始组合对应的融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
本实施例的一示例中,所述设定条件包括以下条件中的任意一种或多种:当前块的尺寸大于NxM,N,M为设定的正整数;及,当前块选中的帧内预测模式不为整数斜率的角度模式;所述预定角度模式包括全部角度模式,或包括除整数斜率的角度模式之外的其他角度模式。
本实施例的第一预测结果和第二预测结果的加权和可以按照上述公式一或公式二计算。
本实施例不需要在模板上按照误差代价排序同时排序参考行融合和不融合。当满足下面条件的一个或多个,即使用相邻参考行的融合模式参与模板排序和生成组合{参考行,预测模式}列表;否则则不采用融合模式:
当当前块的尺寸大于NxM;
当当前帧内预测模式不是PLANAR,DC,水平角度,垂直角度,-45度角度,45度角度,135度角度模式
在预测阶段,当前块使用TMRL时,根据解码获得的tmrl_idx与基于模板排序生成的候选列表确定选中的参考行和预测模式,并根据相同的条件确定是否利用相邻参考行生成融合的预测结果。即在满足下面条件的一个或多个时,对选中的角度模式进行IPF融合:
当当前块的尺寸大于NxM;
当当前帧内预测模式不是PLANAR,DC,水平角度,垂直角度,-45度角度,45度角度,135度角度模式。
本公开又一实施例提供了一种多参考行帧内预测模式候选列表的构建方法,可应用于编码器,也可应用于解码器,如图16所示,所述方法包括:
步骤1010,根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与角度模式的N×M种原始组合;
步骤1020,根据所述N×M种原始组合中包括预定角度模式的组合得到对应的融合组合;
步骤1030,根据所述N×M种原始组合和得到的融合组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
步骤1040,按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,其中,K,N,M为设定的正整数,1≤K≤N×M。
本实施例根据所述N×M种原始组合中包括预定角度模式的组合得到对应的融合组合,包括:根据每一种原始组合得到一种融合组合;基于该融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和 角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行;
所述预定角度模式包括全部角度模式,或包括除整数斜率的角度模式之外的其他角度模式。
本实施例在参与组合的帧内预测模式均为角度模式时,需要进行5xMx2种组合在模板区域上的预测和误差值计算,
本实施例的第一预测结果和第二预测结果的加权和可以按照上述公式一或公式二计算。
本公开上述实施例中,以参与组合的参考行为索引为{1,3,5,7,12}的参考行为例,原始组合对应的融合组合例如可以是,原始组合为索引为1的参考行与某一角度模式的组合时,对应的融合组合可以为索引为1和2的两个参考行与该角度模式的组合,基于该融合组合对当前块预测时,会对根据索引为1的参考行和该角度模式的预测结果,以及索引为2的参考行和该角度模式的预测结果进行加权,得到当前块最终的预测结果。在另一示例中,第二参考行选择内侧的相邻参考行,则对应的融合组合可以为索引为0和1的两个参考行与该角度模式的组合。
依照以上实施例构建TMRL模式的候选列表的方法应用于视频编码和视频解码时,如果当前块选中的是候选列表中的融合组合,则按照该融合组合的预测方式对当前块进行融合预测,如果选中的是候选列表中的原始组合,则按照原始组合的预测方式对当前块进行预测,不使用IPF。
本公开一实施例还提供了一种码流,其中,所述码流通过本公开任一实施例所述的视频编码方法生成。
本公开一实施例还提供了一种多参考行帧内预测模式候选列表的构建装置,如图17所示,包括处理器71以及存储有计算机程序的存储器73,其中,所述处理器71执行所述计算机程序时能够实现本文任一实施例所述的多参考行帧内预测模式候选列表的构建方法。
本公开一实施例还提供了一种帧内预测融合装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现如本公开任一实施例所述的帧内预测融合方法。
本公开一实施例还提供了一种视频解码装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现本公开任一实施例所述的视频解码方法。
本公开一实施例还提供了一种视频编码装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现本公开任一实施例所述的视频编码方法。
本公开一实施例还提供了一种视频编解码系统,其中,包括本公开任一实施例所述的视频编码装置和本公开任一实施例所述的视频解码装置。
本公开一实施例还提供了一种非瞬态计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,其中,所述计算机程序时被处理器执行时实现本公开任一实施例所述的帧内预测融合方法,本公开任一实施例所述的多参考行帧内预测模式候选列表的构建方法,或实现本公开任一实施例所述的视频解码方法,或实现本公开任一实施例所述的视频编码方法。
本公开上述实施例的处理器可以是通用处理器,包括中央处理器(CPU)、网络处理器(Network Processor,简称NP)、微处理器等等,也可以是其他常规的处理器等;所述处理器还可以是数字信号处理器(DSP)、专用集成电路(ASIC)、现成可编程门阵列(FPGA)、离散逻辑或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件,或其它等效集成或离散的逻辑电路,也可以是上述器件的组合。即上述实施例的处理器可以是实现本发明实施例中公开的各方法、步骤及逻辑框图的任何处理器件或器件组合。如果部分地以软件来实施本公开实施例,那么可将用于软件的指令存储在合适的非易失性计算机可读存储媒体中,且可使用一个或多个处理器在硬件中执行所述指令从而实施本公开实施例的方法。本文中所使用的术语“处理器”可指上述结构或适合于实施本文中所描述的技术的任意其它结构。
在以上一个或多个示例性实施例中,所描述的功能可以硬件、软件、固件或其任一组合来实施。如果以软件实施,那么功能可作为一个或多个指令或代码存储在计算机可读介质上或经由计算机可读介质传输,且由基于硬件的处理单元执行。计算机可读介质可包含对应于例如数据存储介质等有形介质的计算机可读存储介质,或包含促进计算机程序例如根据通信协议从一处传送到另一处的任何介质 的通信介质。以此方式,计算机可读介质通常可对应于非暂时性的有形计算机可读存储介质或例如信号或载波等通信介质。数据存储介质可为可由一个或多个计算机或者一个或多个处理器存取以检索用于实施本公开中描述的技术的指令、代码和/或数据结构的任何可用介质。计算机程序产品可包含计算机可读介质。
举例来说且并非限制,此类计算机可读存储介质可包括RAM、ROM、EEPROM、CD-ROM或其它光盘存储装置、磁盘存储装置或其它磁性存储装置、快闪存储器或可用来以指令或数据结构的形式存储所要程序代码且可由计算机存取的任何其它介质。而且,还可以将任何连接称作计算机可读介质举例来说,如果使用同轴电缆、光纤电缆、双绞线、数字订户线(DSL)或例如红外线、无线电及微波等无线技术从网站、服务器或其它远程源传输指令,则同轴电缆、光纤电缆、双纹线、DSL或例如红外线、无线电及微波等无线技术包含于介质的定义中。然而应了解,计算机可读存储介质和数据存储介质不包含连接、载波、信号或其它瞬时(瞬态)介质,而是针对非瞬时有形存储介质。如本文中所使用,磁盘及光盘包含压缩光盘(CD)、激光光盘、光学光盘、数字多功能光盘(DVD)、软磁盘或蓝光光盘等,其中磁盘通常以磁性方式再生数据,而光盘使用激光以光学方式再生数据。上文的组合也应包含在计算机可读介质的范围内。
在一些方面中,本文描述的功能性可提供于经配置以用于编码和解码的专用硬件和/或软件模块内,或并入在组合式编解码器中。并且,可将所述技术完全实施于一个或多个电路或逻辑元件中。
本公开实施例的技术方案可在广泛多种装置或设备中实施,包含无线手机、集成电路(IC)或一组IC(例如,芯片组)。本公开实施例中描各种组件、模块或单元以强调经配置以执行所描述的技术的装置的功能方面,但不一定需要通过不同硬件单元来实现。而是,如上所述,各种单元可在编解码器硬件单元中组合或由互操作硬件单元(包含如上所述的一个或多个处理器)的集合结合合适软件和/或固件来提供。

Claims (45)

  1. 一种帧内预测融合方法,包括:
    在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用帧内预测融合IPF的限制条件是否成立;
    在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF。
  2. 如权利要求1所述的方法,其中:
    所述限制条件包括模式数量限制条件:使用IPF会导致当前块的预测需要融合N种以上的帧内预测模式,N为大于或等于3的整数。
  3. 如权利要求2所述的方法,其中:
    在N=3的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
    当前块选中基于模板的帧内模式推导TIMD融合模式且TIMD融合了满足IPF使用条件的两种角度模式的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时只允许对该两种角度模式中的一种进行IPF融合;
    当前块选中TIMD融合模式且TIMD融合的两种帧内预测模式中只有一种是满足IPF使用条件的角度模式的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对该角度模式进行IPF融合。
  4. 如权利要求1所述的方法,其中:
    所述限制条件包括模式数量限制条件:使用IPF会导致当前块的预测需要融合M种以上的角度模式,M为大于等于2的整数。
  5. 如权利要求4所述的方法,其中:
    在M=3的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
    当前块选中TIMD融合模式且TIMD融合了满足IPF使用条件的两种角度模式的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时只允许对该两种角度模式中的一种进行IPF融合;
    当前块选中TIMD融合模式且TIMD融合的两种帧内预测模式中只有一种是满足IPF使用条件的角度模式的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对该角度模式进行IPF融合。
  6. 如权利要求3或5所述的方法,其中:
    所述只允许对该两种角度模式中的一种进行IPF融合,包括:只允许对该两种角度模式中代价最小或代价次小的角度模式进行IPF融合。
  7. 如权利要求4所述的方法,其中:
    在M=3的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
    当前块选中解码端帧内预测模式导出DIMD融合模式且DIMD融合的两种角度模式均满足IPF使用条件的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时只允许对该两种角度模式中的一种进行IPF融合;
    在当前块使用DIMD融合模式且DIMD融合的两种角度模式中只有一种满足IPF使用条件的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对满足IPF使用条件的该角度模式进行IPF融合。
  8. 如权利要求4所述的方法,其中:
    在M=2的情况下,所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用IPF的限制条件是否成立;在所述限制条件中的至少一种成立时,对当前块进行帧内预测时限制使用IPF,包括:
    当前块选中TIMD融合模式且TIMD融合了两种角度模式的情况下,确定所述模式数量限制条件成立,对当前块进行帧内预测时不允许使用IPF;
    当前块选中TIMD融合模式、TIMD融合的两种帧内预测模式中只有一种是角度模式,且该角度模式满足IPF使用条件的情况下,确定所述模式数量限制条件不成立,对当前块进行帧内预测时允许对满足IPF使用条件的该角度模式进行IPF融合。
  9. 如权利要求3、5、7或8所述的方法,其中:
    所述满足IPF使用条件的角度模式不为整数斜率的角度模式。
  10. 如权利要求3、5、7或8所述的方法,其中:
    所述对当前块进行帧内预测时允许对一角度模式进行IPF融合,包括:将二种以上的预测结果的加权和作为当前块最终的预测结果;所述二种以上的预测结果包括:根据当前块选中的第一参考行和该角度模式对当前块进行预测得到的预测结果,及根据不同于所述第一参考行的第二参考行和该角度模式对当前块进行预测得到的预测结果。
  11. 如权利要求1所述的方法,其中:
    所述限制条件包括以下限制条件中的任意一种或多种:
    限制条件一:当前块选中TIMD融合模式;
    限制条件二:当前块选中DIMD融合模式;
    限制条件三:当前块使用多参考行MRL;
    限制条件四:当前块选中的参考行的索引大于等于K,K为大于等于3的整数;
    限制条件五:当前块的宽小于或等于设定值;
    限制条件五:当前块的高小于或等于设定值;
    限制条件七:当前块的宽乘高小于或等于设定值;
    限制条件八:当前块使用帧内子块划分ISP模式;
    限制条件九:当前块选中的角度模式的角度为-45°、0°、45°、90°、135°中的任意一种,或当前块选中的角度模式为整数斜率的角度模式;
    限制条件十:当前块所属的当前帧为帧间帧;
    限制条件十一:当前块所属的当前帧不为亮度帧;
    在所述限制条件一至限制条件十一中的至少一种成立的情况下,在当前块进行帧内预测时不允许使用IPF。
  12. 如权利要求11所述的方法,其中:
    所述限制条件包括限制条件四,其中,K=3或5或7或12。
  13. 如权利要求1所述的方法,其中:
    所述方法还包括:在所述限制条件均不成立的情况下,对当前块进行帧内预测时使用IPF;所述 对当前块进行帧内预测时使用IPF,包括:
    在当前块选中的帧内预测模式只有一种且为角度模式的情况下,将第一预测结果和第二预测结果的加权和作为当前块最终的预测结果;其中,所述第一预测结果是根据当前块选中的第一参考行和当前块选中的第一角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和所述第一角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
  14. 如权利要求1所述的方法,其中:
    所述方法还包括:在所述限制条件均不成立的情况下,对当前块进行帧内预测时使用IPF;所述对当前块进行帧内预测时使用IPF,包括:
    在当前块选中的帧内预测模式为融合模式的情况下,将三种以上预测结果的加权和作为当前块最终的预测结果;其中,所述三种以上预测结果包括:根据当前块选中的第一参考行和所述融合模式中的每一种帧内预测模式对当前块进行预测得到的多种预测结果,以及根据第二参考行和所述融合模式中允许进行IPF融合的一种或多种角度模式对当前块进行预测得到的一种或多种预测结果;所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
  15. 如权利要求13或14所述的方法,其中:
    在所述第一参考行的索引大于或等于K的情况下,所述第二参考行与所述第一参考行相邻且比所述第一参考行更为靠近当前块,K为大于或等于1的整数。
  16. 如权利要求13或14所述的方法,其中:
    所述第一预测结果和第二预测结果的加权和在计算时,为所述第一预测结果赋予的权重大于为第二预测结果赋予的权重。
  17. 如权利要求13所述的方法,其中:
    所述第一预测结果和第二预测结果的加权和按照以下公式计算:
    p fusion=(w ap a+w bp b)>>shift
    或者
    p fusion=(w ap a+w bp b+offset)>>shift
    其中,p a为所述第一预测结果,p b为所述第二预测结果,p fusion当前块最终的预测结果,w a为p a的权重,w b为p b的权重;
    其中,w a+w b=(1<<shift),offset=1<<(shift-1),shift>=1,shift和offset为设定的参数。
  18. 如权利要求13所述的方法,其中:
    所述在当前块选中的帧内预测模式包括角度模式的情况下,确定当前块使用帧内预测融合IPF的限制条件是否成立;在所述限制条件均不成立的情况下,对当前块进行帧内预测时使用IPF;包括:
    在当前块选中TMRL模式的候选列表中的一扩展参考行和一角度模式,且所述限制条件均不成立的情况下,计算根据该扩展参考行和该角度模式对当前块进行预测的第一预测结果,及根据另一参考行和该角度模式对当前块进行预测的第二预测结果,将所述第一预测结果和第二预测结果的加权和作为当前块最终的预测结果,其中,该另一参考行为索引为0的参考行或该扩展参考行的相邻行。
  19. 一种视频解码方法,包括:
    解码码流,确定当前块选中的参考行和帧内预测模式;
    按照如权利要求1至18中任一所述的帧内预测融合方法对当前块进行预测,得到当前块的预测值;
    根据当前块的预测值确定当前块的重建值。
  20. 一种视频编码方法,包括:
    通过模式选择确定当前块选中的参考行和帧内预测模式;
    按照如权利要求1至18中任一所述的帧内预测融合方法对当前块进行预测,得到当前块的预测值;
    根据当前块的原始值和预测值确定当前块的残差。
  21. 一种视频解码方法,包括:
    解码确定当前块使用基于模板的多参考行帧内预测TMRL模式的情况下,继续解码当前块的TMRL模式索引和TMRL融合标志;
    构建当前块的TMRL模式的候选列表,根据所述候选列表和TMRL模式索引确定当前块选中的扩展参考行和帧内预测模式;
    在所述TMRL融合标志表示使用帧内预测融合IPF的情况下,将第一预测结果和第二预测结果的加权和作为当前块最终的预测结果;
    其中,所述第一预测结果是根据该扩展参考行和该帧内预测模式对当前块进行预测得到的,所述第二预测结果是根据另一参考行和该帧内预测模式对当前块进行预测得到的,该帧内预测模式为角度模式。
  22. 如权利要求21所述的方法,其中:
    所述另一参考行为与该扩展参考行相邻且位于该扩展参考行内侧的参考行,或者为与该扩展参考行相邻且全于该扩展参考行外侧的参考行,或者为索引为0的参考行。
  23. 如权利要求21所述的方法,其中:
    所述方法还包括:在所述TMRL融合标志表示不使用IPF时,根据该扩展参考行和该帧内预测模式对当前块进行预测,得到当前块最终的预测结果。
  24. 如权利要求21所述的方法,其中:
    所述继续解码当前块的TMRL模式索引和TMRL融合标志,包括:在解码当前块的TMRL模式索引之后,判断当前块的尺寸是否满足设定条件,在满足设定条件时,再解码当前块的TMRL融合标志,在不满足设定条件时,跳过对当前块的TMRL融合标志的解码;
    其中,所述设定条件包括以下条件中的任意一种或多种:
    当前块的宽大于设定值;
    当前块的高大于设定值;
    当前块的宽乘高大于设定值;
    当前块所属的当前帧非帧间帧。
  25. 如权利要求21所述的方法,其中:
    所述构建当前块的TMRL模式的候选列表,包括:
    根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种组合,N≥1,M≥1,N×M≥2;
    根据所述N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
    按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,1≤K≤N×M;
    其中,所述M个帧内预测模式从角度为-45°、0°、45°、90°、135°的角度模式之外的其他角度模式中选出。
  26. 如权利要求21所述的方法,其中:
    所述第一预测结果和第二预测结果的加权和按照以下公式计算:
    p fusion=(w ap a+w bp b)>>shift
    或者
    p fusion=(w ap a+w bp b+offset)>>shift
    其中,p a为所述第一预测结果,p b为所述第二预测结果,p fusion当前块最终的预测结果,w a为p a的权重,w b为p b的权重;
    其中,w a+w b=(1<<shift),offset=1<<(shift-1),shift>=1,shift和offset为设定的参数。
  27. 一种视频编码方法,包括:
    构建当前块基于模板的多参考行帧内预测TMRL模式的候选列表,所述候选列表中填入有当前块候选的扩展参考行与帧内预测模式的组合;通过率失真优化,为当前块选中一种参考行和帧内预测模式的组合;
    在当前块的TMRL模式的编码条件满足时,编码当前块的TMRL模式标志以表示当前块使用TMRL模式,编码当前块的TMRL模式索引以表示所述选中的组合在所述候选列表中的位置;
    编码当前块的TMRL融合标志以表示当前块使用帧内预测融合IPF或者不使用IPF;
    其中,所述编码条件至少包括:所述选中的组合在所述候选列表中。
  28. 如权利要求27所述的方法,其中:
    所述编码当前块的TMRL融合标志以表示当前块使用帧内预测融合IPF或者不使用IPF,包括:在当前块选中所述候选列表中包括角度模式的组合的情况下,如设定的限制条件中的至少一种成立时,编码当前块的TMRL融合标志以表示不使用IPF;如所述设定的限制条件均不成立时,编码当前块的TMRL融合标志以表示使用IPF,其中:
    所述设定的限制条件还包括以下限制条件中的任意一种或多种:
    当前块选中TIMD融合模式;
    当前块选中DIMD模式;
    当前块选中的参考行的索引大于等于K,K为大于等于3的整数;
    当前块的宽小于或等于设定值;
    当前块的高小于或等于设定值;
    当前块的宽乘高小于或等于设定值;
    当前块使用帧内子块划分ISP模式;
    当前块选中的角度模式的角度为-45°、0°、45°、90°、135°中的任意一种,或当前块选中的角度模式为整数斜率的角度模式;
    当前块所属的当前帧为帧间帧;
    当前块所属的当前帧为色度帧。
  29. 如权利要求27所述的方法,其中:
    所述构建当前块的TMRL模式的候选列表时,参与组合的帧内预测模式排除角度为-45°、0°、45°、90°、135°的角度模式。
  30. 一种多参考行帧内预测模式候选列表的构建方法,包括:
    根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与角度模式的N×M种原始组合;
    根据所述N×M种原始组合中包括预定角度模式的组合得到对应的融合组合;
    根据所述N×M种原始组合和得到的融合组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
    按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,其中,K,N,M为设定的正整数,1≤K≤N×M。
  31. 如权利要求30所述的方法,其中:
    所述根据所述N×M种原始组合中包括预定角度模式的组合得到对应的融合组合,包括:
    根据每一种包括预定角度模式的原始组合得到一种融合组合;基于该融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行;
    所述预定角度模式包括全部角度模式,或包括除整数斜率的角度模式之外的其他角度模式。
  32. 如权利要求30所述的方法,其中:
    所述第一预测结果和第二预测结果的加权和按照以下公式计算:
    p fusion=(w ap a+w bp b)>>shift
    或者
    p fusion=(w ap a+w bp b+offset)>>shift
    其中,p a为所述第一预测结果,p b为所述第二预测结果,p fusion当前块最终的预测结果,w a为p a的权重,w b为p b的权重;
    其中,w a+w b=(1<<shift),offset=1<<(shift-1),shift>=1,shift和offset为设定的参数。
  33. 一种多参考行帧内预测模式候选列表的构建方法,包括:
    根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种原始组合;
    根据所述N×M种原始组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
    对误差最小的K种原始组合中每一种包括预定角度模式的组合,分别确定是否需要融合,如需要融合,将该原始组合对应的融合组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表,如不需要融合,将该原始组合填入所述候选列表,其中,K,N,M为设定的正整数,1≤K≤N×M。
  34. 如权利要求33所述的方法,其中:
    所述对误差最小的K种原始组合中每一种包括预定角度模式的组合,分别确定是否需要融合,如需要融合,将该原始组合对应的融合组合填入当前块的TMRL模式的候选列表,包括:
    根据该原始组合对应的融合组合对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差,在该原始组合对应的误差大于所述融合组合对应的误差时,确定需要融合;在该原始组合对应的误差小于等于所述融合组合对应的误差时,确定不需要融合;
    其中,基于该融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行;
    其中,所述预定角度模式包括全部角度模式,或包括除整数斜率的角度模式之外的其他角度模式。
  35. 如权利要求33所述的方法,其中:
    所述第一预测结果和第二预测结果的加权和按照以下公式计算:
    p fusion=(w ap a+w bp b)>>shift
    或者
    p fusion=(w ap a+w bp b+offset)>>shift
    其中,p a为所述第一预测结果,p b为所述第二预测结果,p fusion当前块最终的预测结果,w a为p a的权重,w b为p b的权重;
    其中,w a+w b=(1<<shift),offset=1<<(shift-1),shift>=1,shift和offset为设定的参数。
  36. 一种多参考行帧内预测模式候选列表的构建方法,包括:
    根据当前块的N个扩展参考行和M个帧内预测模式得到扩展参考行与帧内预测模式的N×M种原始组合;
    对所述N×M种原始组合进行融合处理,所述融合处理包括:对每一种包括预定角度模式的原始组合,在满足设定条件时,将该原始组合替换为对应的融合组合;
    根据融合处理后的N×M种组合分别对当前块的模板区域进行预测,并计算所述模板区域的重建值和预测得到的预测值之间的误差;
    按照所述误差从小到大的顺序,将所述误差对应的K种组合填入当前块基于模板的多参考行帧内预测TMRL模式的候选列表;
    其中,K,N,M为设定的正整数,1≤K≤N×M;
    其中,基于该原始组合对应的融合组合对当前块预测时,是将第一预测结果和第二预测结果的加权和作为当前块的预测值;其中,所述第一预测结果是根据该原始组合中的扩展参考行和角度模式对当前块进行预测的结果,所述第二预测结果是根据第二参考行和该原始组合中的角度模式对当前块进行预测的结果,所述第二参考行是所述第一参考行的相邻行或者是索引为0的参考行。
  37. 如权利要求36所述的方法,其中:
    所述设定条件包括以下条件中的任意一种或多种:
    当前块的尺寸大于NxM,N,M为设定的正整数;
    当前块选中的帧内预测模式不为整数斜率的角度模式;
    所述预定角度模式包括全部角度模式,或包括除整数斜率的角度模式之外的其他角度模式。
  38. 如权利要求36所述的方法,其中:
    所述第一预测结果和第二预测结果的加权和按照以下公式计算:
    p fusion=(w ap a+w bp b)>>shift
    或者
    p fusion=(w ap a+w bp b+offset)>>shift
    其中,p a为所述第一预测结果,p b为所述第二预测结果,p fusion当前块最终的预测结果,w a为p a的权重,w b为p b的权重;
    其中,w a+w b=(1<<shift),offset=1<<(shift-1),shift>=1,shift和offset为设定的参数。
  39. 一种码流,其中,所述码流通过如权利要求20、27至29中任一所述的视频编码方法生成。
  40. 一种帧内预测融合装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现如权利要求1至18中任一所述的帧内预测融合方法。
  41. 一种多参考行帧内预测模式候选列表的构建装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现如权利要求30至38中任一所述的多参考行帧内预测模式候选列表的构建方法。
  42. 一种视频解码装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现如权利要求19、21至26中任一所述的视频解码方法。
  43. 一种视频编码装置,包括处理器以及存储有计算机程序的存储器,其中,所述处理器执行所述计算机程序时能够实现如权利要求20、27至29中任一所述的视频编码方法。
  44. 一种视频编解码系统,其中,包括如权利要求36所述的视频编码装置和如权利要求35所述的视频解码装置。
  45. 一种非瞬态计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,其中,所述计算机程序时被处理器执行时能够实现如权利要求1至18中任一所述的帧内预测融合方法,或实现如权利要求19、21至26中任一所述的视频解码方法,或实现如权利要求20、27至29中任一所述的视频编码方法,或实现如权利要求30至38中任一所述的多参考行帧内预测模式候选列表的构建方法。
PCT/CN2022/106337 2022-07-08 2022-07-18 一种帧内预测融合方法、视频编解码方法、装置和系统 Ceased WO2024007366A1 (zh)

Priority Applications (6)

Application Number Priority Date Filing Date Title
KR1020257003534A KR20250035555A (ko) 2022-07-08 2022-07-18 인트라 예측 융합 방법, 비디오 인코딩 및 디코딩 방법, 장치 및 시스템
JP2024577022A JP2025521765A (ja) 2022-07-08 2022-07-18 イントラ予測融合方法、ビデオエンコーディング方法及び装置、ビデオデコーディング方法及び装置、並びにビデオコーディングシステム
CN202280097794.0A CN119487831A (zh) 2022-07-08 2022-07-18 一种帧内预测融合方法、视频编解码方法、装置和系统
TW112125248A TW202406347A (zh) 2022-07-08 2023-07-06 一種幀內預測融合方法、視訊編解碼方法、碼流、裝置、系統和儲存媒介
MX2024016089A MX2024016089A (es) 2022-07-08 2024-12-18 Metodo y aparato para la fusion de intraprediccion, metodo y aparato codificador de video, metodo y aparato decodificador de video, sistema codificador y decodificador de video
US18/990,493 US20250126291A1 (en) 2022-07-08 2024-12-20 Method for intra prediction fusion and non-transitory computer-readable storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202210806435 2022-07-08
CN202210806435.X 2022-07-08

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US18/990,493 Continuation US20250126291A1 (en) 2022-07-08 2024-12-20 Method for intra prediction fusion and non-transitory computer-readable storage medium

Publications (1)

Publication Number Publication Date
WO2024007366A1 true WO2024007366A1 (zh) 2024-01-11

Family

ID=89454664

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/106337 Ceased WO2024007366A1 (zh) 2022-07-08 2022-07-18 一种帧内预测融合方法、视频编解码方法、装置和系统

Country Status (7)

Country Link
US (1) US20250126291A1 (zh)
JP (1) JP2025521765A (zh)
KR (1) KR20250035555A (zh)
CN (1) CN119487831A (zh)
MX (1) MX2024016089A (zh)
TW (1) TW202406347A (zh)
WO (1) WO2024007366A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240357078A1 (en) * 2023-04-14 2024-10-24 Qualcomm Incorporated Fusion for template matching based on filtering or position in video coding

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190215521A1 (en) * 2016-09-22 2019-07-11 Mediatek Inc. Method and apparatus for video coding using decoder side intra prediction derivation
US20190379891A1 (en) * 2016-10-14 2019-12-12 Industry Academy Cooperation Foundation Of Sejong University Method and apparatus for encoding/decoding an image
CN111295881A (zh) * 2017-11-13 2020-06-16 联发科技(新加坡)私人有限公司 用于图像和视频编解码的画面内预测融合的方法和装置
CN112956192A (zh) * 2018-10-31 2021-06-11 交互数字Vc控股公司 多参考行帧内预测和最可能的模式
CN113170126A (zh) * 2018-11-08 2021-07-23 Oppo广东移动通信有限公司 视频信号编码/解码方法以及用于所述方法的设备
WO2022116317A1 (zh) * 2020-12-03 2022-06-09 Oppo广东移动通信有限公司 帧内预测方法、编码器、解码器以及存储介质

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116962721A (zh) * 2016-05-04 2023-10-27 微软技术许可有限责任公司 利用样本值的非相邻参考线进行帧内图片预测的方法
EP3567860A1 (en) * 2018-05-09 2019-11-13 InterDigital VC Holdings, Inc. Method and apparatus for blended intra prediction
US11838498B2 (en) * 2021-06-28 2023-12-05 Tencent America LLC Harmonized design for intra bi-prediction and multiple reference line selection
EP4548577A1 (en) * 2022-06-30 2025-05-07 InterDigital CE Patent Holdings, SAS Methods and apparatuses for encoding and decoding an image or a video using combined intra modes

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190215521A1 (en) * 2016-09-22 2019-07-11 Mediatek Inc. Method and apparatus for video coding using decoder side intra prediction derivation
US20190379891A1 (en) * 2016-10-14 2019-12-12 Industry Academy Cooperation Foundation Of Sejong University Method and apparatus for encoding/decoding an image
CN111295881A (zh) * 2017-11-13 2020-06-16 联发科技(新加坡)私人有限公司 用于图像和视频编解码的画面内预测融合的方法和装置
CN112956192A (zh) * 2018-10-31 2021-06-11 交互数字Vc控股公司 多参考行帧内预测和最可能的模式
CN113170126A (zh) * 2018-11-08 2021-07-23 Oppo广东移动通信有限公司 视频信号编码/解码方法以及用于所述方法的设备
WO2022116317A1 (zh) * 2020-12-03 2022-06-09 Oppo广东移动通信有限公司 帧内预测方法、编码器、解码器以及存储介质

Also Published As

Publication number Publication date
KR20250035555A (ko) 2025-03-12
JP2025521765A (ja) 2025-07-10
TW202406347A (zh) 2024-02-01
MX2024016089A (es) 2025-02-10
US20250126291A1 (en) 2025-04-17
CN119487831A (zh) 2025-02-18

Similar Documents

Publication Publication Date Title
CN110393010B (zh) 视频译码中的帧内滤波旗标
TWI745594B (zh) 與視訊寫碼中之變換處理一起應用之內部濾波
CN117956149A (zh) 自适应环路滤波器
CN110650337B (zh) 一种图像编码方法、解码方法、编码器、解码器及存储介质
JP2020537468A (ja) ビデオコーディングのための空間変動変換
US20250126291A1 (en) Method for intra prediction fusion and non-transitory computer-readable storage medium
WO2024145857A1 (zh) 帧内模板匹配预测方法、视频编解码方法、装置和系统
US20250159252A1 (en) Video coding method, apparatus and system
US20250260807A1 (en) Video encoding method, and video decoding method and apparatus
CN112470471A (zh) 用于视频译码的受约束编码树
TW202404366A (zh) 多參考行索引列表排序方法、視訊編解碼方法、裝置、系統和儲存媒介
WO2024174253A1 (zh) 基于插值滤波的帧内预测、视频编解码方法、装置和系统
CN119213776A (zh) 一种环路滤波方法、视频编解码方法、装置和系统
RU2858702C2 (ru) Способ слияния внутреннего предсказания (ipf) (варианты) и энергонезависимый машиночитаемый носитель данных
TW202005370A (zh) 視訊編碼及解碼
WO2024145851A1 (zh) 帧内模板匹配预测方法、视频编解码方法、装置和系统
CN121099052A (zh) 跨分量预测方法、视频编解码方法、装置、介质及设备
HK40009762B (zh) 視頻譯碼中的幀內濾波旗標
HK40009762A (zh) 視頻譯碼中的幀內濾波旗標

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22949946

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: MX/A/2024/016089

Country of ref document: MX

WWE Wipo information: entry into national phase

Ref document number: 2024577022

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 202280097794.0

Country of ref document: CN

WWE Wipo information: entry into national phase

Ref document number: 202517008707

Country of ref document: IN

ENP Entry into the national phase

Ref document number: 20257003534

Country of ref document: KR

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 1020257003534

Country of ref document: KR

WWE Wipo information: entry into national phase

Ref document number: 2025102480

Country of ref document: RU

WWP Wipo information: published in national office

Ref document number: MX/A/2024/016089

Country of ref document: MX

NENP Non-entry into the national phase

Ref country code: DE

WWP Wipo information: published in national office

Ref document number: 202280097794.0

Country of ref document: CN

WWP Wipo information: published in national office

Ref document number: 202517008707

Country of ref document: IN

WWP Wipo information: published in national office

Ref document number: 2025102480

Country of ref document: RU

WWP Wipo information: published in national office

Ref document number: 1020257003534

Country of ref document: KR

122 Ep: pct application non-entry in european phase

Ref document number: 22949946

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 2025102480

Country of ref document: RU