WO2025007250A1 - 视频解码方法、装置、设备、及存储介质 - Google Patents

视频解码方法、装置、设备、及存储介质 Download PDF

Info

Publication number
WO2025007250A1
WO2025007250A1 PCT/CN2023/105570 CN2023105570W WO2025007250A1 WO 2025007250 A1 WO2025007250 A1 WO 2025007250A1 CN 2023105570 W CN2023105570 W CN 2023105570W WO 2025007250 A1 WO2025007250 A1 WO 2025007250A1
Authority
WO
WIPO (PCT)
Prior art keywords
motion information
block
sub
reference image
current block
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/105570
Other languages
English (en)
French (fr)
Inventor
王凡
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Guangdong Oppo Mobile Telecommunications Corp Ltd
Original Assignee
Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Guangdong Oppo Mobile Telecommunications Corp Ltd filed Critical Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority to CN202380100003.XA priority Critical patent/CN121511598A/zh
Priority to PCT/CN2023/105570 priority patent/WO2025007250A1/zh
Publication of WO2025007250A1 publication Critical patent/WO2025007250A1/zh
Priority to MX2025015514A priority patent/MX2025015514A/es
Priority to US19/428,443 priority patent/US20260113474A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/44Decoders specially adapted therefor, e.g. video decoders which are asymmetric with respect to the encoder
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • H04N19/137Motion inside a coding unit, e.g. average field, frame or block difference
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation
    • H04N19/513Processing of motion vectors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation
    • H04N19/513Processing of motion vectors
    • H04N19/517Processing of motion vectors by encoding
    • H04N19/52Processing of motion vectors by encoding by predictive encoding

Definitions

  • the present application relates to the field of video encoding and decoding technology, and in particular to a video decoding method, device, equipment, and storage medium.
  • Digital video technology can be incorporated into a variety of video devices, such as digital televisions, smart phones, computers, e-readers or video players, etc. With the development of video technology, the amount of data included in video data is large. In order to facilitate the transmission of video data, video devices implement video compression technology to make video data more efficiently transmitted or stored.
  • prediction can eliminate or reduce the redundancy in the video and improve the compression efficiency.
  • the decoding end improves the motion information determined by the decoding.
  • the current motion information improvement method has a poor improvement effect, resulting in inaccurate prediction at the decoding end, which in turn affects the decoding performance of the video.
  • the embodiments of the present application provide a video decoding method, apparatus, device, and storage medium, which can improve the improvement effect of motion information, the prediction accuracy of the current block and the decoding performance of the video.
  • the present application provides a video decoding method, applied to a decoder, comprising:
  • a prediction value of the current block is determined.
  • the present application provides a video decoding device, which is used to execute the method in the first aspect or its respective implementations.
  • the device includes a functional unit for executing the method in the first aspect or its respective implementations.
  • a video decoder comprising a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its implementations.
  • a video coding and decoding system including a video encoder and a video decoder.
  • the video decoder is used to execute the method in the first aspect or its implementation manners.
  • a chip for implementing the method of the first aspect.
  • the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes the method of the first aspect.
  • a computer-readable storage medium for storing a computer program, wherein the computer program enables a computer to execute the method of the first aspect.
  • a computer program product comprising computer program instructions, wherein the computer program instructions enable a computer to execute the method of the first aspect.
  • the decoding end when decoding the current block, the decoding end first determines the first motion information of the current block, and then improves the first motion information based on the motion information of the reference image of the current block to obtain the second motion information. That is to say, in the embodiment of the present application, when improving the first motion information, the motion information of the reference image is taken into account to achieve effective improvement of the first motion information and obtain accurate second motion information. Then, when determining the prediction value of the current block based on the accurate second motion information, the prediction accuracy of the current block can be improved, thereby improving the decoding performance of the video.
  • FIG1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present application.
  • FIG2 is a schematic block diagram of a video encoder according to an embodiment of the present application.
  • FIG3 is a schematic block diagram of a video decoder according to an embodiment of the present application.
  • FIG4 is a schematic diagram of a GOP structure
  • FIG5 is a schematic diagram of CU division
  • FIG6 is a schematic diagram of a spatial domain and a temporal domain block
  • FIG7 is a schematic diagram of deriving time domain motion information
  • FIG8 is a schematic diagram of the principle of SbTMVP
  • FIG9 is a schematic diagram of the principle of MMVD
  • FIG10A is a schematic diagram of a principle of affine
  • FIG10B is another schematic diagram of the principle of affine
  • FIG11 is a schematic diagram of weight allocation
  • FIG12 is a schematic diagram of the principle of DMVR
  • FIG13 is a schematic diagram of template matching
  • FIG14 is a schematic diagram of a video decoding method flow chart provided by an embodiment of the present application.
  • FIG15 is a schematic diagram of a current block and a reference block
  • FIG16 is a schematic diagram showing the motion of the current block and two reference blocks
  • FIG17A is a schematic diagram of a motion when the motion information of the current block is greater than the motion information of the reference block;
  • FIG17B is a schematic diagram of a motion when the motion information of the current block is smaller than the motion information of the reference block;
  • FIG18A is another motion diagram when the motion information of the current block is smaller than the motion information of the reference block
  • FIG18B is another motion diagram when the motion information of the current block is greater than the motion information of the reference block
  • FIG19A is a schematic diagram of a motion information search
  • FIG19B is another schematic diagram of motion information search
  • FIG19C is a schematic diagram of another motion information search
  • FIG20 is a schematic diagram showing different motion information in a reference block
  • FIG21 is a schematic block diagram of a video decoding device provided by an embodiment of the present application.
  • FIG. 22 is a schematic block diagram of an electronic device provided in an embodiment of the present application.
  • the present application can be applied to the field of image coding and decoding, the field of video coding and decoding, the field of hardware video coding and decoding, the field of dedicated circuit video coding and decoding, the field of real-time video coding and decoding, etc.
  • AVC H.264/audio and video coding
  • HEVC H.265/high efficiency video coding
  • VVC VVC
  • the solution of the present application may be combined with other proprietary or industry standards and operate, the standards include ITU-TH.261, ISO/IEC MPEG-1 Visual, ITU-TH.262 or ISO/IEC MPEG-2 Visual, ITU-TH.263, ISO/IEC MPEG-4 Visual, ITU-TH.264 (also known as ISO/IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions.
  • SVC scalable video coding
  • MVC multi-view video coding
  • FIG1 is a schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application. It should be noted that FIG1 is only an example, and the video encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in FIG1.
  • the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120.
  • the encoding device is used to encode (which can be understood as compression) the video data to generate a code stream, and transmit the code stream to the decoding device.
  • the decoding device decodes the code stream generated by the encoding device to obtain decoded video data.
  • the encoding device 110 of the embodiment of the present application can be understood as a device with a video encoding function
  • the decoding device 120 can be understood as a device with a video decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, vehicle-mounted computers, etc.
  • the encoding device 110 may transmit the encoded video data (eg, a code stream) to the decoding device 120 via the channel 130.
  • the channel 130 may include one or more media and/or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.
  • the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded video data directly to the decoding device 120 in real time.
  • the encoding device 110 can modulate the encoded video data according to the communication standard and transmit the modulated video data to the decoding device 120.
  • the communication medium includes a wireless communication medium, such as a radio frequency spectrum, and optionally, the communication medium may also include a wired communication medium, such as one or more physical transmission lines.
  • the channel 130 includes a storage medium, which can store the video data encoded by the encoding device 110.
  • the storage medium includes a variety of locally accessible data storage media, such as optical disks, DVDs, flash memories, etc.
  • the decoding device 120 can obtain the encoded video data from the storage medium.
  • the channel 130 may include a storage server that can store the video data encoded by the encoding device 110.
  • the decoding device 120 can download the stored encoded video data from the storage server.
  • the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.
  • FTP file transfer protocol
  • the encoding device 110 includes a video encoder 112 and an output interface 113.
  • the output interface 113 may include a modulator/demodulator (modem) and/or a transmitter.
  • the encoding device 110 may further include a video source 111 in addition to the video encoder 112 and the input interface 113 .
  • the video source 111 may include at least one of a video acquisition device (eg, a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.
  • a video acquisition device eg, a video camera
  • a video archive e.g., a video archive
  • a video input interface e.g., a computer graphics system
  • the video input interface is used to receive video data from a video content provider
  • the computer graphics system is used to generate video data.
  • the video encoder 112 encodes the video data from the video source 111 to generate a bitstream.
  • the video data may include one or more pictures or a sequence of pictures.
  • the bitstream contains the encoding information of the picture or the sequence of pictures in the form of a bitstream.
  • the encoding information may include the encoded picture data and associated data.
  • the associated data may include a sequence parameter set (SPS for short), a picture parameter set (PPS for short) and other syntax structures.
  • SPS sequence parameter set
  • PPS picture parameter set
  • the syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
  • the video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113.
  • the encoded video data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.
  • the decoding device 120 includes an input interface 121 and a video decoder 122 .
  • the decoding device 120 may include a display device 123 in addition to the input interface 121 and the video decoder 122 .
  • the input interface 121 includes a receiver and/or a modem.
  • the input interface 121 can receive the encoded video data through the channel 130 .
  • the video decoder 122 is used to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to the display device 123 .
  • the display device 123 displays the decoded video data.
  • the display device 123 may be integrated with the decoding device 120 or external to the decoding device 120.
  • the display device 123 may include a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
  • FIG1 is only an example, and the technical solution of the embodiment of the present application is not limited to FIG1 .
  • the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.
  • FIG2 is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 can be used to perform lossy compression on an image, or can be used to perform lossless compression on an image.
  • the lossless compression can be visually lossless compression or mathematically lossless compression.
  • the video encoder 200 can be applied to image data in luminance and chrominance (YCbCr, YUV) format.
  • the YUV ratio can be 4:2:0, 4:2:2 or 4:4:4, Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation.
  • 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr)
  • 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr)
  • 4:4:4 means full pixel display (YYYYCbCrCbCrCbCrCbCr).
  • the video encoder 200 reads video data, and for each frame of the video data, divides the frame into a number of coding tree units (CTUs).
  • CTB may be referred to as a "tree block", “largest coding unit” (LCU) or “coding tree block” (CTB).
  • Each CTU may be associated with a pixel block of equal size within the image.
  • Each pixel may correspond to a luminance (luminance or luma) sample and two chrominance (chrominance or chroma) samples. Therefore, each CTU may be associated with a luminance sample block and two chrominance sample blocks.
  • the size of a CTU is, for example, 128 ⁇ 128, 64 ⁇ 64, 32 ⁇ 32, etc.
  • a CTU may be further divided into a number of coding units (CUs) for encoding, and a CU may be a rectangular block or a square block.
  • CU can be further divided into prediction unit (PU) and transform unit (TU), which makes encoding, prediction and transform separated and more flexible in processing.
  • PU prediction unit
  • TU transform unit
  • CTU is divided into CU in quadtree mode
  • CU is divided into TU and PU in quadtree mode.
  • the video encoder and video decoder may support various PU sizes. Assuming that the size of a particular CU is 2N ⁇ 2N, the video encoder and video decoder may support PU sizes of 2N ⁇ 2N or N ⁇ N for intra-frame prediction, and support symmetric PUs of 2N ⁇ 2N, 2N ⁇ N, N ⁇ 2N, N ⁇ N or similar sizes for inter-frame prediction. The video encoder and video decoder may also support asymmetric PUs of 2N ⁇ nU, 2N ⁇ nD, nL ⁇ 2N, and nR ⁇ 2N for inter-frame prediction.
  • the video encoder 200 may include: a prediction unit 210, a residual unit 220, a transform/quantization unit 230, an inverse transform/quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy coding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.
  • the current block may be referred to as a current coding unit (CU) or a current prediction unit (PU), etc.
  • a prediction block may also be referred to as a prediction image block or an image prediction block, and a reconstructed image block may also be referred to as a reconstructed block or an image reconstructed image block.
  • the prediction unit 210 includes an inter-frame prediction unit 211 and an intra-frame estimation unit 212. Since there is a strong correlation between adjacent pixels in a frame of a video, an intra-frame prediction method is used in the video coding and decoding technology to eliminate spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent frames in a video, an inter-frame prediction method is used in the video coding and decoding technology to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency.
  • the inter-frame prediction unit 211 can be used for inter-frame prediction.
  • Inter-frame prediction can include motion estimation and motion compensation. It can refer to the image information of different frames.
  • Inter-frame prediction uses motion information to find reference blocks from reference frames, and generates prediction blocks based on the reference blocks to eliminate temporal redundancy.
  • the frames used for inter-frame prediction can be P frames and/or B frames. P frames refer to forward prediction frames, and B frames refer to bidirectional prediction frames.
  • Inter-frame prediction uses motion information to find reference blocks from reference frames, and generates prediction blocks based on the reference blocks.
  • Motion information includes a reference frame list where the reference frame is located, a reference frame index, and a motion vector.
  • the motion vector can be an integer pixel or a sub-pixel.
  • the motion vector is a sub-pixel
  • the integer pixel or sub-pixel block in the reference frame found according to the motion vector is called a reference block.
  • Some technologies will directly use the reference block as a prediction block, and some technologies will generate a prediction block based on the reference block. Reprocessing the reference block to generate a prediction block can also be understood as taking the reference block as a prediction block and then processing the prediction block to generate a new prediction block.
  • the intra-frame estimation unit 212 only refers to the information of the same frame image to predict the pixel information in the current code image block to eliminate spatial redundancy.
  • the frame used for intra-frame prediction can be an I frame.
  • the intra-frame prediction modes used by HEVC are Planar, DC, and 33 angle modes, for a total of 35 prediction modes.
  • the intra-frame modes used by VVC are Planar, DC, and 65 angle modes, for a total of 67 prediction modes.
  • the residual unit 220 may generate a residual block of the CU based on the pixel blocks of the CU and the prediction blocks of the PUs of the CU. For example, the residual unit 220 may generate a residual block of the CU so that each sample in the residual block has a value equal to the difference between the following two: a sample in the pixel blocks of the CU and a corresponding sample in the prediction blocks of the PUs of the CU.
  • the transform/quantization unit 230 may quantize the transform coefficients.
  • the transform/quantization unit 230 may quantize the transform coefficients associated with the TUs of the CU based on a quantization parameter (QP) value associated with the CU.
  • QP quantization parameter
  • the video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
  • the inverse transform/quantization unit 240 may apply inverse quantization and inverse transform to the quantized transform coefficients, respectively, to reconstruct a residual block from the quantized transform coefficients.
  • the reconstruction unit 250 may add the samples of the reconstructed residual block to the corresponding samples of one or more prediction blocks generated by the prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of the CU in this manner, the video encoder 200 may reconstruct the pixel blocks of the CU.
  • the loop filter unit 260 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent coded pixels. For example, a deblocking filter operation may be performed to reduce the blocking effect of the pixel blocks associated with the CU.
  • the loop filter unit 260 includes a deblocking filter unit and a sample adaptive offset/adaptive loop filter (SAO/ALF) unit, wherein the deblocking filter unit is used to remove the block effect, and the SAO/ALF unit is used to remove the ringing effect.
  • SAO/ALF sample adaptive offset/adaptive loop filter
  • the decoded image buffer 270 may store the reconstructed pixel blocks.
  • the inter prediction unit 211 may use the reference image containing the reconstructed pixel blocks to perform inter prediction on PUs of other images.
  • the intra estimation unit 212 may use the reconstructed pixel blocks in the decoded image buffer 270 to perform intra prediction on other PUs in the same image as the CU.
  • the entropy coding unit 280 may receive the quantized transform coefficients from the transform/quantization unit 230.
  • the entropy coding unit 280 may perform One or more entropy encoding operations are performed to generate entropy coded data.
  • FIG. 3 is a schematic block diagram of a video decoder according to an embodiment of the present application.
  • the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization/transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded image buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.
  • the video decoder 300 may receive a bitstream.
  • the entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse the syntax elements in the bitstream that have been entropy encoded.
  • the prediction unit 320, the inverse quantization/transformation unit 330, the reconstruction unit 340, and the loop filter unit 350 may decode the video data according to the syntax elements extracted from the bitstream, that is, generate decoded video data.
  • the prediction unit 320 includes an intra estimation unit 322 and an inter prediction unit 321 .
  • the intra estimation unit 322 may perform intra prediction to generate a prediction block for the PU.
  • the intra estimation unit 322 may use an intra prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs.
  • the intra estimation unit 322 may also determine the intra prediction mode for the PU according to one or more syntax elements parsed from the code stream.
  • the inter prediction unit 321 may construct a first reference image list (list 0) and a second reference image list (list 1) according to the syntax elements parsed from the code stream.
  • the entropy decoding unit 310 may parse the motion information of the PU.
  • the inter prediction unit 321 may determine one or more reference blocks of the PU according to the motion information of the PU.
  • the inter prediction unit 321 may generate a prediction block of the PU according to one or more reference blocks of the PU.
  • the inverse quantization/transform unit 330 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU.
  • the inverse quantization/transform unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.
  • the inverse quantization/transform unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.
  • the reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.
  • the loop filtering unit 350 may perform a deblocking filtering operation to reduce blocking effects of pixel blocks associated with a CU.
  • the video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360.
  • the video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
  • the basic process of video encoding and decoding is as follows: at the encoding end, a frame of image is divided into blocks, and for the current block, the prediction unit 210 uses intra-frame prediction or inter-frame prediction to generate a prediction block of the current block.
  • the residual unit 220 can calculate the residual block based on the original block of the prediction block and the current block, that is, the difference between the original block of the prediction block and the current block, and the residual block can also be called residual information.
  • the residual block can remove information that is not sensitive to the human eye through the transformation and quantization process of the transformation/quantization unit 230 to eliminate visual redundancy.
  • the residual block before transformation and quantization by the transformation/quantization unit 230 can be called a time domain residual block, and the time domain residual block after transformation and quantization by the transformation/quantization unit 230 can be called a frequency residual block or a frequency domain residual block.
  • the entropy coding unit 280 receives the quantized change coefficient output by the change quantization unit 230, and can entropy encode the quantized change coefficient and output a bit stream. For example, the entropy coding unit 280 can eliminate character redundancy according to the target context model and the probability information of the binary bit stream.
  • the entropy decoding unit 310 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block.
  • the prediction unit 320 uses intra-frame prediction or inter-frame prediction to generate a prediction block of the current block based on the prediction information.
  • the inverse quantization/transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block.
  • the reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block.
  • the reconstructed blocks constitute a reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or on the block to obtain a decoded image.
  • the encoding end also requires similar operations as the decoding end to obtain a decoded image.
  • the decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction for subsequent
  • the block division information determined by the encoder as well as the mode information or parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc., are carried in the bitstream when necessary.
  • the decoder parses the bitstream and determines the same block division information, prediction, transformation, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder by analyzing the existing information, thereby ensuring that the decoded image obtained by the encoder is the same as the decoded image obtained by the decoder.
  • the above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. The present application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
  • the current block may be a current coding unit (CU) or a current prediction unit (PU), etc.
  • CU current coding unit
  • PU current prediction unit
  • an image may be divided into slices, etc. Slices in the same image may be processed in parallel, that is, there is no data dependency between them.
  • "Frame” is a commonly used term, and it can generally be understood that a frame is an image. In the application, the frame may also be replaced by an image or a slice, etc.
  • inter-frame prediction uses temporal correlation to eliminate redundancy.
  • the frame rate of general videos will be 30 frames per second, 50 frames per second, 60 frames per second, or even 120 frames per second.
  • the correlation between adjacent frames in the same scene is very high.
  • Inter-frame prediction technology uses this correlation to refer to the content of the frames that have been encoded and decoded to predict the current content to be encoded. Inter-frame prediction can greatly improve encoding performance.
  • the most basic inter-frame prediction method is translational prediction.
  • Translational prediction assumes that the content to be predicted is in translational motion between the current image and the reference image.
  • the content of the current block (coding unit or prediction unit) is in translational motion between the current image and the reference image.
  • MV motion vector
  • Translational motion accounts for a large proportion in videos. Stationary backgrounds, objects that move as a whole, and lens translation can all be processed using translational prediction.
  • Bidirectional prediction finds two reference blocks from the reference image and performs weighted average on the two reference blocks to obtain a prediction block that is as similar to the current block as possible. For example, for some scenes, weighted average of a reference block from the front and back of the current frame may be more similar to the current block than a single reference block. Based on this, bidirectional prediction improves compression performance on the basis of unidirectional prediction.
  • POC picture order count
  • P-image P Frame
  • RPL0 reference image list
  • RPL0 The reference image list RPL0 contains all reference images whose POC is before the current image.
  • B-image B Frame was previously available. The image predicted by the reference image with POC before the current image and the reference image with POC after the current image.
  • RPL0 contains all reference images with POC before the current image
  • RPL1 contains all reference images with POC after the current image.
  • RPL0 contains all reference images with POC before the current image
  • RPL1 contains all reference images with POC after the current image.
  • RPL0 contains all reference images with POC before the current image
  • RPL1 contains all reference images with POC after the current image.
  • RPL0 contains all reference images with POC before the current image
  • RPL1 contains all reference images with POC after the current image.
  • RPL0 contains all reference images with POC before the current image
  • RPL1 contains all reference images with POC after the current image.
  • For a current block you can only refer to the reference block of a certain image in RPL0, which is also called forward prediction; you can also only refer to the reference block of a certain image in RPL1, which is also called backward prediction; you can also refer to the reference block of a certain image in RPL0 and the reference block of a certain image in RPL1 at the same
  • a simple way to refer to two reference blocks at the same time is to average the pixels in each corresponding position of the two reference blocks to obtain the prediction block of the current block.
  • B image no longer restricts RPL0 to contain all reference images with POC before the current image, and RPL1 to contain all reference images with POC after the current image. Therefore, RPL0 can also contain reference images with POC after the current image, and RPL1 can also contain reference images with POC before the current image.
  • the current block can also refer to a reference image with a POC before the current image or a reference image with a POC after the current image. This kind of B image is also called a generalized B image.
  • the encoding and decoding order of RA (Random Access) configuration is different from the POC order.
  • the B picture can refer to the information before the current picture and the information after the current picture, which significantly improves the encoding performance.
  • a classic GOP (group of pictures) structure of RA is shown in Figure 4, and the arrows in the figure represent the reference relationship.
  • the I picture does not need a reference picture.
  • the P picture with POC 4 is decoded.
  • the I picture with POC 0 can be referenced.
  • the B picture with POC 2 is decoded.
  • the I picture with POC 0 and the P picture with POC 4 can be referenced, and so on.
  • Low Delay configuration is divided into Low Delay P and Low Delay B.
  • Low Delay P is the traditional Low Delay configuration. Its typical structure is IPPP..., that is, an I image is encoded and decoded first, and the subsequent images are all P images.
  • the typical structure of Low Delay B is IBBB..., the difference from Low Delay P is that each inter-frame image is a B image, that is, using two reference image lists, the current block can simultaneously refer to the reference block of an image in RPL0 and the reference block of an image in RPL1.
  • the compression efficiency of RA configuration is higher than that of LD configuration
  • the compression efficiency of LDB configuration is higher than that of LDP configuration. This is because bidirectional prediction can refer to backward information and can reduce prediction errors through some technologies, such as weighted average.
  • a reference image list of the current image can have at most several reference images, such as 2, 3 or 4.
  • reference images which reference images are in RPL0 and RPL1
  • the same reference image may appear in RPL0 and RPL1 at the same time. That is, the codec allows the current block to refer to two reference blocks of the same reference image at the same time.
  • Codecs usually use the index value index in the reference image list to correspond to the reference image. If a reference image list is 4 in length, then index has four values: 0, 1, 2, and 3. For example, RPL0 of the current frame has four reference images with POCs 5, 4, 3, and 0. Then RPL0 index 0 is the reference image with POC 5, RPL0 index 1 is the reference image with POC 4, RPL0 index 2 is the reference image with POC 3, and RPL0 index 3 is the reference image with POC 0.
  • Inter-frame prediction uses motion information to represent "motion".
  • Basic motion information includes reference picture information and motion vector (MV) information.
  • MV motion vector
  • unidirectional motion information and bidirectional motion information can use the same data structure, but the two sets of reference frame information and motion vector information of the bidirectional motion information are valid, while one set of reference frame information and motion vector information of the unidirectional motion information is invalid.
  • the said validity can also be said to be “used”, and the said invalidity can also be said to be "not used”.
  • VVC supports two reference image lists, denoted as RPL0 and RPL1.
  • RPL0 the reference image index refIdxL0 corresponding to RPL0
  • the motion vector mvL0 corresponding to RPL0 the reference image index refIdxL1 corresponding to RPL1
  • the motion vector mvL0 corresponding to RPL1 the motion vector mvL0 corresponding to RPL1.
  • the reference image index corresponding to RPL0 and the reference image index corresponding to RPL1 here can be understood as the information of the above-mentioned reference image.
  • VVC uses two flags to indicate whether to use the motion information corresponding to RPL0 and whether to use the motion information corresponding to RPL1, respectively denoted as predFlagL0 and predFlagL1.
  • predFlagL0 and predFlagL1 indicate whether the above-mentioned unidirectional motion information is "valid". Therefore, although VVC does not explicitly mention the data structure of motion information, it uses the reference image index, motion vector and "valid" flag corresponding to each reference image list to represent the motion information. In the VVC standard text, motion information does not appear, but motion vectors are used. It can also be considered that the reference image index and the flag of whether to use the corresponding motion information are attached to the motion vector. For the convenience of description, "motion information” is still used in this article, but it should be understood that “motion vector” can also be used to describe it. “Motion information” can also be called "motion parameter".
  • the motion vector can be represented by (x, y), that is, a horizontal component and a vertical component. Since videos are represented by pixels, there is a distance between pixels, and the movement of an object between adjacent images may not always correspond to the whole pixel distance. For example, in a distant video, the distance between 2 pixels is 1 meter for the distant object, and the distance moved by the object in the time between 2 frames is 0.5 meters. This scene cannot be well represented by the whole pixel motion vector. Therefore, the motion vector can be expressed at the sub-pixel level, such as 1/2 pixel accuracy, 1/4 pixel accuracy, 1/8 pixel accuracy, and 1/16 pixel accuracy, to express the motion more finely. And the pixel value of the sub-pixel position in the reference image is obtained by interpolation.
  • the unidirectional prediction and bidirectional prediction in the above-mentioned translation prediction are both based on blocks, such as coding units or prediction units. That is, a pixel matrix is used as a unit for prediction.
  • the most basic block is a rectangular block, such as a square and a rectangle.
  • Video coding standards such as HEVC and VVC allow the encoder to determine the size and division method of the coding unit and prediction unit according to the content of the video. Areas with simple textures or motions tend to use larger blocks, while areas with complex textures or motions tend to use smaller blocks. The deeper the level of block division, the more complex blocks that are closer to the actual texture or motion can be divided, but the corresponding overhead for representing these divisions is also greater. Motion information may also need to be transmitted in the bitstream. And usually, the finer the block division, the greater the overhead of motion information.
  • motion vector prediction MVP motion vector prediction
  • MVD motion vector difference
  • each inter-frame coded block requires a motion information.
  • the CU division is equal to the PU division is equal to the TU division, that is, a coding unit has a prediction unit of the same size and position and a transform unit of the same size and position.
  • the CU The division is more flexible. Compared with HEVC, VVC tends to weaken PU and TU. Differences in any link of prediction, transformation, quantization, and entropy coding may lead to the division of CU. For example, if the motion information of two regions is different, the encoder may divide the two regions into different CUs.
  • the encoder may also divide the two regions into different CUs. How to divide is determined by the overall compression efficiency, not entirely by a single factor. As a result, the same object or regions with the same or similar motion will be divided into different CUs.
  • Figure 5 is an example in HEVC
  • a is the original image, in which there is an iron rod moving in the direction indicated by the arrow, and the background area moves less.
  • Figure b shows the block division of HEVC
  • Figure c removes the boundaries of the blocks with the same motion information in Figure b. It can be seen that many adjacent blocks use the same motion information. In this case, if the motion information is encoded separately for each block, there will be obvious waste.
  • the complete motion information of VVC mentioned above includes the reference image index of RPL0, the flag whether MV is used, the reference image index of RPL1, and the flag whether MV is used.
  • the basic principle of the merge mode is that the current block can inherit the motion information of the adjacent blocks, including the reference image information and the motion vector information.
  • Merge mode can build a merge candidate list. If the current block uses Merge mode, an index can be used to indicate which motion information the current block merges with, so there is no need to encode the complete motion information.
  • an index can be used to indicate which motion information the current block merges with, so there is no need to encode the complete motion information.
  • the adjacent blocks in the spatial domain refer to blocks adjacent to the current block in the same image
  • the non-adjacent blocks in the spatial domain refer to blocks non-adjacent to the current block in the same image.
  • the motion information in the time domain and the motion information of non-adjacent blocks in the time domain refer to motion information at a specified position on a collocated reference image.
  • the large gray block is the current block, wherein positions 1, 2, 3, 4, and 5 are positions of adjacent blocks in the spatial domain used by Merge, and other dark gray positions are positions of non-adjacent blocks in the spatial domain used by Merge.
  • Position 6 is the position used for motion information in the time domain.
  • the position corresponding to the center of the current block is used.
  • Other light gray positions are positions used for motion information of non-adjacent blocks in the time domain.
  • the temporal motion information is derived based on the motion information at the corresponding position of the collocated reference image.
  • Temporal motion information prediction is used as a supplement to spatial motion information prediction.
  • the correlation between adjacent regions on the same image is stronger than that on different images.
  • temporal motion information is more useful. For example, if the current block and the surrounding adjacent blocks in the current image belong to different objects and have completely different motions, the motion of the block that belongs to the same object as the current block in a reference image can provide better motion information prediction for the current block.
  • the motion vector of the co-located block on the co-located reference image (the block that obtains the temporal motion information is called the co-located block here) is the vector from the co-located reference image col_pic to the reference image col_ref of the co-located block.
  • the motion vector it needs is the vector from the current image curr_pic to the reference image curr_ref of the current block.
  • the scaling ratio can be determined based on td and tb.
  • the motion vector of the co-located block be (col_mv_x, col_mv_y)
  • each 4x4 sub-block stores a set of motion information. It is understandable that if the cost of hardware implementation is not considered, the same-position reference image can also store a set of motion information for each pixel.
  • VVC introduces sub-block-based temporal motion vector prediction, namely SbTMVP (Subblock-based temporal motion vector prediction).
  • SbTMVP Motion vector prediction
  • TMVP temporary motion vector prediction
  • SbTMVP is based on sub-blocks, so SbTMVP can obtain an MVP for each sub-block. This is also the essential difference between SbTMVP and TMVP.
  • TMVP uses the position of the lower right corner of the current block or the position of the center of the current block to locate the co-located block
  • SbTMVP finds a motion offset based on the motion of the surrounding blocks to determine the position.
  • the motion displacement is set to the motion vector of A1 using the co-located reference image. Otherwise, the motion displacement is set to (0, 0).
  • the position is found according to the motion displacement, and then the MV corresponding to the position of each sub-block in the "co-located block" is scaled to obtain the MVP of each sub-block.
  • Merge mode directly selects the motion information in the merge candidate list as the motion information of the current block. In actual videos, there are sometimes some differences between the actual motion vector of the current block and the motion vector in the selected merge candidate list.
  • MVD in Merge mode is a special Merge mode in VVC, which encodes MVD in this case in an efficient way. Ordinary Merge does not require encoding and decoding MVD (motion vector difference). Ordinary inter mode requires direct encoding and decoding of MVD. As shown in Figure 9, MMVD takes advantage of the fact that MVD is more distributed in a single horizontal direction or a single vertical direction, with more MVDs for smaller values and fewer MVDs for larger values.
  • MMVD can only represent MVDs of specific values in some specific directions, and it cannot represent any MVD. It uses mmvd_direction_idx to represent the direction of the MVD, which can also be understood as whether the x and y of the MVD are non-zero and the positive and negative signs. mmvd_distance_idx represents the absolute value of the non-zero x and y of the MVD, MmvdDistance.
  • ph_mmvd_fullpel_only_flag is an image header flag that can set 2 different combinations of MMVD.
  • Affine uses a linear model to calculate the motion vector of each sub-block or each pixel in the current block based on the motion vectors of 2 control points (4 parameters, a motion vector includes 2 parameters x and y) or 3 control points (6 parameters).
  • the motion vector at the (x, y) position in the current block is derived according to the following formula (3):
  • the motion vector at the (x, y) position in the current block is derived according to the following formula (4):
  • (mv0x, mv0y) is the motion vector of the control point at the upper left corner of the current block
  • (mv1x, mv1y) is the motion vector of the control point at the upper right corner of the current block
  • (mv2x, mv2y) is the motion vector of the control point at the lower left corner of the current block.
  • Affine used in VVC divides the current block into 4x4 sub-blocks, calculates an MV for each sub-block and performs motion compensation.
  • Figure 10B is an example of Affine deriving motion vectors based on sub-blocks. It can be understood that with the enhancement of hardware processing capabilities, Affine can also perform pixel-based processing. That is, a motion vector is derived for each pixel, and motion compensation is performed on a pixel based on the motion vector.
  • Affine only needs a few control points to derive the motion vector for each sub-block or each pixel, which can achieve more precise predictions than motion compensation based on the entire block. Compared with dividing smaller CUs, Affine has much less overhead.
  • HEVC supports a maximum of 64x64 CTUs and can recursively perform quadtree division.
  • VVC supports a more flexible block division method than HEVC, supporting a maximum of 128x128 CTUs, including quadtree, ternary tree and binary tree divisions. These division methods.
  • block division is becoming more and more flexible, whether it is CU, PU, or TU, it can only be divided into rectangular blocks. It should be noted that VVC has weakened the division of PU and TU.
  • the boundaries of texture or motion in natural videos are diverse. For example, when encountering an oblique object boundary, if you simply use rectangular blocks to approach the boundary, you will divide it into many small blocks, which will significantly increase the overhead.
  • the geometric partitioning prediction mode (Geometric partitioning Mode, GPM) can better handle textures and boundaries in natural videos.
  • GPM uses two prediction blocks with the same size as the current block. Some pixel positions in the prediction block of GPM use 100% of the pixel values of the corresponding positions of the first prediction block, and some pixel positions use 100% of the pixel values of the corresponding positions of the second prediction block. In the boundary area or transition area, the pixel values of the corresponding positions of the two prediction blocks are used in a certain proportion. The weight of the boundary area is also gradually transitioned. Of course, in response to scenes such as screen content encoding, the transition area can also be not used. How these weights are specifically distributed is determined by the "division" mode of GPM. The weight of each pixel position is determined according to the "division" mode of GPM.
  • GPM uses two prediction blocks of different sizes from the current block, that is, each takes a required part. The part with a weight of 0 is eliminated. This is an implementation problem and is not the focus of the present invention.
  • FIG11 is a weight diagram of 64 modes of GPM in VVC on a square block.
  • Black indicates that the weight value of the corresponding position of the first prediction block is 0%
  • white indicates that the weight value of the corresponding position of the first prediction block is 100%
  • the gray area indicates that the weight value of the corresponding position of the first prediction block is a certain weight value greater than 0% and less than 100% according to the different shades of color.
  • the weight value of the corresponding position of the second reference block is 100% minus the weight value of the corresponding position of the first reference block.
  • GPM can be said to be a prediction mode or prediction method because it ultimately generates a prediction block. It can also be said that GPM is a "partitioning" mode that simulates the division of the prediction block, which is similar to the implementation of PU division, but without substantial division.
  • the first prediction block and the second prediction block used by the above GPM can be prediction blocks generated by intra-frame prediction, prediction blocks generated by unidirectional prediction between frames, or prediction blocks generated by bidirectional prediction between frames.
  • the bit rate of general consumer video is limited, so video compression usually seeks a compromise between bit stream overhead and distortion.
  • block division for the same content, within a certain range, the finer the division, the greater the overhead and the less distortion; the coarser the division, the less overhead and the greater the distortion.
  • encoding of motion information for the same content, within a certain range, the more precise the motion information, the greater the overhead and the less distortion; the coarser the motion information, the less overhead and the greater the distortion.
  • Some decoding methods use the information on the decoding side for processing and calculation without occupying overhead, so as to improve motion information, improve prediction effect, and reduce distortion.
  • Not occupying overhead means that there is no instruction made by the encoder based on the original image, and it is automatically processed based on the available information.
  • Two typical decoding methods in VVC are DMVR (Decoder side motion vector refinement) and BDOF (bi-directional optical flow).
  • DMVR in VVC uses bilateral matching BM (bilateral matching), that is, to calculate the matching cost of the reference blocks on both sides, such as SAD (sum of absolute difference). DMVR searches for the matching cost of the MVs around the original MV.
  • the MVs of the two reference images are moved in a mirrored manner, that is, one side moves MVdiff and the other side moves -MVdiff based on their respective original MVs, as shown in Figure 12.
  • the search also supports sub-pixel search, so DMVR may find MVs with higher accuracy than the original MVs.
  • the search is performed according to certain rules. Generally, the integer pixel MVs within a certain range are searched first to find the integer pixel MV with the lowest matching cost, and then the sub-pixel MVs are searched based on the integer pixel MVs. If an MV with a lower matching cost than the original MV is found, the MV with a lower matching cost is used for motion compensation prediction.
  • the MV improved by DMVR can theoretically be used to store MVs and surrounding blocks. For example, when the merge candidate list is constructed for the current block, if the surrounding blocks use DMVR to improve the MV, using the improved MV to construct the merge candidate list can achieve better compression effect, but for hardware implementation considerations, VVC does not do this.
  • DMVR can be processed based on sub-blocks.
  • the sub-blocks will be divided into 16-pixel blocks.
  • this is based on the consideration of hardware implementation complexity, because DMVR needs to be searched at the decoding end, and limiting the size of the sub-block can reduce the cost of cache.
  • dividing into sub-blocks provides better flexibility, and each sub-block can improve the MV independently, which to a certain extent achieves the effect of improving the division accuracy, which also improves the compression efficiency.
  • Bidirectional optical flow is also a typical decoding method.
  • BDOF improves MV and prediction based on the principle of optical flow.
  • Optical flow is the instantaneous speed of the pixel movement of a moving object in space on the observation imaging plane.
  • Optical flow has some basic assumptions, such as constant brightness, that is, the brightness of the same target does not change when it moves between different images. Time is continuous or the movement is small. That is, the change of time will not cause a drastic change in the target position.
  • BDOF in VVC will derive a motion vector deviation (v x ,v y ), which is calculated by minimizing the difference between the predicted values in the two directions. This motion vector deviation is also used to adjust the predicted value in the corresponding sub-block.
  • the process of deriving the predicted value includes:
  • I (k) (i,j) is the predicted value of the coordinate (i,j) of the reference image list k
  • k 0,1
  • (v x ,v y ) is determined by the following formula (6):
  • ⁇ (i,j) (I (1) (i,j)>>n b )-(I (0) (i,j)>>n b ).
  • is a 6x6 window around the current 4x4 sub-block, na is min(1, bitDepth-11), and n b is min(4, bitDepth-8).
  • each prediction value within the 4x4 sub-block is adjusted based on the motion vector deviation and gradient.
  • the adjustment value of the predicted value is determined by the following formula (8):
  • the prediction value of the current block is adjusted based on the adjustment value of the prediction value to obtain the prediction value of BDOF.
  • Ooffset and shift are calculated based on the bit depth of brightness.
  • n a , n b and shift are all processed to reduce the bit width during the calculation process.
  • the motion vector deviation of BDOF can achieve high precision, making the prediction more accurate, and the sub-block-based processing also improves flexibility. It is similar to DMVR.
  • DMVR is based on block matching
  • BDOF is based on the principle of optical flow. They can be used in combination.
  • An example is as follows, which can be called multi-pass decoder-side motion vector refinement (MDMVR for short).
  • MDMVR may include: Step 1, motion vector improvement based on bidirectional matching of the whole block. Step 2, motion vector improvement based on bidirectional matching of sub-blocks. The sub-block size of this step may be 16x16. Step 3, motion vector improvement based on bidirectional optical flow of sub-blocks. The sub-block size of this step may be 8x8. Currently, further steps may be added on this basis, such as a fourth step, motion vector improvement based on bidirectional optical flow of 4x4 sub-blocks. Or a motion vector improvement based on bidirectional optical flow of points, etc.
  • the template matching method was first used in inter-frame prediction. It uses the correlation between adjacent pixels and takes some areas around the current block as templates.
  • the current block is encoded and decoded, its left and upper sides have been encoded and decoded according to the encoding order.
  • the existing hardware decoder is implemented, it is not necessarily guaranteed that when the current block starts to be decoded, its left and upper sides have been decoded.
  • this refers to inter-frame blocks.
  • the inter-frame coded block when the inter-frame coded block generates a prediction block, the surrounding reconstructed pixels are not required, so the prediction process of the inter-frame block can be carried out in parallel.
  • the intra-frame coded block must use the reconstructed pixels on the left and upper sides as reference pixels.
  • the left and upper sides are available, which means that the hardware design can be adjusted accordingly.
  • the right and lower sides are not available under the current standard such as VVC encoding order.
  • the rectangular areas on the left and upper sides of the current block are set as templates.
  • the height of the template part on the left is generally the same as the height of the current block, and the width of the template part on the upper side is generally the same as the width of the current block, but of course they can be different.
  • the best matching position of the template is found in the reference image to determine the motion information or motion vector of the current block. This process can be roughly described as starting from a starting position in a certain reference image and searching within a certain range around it.
  • the search rules such as the search range and search step length, can be pre-set. Each time a position is moved to, the matching degree between the template corresponding to the position and the template around the current block is calculated.
  • the so-called matching degree can be measured by some distortion costs, such as SAD (sum of absolute difference), SATD (sum of absolute transformed difference).
  • SAD sum of absolute difference
  • SATD sum of absolute transformed difference
  • MSE mean-square error
  • the cost is calculated using the predicted block of the template corresponding to the position and the reconstructed block of the template around the current block.
  • the sub-pixel position can also be searched, and the motion information of the current block can be determined based on the position with the highest degree of matching.
  • the motion information suitable for the template may also be the appropriate motion information for the current block.
  • the template matching method may not necessarily be applicable to all blocks, so some methods can be used to determine whether the current block uses the above template matching method, such as using a control switch in the current block to indicate whether the template matching method is used.
  • This template matching method is called DMVD (decoder side motion vector derivation).
  • DMVD decoder side motion vector derivation
  • Both the encoder and the decoder can use the template to search to derive motion information or find better motion information based on the original motion information. It does not need to transmit specific motion vectors or motion vector differences, but the encoder and decoder perform the same regular search to ensure the consistency of encoding and decoding.
  • the template matching method can improve compression performance, but it also requires “searching" at the decoding end, which brings a certain degree of decoding complexity.
  • the encoder can decide which prediction mode or model to use for the current block, such as whether to use the merge mode, whether to use the MMVD mode, whether to predict by the whole block or the sub-block. If it is predicted by the whole block, the merge candidate list also includes motion vector prediction in the spatial domain and motion vector prediction in the temporal domain. If it is predicted by the sub-block, there are also modes such as SbTMVP and Affine in the sub-block candidates. Whether to use the GPM mode, etc. On the one hand, the more detailed and accurate information is told to the decoder through the bitstream, the decoder can make better predictions, but the corresponding overhead is greater.
  • the encoder needs to make a trade-off between bitrate and distortion.
  • the above methods such as merge, MMVD, GPM, SbTMVP, Affine, etc., tell the decoder as much information as possible in a more efficient way, and the decoder executes according to the instructions of the encoder.
  • the decoding algorithms such as DMVR and BDOF use block matching or optical flow methods to compensate for the distortion caused by inaccurate motion vectors, which to a certain extent also gives the encoder more room to transmit less information to save bitstream overhead. It can be said that the "smarter" the decoder is, the smaller the distortion it can make when the encoder gives the same instructions.
  • the decoder obtains the indication of motion information from the bitstream to obtain the initial motion information, finds the reference block or the area around the reference block according to the initial motion information, and improves the initial motion information and/or improves the prediction value according to the pixel value information of the reference block or the area around the reference block.
  • DMVR searches around the initial MV, and its search process and the final selection of which MV depends on the result of block matching.
  • BDOF also calculates information such as gradient based on pixel values, and calculates the instantaneous motion vector difference through the optical flow method.
  • the present application improves the first motion information based on the motion information of the reference image of the current block to obtain the second motion information, wherein the motion information of the reference image is used to improve the first motion information and/or for block division. That is to say, in the embodiment of the present application, when improving the first motion information, the motion information of the reference image is taken into account to achieve effective improvement of the first motion information and obtain accurate second motion information, and then based on the accurate second motion information, when determining the prediction value of the current block, the prediction accuracy of the current block can be improved, thereby improving the decoding performance of the video.
  • the video decoding method provided in the embodiment of the present application is introduced by taking the decoding end as an example.
  • FIG14 is a schematic diagram of a video decoding method flow chart provided by an embodiment of the present application, and the embodiment of the present application is applied to the video decoders shown in FIG1 and FIG3. As shown in FIG14, the method of the embodiment of the present application includes:
  • S101 Determine first motion information of a current block.
  • the decoding method provided in the embodiment of the present application is applied to inter-frame prediction to improve the motion information of the current block.
  • the encoder carries less and less prediction-related information in the bitstream for the sake of bitrate considerations, which makes the motion information of the current block obtained by the decoder based on the prediction-related information carried in the bitstream inaccurate. Therefore, the decoder can improve the motion information determined by the decoding to improve the prediction effect.
  • the motion information of the reference image is not considered, which makes the improvement of the motion information unsatisfactory.
  • the motion information of the reference image is taken into consideration, thereby achieving effective improvement of the motion information of the current block, thereby improving the prediction accuracy of the current block and enhancing the decoding performance of the video.
  • the first motion information of the current block can be understood as the initial motion information of the current block, that is, the decoding end obtains the prediction related information carried by the bitstream by decoding the bitstream, and determines the motion information based on the prediction related information.
  • the prediction related information may include information such as the prediction mode.
  • the first motion information of the current block can be understood as motion information after the initial motion information of the current block has been improved once or multiple times. Through the method of the embodiment of the present application, the first motion information is improved again based on the motion information of the reference image.
  • inter-frame prediction uses motion information to represent "motion".
  • Basic motion information includes reference picture information and motion vector (MV) information.
  • MV motion vector
  • a block uses bidirectional prediction, it is necessary to find two reference blocks, so two sets of reference image information and motion vector information are required. Each set can be understood as a unidirectional motion information, and these two sets are combined to form a bidirectional motion information.
  • the motion information of the embodiments of the present application may refer to unidirectional motion information, that is, information including a set of reference image information and motion vector information.
  • the motion information of the embodiments of the present application may refer to bidirectional motion information, that is, information including two sets of reference image information and motion vector information.
  • the motion information of the embodiments of the present application may refer to multi-directional motion information, that is, information including multiple sets of reference image information and motion vector information.
  • the reference image index corresponding to each reference image list, the motion vector and the "valid" flag may be used together to represent the motion information.
  • the embodiment of the present application does not limit the specific method for the decoding end to determine the first motion information of the current block.
  • the encoder carries the prediction mode of the current block in the bitstream, so that the decoder obtains the prediction mode of the current block by decoding the bitstream, and further determines the first motion information of the current block based on the prediction mode.
  • the decoding end obtains initial motion information of the current block based on the prediction mode, and determines the initial motion information as the first motion information.
  • the decoding end obtains the initial motion information of the current block based on the prediction mode, and then improves the initial motion information, and determines the improved initial motion information as the first motion information.
  • the method for the decoding end to improve the initial motion information may be to improve it by using the above-mentioned DMVR and/or BDOF.
  • the decoding end uses the DMVR improvement method to improve the initial motion information of the current block to obtain the first motion information.
  • the decoding end uses the BDOF improvement method to improve the initial motion information of the current block to obtain the first motion information.
  • the decoding end first uses the DMVR improvement method to improve the initial motion information, and then uses the BDOF improvement method to improve it to obtain the first motion information.
  • the decoding end first uses the BDOF improvement method to improve the initial motion information, and then uses the BDOF improvement method to improve it to obtain the first motion information.
  • the specific improvement methods of DMVR and BDOF refer to the description of the above-mentioned embodiments and will not be repeated here.
  • S102 Based on the motion information of the reference image of the current block, improve the first motion information to obtain the second motion information of the current block.
  • the motion information of the reference image is used to improve the first motion information and/or for block division.
  • the motion information of the embodiment of the present application includes information of the reference image and information of the motion vector. Based on this, the decoding end can obtain the information of the reference image of the current block from the first motion information of the current block determined above, such as obtaining the index of the reference image, and then obtain the reference image of the current block from the reference image list based on the index.
  • the current block corresponds to a reference image list, recorded as RPL0.
  • the index of the reference image of the current block in the reference image list RPL0 is determined, and then based on the index, the reference image corresponding to the index in the reference image list RPL0 is determined as the reference image of the current block.
  • the decoding end determines the index of the reference image of the current block in the reference image list RPL0 in at least the following ways: Method 1, the encoding end and the decoding end determine each reference image in the reference image list RPL0, such as the first reference image, as the reference image of the current block by default, so that the encoding end does not need to indicate the index of the reference image of the current block in the bitstream. Method 2, the encoding end writes the index of the reference image of the current block in the reference image list RPL0 into the bitstream, so that the decoding end obtains the index of the reference image of the current block in the reference image list RPL0 by decoding the bitstream, and then obtains the reference image of the current block based on the index.
  • Method 1 the encoding end and the decoding end determine each reference image in the reference image list RPL0, such as the first reference image, as the reference image of the current block by default, so that the encoding end does not need to indicate the index of the reference image of the current block in the bitstream.
  • the current block corresponds to two reference image lists, which are respectively recorded as RPL0 and RPL1.
  • the decoding end determines that one reference image of the current block is indexed refIdxL0 in the reference image list RPL0, and the other reference image is indexed refIdxL1 in the reference image list RPL1, and then based on these two indexes, the two reference images of the current block are determined in the reference image list RPL0 and the reference image list RPL1.
  • the decoding end determines the index of the reference image of the current block in the reference image list RPL0 in at least the following ways: Method 1, the encoding end and the decoding end determine each reference image in the reference image list RPL0, such as the first reference image, as the reference image of the current block by default, so that the encoding end does not need to indicate the index of the reference image of the current block in the bitstream. Method 2: The encoder writes the index of the reference image of the current block in the reference image list RPL0 into the bitstream, so that the decoder obtains the index of the reference image of the current block in the reference image list RPL0 by decoding the bitstream, and then obtains the reference image of the current block based on the index.
  • Method 1 the encoding end and the decoding end determine each reference image in the reference image list RPL0, such as the first reference image, as the reference image of the current block by default, so that the encoding end does not need to indicate the index of the reference image of the current block in the bitstream.
  • the decoder determines the reference image of the current block. Since the reference images are all decoded images, their motion information is known. Therefore, the decoder can directly obtain the motion information of the reference image, and then determine the second motion information of the current block based on the motion information of the reference image and the first motion information of the current block determined in the above steps, where the second motion information can be understood as more accurate motion information obtained by improving the first motion information.
  • the motion information of the reference image plays at least two roles in improving the first motion information.
  • One is that the motion information of the reference image directly participates in the improvement of the first motion information, for example, it is used to guide the search process of the second motion information.
  • the second is that, in the process of improving the first motion information, the motion information of the reference image is used to indicate the division of blocks.
  • the following introduces a process in which the decoding end improves the first motion information based on the motion information of the reference image of the current block to obtain the second motion information of the current block.
  • the embodiment of the present application does not limit the specific process of improving the first motion information based on the motion information of the reference image at the decoding end to obtain the second motion information of the current block.
  • the decoding end improves the first motion information by at least several methods shown in the following embodiments to obtain the second motion information.
  • a template reference area corresponding to the template of the current block is determined in the reference image, wherein the motion information of the template reference area is known, and the motion information of the template of the current block is also known, so the first motion information of the current block can be improved based on the motion information of the template reference area and the motion information of the template of the current block to obtain the second motion information of the current block.
  • the difference value between the motion information of the template reference area and the motion information of the template of the current block is determined, and the difference value is added to the first motion information to obtain the second motion information of the current block.
  • the above S102 includes the following steps S102-A to S102-C:
  • the decoding end first determines the reference block corresponding to the current block in the reference image based on the first motion information. Since the motion information of the reference block is known, the decoding end can determine whether the first motion information of the current block is accurate based on the motion information of the reference block. For example, assuming that the reference block and the current block belong to the same object moving as a whole in the image, and the motion of the entire object is a uniform motion, the reference block can be used as the same block as the current block, and the time domain motion information of the current block can be determined based on the motion information of the reference block. For the convenience of description, the time domain motion information is recorded as the third motion information, and then the first motion information is improved based on the third motion information to achieve accurate improvement of the first motion information.
  • the first motion information includes information of the motion vector of the current block.
  • the improvement of the first motion information can be understood as the improvement of the motion vector included in the first motion information.
  • the decoding end first determines the reference block corresponding to the current block in the reference image based on the first motion information, wherein the first motion information of the current block can be unidirectional motion information or bidirectional motion information, and these two cases are introduced respectively below.
  • the first motion information of the current block includes unidirectional motion information, that is, the current block corresponds to a reference image and a motion vector.
  • the reference image included in the first motion information of the current block curr_block in the current image curr_pic is the reference image ref_pic_0 in the reference image sequence RPL0
  • the motion vector is the first motion vector mv_0.
  • the playback order of the reference image ref_pic_0 is before the current image curr_pic.
  • the decoding end can locate the corresponding reference block ref_block_0 in the reference image ref_pic_0 based on the position of the current block curr_block and the first vector mv_0 of the current block.
  • the current block represented by the dotted box in the reference image in Figure 15 can be understood as the same-position block of the current block in the reference image.
  • the first motion information of the current block includes bidirectional motion information, that is, the current block corresponds to two reference images, which are recorded as the first reference image and the second reference image, and the two motion vectors are recorded as the first motion vector and the second motion vector.
  • the first reference image included in the first motion information of the current block curr_block in the current image curr_pic is the reference image ref_pic_0 of the reference image sequence RPL0
  • the first motion vector is mv_0
  • the second reference image of the current block is the reference image ref_pic_1 in the reference image sequence RPL1
  • the second motion vector is mv_1.
  • the decoding end can locate the corresponding first reference block ref_block_0 in the first reference image ref_pic_0 according to the position of the current block curr_block and the first motion vector mv_0, and can locate the corresponding second reference block ref_block_1 in the second reference image ref_pic_1 according to the position of the current block curr_block and the second motion vector mv_1.
  • the decoding end determines the reference block of the current block in the reference image based on the first motion information of the current block. Since the reference images are all decoded images and their motion information is known, the motion information of the reference block can also be obtained.
  • the decoding end can infer the direction in which the first motion vector of the current block is more likely to deviate based on the motion information of the reference block. Specifically, based on the motion information of the reference block, the temporal motion information of the current block is determined, recorded as the third motion information, and the third motion information is compared with the first motion information of the current block to determine the direction in which the first motion information is more likely to deviate.
  • the implementation methods for determining the third motion information in the above S102-B include but are not limited to the following:
  • the decoding end uses the current block as the co-located block of the reference block, determines the time domain motion information of the reference block according to the motion information of the reference block, and then determines the opposite number of the time domain motion information as the third motion information.
  • the decoder can infer the vector information in the time domain motion information when the reference block moves from the current position in the reference image to the current block in the current image based on the motion information of the reference block, which is recorded as mv_t. For example, the decoder uses the method for deriving the time domain motion information to derive mv_t. Then, the opposite vector of mv_t -mv_t is determined as the motion vector in the third motion information.
  • (ref_mv_x and ref_mv_y) are the motion vectors of the reference block on the x-axis and y-axis
  • td is the POC distance between the co-located reference image col_pic and the reference image col_ref of the co-located block
  • tb is the POC distance between the current image curr_pic and the reference image curr_ref.
  • the absolute value of the motion vector -mv_t in the third motion information determined above is compared with the absolute value of the motion vector mv_0 in the first motion information of the current block to determine in which direction the first motion information is more likely to shift.
  • the absolute value of the motion vector -mv_t in the third motion information is less than the absolute value of the motion vector mv_0 in the first motion information of the current block, it means that the reference block can only reach the position of the dotted box shown in Figure 17A according to its existing motion, so it can be inferred that the motion vector mv_0 of the current block in the first motion information is too large.
  • the absolute value of the motion vector -mv_t in the third motion information is greater than the absolute value of the motion vector mv_0 in the first motion information of the current block, it means that the reference block can reach the position of the dotted box shown in Figure 17B according to its existing motion, so it can be inferred that the motion vector mv_0 of the current block in the first motion information is too small.
  • the decoder can infer the motion vector 0 of the first reference block from the first reference image to the current image based on the motion information of the first reference block, which is recorded as mv_0_t. And the decoder can infer the motion vector 1 of the second reference block from the second reference image to the current image based on the motion information of the second reference block, which is recorded as mv_1_t.
  • motion vector 0 and motion vector 1 can be derived using the method of deriving temporal motion information.
  • ref_mv_0_x and ref_mv_0_y are motion vectors of the first reference block on the x-axis and y-axis
  • td0 is the POC distance between the first co-located reference image col_pic0 and the first reference image col_ref0 of the co-located block
  • tb0 is the POC distance between the current image curr_pic and the first reference image curr_ref0.
  • ref_mv_1_x and ref_mv_1_y are motion vectors of the second reference block on the x-axis and y-axis
  • td1 is the POC distance between the second co-located reference image col_pic1 and the second reference image col_ref1 of the co-located block
  • tb1 is the POC distance between the current image curr_pic and the second reference image curr_ref1.
  • the decoding end determines the opposite vector -mv_0_t of mv_0_t as the first prediction direction motion vector in the third motion information, and determines the phase vector -mv_1_t of mv_1_t as the second prediction direction motion vector in the third motion information.
  • the absolute value of the motion information in the third motion information determined above is compared with the absolute value of the motion vector in the first motion information of the current block to determine in which direction the first motion information is more likely to shift.
  • the first motion information and the third motion information both include a first prediction direction motion vector and a second prediction direction motion vector, so that the motion information in each direction is compared separately.
  • the absolute value of the first prediction direction motion vector -mv_0_t in the third motion information is compared with the absolute value of the first prediction direction motion example mv_0 in the first motion information. For example, if the absolute value of the first prediction direction motion vector -mv_0_t in the third motion information is smaller than the absolute value of the first prediction direction motion vector mv_0 in the first motion information, it means that when the first reference block moves from the current position in the first reference image to the current image according to its existing motion vector, it cannot reach the position of the current block in the current image, so it can be inferred that the first prediction direction motion vector mv_0 in the first motion information is too large.
  • the absolute value of the first prediction direction motion vector -mv_0_t in the third motion information is larger than the absolute value of the first prediction direction motion vector mv_0 in the first motion information, it means that when the first reference block moves from the current position of the first reference block to the current image according to its existing motion vector, its position in the current image exceeds the position of the current block in the current image, so it can be inferred that the first prediction direction motion vector mv_0 in the first motion information is too small.
  • the absolute value of the second prediction direction motion vector -mv_1_t in the third motion information is compared with the absolute value of the second prediction direction motion vector mv_1 in the first motion information.
  • the absolute value of the second prediction direction motion vector -mv_1_t in the third motion information is smaller than the absolute value of the second prediction direction motion vector mv_1 in the first motion information, it means that when the second reference block moves from the current position in the second reference image to the current image according to its existing motion vector, it cannot reach the current block in the current image, so it can be inferred that the second prediction direction motion vector mv_0 in the first motion information is too large.
  • the absolute value of the second prediction direction motion vector -mv_1_t in the third motion information is larger than the absolute value of the second prediction direction motion vector mv_1 in the first motion information, it means that after the second reference block moves from the current position in the second reference image to the current image according to its existing motion vector, its position in the current image exceeds the position of the current block in the current image, so it can be inferred that the second prediction direction motion vector mv_1 in the first motion information is too small.
  • the decoding end determines the time domain motion information of the current block as the third motion vector based on the motion information of the reference block.
  • the decoding end determines the time domain motion information of the current block as the third motion information according to the motion information of the reference block.
  • ref_mv_x and ref_mv_y are the motion vectors of the reference block on the x-axis and y-axis
  • mv_t’_x and mv_t’_y are the first prediction direction motion vector and the second prediction direction motion vector in the third motion information
  • td is the POC distance between the co-located reference image col_pic and the reference image col_ref of the co-located block
  • tb is the POC distance between the current image curr_pic and the reference image curr_ref.
  • the absolute value of the motion vector mv_t' in the third motion information determined above is compared with the absolute value of the motion vector mv_0 in the first motion information of the current block to determine in which direction the first motion information is more likely to shift.
  • the absolute value of the motion vector mv_t' in the third motion information is smaller than the absolute value of the motion vector mv_0 in the first motion information of the current block, it means that when the current block moves from the current position of the current image to the reference image according to its existing motion vector, it can only reach the dotted box position shown in Figure 18A, then it can be inferred that the motion vector mv_0 of the current block in the first motion information is too small.
  • the absolute value of the motion vector mv_t’ in the third motion information is greater than the motion vector mv_0 in the first motion information of the current block, it means that when the current block moves from the current position of the current image to the reference image according to its existing motion vector, it can reach the position of the dotted box shown in Figure 18B. Then it can be inferred that the motion vector mv_0 of the current block in the first motion information is too large.
  • the first motion information and the third motion information both include a first prediction direction motion vector and a second prediction direction motion vector.
  • the decoder can infer the motion vector mv_0_t’ of the current block when it moves from the current position of the current image to the first reference image based on the motion information of the first reference block.
  • the decoder can infer the motion vector mv_1_t’ of the current block when it moves from the current position of the current image to the second reference image based on the motion information of the second reference block.
  • mv_0_t’ is the first prediction direction motion vector in the third motion information
  • mv_1_t’ is the second prediction direction motion vector in the third motion information.
  • mv_0_t’ and mv_1_t’ can be derived using the method for deriving time domain motion information.
  • ref_mv_0_x and ref_mv_0_y are motion vectors of the first reference block on the x-axis and y-axis
  • td0 is the POC distance between the first co-located reference image col_pic0 and the first reference image col_ref0 of the co-located block
  • tb0 is the POC distance between the current image curr_pic and the first reference image curr_ref0.
  • ref_mv_1_x and ref_mv_1_y are motion vectors of the second reference block on the x-axis and y-axis
  • td1 is the POC distance between the second co-located reference image col_pic1 and the second reference image col_ref1 of the co-located block
  • tb1 is the POC distance between the current image curr_pic and the second reference image curr_ref1.
  • the absolute value of the motion vector in the third motion information determined above is compared with the absolute value of the motion vector in the first motion information of the current block to determine in which direction the first motion information is more likely to shift.
  • the absolute value of the first predicted direction motion vector mv_0_t' in the third motion information is compared with the absolute value of the first predicted direction motion vector mv_0 in the first motion information.
  • the first predicted direction motion vector mv_0_t' is less than the absolute value of the first predicted direction motion vector mv_0 in the first motion information, it means that when the current block moves from the current position in the current image to the first reference image according to the first predicted direction motion vector mv_0, it cannot reach the position of the current first reference block in the first reference image. It can be inferred that the first predicted direction motion vector mv_0 in the first motion information is too large.
  • the absolute value of the first predicted direction motion vector mv_0_t' in the third motion information is greater than the absolute value of the first predicted direction motion vector mv_0 in the first motion information, it means that the position of the current block in the first reference image after moving from the current position in the current image to the first reference image according to the first predicted direction motion vector mv_0 exceeds the current position of the first reference block in the first reference image. It can be inferred that the first predicted direction motion vector mv_0 in the first motion information is too small.
  • the absolute value of the second predicted direction motion vector mv_1_t' in the third motion information is compared with the absolute value of the second predicted direction motion vector mv_1 in the first motion information.
  • the absolute value of the second prediction direction motion vector mv_1_t' in the third motion information is smaller than the absolute value of the second prediction direction motion vector mv_1 in the first motion information, it means that the current block moves from the current position in the current image to the second reference image according to the second prediction direction motion vector mv_0, and cannot reach the current position of the second reference block in the second reference image, so it can be inferred that the second prediction direction motion vector mv_0 in the first motion information is too large.
  • the absolute value of the second prediction direction motion vector mv_1_t' in the third motion information is larger than the absolute value of the second prediction direction motion vector mv_1 in the first motion information, it means that the position of the current block in the second reference image after moving from the current position in the current image to the second reference image according to the second prediction direction motion vector mv_0 exceeds the current position of the second reference block in the second reference image, so it can be inferred that the second prediction direction motion vector mv_1 in the first motion information is too small.
  • the decoding end determines the third motion information based on the above steps, and then executes the above step S102-C.
  • using the third motion information to improve the first motion information may result in a greater error, for example, if the object does not move at a uniform speed. Therefore, in an embodiment of the present application, before the decoder improves the first motion information based on the third motion information to obtain the second motion information, it first determines the difference between the first motion information and the third motion information.
  • the embodiment of the present application does not limit the specific manner in which the decoding end determines the difference value between the first motion information and the third motion information.
  • an absolute value of a difference between a motion vector in the first motion information and a motion vector in the third motion information is determined as a difference value between the first motion information and the third motion information.
  • mv_0_x and mv_0_y are motion vectors in the first motion information
  • mv_0_t’_x and mv_0_t’_y are motion vectors in the third motion information
  • diff is the difference value between the first motion information and the third motion information.
  • the step of improving the first motion information based on the third motion information to obtain the second motion information is skipped, and the first motion information is directly improved by a method such as DMVR to obtain the second motion information. If the difference value is less than or equal to the preset threshold value, the above step S102-C is executed to improve the first motion information based on the third motion information to obtain the second motion information.
  • the embodiment of the present application does not limit the specific manner in which the decoding end improves the first motion information based on the third motion information to obtain the second motion information.
  • the decoding end compares the third motion information with the first motion information to adjust the first motion information based on the third motion information to obtain improved second motion information.
  • the first motion information can be adaptively increased to obtain the second motion information. For example, when searching for the second motion information around the first motion information, a larger motion vector can be searched for.
  • the first motion information can be adaptively reduced to obtain the second motion information. For example, when searching for the second motion information around the first motion information, a smaller motion vector can be searched for.
  • the above S102-C includes the following steps S102-C1 and S102-C2:
  • the decoding end determines the fourth motion information based on the third motion information and the first motion information in the following implementations, but is not limited to:
  • an average value of the third motion information and the first motion information is determined as the fourth motion information.
  • mv_0_c_x and mv_0_c_y in formula (17) are motion vectors in the fourth motion information
  • mv_0_x and mv_0_y are motion vectors in the first motion information
  • mv_0_t’_x and mv_0_t’_y are motion vectors in the third motion information.
  • mv_0_c_x and mv_0_c_y in formula (18) are the first predicted direction motion vector in the fourth motion information
  • mv_1_c_x and mv_1_c_y are the second predicted direction motion vector in the fourth motion information
  • mv_0_x and mv_0_y are the first predicted direction motion vector in the first motion information
  • mv_1_x and mv_1_y are the second predicted direction motion vector in the first motion information
  • mv_0_t’_x and mv_0_t’_y are the first predicted direction motion vector in the third operation information
  • mv_1_t’_x and mv_1_t’_y are the second predicted direction motion vector in the third operation information.
  • Method 2 determine the weights corresponding to the third motion information and the first motion information, determine a weighted average of the third motion information and the first motion information based on the weights, and then determine the weighted average as the fourth motion information.
  • the embodiment of the present application does not limit the specific manner in which the decoding end determines the weights corresponding to the third motion information and the first motion information.
  • the weight corresponding to the third motion information is greater than the weight corresponding to the first motion information.
  • the weight corresponding to the third motion information is less than the weight corresponding to the first motion information.
  • the first motion information is a hypothesis (or prediction) of motion information determined based on relevant prediction information provided by the encoder, and the relevant prediction information provided by the encoder includes the encoder from a certain candidate list. A motion information that the encoder considers appropriate is selected, or a prediction mode that the encoder selects that it considers appropriate.
  • the third motion information is a hypothesis (or prediction) of motion information inferred by the decoder, such as the motion information of the current block derived by the decoder based on the motion information on the reference image.
  • the first motion information is obtained through the selection of the encoder, and the third motion information is derived based on the motion information of the reference image, but the reference image is not the current image after all. Therefore, in this example, a higher weight can be set for the first motion information and a lower weight can be set for the third motion information.
  • the decoding end determines the weights corresponding to the third motion information and the first motion information, the third motion information and the first motion information are weighted to obtain fourth motion information.
  • a0* is the weight corresponding to the first motion information
  • b0* is the weight corresponding to the third motion information
  • a0 is greater than b0.
  • b0 1-a0.
  • a0 may be 3/4, 5/8, etc.
  • b0 may be 1/4, 2/8, etc.
  • the division in the above formula can also be replaced by a right shift >>.
  • a0* in formula (21) is the weight corresponding to the first prediction direction motion information (or the first prediction direction motion vector) in the first motion information
  • b0* is the weight corresponding to the second prediction direction motion information (or the second prediction direction motion vector) in the first motion information
  • a1* is the weight corresponding to the first prediction direction motion information (or the first prediction direction motion vector) in the third motion information
  • b1* is the weight corresponding to the second prediction direction motion information (or the second prediction direction motion vector) in the third motion information.
  • different weights may be set according to specific circumstances, such as setting different weights according to different prediction modes.
  • the decoding end After the decoding end determines the fourth motion information based on the above steps, it executes the above S102-C2 to determine the second motion information based on the fourth motion information.
  • the specific implementation methods of the decoding end in the above S102-C2 determining the second motion information based on the fourth motion information include but are not limited to the following:
  • Method 1 Positioning the search center of the second motion information at the position corresponding to the fourth motion information.
  • the above S102-C2 includes the following step S102-C2-a:
  • the second motion information is obtained by searching around the first motion information.
  • the decoding end determines the positioning point of the reference block of the current block in the reference image based on the first motion information.
  • the positioning point can be the position of the upper left corner of the reference block or the center position of the reference block, etc.
  • a search is performed near the positioning point specified by the first motion information, for example, the search is performed with the positioning point as the center, or the search is performed in the same range up, down, left, and right around the positioning point. Exemplarily, in FIG.
  • a square is a pixel position
  • the black square is the positioning point specified by the first motion information in the reference image
  • the white square is the search position of the second motion information.
  • the motion information corresponding to the position with the lowest cost among these square search points is determined as the second motion information.
  • the first motion information in the embodiment of the present application may be inaccurate. For example, based on the motion information of the reference image, it is determined that the first motion information is too large or too small. At this time, when the second motion information is searched with the inaccurate first motion information as the search center, the search of the second motion information is inaccurate.
  • the embodiment of the present application determines the third motion information based on the motion information of the reference image, and then modifies the search center of the second motion information based on the third motion information and the first motion information of the current block.
  • the fourth motion information is determined, and then the corresponding position of the fourth motion information in the reference image is used as the search center point of the second motion information.
  • the second motion information is searched in the reference image to obtain the second motion information. Since the fourth motion information takes into account the motion information of the reference image and the first motion information of the current block, the position specified by the fourth motion information is used as the search center of the second motion information.
  • the accuracy of the search center can be improved, thereby improving the search accuracy of the second motion information and improving the decoding prediction effect.
  • the black square is the first motion information specified as a position in the reference image
  • the gray square is the third motion information specified as a position in the reference image.
  • the position corresponding to the average value of the third motion information and the first motion information (i.e., the fourth motion information) in the reference image is used as the search center of the second motion information.
  • the resulting search range is shown in FIG19B.
  • the search range is shifted to the right, and the search is biased towards smaller motion information, thereby achieving accurate search for the second motion information.
  • the first motion information is less than the third motion information, it means that the first operation information is too small.
  • searching for the second motion information it is more inclined to search for larger motion information.
  • the search range of the second motion information is shifted to the left compared to FIG19A, and the search is biased towards larger motion information, thereby achieving accurate search for the second motion information.
  • the specific process of determining the search range of the second motion information is basically the same.
  • the position corresponding to the fourth motion information in the reference image is used as the search center point of the second motion information
  • the process of searching for the second motion information in the reference image can refer to the schemes shown in Figures 19B and 19C.
  • the position corresponding to the fourth motion information in the reference image is used as the search center point of the second motion information
  • the motion information search is performed within a preset search range near the search center point in the reference image, and the cost of the motion information corresponding to each position point searched is determined, and then the motion information corresponding to the position point with the smallest cost is determined as the current block.
  • the second motion information of the block In unidirectional prediction, the cost of the motion information corresponding to each position point can be represented by the matching cost of the motion information of the template of the position point and the motion information of the template of the current block.
  • the current block includes a first reference image and a second reference image
  • the second motion information and the fourth motion information both include first prediction direction motion information and second prediction direction motion information.
  • the decoding end uses the position corresponding to the first prediction direction motion information in the fourth motion information in the first reference image as the search center point of the first prediction direction motion information in the second motion information, and uses the position corresponding to the second prediction direction motion information in the fourth motion information in the second reference image as the search center point of the second prediction direction motion information in the second motion information, and performs motion information search within the preset search range of the first reference image and the second reference image, and determines the bilateral matching cost of each pair of bilateral motion information searched, wherein each pair of motion information includes a first prediction direction motion information and a second prediction direction motion information; and then determines the second motion information from the multiple pairs of bilateral motion information searched based on the bilateral matching cost.
  • the decoding end may combine n possible MVs corresponding to the first reference image side with n possible MVs corresponding to the second reference image side in pairs to obtain n 2 pairs of bilateral motion information.
  • the MVs of the two reference images are moved in a mirror image, that is, MVdiff is moved on one side and -MVdiff is moved on the other side based on the MVs corresponding to the respective search center points.
  • n pairs of bilateral motion information can be obtained.
  • the embodiment of the present application does not limit the specific method of determining the bilateral matching cost between bilateral motion information.
  • the i-th pair of bilateral motion information includes first prediction direction motion information (for example, first prediction direction motion vector MV0) and second prediction direction motion information (for example, second prediction direction motion vector MV1). Since MV0 and MV1 are both vectors, the distance between MV0 and MV1 can be determined based on the vector distance, and then the distance can be determined as the bilateral matching cost corresponding to the i-th pair of bilateral motion information.
  • first prediction direction motion information for example, first prediction direction motion vector MV0
  • second prediction direction motion information for example, second prediction direction motion vector MV1
  • the decoding end may determine a first prediction block in a first reference image based on the first prediction direction motion information in the i-th pair of bilateral motion information, and determine a second prediction block in a second reference image based on the second prediction direction motion information in the i-th pair of bilateral motion information; determine a matching cost of the first prediction block and the second prediction block; and determine a bilateral matching cost of the i-th pair of bilateral motion information based on the matching cost of the first prediction block and the second prediction block.
  • the SAD cost of the first prediction block and the second prediction block is determined as the matching cost of the first prediction block and the second prediction block, and then the matching cost is determined as the bilateral matching cost of the i-th pair of bilateral motion information, or the matching cost is multiplied or divided by a preset coefficient to obtain the bilateral matching cost of the i-th pair of bilateral motion information.
  • the decoding end can determine the bilateral matching cost of each pair of bilateral motion information among the multiple pairs of bilateral motion information searched, and then determine the pair of bilateral motion information with the smallest bilateral matching cost among the multiple pairs of bilateral motion information searched as the second motion information, and the obtained second motion information is the bilateral motion information.
  • the position corresponding to the fourth motion information in the reference image is used as the search center point of the second motion information, and the specific process of searching for the second motion information in the reference image is introduced.
  • Method 2 modify the cost in the second motion information search process through the fourth motion information.
  • the above S102-C2 includes the following steps S102-C2-b1 to S102-C2-b4:
  • the first cost of each candidate motion information can be searched and determined based on the current first motion information, and then the cost coefficient determined by the fourth motion information can be used to correct the first cost of the candidate motion information to obtain the second cost, and then based on the second cost of each candidate motion information, the second motion information can be selected from each candidate motion information.
  • the decoding end first uses the position corresponding to the first motion information in the reference image as the search center point of the second motion information, performs motion information search in the reference image, and determines the first cost of each candidate motion information searched.
  • the black square is the positioning point specified by the first motion information in the reference image, and the positioning point is used as the search center of the second motion information.
  • the motion information at the position corresponding to the white square is recorded as the candidate motion information of the second motion information, and the first cost of each of these candidate motion information is determined.
  • a cost coefficient corresponding to the candidate motion information is determined based on the candidate motion information and the fourth motion information. For example, the absolute value of the difference between the candidate motion information and the fourth motion information is determined as the cost coefficient corresponding to the candidate motion information. Alternatively, the sum of the absolute value of the difference between the candidate motion information and the fourth motion information and a preset value is determined as the cost coefficient corresponding to the candidate motion information.
  • the first cost of the candidate motion information can be modified by the above-determined cost coefficient to obtain the second cost.
  • the product of the cost coefficient and the first cost of the candidate motion information is determined as the second cost of the candidate motion information.
  • the specific process of determining the search range of the second motion information is basically the same.
  • the decoding end first uses the position corresponding to the first motion information in the reference image as the search center point of the second motion information, performs motion information search within a preset search range of the reference image, and determines the first cost of each candidate motion information searched.
  • the first cost of each candidate motion information can be represented by the matching cost of the motion information of the template at the location of the candidate motion information and the motion information of the template of the current block.
  • the cost coefficient corresponding to each candidate motion information is determined.
  • the absolute value of the difference between the candidate motion information and the fourth motion information is determined as the cost coefficient corresponding to the candidate motion information.
  • the absolute value of the difference between the candidate motion information and the fourth motion information is added to the preset value.
  • the value is determined as the cost coefficient corresponding to the candidate motion information.
  • the first cost is modified based on the cost coefficient corresponding to the candidate motion information to obtain the second cost of the candidate motion information.
  • the product of the cost coefficient and the first cost of the candidate motion information is determined as the second cost of the candidate motion information.
  • the second costs of the multiple candidate motion information searched can be determined, and then the candidate motion information with the smallest second cost among the multiple candidate motion information is determined as the second motion information, and the second motion information is unidirectional motion information.
  • the current block includes a first reference image and a second reference image
  • the first motion information, the second motion information, the fourth motion information and the candidate motion information all include first prediction direction motion information and second prediction direction motion information.
  • the decoding end when searching, uses the position corresponding to the first prediction direction motion information in the first motion information in the first reference image as the search center point of the first prediction direction motion information in the second motion information, and can search for n possible MVs on the first reference image side, and uses the position corresponding to the second prediction direction motion information in the first motion information in the second reference image as the search center point of the second prediction direction motion information in the second motion information, and can search for n possible MVs on the second reference image side.
  • the n possible MVs on both sides are combined in pairs to obtain n 2 pairs of bilateral motion information, and then obtain n 2 candidate motion information, each of which is bilateral motion information.
  • the position corresponding to the first predicted direction motion information in the first motion information in the first reference image is used as the search center point of the first predicted direction motion information in the second motion information
  • the position corresponding to the second predicted direction motion information in the first motion information in the second reference image is used as the search center point of the second predicted direction motion information in the second motion information.
  • the first cost of the candidate motion information is determined, and the first cost may be a bilateral matching cost.
  • the process of determining the first cost of each candidate motion information is the same, and one candidate motion information is used as an example for explanation.
  • the candidate motion information includes first prediction direction motion information (e.g., first prediction direction motion vector MV0) and second prediction direction motion information (e.g., second prediction direction motion vector MV1). Since MV0 and MV1 are both vectors, the distance between MV0 and MV1 can be determined based on the vector distance, and then the distance can be determined as the first cost of the candidate motion information.
  • first prediction direction motion information e.g., first prediction direction motion vector MV0
  • second prediction direction motion information e.g., second prediction direction motion vector MV1
  • the decoding end may determine a first prediction block in a first reference image based on the first prediction direction motion information in the candidate motion information, and determine a second prediction block in a second reference image based on the second prediction direction motion information in the candidate motion information; determine the matching cost of the first prediction block and the second prediction block; and determine the first cost of the candidate motion information based on the matching cost of the first prediction block and the second prediction block.
  • the SAD cost of the first prediction block and the second prediction block is determined as the matching cost of the first prediction block and the second prediction block, and then the matching cost is determined as the first cost of the candidate motion information, or the matching cost is multiplied or divided by a preset coefficient to obtain the first cost of the candidate motion information.
  • a cost coefficient corresponding to the candidate motion information is determined.
  • the second method does not change the search range of the second motion information, but is more inclined to select candidate motion information close to the fourth motion information, and then the first cost of each candidate motion information can be multiplied by a coefficient. For example, a smaller cost coefficient can be set for the candidate motion information close to the fourth motion information, and a larger cost coefficient can be set for the candidate motion information far from the fourth motion information.
  • the corresponding cost coefficient may be one or two.
  • the absolute value of the difference between the first prediction direction motion information in the candidate motion information and the first prediction direction motion information in the fourth motion information can be determined, recorded as difference 1
  • the absolute value of the difference between the second prediction direction motion information in the candidate motion information and the second prediction direction motion information in the fourth motion information can be determined, recorded as difference 2.
  • difference 3 is determined, for example, the sum of difference 1 and difference 2, or the average value, is used to determine difference 3.
  • the candidate motion information corresponds to a cost coefficient
  • the difference 3 is determined as the cost coefficient corresponding to the candidate motion information, or the sum of difference 3 and a preset value is determined as the cost coefficient corresponding to the candidate motion information.
  • a cost coefficient determined and the above-mentioned first cost can be used to determine the second cost of the candidate motion information, for example, the product of the first cost of the candidate motion information and a cost coefficient corresponding to the candidate motion information is determined as the second cost of the candidate motion information.
  • the above S102-C2-b2 includes the following steps:
  • the decoding end determines the first cost coefficient corresponding to the first predicted direction motion information in the candidate motion information based on the first predicted direction motion information in the candidate motion information and the first predicted direction motion information in the fourth motion information, and determines the second cost coefficient corresponding to the second predicted direction motion information in the candidate motion information based on the second predicted direction motion information in the candidate motion information and the second predicted direction motion information in the fourth motion information.
  • the process of determining the first cost coefficient and the process of determining the second cost coefficient at the decoding end are basically the same.
  • the absolute value of the difference between the first prediction direction motion information in the candidate motion information and the first prediction direction motion information in the fourth motion information is determined as the first cost coefficient corresponding to the first prediction direction motion information in the candidate motion information.
  • An absolute value of a difference between the second predicted direction motion information and the second predicted direction motion information in the fourth motion information is determined as a second cost coefficient corresponding to the second predicted direction motion information in the candidate motion information.
  • the absolute value of the difference between the i-th predicted direction motion information in the candidate motion information and the i-th predicted direction motion information in the fourth motion information is determined, i being one or two; based on the absolute value of the difference, the i-th cost coefficient is determined, wherein the i-th cost coefficient is negatively correlated with the absolute value of the difference. That is, the absolute value 1 of the difference between the first predicted direction motion information in the candidate motion information and the first predicted direction motion information in the fourth motion information is determined, and based on the absolute value 1 of the difference, the first cost coefficient is determined, and the first cost coefficient is negatively correlated with the absolute value 1 of the difference, i.e., the larger the distance 1 is, the smaller the first cost coefficient is.
  • the absolute value 2 of the difference between the second predicted direction motion information in the candidate motion information and the second predicted direction motion information in the fourth motion information is determined, and based on the absolute value 2 of the difference, the second cost coefficient is determined, and the second cost coefficient is negatively correlated with the absolute value 2 of the difference, i.e., the larger the absolute value 2 of the difference is, the smaller the second cost coefficient is.
  • the embodiment of the present application does not limit the specific method of determining the i-th cost coefficient based on the absolute value of the difference.
  • the absolute value of the difference is determined as the i-th cost coefficient. That is, the absolute value 1 of the difference between the first predicted direction motion information in the candidate motion information and the first predicted direction motion information in the fourth motion information is used to determine the first cost coefficient, and the absolute value 2 of the difference between the second predicted direction motion information in the candidate motion information and the second predicted direction motion information in the fourth motion information is used to determine the second cost coefficient.
  • the absolute value of the difference and the minimum value in the first preset value are determined; based on the minimum value, the i-th cost coefficient is determined. That is, the absolute value 1 of the difference between the first predicted direction motion information in the candidate motion information and the first predicted direction motion information in the fourth motion information is compared with the first preset value, and the absolute value 1 of the difference and the minimum value 1 in the first preset value are determined, and then the first cost coefficient is determined based on the minimum value 2.
  • the absolute value 2 of the difference between the second predicted direction motion information in the candidate motion information and the second predicted direction motion information in the fourth motion information is compared with the first preset value, and the absolute value 2 of the difference and the minimum value 2 in the first preset value are determined, and then the second cost coefficient is determined based on the minimum value 2.
  • the embodiment of the present application does not limit the specific value of the first preset value.
  • the first preset value is a value greater than 0.
  • the first preset value is 4.
  • the embodiment of the present application does not limit the specific method of determining the i-th cost coefficient based on the minimum value.
  • the minimum value is determined as the i-th cost coefficient.
  • the minimum value 1 is determined as the first cost coefficient, and the minimum value 2 is determined as the second cost coefficient.
  • the sum of the minimum value and the second preset value is determined as the i-th cost coefficient.
  • the sum of the minimum value 1 and the second preset value is determined as the first cost coefficient, and the sum of the minimum value 2 and the second preset value is determined as the second cost coefficient.
  • coef_0 is the first cost coefficient of the candidate motion information
  • coef_0 is the second cost coefficient of the candidate motion information
  • mv_0_x and mv_0_y are the first predicted direction motion information (i.e., the first predicted direction motion vector) in the candidate motion information
  • mv_1_x and mv_1_y are the second predicted direction motion information (i.e., the second predicted direction motion vector) in the candidate motion information.
  • mv_0_c_x and mv_0_c_y are the first predicted direction motion information (i.e., the first predicted direction motion vector) in the fourth motion information
  • mv_0_c_x and mv_0_c_y are the second predicted direction motion information (i.e., the second predicted direction motion vector) in the fourth motion information.
  • a is the first preset value
  • b is the second preset value
  • min() is the operation of taking the minimum value.
  • the division in the above formula (21) can also be replaced by right shift.
  • the embodiment of the present application does not limit the specific values of the first preset value a and the second preset value b.
  • the first preset value a is 4.
  • the second preset value b is 32.
  • the first cost of the candidate motion information is corrected based on the first cost coefficient and the second cost coefficient to obtain the second cost of the candidate motion information.
  • the embodiment of the present application does not limit the specific manner in which the decoding end corrects the first cost of the candidate motion information based on the first cost coefficient and the second cost coefficient to obtain the second cost of the candidate motion information.
  • the first cost coefficient and the second cost coefficient are added together and then multiplied with the first cost of the candidate motion information to obtain the second cost of the candidate motion information.
  • the first cost of the candidate motion information is multiplied by the first cost coefficient and the second cost coefficient to obtain the second cost of the candidate motion information.
  • SAD_c is the second cost of the candidate motion information
  • SAD is the first cost of the candidate motion information
  • coef_0 is the first cost coefficient of the candidate motion information
  • coef_1 is the second cost coefficient of the candidate motion information.
  • the decoding end can obtain the second cost of each of the multiple candidate motion information searched, and the second cost is a cost corrected based on the motion information of the reference image, and its accuracy is higher than the first cost. Then, based on the second cost of each of the multiple candidate motion information, the second motion information of the current block is determined from the multiple candidate motion information. For example, the candidate motion information with the smallest second cost among the multiple candidate motion information searched is determined as the second motion information, so as to achieve accurate determination of the second motion information, thereby improving the prediction accuracy of the current block and improving the decoding effect of the decoding end.
  • the above embodiments all take the current block as a whole block and introduce the process of improving the first motion information of the entire current block as a whole block.
  • the method of the above embodiments is also used for sub-block-based processing, for example, the current block is divided into multiple sub-blocks, and the first motion information of each sub-block is improved separately in the same way as the first motion information of the current block, so as to obtain the second motion information of each sub-block, and then based on the second motion information of each sub-block, the predicted value of each sub-block is obtained, and the predicted value of each sub-block constitutes the predicted value of the current block.
  • the above embodiment introduces the specific process of obtaining the second motion information by improving the first motion information at the decoding end in case 1 when the motion information of the reference image is used to directly participate in improving the first motion information.
  • Case 2 The motion information of the reference image in the embodiment of the present application can be used to guide the division method of the related blocks in the process of improving the first motion information.
  • the above S102 includes the following steps S102-D to S102-F:
  • the reference block of the current block is determined, and the motion information of the reference block is obtained.
  • the motion vector of the gray area in the upper right corner of the reference block is significantly different from that of other areas.
  • a threshold can be set, and the difference between the two motion vectors is considered to be significantly different when it exceeds the threshold. Since the current block has a strong correlation with the reference block, the motion vector of the upper left corner area of the current block is also significantly different from that of other areas. If the current block is improved as a whole, the improvement effect is not obvious.
  • the current block when improving the first motion information of the current block, based on the motion information of the reference image block, the current block is divided into at least one sub-block, and the motion information is improved separately for each sub-block.
  • This not only reduces the hardware implementation cost, but also provides better flexibility when divided into sub-blocks for improvement.
  • the MV can be improved independently for each sub-block, which achieves the effect of improving the accuracy of motion information improvement to a certain extent.
  • the decoding end indicates the division of the current block based on the motion information of the reference image.
  • the embodiment of the present application does not limit the specific manner in which the decoding end divides the current block into at least one sub-block based on the motion information of the reference image.
  • the decoding end first determines the reference block corresponding to the current block in the reference image, and then randomly samples the motion information of several points in the reference block, and then compares the motion information of these points, and divides the points of motion information in these points into a sub-block, and then divides the reference block into at least one sub-block. Then, the decoding end determines the sub-blocks corresponding to each sub-block in the reference block in the current block, and then divides the current block into at least one sub-block.
  • the above S102-D includes the following steps S102-D1 to S102-D4:
  • S102-D3 classify the acquired motion information of the M sub-blocks to obtain P classification results, where P is a positive integer less than or equal to M;
  • the decoding end first determines the reference block of the current block in the reference image based on the first motion information of the current block, and then pre-divides the current block into M sub-blocks, and the sizes of the M sub-blocks can be the same or different. For example, when the M sub-blocks are all 16X16, or 8X8, or 4X4, or when the width and height of the current block are both 2N, the current block can be divided into 4NxN sub-blocks, or when the size of the current block is Nx2N or 2N XN, the current block can be divided into 2NxN sub-blocks. Then, the M sub-blocks corresponding to the M sub-blocks in the reference block in the current block are determined.
  • the decoding end can obtain the motion information of the M sub-blocks corresponding to the M sub-blocks of the current block in the reference block.
  • the correlation between the current block and the reference block is strong, so the decoding end clusters the motion information of the M sub-blocks in the reference block to obtain P classification results, and then divides the current block into at least one sub-block based on the P classification results.
  • the current block is divided into P sub-blocks, one of which corresponds to one classification result.
  • the decoding end divides the current block into at least one sub-block, and then improves each sub-block in the at least one sub-block separately.
  • the improvement process of each sub-block is basically the same.
  • the i-th sub-block is taken as an example for explanation.
  • the embodiment of the present application improves the first motion information of the i-th sub-block, and the specific method of obtaining the second motion information of the i-th sub-block is not limited.
  • the decoding end adopts the above-mentioned DMVR and/or BDOF and other methods to improve the first motion information of the i-th sub-block to obtain the second motion information of the i-th sub-block.
  • the decoding end improves the first motion information of the i-th sub-block by using the motion information of the reference image. Specifically, the decoding end determines the reference block corresponding to the i-th sub-block in the reference image based on the first motion information of the i-th sub-block; determines the time domain motion information of the i-th sub-block as the third motion information corresponding to the i-th sub-block based on the motion information of the reference block of the i-th sub-block; improves the first motion information of the i-th sub-block based on the third motion information to obtain the second motion information of the i-th sub-block.
  • the specific process of this implementation can refer to the specific description of improving the first motion information of the current block in the above situation 1, and it only needs to replace the above current block with the i-th sub-block, which will not be repeated here.
  • the decoding end may determine the second motion information of each sub-block in the current block based on the above steps, and further determine the second motion information of the current block based on the second motion information of at least one sub-block.
  • the second motion information of the at least one sub-block is determined as the second motion information of the current block.
  • the second motion information of the current block includes the second motion information of each sub-block in the at least one sub-block.
  • the prediction value of each sub-block in the at least one sub-block can be determined based on the second motion information of the at least one sub-block, and then the prediction value of the at least one sub-block constitutes the prediction value of the current block.
  • the average value of the second motion information of the at least one sub-block is determined as the second motion information of the current block.
  • the second motion information of the current block includes the motion information of a whole block.
  • the above embodiment introduces the process of improving the motion information of each sub-block in the current block separately.
  • the decoding end when the decoding end improves the first motion information of the current block, it may perform multiple improvements.
  • the above S102 includes the following steps S102-G:
  • S102-G Based on the motion information of the reference image of the current block, perform N rounds of improvement on the first motion information to obtain second motion information, where N is a positive integer greater than 1.
  • the embodiment of the present application does not limit the specific improvement methods used in these N rounds of improvements.
  • the specific improvement methods used in the N rounds of improvement are all different.
  • the blocks of the previous round are divided in the next round of improvement.
  • the decoding end improves the first motion information of the current block based on the motion vector improvement method of the bidirectional matching of the whole block.
  • the current block is divided into at least one sub-block, and the motion information of each sub-block of the current block after the first round of improvement is improved based on the motion vector improvement of the bidirectional matching of the sub-block.
  • the size of the sub-block in the second round can be 16x16.
  • the sub-block of the second round is divided into at least one sub-block, and the motion information after the second round of improvement is improved based on the motion vector improvement of the bidirectional optical flow of the sub-block.
  • the size of the sub-block in this round can be 8x8.
  • further steps can be enriched on this basis, such as a fourth round of improvement, for example, based on the motion vector improvement of the bidirectional optical flow of the 4x4 sub-block, to improve the motion information after the third round of improvement.
  • further improvement can be performed by using improvement methods such as motion vector improvement based on the bidirectional optical flow of the point. Multiple rounds of improvement optimize the motion vector from top to bottom in multiple layers.
  • at least one round of sub-block division is performed based on the motion information of the reference image.
  • the specific division process can refer to the relevant description of S102-D above, thereby improving the rationality and accuracy of the sub-block division.
  • the above S102-G includes the following steps S102-G1 to S102-G4:
  • S102-G4 improve the motion information of each sub-block in at least one sub-block corresponding to the j+1th round, repeat N rounds, and obtain second motion information.
  • the decoding end first performs a first round of improvement on the first motion information of the current block to obtain the first round of improved motion information 1 corresponding to the current block. Then, based on the motion information of the reference image, the current block is divided into blocks to obtain at least one sub-block 2 corresponding to the second round. The motion information of each sub-block 2 in the at least one sub-block 2 corresponding to the second round is improved, and the motion information of the sub-block 2 is the motion information after the first round of improvement, thereby obtaining the improved motion information of the at least one sub-block 2 corresponding to the second round. Then, based on the motion information of the reference image, the sub-block 2 is divided into blocks to obtain at least one sub-block 3 corresponding to the third round.
  • the motion information of each sub-block 3 in the at least one sub-block 3 corresponding to the third round is improved, and the motion information of the sub-block 3 is the motion information after the second round of improvement, thereby obtaining the improved motion information of the at least one sub-block 3 corresponding to the third round.
  • the sub-block 3 is divided into blocks to obtain at least one sub-block 4 corresponding to the fourth round.
  • the motion information of each sub-block 4 in the at least one sub-block 4 corresponding to the fourth round is improved.
  • the motion information of the sub-block 4 is the motion information after the third round of improvement, and then the improved motion information of the at least one sub-block 4 corresponding to the fourth round is obtained.
  • the second motion information of the current block is obtained.
  • the method of dividing the current block or sub-block into blocks based on the motion information of the reference image is basically the same as the above S102-D.
  • the decoding end first determines the reference block corresponding to the second block in the reference image, and the second block is the current block or the sub-block corresponding to the jth round. Then, the motion information of the M sub-blocks corresponding to the M sub-blocks of the second block in the reference block is obtained, where M is a positive integer greater than 1. Then, the obtained motion information of the M sub-blocks is classified to obtain P classification results, where P is a positive integer less than or equal to M.
  • the second block is divided into at least one sub-block, for example, based on the sub-blocks corresponding to the P classification results, the second block is divided into P sub-blocks.
  • S102-D the relevant description of S102-D above for details, which will not be repeated here.
  • the decoding end improves the motion information of the sub-block, and the specific method of obtaining the improved motion information of the sub-block is not limited.
  • the decoding end adopts the above-mentioned DMVR and/or BDOF and other methods to improve the motion information of the above-mentioned sub-blocks.
  • the decoding end improves the motion information of the sub-block by using the motion information of the reference image. Specifically, the decoding end determines the reference block corresponding to the sub-block in the reference image based on the motion information of the sub-block; determines the temporal motion information of the sub-block as the third motion information corresponding to the sub-block based on the motion information of the reference block of the sub-block; improves the motion information of the sub-block based on the third motion information to obtain the improved motion information of the sub-block.
  • the specific process of this implementation can refer to the specific description of improving the first motion information of the current block in the above situation 1, and it only needs to replace the above current block with the sub-block, which will not be repeated here.
  • the above embodiment introduces the process of performing multiple improvements on the first motion information of the current block.
  • the decoding end before the decoding end improves the first motion information based on the motion information of the reference image of the current block to obtain the second motion information of the current block, it first determines whether the current block meets the preset motion vector improvement condition based on the entire block. If it is determined that the current block meets the motion vector improvement condition based on the entire block, the first motion information is improved based on the motion information of the reference image of the current block to obtain the second motion information of the current block.
  • the method of the embodiment of the present application includes the following steps:
  • Step 1 Based on the motion information of the reference image, the current block is divided into blocks to obtain a plurality of first sub-blocks;
  • Step 2 for any first sub-block of the plurality of first sub-blocks, determining whether the first sub-block satisfies a motion vector improvement condition based on the entire block;
  • Step 3 if the first sub-block does not meet the motion vector improvement condition based on the entire block, the first sub-block is divided into blocks based on the motion information of the reference image to obtain a plurality of second sub-blocks;
  • Step 4 for any second sub-block among the multiple second sub-blocks, determine whether the second sub-block satisfies the motion vector improvement condition based on the entire block, and repeat the process until the divided sub-block satisfies the motion vector improvement condition based on the entire block, or the size of the divided sub-block satisfies the preset size.
  • the decoding end when the decoding end determines that the current block does not meet the motion vector improvement condition based on the entire block, it divides the current block layer by layer until the sub-blocks after division meet the motion vector improvement condition based on the entire block, or when the sub-blocks after division do not meet the motion vector improvement condition based on the entire block, but the size of the sub-blocks after division meets the preset size, the sub-block division is stopped.
  • the embodiment of the present application does not limit the specific content of the above-mentioned motion vector improvement condition based on the entire block.
  • the whole-block-based motion vector improvement condition may be that the size of the block satisfies a preset threshold size, and the preset threshold may be 8X8 or 4X4, etc.
  • the current block is divided into blocks to obtain multiple first sub-blocks.
  • the size of the first sub-block is 8X8, the size of the first sub-block does not meet the motion vector improvement condition based on the entire block. Then, based on the motion information of the reference image, the first sub-block is divided into blocks to obtain multiple second sub-blocks.
  • the second sub-block meets the motion vector improvement condition based on the entire block, and then based on the motion information of the reference image, the first motion information of each second sub-block is improved to obtain the second motion information of each second sub-block.
  • the decoder determines whether the current block or the first sub-block or the second sub-block satisfies the motion vector improvement condition based on the entire block through the following steps a to d:
  • Step a determining a reference block corresponding to the first block in the reference image, the first block being the current block or the first sub-block or the second sub-block;
  • Step b obtaining motion information of M sub-blocks corresponding to the M sub-blocks of the first block in the reference block, where M is a positive integer greater than 1;
  • Step c classifying the acquired motion information of the M sub-blocks to obtain P classification results, where P is a positive integer less than or equal to M;
  • Step d Based on the P classification results, determine whether the first block meets the motion vector improvement condition based on the entire block.
  • the specific method for the decoding end to determine whether the current block or the first sub-block or the second sub-block meets the motion vector improvement condition based on the entire block is basically the same.
  • the first block is used to replace the current block or the first sub-block or the second sub-block.
  • the decoding end first determines the reference block corresponding to the first block in the reference image, then divides the first block into M sub-blocks, and obtains the motion information of the M sub-blocks corresponding to the M sub-blocks of the first block in the reference block. Then, the obtained motion information of the M sub-blocks is classified to obtain P types of classification results.
  • the specific implementation process of the above steps a to c can refer to the specific description of S102-D1 to S102-D3 above. It only needs to replace the current block in S102-D1 to S102-D3 with the first block, and P types of classification results corresponding to the first block can be obtained.
  • P is equal to 1
  • P is greater than 1, it means that the motion information of the M sub-blocks in the first block is not completely the same, and the first block needs to be divided, and then it is determined that the first block does not meet the motion vector improvement condition based on the entire block.
  • the current block is divided into multiple first sub-blocks based on the motion information of the reference image, and through the above steps a to d, it is determined whether each first sub-block meets the motion vector improvement condition based on the whole block. If part of the first sub-blocks meet the motion vector improvement condition based on the whole block, the motion information of the part of the first sub-blocks is improved based on the motion information of the reference image.
  • each first sub-block in the part of the first sub-blocks is divided based on the motion information of the reference image to obtain multiple second sub-blocks. Then, through the above steps a to d, it is determined whether each second sub-block in the multiple second sub-blocks meets the motion vector improvement condition based on the whole block, and for the second sub-block that meets the condition, the motion information of the second sub-block is improved based on the motion information of the reference image.
  • the second sub-blocks that do not meet the condition are continued to be divided, and the above steps are repeated until all sub-blocks meet the motion vector improvement condition based on the whole block.
  • a motion vector improvement method based on bidirectional optical flow of the sub-block is used by default to improve the motion information of the sub-block.
  • the process of dividing the current block or the first sub-block or the second sub-block by the decoding end based on the motion information of the reference image is basically the same as the above S102-D.
  • the decoding end first determines the reference block corresponding to the second block in the reference image, and the second block is the current block or the first sub-block or the second sub-block.
  • the motion information of the M sub-blocks corresponding to the M sub-blocks of the second block in the reference block is obtained, where M is a positive integer greater than 1.
  • the obtained motion information of the M sub-blocks is classified to obtain P classification results, where P is a positive integer less than or equal to M.
  • the second block is divided into at least one sub-block, for example, based on the sub-blocks corresponding to the P classification results, the second block is divided into P sub-blocks.
  • S102-D the relevant description of S102-D above for details, which will not be repeated here.
  • the above embodiment combines Case 1 and Case 2 to introduce the process of using the motion information of the reference image at the decoding end to improve the first motion information of the current block, and the process of using the motion information of the reference image to guide block division.
  • the decoding end After obtaining the second motion information of the current block based on the above steps, the decoding end performs the following step S103.
  • the decoding end takes into account the motion information of the reference image when improving the first motion information of the current block, thereby achieving effective improvement of the first motion information of the current block and obtaining accurate second motion information.
  • the prediction accuracy of the current block can be improved, thereby improving the decoding performance of the video.
  • the second motion information includes a motion vector in one direction, and based on the motion vector, a prediction block of the current block is determined in a reference image of the current block to obtain a prediction value of the current block.
  • the current block if the current block adopts bidirectional prediction, the current block includes a first reference image and a second reference image, and the second motion information includes a first prediction direction motion vector and a second prediction direction motion vector.
  • the decoding end determines prediction block 1 in the first reference image based on the first prediction direction motion vector in the second motion information, and determines prediction block 2 in the second reference image based on the second prediction direction motion vector in the second motion information, and then obtains the prediction value of the current block based on prediction block 1 and prediction block 2.
  • the average value or weighted average value of prediction block 1 and prediction block 2 is determined as the prediction value of the current block.
  • the decoding end when decoding the current block, the decoding end first determines the first motion information of the current block, and then improves the first motion information based on the motion information of the reference image of the current block to obtain the second motion information. That is to say, in the embodiment of the present application, when improving the first motion information, the motion information of the reference image is taken into account to achieve effective improvement of the first motion information and obtain accurate second motion information. Then, when determining the prediction value of the current block based on the accurate second motion information, the prediction accuracy of the current block can be improved, thereby improving the decoding performance of the video.
  • Figures 14 to 20 are merely examples of the present application and should not be construed as limitations to the present application.
  • the size of the sequence number of each process does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
  • the term "and/or” is merely a description of the association relationship of associated objects, indicating that three relationships may exist. Specifically, A and/or B can represent: A exists alone, A and B exist at the same time, and B exists alone.
  • the character "/" in the present application generally indicates that the associated objects before and after are in an "or" relationship.
  • FIG. 21 is a schematic block diagram of a video decoding device provided in an embodiment of the present application.
  • the video decoding device 10 is applied to the above-mentioned video decoder.
  • the video decoding device 10 includes:
  • a determination unit 11 configured to determine first motion information of a current block
  • An improving unit 12 configured to improve the first motion information based on the motion information of the reference image of the current block to obtain second motion information of the current block;
  • the prediction unit 13 is configured to determine a prediction value of the current block based on the second motion information.
  • the improvement unit 12 is specifically used to determine, in the reference image, a reference block corresponding to the current block based on the first motion information; determine, based on the motion information of the reference block, the time domain motion information of the current block as the third motion information; and improve the first motion information based on the third motion information to obtain the second motion information.
  • the improving unit 12 is specifically configured to determine fourth motion information based on the third motion information and the first motion information; and determine the second motion information based on the fourth motion information.
  • the improving unit 12 is specifically configured to determine an average value of the third motion information and the first motion information as the determined fourth motion information.
  • the improvement unit 12 is specifically used to determine the weights corresponding to the third motion information and the first motion information; based on the weights, determine the weighted average of the third motion information and the first motion information; and determine the weighted average as the determined fourth motion information.
  • the weight of the first motion information is greater than the weight of the third motion information.
  • the improving unit 12 is specifically configured to use a position corresponding to the fourth motion information in the reference image as a search center point for the second motion information, and search the reference image to obtain the second motion information.
  • the reference image includes a first reference image and a second reference image
  • the second motion information and the fourth motion information both include first predicted direction motion information and second predicted direction motion information
  • the improvement unit 12 is specifically used to use the position corresponding to the first predicted direction motion information in the fourth motion information in the first reference image as the search center point of the first predicted direction motion information in the second motion information, and use the position corresponding to the second predicted direction motion information in the fourth motion information in the second reference image as the search center point of the second predicted direction motion information in the second motion information, perform motion information search within a preset search range of the first reference image and the second reference image, determine the bilateral matching cost of each pair of bilateral motion information searched, each pair of bilateral motion information including a first predicted direction motion information and a second predicted direction motion information; and determine the second motion information from the multiple pairs of bilateral motion information searched based on the bilateral matching cost.
  • the improvement unit 12 is specifically used to determine, for the i-th pair of bilateral motion information searched, a first prediction block in the first reference image based on the first prediction direction motion information in the i-th pair of bilateral motion information, and a second prediction block in the second reference image based on the second prediction direction motion information in the i-th pair of bilateral motion information, where i is a positive integer; determine a matching cost between the first prediction block and the second prediction block; and determine the bilateral matching cost of the i-th pair of bilateral motion information based on the matching cost between the first prediction block and the second prediction block.
  • the improving unit 12 is specifically configured to determine a pair of bilateral motion information with the minimum bilateral matching cost among the searched multiple pairs of bilateral motion information as the second motion information.
  • the improvement unit 12 is specifically used to use the position corresponding to the first motion information in the reference image as the search center point of the second motion information, perform motion information search in the reference image, and determine the first cost of each candidate motion information searched; determine the cost coefficient corresponding to the candidate motion information based on the candidate motion information and the fourth motion information; correct the first cost based on the cost coefficient corresponding to the candidate motion information to obtain the second cost of the candidate motion information; and determine the second motion information based on the second costs of the multiple candidate motion information searched.
  • the reference image includes a first reference image and a second reference image
  • the first motion information, the second motion information, the fourth motion information and the candidate motion information all include first predicted direction motion information and second predicted direction motion information
  • the improvement unit 12 is specifically used to use the position corresponding to the first predicted direction motion information in the first motion information in the first reference image as the search center point of the first predicted direction motion information in the second motion information, and use the position corresponding to the second predicted direction motion information in the first motion information in the second reference image as the search center point of the second predicted direction motion information in the second motion information, perform motion information search within a preset search range of the first reference image and the second reference image, and determine the first cost of each candidate motion information searched.
  • the improvement unit 12 is specifically used to determine the first cost coefficient corresponding to the first predicted direction motion information in the candidate motion information based on the first predicted direction motion information in the candidate motion information and the first predicted direction motion information in the fourth motion information; and determine the second cost coefficient corresponding to the second predicted direction motion information in the candidate motion information based on the second predicted direction motion information in the candidate motion information and the second predicted direction motion information in the fourth motion information.
  • the improvement unit 12 is specifically used to determine the absolute value of the difference between the i-th predicted direction motion information in the candidate motion information and the i-th predicted direction motion information in the fourth motion information, where i is one or two; based on the absolute value of the difference, determine the i-th cost coefficient, wherein the i-th cost coefficient is negatively correlated with the absolute value of the difference.
  • the improving unit 12 is specifically configured to determine a minimum value between the absolute value of the difference and a first preset value; and determine the i-th cost coefficient based on the minimum value.
  • the improving unit 12 is specifically configured to determine the sum of the minimum value and the second preset value as the i-th cost coefficient.
  • the improving unit 12 is specifically configured to modify the first cost of the candidate motion information based on the first cost coefficient and the second cost coefficient to obtain the second cost of the candidate motion information.
  • the improving unit 12 is specifically configured to multiply the first cost with the first cost coefficient and the second cost coefficient to obtain the second cost of the candidate motion information.
  • the improving unit 12 is specifically configured to determine the candidate motion information with the smallest second cost among the searched multiple candidate motion information as the second motion information.
  • the improvement unit 12 before improving the first motion information based on the third motion information to obtain the second motion information, further determines the difference value between the first motion information and the third motion information; if the difference value is less than or equal to a preset threshold, the first motion information is improved based on the third motion information to obtain the second motion information.
  • the improvement unit 12 is specifically used to divide the current block into at least one sub-block based on the motion information of the reference image; for the i-th sub-block in the at least one sub-block, improve the first motion information of the i-th sub-block to obtain the second motion information of the i-th sub-block, where i is a positive integer; and obtain the second motion information of the current block based on the second motion information of the N sub-blocks.
  • the improvement unit 12 is specifically used to determine, in the reference image, a reference block corresponding to the ith sub-block based on the first motion information of the ith sub-block; determine third motion information of the ith sub-block moving from the current image to the reference image according to the motion information of the reference block of the ith sub-block; and improve the first motion information of the ith sub-block based on the third motion information to obtain second motion information of the ith sub-block.
  • the improving unit 12 is specifically configured to perform N rounds of improvement on the first motion information based on the motion information of the reference image of the current block to obtain the second motion information, where N is a positive integer greater than 1.
  • the improvement unit 12 is specifically used to improve the motion information of each sub-block in the at least one sub-block corresponding to the j-th round to obtain the improved motion information of the at least one sub-block corresponding to the j-th round, and if the j is equal to 1, the sub-block corresponding to the j-th round is the current block; based on the motion information of the reference image, each sub-block in the at least one sub-block corresponding to the j-th round is divided into blocks to obtain at least one sub-block corresponding to the j+1-th round; the motion information of each sub-block in the at least one sub-block corresponding to the j+1-th round is improved, and N rounds are repeated to obtain the second motion information.
  • the improvement unit 12 is further used to determine whether the current block satisfies a preset whole-block-based motion vector improvement condition before improving the first motion information based on the motion information of the reference image of the current block to obtain the second motion information of the current block; if the current block satisfies the whole-block-based motion vector improvement condition, the first motion information is improved based on the motion information of the reference image of the current block to obtain the second motion information of the current block.
  • the improvement unit 12 is further used to divide the current block into blocks based on the motion information of the reference image to obtain multiple first sub-blocks; for any first sub-block of the multiple first sub-blocks, determine whether the first sub-block meets the whole-block-based motion vector improvement condition; if the first sub-block does not meet the whole-block-based motion vector improvement condition, divide the first sub-block into blocks based on the motion information of the reference image to obtain multiple second sub-blocks; for any second sub-block of the multiple second sub-blocks, determine whether the second sub-block meets the whole-block-based motion vector improvement condition, and repeat the execution until the divided sub-blocks meet the whole-block-based motion vector improvement condition, or the size of the divided sub-blocks meets the preset size.
  • the improvement unit 12 is specifically used to determine a reference block corresponding to the first block in the reference image, where the first block is the current block or the first sub-block or the second sub-block; obtain motion information of M sub-blocks corresponding to M sub-blocks of the first block in the reference block, where M is a positive integer greater than 1; classify the obtained motion information of the M sub-blocks to obtain P classification results, where P is a positive integer less than or equal to M; based on the P classification results, determine whether the first block meets the whole-block-based motion vector improvement condition.
  • the improvement unit 12 is specifically configured to determine that the first block satisfies the whole-block-based motion vector improvement condition if P is equal to 1; and determine that the first block does not satisfy the whole-block-based motion vector improvement condition if P is greater than 1.
  • the improvement unit 12 is specifically used to determine a reference block corresponding to the second block in the reference image, where the second block is any one of the current block, the first sub-block, the second sub-block, and the sub-block corresponding to the jth round; obtain motion information of M sub-blocks of the second block corresponding to M sub-blocks in the reference block, where M is a positive integer greater than 1; classify the acquired motion information of the M sub-blocks to obtain P classification results, where P is a positive integer less than or equal to M; and divide the second block into at least one sub-block based on the P classification results.
  • the improving unit 12 is specifically configured to divide the second block into P sub-blocks based on the sub-blocks corresponding to the P classification results.
  • the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, no further description is given here.
  • the device 10 shown in FIG. 21 can perform the decoding method of the decoding end of the embodiment of the present application, and the aforementioned and other operations and/or functions of each unit in the device 10 are respectively for implementing the corresponding processes in each method such as the decoding method of the above-mentioned decoding end, and for the sake of brevity, no further description is given here.
  • the functional unit can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software units.
  • the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and/or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or a combination of hardware and software units in the decoding processor to perform.
  • the software unit can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc.
  • the storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.
  • FIG. 22 is a schematic block diagram of an electronic device provided in an embodiment of the present application.
  • the electronic device 30 may be a video decoder as described in an embodiment of the present application, and the electronic device 30 may include:
  • the memory 33 and the processor 32, the memory 33 is used to store the computer program 34 and transmit the program code 34 to the processor 32.
  • the processor 32 can call and run the computer program 34 from the memory 33 to implement the method in the embodiment of the present application.
  • the processor 32 may be configured to execute the steps in the above method 200 according to the instructions in the computer program 34 .
  • the processor 32 may include but is not limited to:
  • DSP digital signal processor
  • ASIC application-specific integrated circuit
  • FPGA field programmable gate array
  • the memory 33 includes but is not limited to:
  • Non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory.
  • the volatile memory can be random access memory (RAM), which is used as an external cache.
  • RAM random access memory
  • SRAM static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDR SDRAM double data rate synchronous dynamic random access memory
  • ESDRAM enhanced synchronous dynamic random access memory
  • SLDRAM synchronous link DRAM
  • Direct Rambus RAM Direct Rambus RAM
  • the computer program 34 may be divided into one or more units, which are stored in the memory 33 and executed by the processor 32 to complete the method provided by the present application.
  • the one or more units may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 34 in the electronic device 30.
  • the electronic device 30 may further include:
  • the transceiver 33 may be connected to the processor 32 or the memory 33 .
  • the processor 32 may control the transceiver 33 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices.
  • the transceiver 33 may include a transmitter and a receiver.
  • the transceiver 33 may further include an antenna, and the number of antennas may be one or more.
  • bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
  • the present application also provides a computer storage medium on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment.
  • the present application embodiment also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.
  • the present application also provides a code stream, which is generated according to the above encoding method.
  • the computer program product includes one or more computer instructions.
  • the computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
  • the computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
  • the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.
  • the computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations.
  • the available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
  • a magnetic medium e.g., a floppy disk, a hard disk, a magnetic tape
  • an optical medium e.g., a digital video disc (DVD)
  • DVD digital video disc
  • SSD solid state disk
  • the disclosed systems, devices and methods can be implemented in other ways.
  • the device embodiments described above are only schematic.
  • the division of the unit is only a logical function division.
  • Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
  • each functional unit in each embodiment of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

提供一种视频解码方法、装置、设备、及存储介质,解码端在对当前块进行解码时,首先确定当前块的第一运动信息(S101),接着,基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到第二运动信息(S102)。也就是说,在本申请实施例中,在对第一运动信息进行改善时,考虑了参考图像的运动信息,实现对第一运动信息的有效改善,得到准确的第二运动信息,进而基于该准确的第二运动信息,确定当前块的预测值时,可以提高当前块的预测准确性,进而提高视频的解码性能。

Description

视频解码方法、装置、设备、及存储介质 技术领域
本申请涉及视频编解码技术领域,尤其涉及一种视频解码方法、装置、设备、及存储介质。
背景技术
数字视频技术可以并入多种视频装置中,例如数字电视、智能手机、计算机、电子阅读器或视频播放器等。随着视频技术的发展,视频数据所包括的数据量较大,为了便于视频数据的传输,视频装置执行视频压缩技术,以使视频数据更加有效的传输或存储。
由于视频中存在时间或空间冗余,通过预测可以消除或降低视频中的冗余,提高压缩效率。在预测时,为了提高解码准确性,则解码端对解码确定的运动信息进行改善。但是,目前的运动信息改善方法,其改善效果不佳,导致解码端的预测不够准确,进而影响视频的解码性能。
发明内容
本申请实施例提供了一种视频解码方法、装置、设备、及存储介质,可以提高运动信息的改善效果提升,当前块的预测准确性和视频的解码性能。
第一方面,本申请提供了一种视频解码方法,应用于解码器,包括:
确定当前块的第一运动信息;
基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息;
基于所述第二运动信息,确定所述当前块的预测值。
第二方面,本申请提供了一种视频解码装置,用于执行上述第一方面或其各实现方式中的方法。具体地,该装置包括用于执行上述第一方面或其各实现方式中的方法的功能单元。
第三方面,提供了一种视频解码器,包括处理器和存储器。该存储器用于存储计算机程序,该处理器用于调用并运行该存储器中存储的计算机程序,以执行上述第一方面或其各实现方式中的方法。
第四方面,提供了一种视频编解码系统,包括视频编码器和视频解码器。视频解码器用于执行上述第一方面或其各实现方式中的方法。
第五方面,提供了一种芯片,用于实现上述第一方面的方法。具体地,该芯片包括:处理器,用于从存储器中调用并运行计算机程序,使得安装有该芯片的设备执行如上述第一方面的方法。
第六方面,提供了一种计算机可读存储介质,用于存储计算机程序,该计算机程序使得计算机执行上述第一方面的方法。
第七方面,提供了一种计算机程序产品,包括计算机程序指令,该计算机程序指令使得计算机执行上述第一方面的方法。
第八方面,提供了一至第二方面中的任一方面或其各实现方式中的方法。
基于以上技术方案,解码端在对当前块进行解码时,首先确定当前块的第一运动信,接着,基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到第二运动信息。也就是说,在本申请实施例中,在对第一运动信息进行改善时,考虑了参考图像的运动信息,实现对第一运动信息的有效改善,得到准确的第二运动信息,进而基于该准确的第二运动信息,确定当前块的预测值时,可以提高当前块的预测准确性,进而提高视频的解码性能。
附图说明
图1为本申请实施例涉及的一种视频编解码系统的示意性框图;
图2是本申请实施例涉及的视频编码器的示意性框图;
图3是本申请实施例涉及的视频解码器的示意性框图;
图4为一种GOP结构示意图;
图5为一种CU划分示意图;
图6为一种空域和时域块的示意图;
图7为一种时域运动信息导出示意图;
图8为SbTMVP的原理示意图;
图9为MMVD的原理示意图;
图10A为仿射Affine的一种原理示意图;
图10B为仿射Affine的另一种原理示意图;
图11为一种权重分配示意图;
图12为DMVR的原理示意图;
图13为模板匹配示意图;
图14为本申请一实施例提供的视频解码方法流程示意图;
图15为当前块和参考块的示意图;
图16为当前块和两个参考块的运动示意图;
图17A为当前块的运动信息大于参考块的运动信息时的一种运动示意图;
图17B为当前块的运动信息小于参考块的运动信息时的一种运动示意图;
图18A为当前块的运动信息小于参考块的运动信息时的另一种运动示意图;
图18B为当前块的运动信息大于参考块的运动信息时的另一种运动示意图;
图19A为一种运动信息搜索示意图;
图19B为另一种运动信息搜索示意图;
图19C为又一种运动信息搜索示意图;
图20为参考块中存在不同运动信息时的示意图;
图21是本申请一实施例提供的视频解码装置的示意性框图;
图22是本申请实施例提供的电子设备的示意性框图。
具体实施方式
本申请可应用于图像编解码领域、视频编解码领域、硬件视频编解码领域、专用电路视频编解码领域、实时视频编解码领域等。例如,本申请的方案可结合至音视频编码标准(audio video coding standard,简称AVS),例如,H.264/音视频编码(audio video coding,简称AVC)标准,H.265/高效视频编码(high efficiency video coding,简称HEVC)标准以及H.266/多功能视频编码(versatile video coding,简称VVC)标准。或者,本申请的方案可结合至其它专属或行业标准而操作,所述标准包含ITU-TH.261、ISO/IECMPEG-1Visual、ITU-TH.262或ISO/IECMPEG-2Visual、ITU-TH.263、ISO/IECMPEG-4Visual,ITU-TH.264(还称为ISO/IECMPEG-4AVC),包含可分级视频编解码(SVC)及多视图视频编解码(MVC)扩展。应理解,本申请的技术不限于任何特定编解码标准或技术。
为了便于理解,首先结合图1对本申请实施例涉及的视频编解码系统进行介绍。
图1为本申请实施例涉及的一种视频编解码系统的示意性框图。需要说明的是,图1只是一种示例,本申请实施例的视频编解码系统包括但不限于图1所示。如图1所示,该视频编解码系统100包含编码设备110和解码设备120。其中编码设备用于对视频数据进行编码(可以理解成压缩)产生码流,并将码流传输给解码设备。解码设备对编码设备编码产生的码流进行解码,得到解码后的视频数据。
本申请实施例的编码设备110可以理解为具有视频编码功能的设备,解码设备120可以理解为具有视频解码功能的设备,即本申请实施例对编码设备110和解码设备120包括更广泛的装置,例如包含智能手机、台式计算机、移动计算装置、笔记本(例如,膝上型)计算机、平板计算机、机顶盒、电视、相机、显示装置、数字媒体播放器、视频游戏控制台、车载计算机等。
在一些实施例中,编码设备110可以经由信道130将编码后的视频数据(如码流)传输给解码设备120。信道130可以包括能够将编码后的视频数据从编码设备110传输到解码设备120的一个或多个媒体和/或装置。
在一个实例中,信道130包括使编码设备110能够实时地将编码后的视频数据直接发射到解码设备120的一个或多个通信媒体。在此实例中,编码设备110可根据通信标准来调制编码后的视频数据,且将调制后的视频数据发射到解码设备120。其中通信媒体包含无线通信媒体,例如射频频谱,可选的,通信媒体还可以包含有线通信媒体,例如一根或多根物理传输线。
在另一实例中,信道130包括存储介质,该存储介质可以存储编码设备110编码后的视频数据。存储介质包含多种本地存取式数据存储介质,例如光盘、DVD、快闪存储器等。在该实例中,解码设备120可从该存储介质中获取编码后的视频数据。
在另一实例中,信道130可包含存储服务器,该存储服务器可以存储编码设备110编码后的视频数据。在此实例中,解码设备120可以从该存储服务器中下载存储的编码后的视频数据。可选的,该存储服务器可以存储编码后的视频数据且可以将该编码后的视频数据发射到解码设备120,例如web服务器(例如,用于网站)、文件传送协议(FTP)服务器等。
一些实施例中,编码设备110包含视频编码器112及输出接口113。其中,输出接口113可以包含调制器/解调器(调制解调器)和/或发射器。
在一些实施例中,编码设备110除了包括视频编码器112和输入接口113外,还可以包括视频源111。
视频源111可包含视频采集装置(例如,视频相机)、视频存档、视频输入接口、计算机图形系统中的至少一个,其中,视频输入接口用于从视频内容提供者处接收视频数据,计算机图形系统用于产生视频数据。
视频编码器112对来自视频源111的视频数据进行编码,产生码流。视频数据可包括一个或多个图像(picture)或图像序列(sequence of pictures)。码流以比特流的形式包含了图像或图像序列的编码信息。编码信息可以包含编码图像数据及相关联数据。相关联数据可包含序列参数集(sequence parameter set,简称SPS)、图像参数集(picture parameter set,简称PPS)及其它语法结构。SPS可含有应用于一个或多个序列的参数。PPS可含有应用于一个或多个图像的参数。语法结构是指码流中以指定次序排列的零个或多个语法元素的集合。
视频编码器112经由输出接口113将编码后的视频数据直接传输到解码设备120。编码后的视频数据还可存储于存储介质或存储服务器上,以供解码设备120后续读取。
在一些实施例中,解码设备120包含输入接口121和视频解码器122。
在一些实施例中,解码设备120除包括输入接口121和视频解码器122外,还可以包括显示装置123。
其中,输入接口121包含接收器及/或调制解调器。输入接口121可通过信道130接收编码后的视频数据。
视频解码器122用于对编码后的视频数据进行解码,得到解码后的视频数据,并将解码后的视频数据传输至显示装置123。
显示装置123显示解码后的视频数据。显示装置123可与解码设备120整合或在解码设备120外部。显示装置123可包括多种显示装置,例如液晶显示器(LCD)、等离子体显示器、有机发光二极管(OLED)显示器或其它类型的显示装置。
此外,图1仅为实例,本申请实施例的技术方案不限于图1,例如本申请的技术还可以应用于单侧的视频编码或单侧的视频解码。
下面对本申请实施例涉及的视频编码框架进行介绍。
图2是本申请实施例涉及的视频编码器的示意性框图。应理解,该视频编码器200可用于对图像进行有损压缩(lossy compression),也可用于对图像进行无损压缩(lossless compression)。该无损压缩可以是视觉无损压缩(visually lossless compression),也可以是数学无损压缩(mathematically lossless compression)。
该视频编码器200可应用于亮度色度(YCbCr,YUV)格式的图像数据上。例如,YUV比例可以为4:2:0、4:2:2或者4:4:4,Y表示明亮度(Luma),Cb(U)表示蓝色色度,Cr(V)表示红色色度,U和V表示为色度(Chroma)用于描述色彩及饱和度。例如,在颜色格式上,4:2:0表示每4个像素有4个亮度分量,2个色度分量(YYYYCbCr),4:2:2表示每4个像素有4个亮度分量,4个色度分量(YYYYCbCrCbCr),4:4:4表示全像素显示(YYYYCbCrCbCrCbCrCbCr)。
例如,该视频编码器200读取视频数据,针对视频数据中的每帧图像,将一帧图像划分成若干个编码树单元(coding tree unit,CTU),在一些例子中,CTB可被称作“树型块”、“最大编码单元”(Largest Coding unit,简称LCU)或“编码树型块”(coding tree block,简称CTB)。每一个CTU可以与图像内的具有相等大小的像素块相关联。每一像素可对应一个亮度(luminance或luma)采样及两个色度(chrominance或chroma)采样。因此,每一个CTU可与一个亮度采样块及两个色度采样块相关联。一个CTU大小例如为128×128、64×64、32×32等。一个CTU又可以继续被划分成若干个编码单元(Coding Unit,CU)进行编码,CU可以为矩形块也可以为方形块。CU可以进一步划分为预测单元(prediction Unit,简称PU)和变换单元(transform unit,简称TU),进而使得编码、预测、变换分离,处理的时候更灵活。在一种示例中,CTU以四叉树方式划分为CU,CU以四叉树方式划分为TU、PU。
视频编码器及视频解码器可支持各种PU大小。假定特定CU的大小为2N×2N,视频编码器及视频解码器可支持2N×2N或N×N的PU大小以用于帧内预测,且支持2N×2N、2N×N、N×2N、N×N或类似大小的对称PU以用于帧间预测。视频编码器及视频解码器还可支持2N×nU、2N×nD、nL×2N及nR×2N的不对称PU以用于帧间预测。
在一些实施例中,如图2所示,该视频编码器200可包括:预测单元210、残差单元220、变换/量化单元230、反变换/量化单元240、重建单元250、环路滤波单元260、解码图像缓存270和熵编码单元280。需要说明的是,视频编码器200可包含更多、更少或不同的功能组件。
可选的,在本申请中,当前块(current block)可以称为当前编码单元(CU)或当前预测单元(PU)等。预测块也可称为预测图像块或图像预测块,重建图像块也可称为重建块或图像重建图像块。
在一些实施例中,预测单元210包括帧间预测单元211和帧内估计单元212。由于视频的一个帧中的相邻像素之间存在很强的相关性,在视频编解码技术中使用帧内预测的方法消除相邻像素之间的空间冗余。由于视频中的相邻帧之间存在着很强的相似性,在视频编解码技术中使用帧间预测方法消除相邻帧之间的时间冗余,从而提高编码效率。
帧间预测单元211可用于帧间预测,帧间预测可以包括运动估计(motion estimation)和运动补偿(motion compensation),可以参考不同帧的图像信息,帧间预测使用运动信息从参考帧中找到参考块,根据参考块生成预测块,用于消除时间冗余;帧间预测所使用的帧可以为P帧和/或B帧,P帧指的是向前预测帧,B帧指的是双向预测帧。帧间预测使用运动信息从参考帧中找到参考块,根据参考块生成预测块。运动信息包括参考帧所在的参考帧列表,参考帧索引,以及运动矢量。运动矢量可以是整像素的或者是分像素的,如果运动矢量是分像素的,那么需要在参考帧中使用插值滤波做出所需的分像素的块,这里把根据运动矢量找到的参考帧中的整像素或者分像素的块叫参考块。有的技术会直接把参考块作为预测块,有的技术会在参考块的基础上再处理生成预测块。在参考块的基础上再处理生成预测块也可以理解为把参考块作为预测块然后再在预测块的基础上处理生成新的预测块。
帧内估计单元212只参考同一帧图像的信息,预测当前码图像块内的像素信息,用于消除空间冗余。帧内预测所使用的帧可以为I帧。
帧内预测有多种预测模式,以国际数字视频编码标准H系列为例,H.264/AVC标准有8种角度预测模式和1种非角度预测模式,H.265/HEVC扩展到33种角度预测模式和2种非角度预测模式。HEVC使用的帧内预测模式有平面模式(Planar)、DC和33种角度模式,共35种预测模式。VVC使用的帧内模式有Planar、DC和65种角度模式,共67种预测模式。
需要说明的是,随着角度模式的增加,帧内预测将会更加精确,也更加符合对高清以及超高清数字视频发展的需求。
残差单元220可基于CU的像素块及CU的PU的预测块来产生CU的残差块。举例来说,残差单元220可产生CU的残差块,使得残差块中的每一采样具有等于以下两者之间的差的值:CU的像素块中的采样,及CU的PU的预测块中的对应采样。
变换/量化单元230可量化变换系数。变换/量化单元230可基于与CU相关联的量化参数(QP)值来量化与CU的TU相关联的变换系数。视频编码器200可通过调整与CU相关联的QP值来调整应用于与CU相关联的变换系数的量化程度。
反变换/量化单元240可分别将逆量化及逆变换应用于量化后的变换系数,以从量化后的变换系数重建残差块。
重建单元250可将重建后的残差块的采样加到预测单元210产生的一个或多个预测块的对应采样,以产生与TU相关联的重建图像块。通过此方式重建CU的每一个TU的采样块,视频编码器200可重建CU的像素块。
环路滤波单元260用于对反变换与反量化后的像素进行处理,弥补失真信息,为后续编码像素提供更好的参考,例如可执行消块滤波操作以减少与CU相关联的像素块的块效应。
在一些实施例中,环路滤波单元260包括去块滤波单元和样点自适应补偿/自适应环路滤波(SAO/ALF)单元,其中去块滤波单元用于去方块效应,SAO/ALF单元用于去除振铃效应。
解码图像缓存270可存储重建后的像素块。帧间预测单元211可使用含有重建后的像素块的参考图像来对其它图像的PU执行帧间预测。另外,帧内估计单元212可使用解码图像缓存270中的重建后的像素块来对在与CU相同的图像中的其它PU执行帧内预测。
熵编码单元280可接收来自变换/量化单元230的量化后的变换系数。熵编码单元280可对量化后的变换系数执行 一个或多个熵编码操作以产生熵编码后的数据。
图3是本申请实施例涉及的视频解码器的示意性框图。
如图3所示,视频解码器300包含:熵解码单元310、预测单元320、反量化/变换单元330、重建单元340、环路滤波单元350及解码图像缓存360。需要说明的是,视频解码器300可包含更多、更少或不同的功能组件。
视频解码器300可接收码流。熵解码单元310可解析码流以从码流提取语法元素。作为解析码流的一部分,熵解码单元310可解析码流中的经熵编码后的语法元素。预测单元320、反量化/变换单元330、重建单元340及环路滤波单元350可根据从码流中提取的语法元素来解码视频数据,即产生解码后的视频数据。
在一些实施例中,预测单元320包括帧内估计单元322和帧间预测单元321。
帧内估计单元322可执行帧内预测以产生PU的预测块。帧内估计单元322可使用帧内预测模式以基于空间相邻PU的像素块来产生PU的预测块。帧内估计单元322还可根据从码流解析的一个或多个语法元素来确定PU的帧内预测模式。
帧间预测单元321可根据从码流解析的语法元素来构造第一参考图像列表(列表0)及第二参考图像列表(列表1)。此外,如果PU使用帧间预测编码,则熵解码单元310可解析PU的运动信息。帧间预测单元321可根据PU的运动信息来确定PU的一个或多个参考块。帧间预测单元321可根据PU的一个或多个参考块来产生PU的预测块。
反量化/变换单元330可逆量化(即,解量化)与TU相关联的变换系数。反量化/变换单元330可使用与TU的CU相关联的QP值来确定量化程度。
在逆量化变换系数之后,反量化/变换单元330可将一个或多个逆变换应用于逆量化变换系数,以便产生与TU相关联的残差块。
重建单元340使用与CU的TU相关联的残差块及CU的PU的预测块以重建CU的像素块。例如,重建单元340可将残差块的采样加到预测块的对应采样以重建CU的像素块,得到重建图像块。
环路滤波单元350可执行消块滤波操作以减少与CU相关联的像素块的块效应。
视频解码器300可将CU的重建图像存储于解码图像缓存360中。视频解码器300可将解码图像缓存360中的重建图像作为参考图像用于后续预测,或者,将重建图像传输给显示装置呈现。
视频编解码的基本流程如下:在编码端,将一帧图像划分成块,针对当前块,预测单元210使用帧内预测或帧间预测产生当前块的预测块。残差单元220可基于预测块与当前块的原始块计算残差块,即预测块和当前块的原始块的差值,该残差块也可称为残差信息。该残差块经由变换/量化单元230变换与量化等过程,可以去除人眼不敏感的信息,以消除视觉冗余。可选的,经过变换/量化单元230变换与量化之前的残差块可称为时域残差块,经过变换/量化单元230变换与量化之后的时域残差块可称为频率残差块或频域残差块。熵编码单元280接收到变化量化单元230输出的量化后的变化系数,可对该量化后的变化系数进行熵编码,输出码流。例如,熵编码单元280可根据目标上下文模型以及二进制码流的概率信息消除字符冗余。
在解码端,熵解码单元310可解析码流得到当前块的预测信息、量化系数矩阵等,预测单元320基于预测信息对当前块使用帧内预测或帧间预测产生当前块的预测块。反量化/变换单元330使用从码流得到的量化系数矩阵,对量化系数矩阵进行反量化、反变换得到残差块。重建单元340将预测块和残差块相加得到重建块。重建块组成重建图像,环路滤波单元350基于图像或基于块对重建图像进行环路滤波,得到解码图像。编码端同样需要和解码端类似的操作获得解码图像。该解码图像也可以称为重建图像,重建图像可以为后续的帧作为帧间预测的参考帧。
需要说明的是,编码端确定的块划分信息,以及预测、变换、量化、熵编码、环路滤波等模式信息或者参数信息等在必要时携带在码流中。解码端通过解析码流及根据已有信息进行分析确定与编码端相同的块划分信息,预测、变换、量化、熵编码、环路滤波等模式信息或者参数信息,从而保证编码端获得的解码图像和解码端获得的解码图像相同。
上述是基于块的混合编码框架下的视频编解码器的基本流程,随着技术的发展,该框架或流程的一些模块或步骤可能会被优化,本申请适用于该基于块的混合编码框架下的视频编解码器的基本流程,但不限于该框架及流程。
在一些实施例中,当前块(current block)可以是当前编码单元(CU)或当前预测单元(PU)等。由于并行处理的需要,图像可以被划分成片slice等,同一个图像中的片slice可以并行处理,也就是说它们之间没有数据依赖。而“帧”是一种常用的说法,一般可以理解为一帧是一个图像。在申请中所述帧也可以替换为图像或slice等。
由上述可知,帧间预测利用时间上的相关性来消除冗余。为了使人眼看不出卡顿,一般的视频的帧率会有30帧每秒,50帧每秒,60帧每秒,甚至120帧每秒。在这样的视频中,同一个场景下的相邻帧之间的相关性很高,帧间预测技术利用这种相关性参考已经编解码的帧的内容对当前要编码的内容进行预测。帧间预测可以极大的提升编码性能。
最基本的帧间预测方法是平移(translational)预测,平移预测假定当前要预测的内容在当前图像与参考图像之间是平移运动的,比如说当前块(编码单元或预测单元)的内容在当前图像与参考图像之间是平移运动的,那么就可以通过一个运动矢量(MV)从参考图像中找到这个内容,并把它作为当前块的预测块。平移运动在视频中是占很大比例,静止不动的背景,整体平移的物体,以及镜头的平移等都可以用平移预测来处理。
自然视频中的一些内容不是简单的平移,比如说在平移的过程中有一些细微地变化,包括形状,颜色等的变化。双向预测从参考图像中找到两个参考块,并将这两个参考块进行加权平均,以得到与当前块尽可能像的预测块。比如说对某些场景,从当前帧的前面和后面各找一个参考块进行加权平均,可能比单独一个参考块更像当前块。基于此双向预测在单向预测的基础上对压缩性能又有提升。
POC(picture order count,图像顺序序号)可以作为图像的一个标识,在一段视频序列中,每一个图像有唯一的POC,本申请实施例中认为POC的顺序和播放顺序是相同的。P图像(P Frame)是只能使用POC在当前图像之前的参考图像进行预测的图像。当前参考图像只有一个参考图像列表,记为RPL0。这里RPL可以理解为Reference Picture List的缩写。参考图像列表RPL0中都是POC在当前图像之前的参考图像。B图像(B Frame)早些时候是可以使用 POC在当前图像之前的参考图像及POC在当前图像之后的参考图像进行预测的图像。B图像有两个参考图像列表,记为RPL0和RPL1。一种配置方法是RPL0中都是POC在当前图像之前的参考图像,RPL1中都是POC在当前图像之后的参考图像。对于一个当前块,可以只参考RPL0中的某一图像的参考块,这样也叫做前向预测;也可以只参考RPL1中的某一图像的参考块,这样也叫做后向预测;也可以同时参考RPL0中的某一图像的参考块和RPL1中的某一图像的参考块,这样也叫做双向预测。同时参考两个参考块的一种简单的方法是将两个参考块每一个对应位置的像素进行平均得到当前块的预测块。后来B图像不再限制RPL0中都是POC在当前图像之前的参考图像,RPL1中都是POC在当前图像之后的参考图像。所以RPL0中也可以有POC在当前图像之后的参考图像,RPL1中也可以有POC在当前图像之前的参考图像。当前块也就可以同时参考POC在当前图像之前的参考图像或同时参考POC在当前图像之后的参考图像。这种B图像也叫广义B图像。
RA(Random Access,随机访问)配置的编解码顺序与POC顺序不同。这样B图像可以参考当前图像之前的信息和当前图像之后的信息从而明显提高了编码性能。RA的一种经典的GOP(group of pictures)结构如图4所示,图中的箭头表示参考关系。I图像不需要参考图像,在POC为0的I图像解码出来后,解码POC为4的P图像,解码POC为4的P图像时可以参考POC为0的I图像。然后解码POC为2的B图像,解码POC为2的B图像时可以参考POC为0的I图像和POC为4的P图等等。
LD(Low Delay,低延迟)配置的编解码顺序与POC顺序相同。所以当前图像只能参考当前图像之前的信息。Low Delay配置又分Low Delay P和Low Delay B。Low Delay P即传统的Low Delay配置。其典型的结构是IPPP……,即先编解码一个I图像,之后的图像都是P图像。Low Delay B的典型结构是IBBB……,与Low Delay P的区别在于每个帧间图像都是B图像,即使用两个参考图像列表,当前块可以同时参考RPL0中的某一图像的参考块和RPL1中的某一图像的参考块。
通常情况下RA配置的压缩效率高于LD配置,LDB配置的压缩效率高于LDP配置。一方面是因为双向预测可以参考到后向的信息,另一方面是因为双向预测可以通过一些技术减少预测误差,比如说加权平均等。
当前图像的一个参考图像列表最多可以有几个参考图像,如2个、3个或4个。当编码某一个当前图像时,RPL0和RPL1中各有哪几个参考图像是由某种配置或算法决定的,不是本发明所讨论的重点。但是同一个参考图像可能同时出现在RPL0和RPL1中。即编解码器允许当前块同时参考同一个参考图像的两个参考块。
编解码器通常使用参考图像列表里的索引值index来对应参考图像。如果一个参考图像列表长度为4,则index有0,1,2,3四个值。举例如当前帧的RPL0有POC为5,4,3,0的4个参考图像。则RPL0的index 0为POC 5的参考图像,RPL0的index 1为POC 4的参考图像,RPL0的index 2为POC 3的参考图像,RPL0的index 3为POC 0的参考图像。
帧间预测使用运动信息(motion information)来表示“运动”。基本的运动信息包含参考图像(reference picture)的信息和运动矢量(MV,motion vector)的信息。一个块为了能使用双向预测,自然需要能找到2个参考块,那么就需要2组参考图像的信息和运动矢量的信息。可以把它们每一组理解为一个单向运动信息,而把这2组组合到一起就形成了一个双向运动信息。在具体实现时,单向运动信息和双向运动信息可以使用相同的数据结构,只是双向运动信息的2组参考帧的信息和运动矢量的信息都有效,而单向运动信息的其中一组参考帧的信息和运动矢量的信息是无效的。所述有效也可以说是“使用”,所述无效也可以说是“不使用”。
VVC支持2个参考图像列表,记为RPL0,RPL1,对上述的双向运动信息,VVC使用RPL0对应的参考图像索引refIdxL0,以及参RPL0对应的运动矢量mvL0,RPL1对应的参考图像索引refIdxL1,以及RPL1对应的运动矢量mvL0。这里的RPL0对应的参考图像索引,RPL1对应的参考图像索引可以理解为上述的参考图像的信息。VVC用两个标志来分别表示是否使用RPL0对应的运动信息以及是否使用RPL1对应的运动信息,分别记为predFlagL0和predFlagL1。也可以理解为predFlagL0和predFlagL1表示上述单向运动信息“是否有效”。所以VVC中虽然没有明确地提到运动信息这种数据结构,但是它用每个参考图像列表对应的参考图像索引,运动矢量以及“是否有效”的标志位一起来表示运动信息。在VVC的标准文本中不出现运动信息,而是使用的运动矢量,也可以认为参考图像索引和是否使用对应运动信息的标志是运动矢量的附属。本文中为了描述方便仍然用“运动信息”,但是应当理解,也可以用“运动矢量”来描述。“运动信息”也可以叫做“运动参数”。
对一个二维的图像来说,运动矢量可以用(x,y)来表示,即一个水平方向的分量和一个竖直方向的分量。由于视频都是以像素来表示的,像素之间是有距离的,一个物体的运动在相邻图像之间不一定总是能对应到整像素距离。举个例子,一个远景的视频,2个像素之间的距离在远景的物体上有1米,而这个物体在2帧之间的时间内运动的距离是0.5米,这种场景用整像素的运动矢量就无法很好地表示。因而运动矢量可以做到分像素级别,如1/2像素精度,1/4像素精度,1/8像素精度,1/16像素精度,来把运动表示得更精细。并通过插值的方法来得到参考图像中的分像素位置的像素值。
上述的平移预测中的单向预测和双向预测都是基于块的,如编码单元或预测单元。即把一个像素矩阵作为单位进行预测。最基本的块就是矩形块,如正方形和长方形。视频编解码标准如HEVC、VVC允许编码器根据视频的内容来确定编码单元、预测单元的大小及划分方式。纹理或运动简单的区域倾向于使用较大的块,纹理或运动复杂的区域倾向于使用较小的块。块划分的层次越深,越能划分出复杂的、更贴近于实际纹理或运动的块,但相应地用于表征这些划分的开销也就更大。运动信息也可能需要在码流中传输。而且通常情况下,块划分地越细,通常运动信息的开销越大。
最原始的运动信息表示方法是直接写入完整的运动信息。后来,专家发现可以使用运动矢量预测MVP(motion vector prediction)加运动矢量差MVD(motion vector difference)来表示运动矢量,即MV=MVP+MVD。MVP越精准,MVD就越小,这样在码流中占用的开销也就更小。
可以理解的是每一个帧间编码的块都需要一个运动信息。为了简化问题,假设CU的划分等于PU的划分等于TU的划分,也就是一个编码单元有一个相同大小相同位置的预测单元一个相同大小相同位置的变换单元。实际上随着CU 划分更加灵活,VVC相对于HEVC就有弱化PU和TU的倾向。预测、变换、量化、熵编码中某一个环节的差异都可能导致CU的划分。例如2个区域的运动信息不同,那么编码器可能把这2个区域划分为不同的CU。又例如2个区域运动信息相同或相似,但是残差特性差别很大,那么编码器也可能把这2个区域划分为不同的CU。如何划分是根据整体的压缩效率决定的,并不完全取决于某一个因素。因而会出现同一个物体或者说运动相同或相似的区域被划分成不同的CU。
示例性的,如图5所示,图5是HEVC中的一个例子,a为原图,图中有一根铁杆在按箭头所指的方向运动,背景区域运动较小。b图中为HEVC的块划分情况,c图中去掉了b图中运动信息相同的块的边界。可以看出很多相邻的块使用了相同的运动信息。这种情况下,如果对每一个块都单独编码运动信息会产生明显的浪费。上面提到VVC的完整的运动信息包括RPL0的参考图像索引,MV是否使用的标志,RPL1的参考图像索引,MV是否使用的标志。合并merge模式的基本原理就是当前块可以继承相邻块的运动信息,包括参考图像的信息和运动矢量的信息。
Merge模式可以构建一个合并候选列表,如果当前块使用Merge模式,可以用一个index指示当前块合并哪一个运动信息,从而不需要编码完整的运动信息。构建合并候选列表时,可以加入当前块空域(spatial)上相邻块的运动信息,时域(temporal)上的运动信息,空域上不相邻块的运动信息,时域上不相邻的运动信息,基于历史的运动信息,合成的运动信息等。
所述空域上的相邻块指同一个图像中与当前块相邻的块,所述空域上不相邻块指同一个图像中与当前块不相邻的块。所述时域上的运动信息以及时域上不相邻块的运动信息指在同位(collocated)参考图像上指定位置的运动信息。示例性的,如图6所示,其中大的灰色块为当前块,其中1,2,3,4,5位置为Merge使用的空域上相邻的块的位置,其他深灰色位置为Merge使用的空域上不相邻的块的位置。6位置为时域上的运动信息使用的位置,如果当前块右下角的对应的位置不可用,则使用当前块中心对应的位置。其他浅灰色位置为时域上的不相邻块的运动信息使用的位置。根据同位参考图像的对应位置上的运动信息导出时域运动信息。
时域运动信息预测被当作空域运动信息预测的补充来使用。通常来说同一个图像上的相邻区域的相关性比不同图像上的相关性更强。但是也有一些情况时域运动信息更好用。举一个简单的例子,比如当前图像中当前块和周边的相邻块属于不同的物体,它们有着截然不同的运动,而在某一个参考图像上与当前块属于同一个物体的块的运动,就能为当前块提供较好的运动信息预测。
示例性的,如图7所示,同位参考图像上的同位块(这里将获取时域运动信息的那个块叫同位块)的运动矢量是从该同位参考图像col_pic到该同位块的参考图像col_ref之间的矢量。而对当前块,它需要的运动矢量是从当前图像curr_pic到当前块的参考图像curr_ref之间的矢量。设col_pic和col_ref之间的POC距离是td,curr_pic和curr_ref之间的POC距离是tb。假设同位块上的运动到当前块上的运动是不变的,那么可以根据td,tb来确定缩放的比例。设同位块的运动矢量为(col_mv_x,col_mv_y),则时域运动矢量预测(tmvp_x,tmvp_y)可以按如下公式(1)推导:
tmvp_x=col_mv_x*tb/td,tmvp_y=col_mv_y*tb/td    (1)
在VVC中,同位参考图像上的存储运动信息的最小单元是4x4。也就是每个4x4的子块存储一组运动信息。可以理解的是如果不考虑硬件实现的代价,同位参考图像也可以做到每个像素存储一组运动信息。
VVC中引入了基于子块的时域运动矢量预测,即SbTMVP(Subblock-based temporal motion vector prediction,基于子块的时间运动矢量预测)。常规上说MVP(motion vector prediction,运动矢量预测),TMVP(temporal motion vector prediction,时间运动矢量预测)都是对整个块来说的,也就是整个块共享同一个MVP。而SbTMVP是基于子块的,这样SbTMVP就能对每一个子块获得一个MVP。这也是SbTMVP和TMVP的本质区别。
另一方面,TMVP用当前块右下角的位置或当前块中心的位置来定位同位块,而SbTMVP根据周边块的运动找一个运动偏移来确定位置。VVC中,如果A1位置的块参考了同位参考图像,那么所述运动位移设为A1使用同位参考图像的运动矢量。否则,所述运动位移设为(0,0)。如图8所示,按运动位移找到位置,然后将“同位块”中的对应于每一个子块的位置的MV进行缩放得到每一个子块的MVP。
Merge模式直接选中的merge候选列表中的运动信息作为当前块的运动信息。在实际的视频中,当前块实际的运动矢量和选中的merge候选列表中的运动矢量有时会有一些差别。Merge模式下的MVD(Merge mode with MVD,简称MMVD)是VVC中的一种特殊的Merge模式,它用一种高效的方法编码这种情况下的MVD。普通的Merge不需要编解码MVD(motion vector difference)。普通的inter模式需要直接编解码MVD。如图9所示,而MMVD利用了MVD更多地分布在单水平方向或单竖直方向的特性,数值小的MVD多,数值越大的MVD越少的特性。
如图9所示,MMVD只能表示一些特定的方向上的特定数值的MVD,它并不能表示任意的MVD。它用mmvd_direction_idx表示MVD的方向,当然也可以理解为MVD的x和y是否非零以及正负号,用mmvd_distance_idx表示MVD的x和y非零的那一个的绝对值的大小MmvdDistance。
示例性的,mmvd_distance_idx[x0][y0]和MmvdDistance[x0][y0]的关系,如表1所示:
表1
其中ph_mmvd_fullpel_only_flag是一个图像头flag,可以设置MMVD的2个不同组合。
示例性的,mmvd_direction_idx[x0][y0]和MmvdSign[x0][y0]的关系,如表2所示:
表2
在一些实施例中,MMVD的MVD按如下公式(2)得到:
MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0]
MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1]    (2)
前面介绍了最简单常用的平移运动,在现实世界中,运动不只有平移,还有很多种形式如缩小,放大,旋转,运动的透视(透视:离镜头近的物体显得大,离镜头远的物体显得小),还有很多不规则的运动形式。Affine(仿射)可以用来表示比平移更复杂的运动。如图10所示,Affine根据2个控制点(4个参数,一个运动矢量包括x,y 2个参数)或者3个控制点(6个参数)的运动矢量,利用线性模型计算出当前块中每一个子块或者每一个像素的运动矢量。
在一种示例中,对4参数Affine模型,当前块内(x,y)位置的运动矢量按如下公式(3)推导:
在一种示例中,对6参数Affine模型,当前块内(x,y)位置的运动矢量按如下公式(4)推导:
其中(mv0x,mv0y)为当前块的左上角的控制点的运动矢量,(mv1x,mv1y)为当前块的右上角的控制点的运动矢量,(mv2x,mv2y)为当前块的左下角的控制点的运动矢量。
VVC中使用的Affine为了简化硬件实现的复杂度,将当前块划分为4x4的子块,为每一个子块计算一个MV并进行运动补偿。示例性的,图10B是Affine基于子块推导运动矢量的一个例子,可以理解的是,随着硬件处理能力的增强,Affine也可以做到基于像素的处理。即为每一个像素推导一个运动矢量,根据该运动矢量对一个像素进行运动补偿。
Affine只需要几个控制点就可以为每一个子块或每一个像素推导出各自的运动矢量,相对于基于整块的运动补偿,它可以做到更精细的预测。而相对于划分更细小的CU,Affine的开销要小得多。
在一些实施例中,HEVC支持最大64x64的CTU,可以递归地进行四叉树的划分。VVC支持一套比HEVC更灵活的块划分方法,支持最大128x128的CTU,包括四叉树,三叉树和二叉树的划分。这些划分方法。虽然块划分越来越灵活,但不管是CU,PU,还是TU,始终只能划分成矩形块。需要注意的是VVC已经弱化了PU和TU的划分。自然视频中纹理或者运动的边界是多种多样的。比如遇到一个斜向的物体边界,如果单纯使用矩形块想要逼近边界就会划分出很多的小块,这样会明显地增加开销。几何划分预测预测模式(GeometricpartitioningMode,GPM)可以更好地处理自然视频中的纹理和边界。
GPM使用两个与当前块大小相同的预测块,GPM的预测块中某些像素位置100%使用第一个预测块对应位置的像素值,某些像素位置100%使用第二个预测块对应位置的像素值,而在交界区域或者称过渡区域,按一定比例使用这两个预测块对应位置的像素值。交界区域的权重也是逐渐过渡的。当然应对诸如屏幕内容编码地场景,也可以不使用过渡区域。具体这些权重如何分配,由GPM的“划分”模式决定。根据GPM的“划分”模式确定每个像素位置的权重。当然在某些情况下,比如说块尺寸很小的情况,可能某些GPM的模式下不能保证一定有某些像素位置100%使用第一个预测块对应位置的像素值,某些像素位置100%使用第二个预测块对应位置的像素值。也可以认为GPM使用两个与当前块大小不相同的预测块,即各取所需的一部分。将权重为0的部分剔除出来。这是实现的问题,不是本发明讨论的重点。
示例性地,图11为VVC中的GPM在正方形的块上的64种模式的权重图。黑色表示第一个预测块对应位置的权重值为0%,白色表示第一个预测块对应位置的权重值为100%,灰色区域则按颜色深浅的不同表示第一个预测块对应位置的权重值为大于0%小于100%的某一个权重值。第二个参考块对应位置的权重值则为100%减去第一个参考块对应位置的权重值。
GPM可以说是一种预测模式或预测方法,因为它最终产生一个预测块。也可以说GPM是一种“划分”模式,它把预测块进行了模拟划分,它类似于实现了PU划分,但又没有实质地划分。上述GPM使用的第一个预测块和第二个预测块既可以是帧内预测产生的预测块,也可以是帧间的单向预测产生的预测块,还可以是帧间的双向预测产生的预测块。
在一些实施例中,一般消费类的视频码率是受限的,因而视频压缩通常要去寻求码流开销和失真的折衷。以块划分为例,对相同的内容,在一定范围内,划分地越细开销越大失真越少;划分越粗开销越少失真越大。以运动信息的编码为例,对相同的内容,在一定范围内,运动信息越精准开销越大失真越小;运动信息越粗放开销越小失真越大。一些解码端的方法在不占用开销情况下,利用解码端侧的信息进行处理,计算,达到改善运动信息,改善预测效果,减少失真的效果。不占用开销也就意味着没有编码器根据原始图像做出的指示,它是根据可以得到的信息自动处理的。VVC中2个典型的解码端方法是DMVR(Decoder side motion vector refinement)和BDOF(bi-directional optical flow)。
VVC中的DMVR启动的一个条件是当前块的2个参考图像分别来自于当前图像的一前一后,而且2个参考图像与当前图像的距离相等。另一个启动条件是当前CU使用的是整块的merge模式(包括skip),所述整块的即不包含基于子块的merge如SbTMVP和affine merge,因为merge模式下运动矢量容易出现不够精确的情况。还有一些其他的 条件这里不再赘述。VVC中的DMVR利用双边匹配BM(bilateral matching)也就是对两边的参考块计算匹配代价,如SAD(sum of absolute difference)。DMVR搜索原始MV周边的MV的匹配代价,在移动时,两个参考图像的MV是按镜像移动的,也就是在各自原始MV的基础上一边移动MVdiff,另一边移动-MVdiff,如图12所示。搜索时也支持分像素的搜索,所以DMVR有可能找到比原有MV精度更高的MV。按一定的规则进行搜索,一般先搜索一定范围内的整像素MV,找到匹配代价最小的整像素MV,再在该整像素MV的基础上搜索分像素MV。如果找到了比原始MV匹配代价更小的MV,则用匹配代价更小的MV做运动补偿预测。DMVR改善后的MV理论上可以用于存储MV及周边块使用,比如在当前块构建merge候选列表时,如果周边块用DMVR改善了MV,使用改善后的MV用于构建merge候选列表能达到更好的压缩效果,但是出于硬件实现的考虑,VVC并没有这么做。
DMVR是可以基于子块处理的。实际上在VVC中,如果块的水平方向或竖直方向大于16像素,会以16像素的大小分割子块。这一方面是基于硬件实现复杂度的考虑,因为DMVR需要在解码端进行搜索,而限制子块的大小可以减少缓存的代价。另一方面,划分成子块处理提供了更好的灵活性,每个子块可以独立改善MV,从一定程度上达到了提高划分精度的效果,这也提升了压缩效率。
双向光流(bi-directional optical flow,简称BDOF)也是一个典型的解码端方法。BDOF基于光流原理改善MV和预测。光流是空间运动物体在观察成像平面上的像素运动的瞬时速度。光流有一些基本的假设,如亮度恒定不变,即同一目标在不同图像之间运动时,其亮度不发生变化。时间连续或运动是小运动。即时间的变化不会引起目标位置的剧烈变化。
VVC中的BDOF启动的一个条件是当前块的2个参考图像分别来自于当前图像的一前一后,而且2个参考图像与当前图像的距离相等。VVC对每一个4x4子块,BDOF会推导出一个运动矢量偏差(vx,vy),这个偏差是通过最小化两个方向的预测值的差值计算得到的。这个运动矢量偏差也被用于调整对应子块里的预测值。
示例性的,预测值的推导过程包括:
首先计算2个预测块的水平方向和竖直方向的梯度
示例性的,通过如下公式(5)确定
其中I(k)(i,j)是参考图像列表k,k=0,1的坐标(i,j)的预测值,shift1根据亮度的比特深度bitDepth计算,shift1=max(6,bitDepth-6)。
接着,计算运动矢量偏差(vx,vy)。
示例性的,通过如下公式(6)确定(vx,vy):
其中,th′BIO=2max(5,BD-7).是向下取整,
示例性的,上述S1,S2,S3,S5和S6按如下公式(7)计算:
其中,θ(i,j)=(I(1)(i,j)>>nb)-(I(0)(i,j)>>nb)。Ω是一个当前4x4子块周围的6x6的窗,na为min(1,bitDepth-11),nb为min(4,bitDepth-8)。
然后,根据运动矢量偏差和梯度,对该4x4子块内的每一个预测值进行调整。
示例性的,通过如下公式(8),确定预测值的调整值:
最后,基于上述预测值的调整值对当前块的预测值进行调整,得到BDOF的预测值。
示例性的,通过如下公式(9),确定BDOF的预测值:
predBDOF(x,y)=(I(0)(x,y)+I(1)(x,y)+b(x,y)+ooffset)>>shift      (9)
其中Ooffset和shift根据亮度的比特深度计算得到。na,nb和shift都是为了降低计算过程中的位宽做的处理。
BDOF的运动矢量偏差可以做到很高的精度从而使预测更准确,而基于子块的处理也提高了灵活性,这两个方面 和DMVR是类似的。
DMVR和BDOF都有改善运动矢量的效果。DMVR基于块匹配,BDOF基于光流原理。它们是可以组合使用的。一个例子如下,可以称之为多轮解码端运动矢量改善(Multi-pass decoder-side motion vector refinement,简称MDMVR)。
示例性的,MDMVR可以包括:第一步,基于整块的双向匹配的运动矢量改善。第二步,基于子块的双向匹配的运动矢量改善。这一步的子块大小可以是16x16。第三步,基于子块的双向光流的运动矢量改善。这一步的子块大小可以是8x8。当前在此基础上还可以进一步丰富步骤,如再来一个第四步,基于4x4子块的双向光流的运动矢量改善。或者再来一个基于点的双向光流的运动矢量改善等。
模板匹配(template matching)的方法最早用在帧间预测中,它利用相邻像素之间的相关性,把当前块周边的一些区域作为模板。在当前块进行编解码时,按照编码顺序其左侧及上侧已经编解码完成。当然在现有的硬件解码器实现时,不一定能保证当前块开始解码时,其左侧和上侧已经解码完成,当然这里说的是帧间块,比如在HEVC中帧间编码的块产生预测块时是不需要周边的重建像素的,因而帧间块的预测过程可以并行进行。但是帧内编码的块是一定需要左侧和上侧的重建像素作为参考像素的。理论上左侧和上侧是可得的,也就是说硬件设计做相应的调整是可以实现的。相对来说右侧和下侧在现在标准如VVC的编码顺序下是不可得的。
示例性的,如图13所示,把当前块的左侧和上侧的矩形区域设为模板,左侧的模板部分的高度一般和当前块的高度相同,上侧的模板的部分的宽度一般和当前块的宽度相同,当然也可以不同。在参考图像中寻找模板的最佳匹配位置从而确定当前块的运动信息或者说运动矢量。这个过程大致可以描述为,在某一个参考图像中,从一个起始位置开始,在周边一定范围内进行搜索。可以预先设定好搜索的规则,如搜索范围搜索步长等。每移动一个到位置,计算该位置对应的模板和当前块周边的模板的匹配程度,所谓匹配程度可以用一些失真代价来衡量,比如说SAD(sum of absolute difference),SATD(sum of absolute transformed difference),一般SATD使用的变换是Hadamard变换,MSE(mean-square error)等,SAD,SATD,MSE等的值越小代表匹配程度越高。用该位置对应的模板的预测块和当前块周边的模板的重建块计算代价。除了整像素位置的搜索还可以进行分像素位置的搜索,根据搜索到的匹配程度最高的位置来确定当前块的运动信息。利用相邻像素之间的相关性,对模板合适的运动信息可能也是当前块合适的运动信息。当然模板匹配的方法可能并不一定对所有的块都适用,因而可以使用一些方法确定当前块是否使用上述模板匹配的方法,比如在当前块用一个控制开关表示是否使用模板匹配的方法。这种模板匹配的方法的一个名字叫DMVD(decoder side motion vector derivation)。编码器和解码器都可以利用模板进行搜索从而导出运动信息或者在原有的运动信息的基础上找到更好的运动信息。而它不需要传输具体的运动矢量或运动矢量差,而是由编码器和解码器都进行同样规则的搜索从而保证编码和解码的一致。模板匹配的方法可以提高压缩性能,但是它需要在解码端也进行“搜索”,从而带来了一定的解码端复杂度。
由上述可知,预测方式各种各样,编码器可以决定当前块使用哪一种预测模式或者说模型,比如是否使用merge模式,是否使用MMVD模式,是按整块预测还是子块预测,如果是按整块预测,merge候选列表里面还有空域的运动矢量预测和时域的运动矢量预测,如果是按子块预测,子块的候选中还有诸如SbTMVP和Affine之类的模式。是否使用GPM模式等。一方面越详细越准确的信息通过码流告诉解码器,解码器就能做出更好的预测,但相应的开销越大。编码器需要在码率和失真中间做出权衡。上述这些方式诸如merge,MMVD,GPM,SbTMVP,Affine等,通过更高效的方式把尽可能多的信息告诉解码端,解码端按照编码端的指示执行。而另一方面,解码端算法诸如DMVR和BDOF,它们通过块匹配或者光流的方法弥补运动矢量不够准确带来的失真,这一定程度上也给了编码器更大的空间少传信息以节省码流开销。可以说解码端越“聪明”,在编码端给出相同的指示的情况下,解码端就可以做出更小的失真。
在这些现有技术构建的编解码框架中,解码器从码流中获得运动信息的指示从而获得初始运动信息,根据初始运动信息找到参考块或者参考块周围的区域,并根据参考块或参考块周围的区域的像素值信息改善初始运动信息和/或改善预测值。例如DMVR在初始MV的周边进行搜索,其搜索过程以及最终选择哪一个MV依赖于块匹配的结果,再例如BDOF也是根据像素值计算梯度等信息,通过光流法算出瞬时的运动矢量差。
但是,目前解码端对解码确定的初始运动信息进行改善时,存在改善效果不佳,导致解码确定的当前块的预测值不够准确,进而影响视频的编解码性能的问题。
为了解决上述技术问题,本申请在确定出当前块的第一运动信息后,基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到第二运动信息,其中该参考图像的运动信息用于改善第一运动信息,和/或用于块划。也就是说,在本申请实施例中,在对第一运动信息进行改善时,考虑了参考图像的运动信息,实现对第一运动信息的有效改善,得到准确的第二运动信息,进而基于该准确的第二运动信息,确定当前块的预测值时,可以提高当前块的预测准确性,进而提高视频的解码性能。
下面结合图14,以解码端为例,对本申请实施例提供的视频解码方法进行介绍。
图14为本申请一实施例提供的视频解码方法流程示意图,本申请实施例应用于图1和图3所示视频解码器。如图14所示,本申请实施例的方法包括:
S101、确定当前块的第一运动信息。
本申请实施例提供的解码方法应用于帧间预测,用于对当前块的运动信息进行改善。
由上述可知,编码端出于码率的考量,在码流中携带的预测相关信息越来越少,这样使得解码端基于码流中所携带的预测相关信息,得到的当前块的运动信息不够准确,因此,解码端可以对该解码确定的运动信息进行改善,以提高预测效果。但是,目前的运动信息改善过程中,未考虑参考图像的运动信息,进而使得运动信息的改善不尽如意。
在本申请实施例中,在对当前块的运动信息进行改善时,考虑了参考图像的运动信息,实现对当前块的运动信息的有效改善,进而提高当前块的预测准确性,提升视频的解码性能。
在一些实施例中,上述当前块的第一运动信息可以理解为当前块的初始运动信息,即解码端通过解码码流,得到码流携带的预测相关信息,并基于该预测相关信息,确定得到运动信息。该预测相关信息可以包括预测模式等信息。
在一些实施例中,上述当前块的第一运动信息可以理解为对当前块的初始运动信息已一次或多次改善后的运动信息,通过本申请实施例的方法,基于参考图像的运动信息,对该第一运动信息进行再次改善。
由上述可知,帧间预测使用运动信息(motion information)来表示“运动”。基本的运动信息包含参考图像(reference picture)的信息和运动矢量(MV,motion vector)的信息。在一些实施例中,若一个块使用双向预测时,则需要能找到2个参考块,那么就需要2组参考图像的信息和运动矢量的信息,可以每一组理解为一个单向运动信息,而把这2组组合到一起就形成了一个双向运动信息。
在一些实施例中,本申请实施例运动信息可以是指单向运动信息,即包括一组参考图像的信息和运动矢量的信息。
在一些实施例中,本申请实施例的运动信息可以指双向运动信息,即包括2组参考图像的信息和运动矢量的信息。
在一些实施例中,本申请实施例的运动信息可以指多向运动信息,即包括多组参考图像的信息和运动矢量的信息。
在一些实施例中,可以使用每个参考图像列表对应的参考图像索引,运动矢量以及“是否有效”的标志位一起来表示运动信息。
本申请实施例对解码端确定当前块的第一运动信息的具体方式不做限制。
在一种可能的实现方式中,编码端将当前块的预测模式携带在码流中。这样解码端通过解码码流,得到当前块的预测模式,进而基于该预测模式,确定当前块的第一运动信息。
例如,解码端基于该预测模式,得到当前块的初始运动信息,将该初始运动信息确定为第一运动信息。
再例如,解码端基于该预测模式,得到当前块的初始运动信息,接着对该初始运动信息进行改善,将改善后的初始运动信息确定为第一运动信息。示例性的,解码端对初始运动信息进行改善的方法可以是采用上述DMVR和/或BDOF的方式进行改善。例如解码端使用DMVR的改善方法,对当前块的初始运动信息进行改善,得到第一运动信息。再例如,解码端使用BDOF的改善方法,对当前块的初始运动信息进行改善,得到第一运动信息。再例如,解码端先使用DMVR改善方法对初始运动信息进行改善后,再使用BDOF改善方法进行改善,得到第一运动信息。再例如,解码端先使用BDOF改善方法对初始运动信息进行改善后,再使用BDOF改善方法进行改善,得到第一运动信息。其中DMVR和BDOF的具体改善方法,参照上述实施例的描述,在此不再赘述。
S102、基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到当前块的第二运动信息。
其中,参考图像的运动信息用于改善第一运动信息,和/或用于块划分。
由上述可知,本申请实施例的运动信息包括参考图像的信息和运动矢量的信息,基于此,解码端可以从上述确定的当前块的第一运动信息,获得当前块的参考图像的信息,例如获得参考图像的索引,进而基于该索引从参考图像列表中,得到当前块的参考图像。
在一些实施例中,若本申请实施例中当前块的预测为单向预测时,则该当前块对应一个参考图像列表,记为RPL0。接着,确定当前块的参考图像在该参考图像列表RPL0中的索引,进而基于该索引,将该参考图像列表RPL0中该索引对应的参考图像,确定为当前块的参考图像。示例性的,解码端确定当前块的参考图像在该参考图像列表RPL0中的索引的方式至少包括如下几种:方式1,编码端和解码端默认将该参考图像列表RPL0中的每一个参考图像,例如第一个参考图像,确定为当前块的参考图像,这样编码端无需在码流中指示当前块的参考图像的索引。方式2,编码端将当前块的参考图像在该参考图像列表RPL0中索引写入码流,这样解码端通过解码码流,得到该当前块的参考图像在该参考图像列表RPL0中索引,进而基于该索引,得到当前块的参考图像。
在一些实施例中,若本申请实施例中当前块的预测为双向预测时,则该当前块对应两个参考图像列表,分别记为RPL0和RPL1,解码端确定当前块的一个参考图像在参考图像列表RPL0中索引refIdxL0,另一个参考图像在参考图像列表RPL1中索引refIdxL1,进而基于这两个索引,在参考图像列表RPL0和参考图像列表RPL1中,确定出当前块的两个参考图像。示例性的,解码端确定当前块的参考图像在该参考图像列表RPL0中的索引的方式至少包括如下几种:方式1,编码端和解码端默认将该参考图像列表RPL0中的每一个参考图像,例如第一个参考图像,确定为当前块的参考图像,这样编码端无需在码流中指示当前块的参考图像的索引。方式2,编码端将当前块的参考图像在该参考图像列表RPL0中索引写入码流,这样解码端通过解码码流,得到该当前块的参考图像在该参考图像列表RPL0中索引,进而基于该索引,得到当前块的参考图像。
解码端基于上述步骤,确定出当前块的参考图像,由于参考图像均为已解码的图像,其运动信息已知。因此,解码端可以直接得到参考图像的运动信息,进而基于参考图像的运动信息和上述步骤确定的当前块的第一运动信息,确定当前块的第二运动信息,其中第二运动信息可以理解为对第一运动信息进行改善后得到的更加准确的运动信息。
在本申请实施例中,参考图像的运动信息在改善第一运动信息中的作用至少包括两种,一种是参考图像的运动信息直接参与第一运动信息的改善,例如用于对第二运动信息的搜索过程做指导,第二种是,在第一运动信息的改善过程中,参考图像的运动信息用于指示块的划分。
下面对解码端基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到当前块的第二运动信息的过程进行介绍。
本申请实施例对解码端基于参考图像的运动信息,对第一运动信息进行改善,得到当前块的第二运动信息的具体过程不做限制。
情况1中,若参考图像的运动信息用于直接参与第一运动信息的改善时,则解码端至少通过如下实施例所示的几种方式,实现对第一运动信息的改善,得到第二运动信息。
在一些实施例中,在参考图像中确定当前块的模板对应的模板参考区域,其中模板参考区域的运动信息已知,当前块的模板的运动信息也已知,因此可以基于模板参考区域的运动信息和当前块的模板的运动信息,对当前块的第一运动信息进行改善,得到当前块的第二运动信息。例如,确定模板参考区域的运动信息和当前块的模板的运动信息之间的差异值,在第一运动信息上增加该差异值,得到当前块的第二运动信息。
在一些实施例中,上述S102包括如下S102-A至S102-C的步骤:
S102-A、基于第一运动信息,在参考图像中确定当前块对应的参考块;
S102-B、基于当前块按照参考块的运动信息,确定当前块的时域运动信息作为第三运动信息;
S102-C、基于第三运动信息对第一运动信息进行改善,得到第二运动信息。
在该实施例中,解码端首先基于第一运动信息,在参考图像中确定当前块对应的参考块。由于参考块的运动信息已知,因此解码端可以基于参考块的运动信息,判断当前块的第一运动信息是否准确。例如,假设参考块和当前块属于图像中的同一个整体运动的物体,且整个物体的运动是一个均速运动,这样可以将参考块作为当前块的同位块,基于参考块的运动信息,确定出当前块的时域运动信息,为了便于描述,将该时域运动信息记为第三运动信息,进而基于该第三运动信息对第一运动信息进行改善,实现对第一运动信息的准确改善。
由上述可知,第一运动信息包括当前块的运动矢量的信息,在本申请实施例中,对第一运动信息的改善可以理解为对第一运动信息所包括的运动矢量的改善。
具体的,解码端首先基于第一运动信息,在参考图像中确定当前块对应的参考块,其中当前块的第一运动信息可以为单向运动信息,也可以为双向运动信息,下面对这两种情况分别进行介绍。
在一种示例中,若当前块采用单向预测时,则当前块的第一运动信息包括单向运动信息,即当前块对应一个参考图像以及一个运动矢量。假设当前图像curr_pic中的当前块curr_block的第一运动信息所包括的参考图像为参考图像序列RPL0中的参考图像ref_pic_0,运动矢量为第一运动矢量mv_0。假设参考图像ref_pic_0的播放顺序在当前图像curr_pic之前。这样,如图15所示,解码端根据当前块curr_block的位置和当前块的第一矢量mv_0,可以在参考图像ref_pic_0中定位到对应的参考块ref_block_0。需要说明的是,图15中参考图像中虚线框表示的当前块可以理解为当前块在参考图像中的同位块。
在一种示例中,若当前块采用双向预测时,则当前块的第一运动信息包括双向运动信息,即当前块对应两个参考图像,记为第一参考图像和第二参考图像,以及两个运动矢量记为第一运动矢量和第二运动矢量。假设当前图像curr_pic中的当前块curr_block的第一运动信息所包括的第一参考图像为参考图像序列RPL0的参考图像ref_pic_0,第一运动矢量为mv_0,当前块的第二参考图像为参考图像序列RPL1中的参考图像ref_pic_1,第二运动矢量为mv_1。假设第一参考图像ref_pic_0的播放顺序在当前图像curr_pic之前,第二参考图像ref_pic_1的播放顺序在当前图像curr_pic之后。这样,如图16所示,解码端根据当前块curr_block的位置和第一运动矢量mv_0,可以在第一参考图像ref_pic_0中定位到对应的第一参考块ref_block_0,以及根据当前块curr_block的位置和第二运动矢量mv_1,可以在第二参考图像ref_pic_1中定位到对应的第二参考块ref_block_1。
基于上述步骤,解码端基于当前块的第一运动信息,确定出当前块在参考图像中的参考块,由于参考图像均为已解码图像,其运动信息已知,因此参考块的运动信息也可以得到。
在本申请实施例中,解码端可以根据参考块的运动信息,来推测当前块的第一运动矢量更有可能往哪一个方向偏移。具体的,基于参考块的运动信息,确定当前块的时域运动信息,记为第三运动信息,将该第三运动信息与当前块的第一运动信息进行比较,以确定第一运动信息更有可能往哪一个方向偏移。
上述S102-B中确定第三运动信息的实现方式包括但不限于如下几种:
方式一,解码端将当前块作为参考块的同位块,根据参考块的运动信息,确定参考块的时域运动信息,进而将该时域运动信息的相反数,确定为第三运动信息。
在一种示例中,若当前块采用单向预测时,假设参考块内部的运动都是相同的,解码端根据参考块的运动信息可以推测出参考块从参考图像中的当前位置出发向当前图像中的当前块运动时的时域运动信息中的矢量信息,记为mv_t。例如,解码端利用时域运动信息的导出方法来推导mv_t。进而将该mv_t的相反向量-mv_t,确定为第三运动信息中的运动矢量。
示例性的,通过如下公式(10),确定参考块的时域运动信息的运动矢量mv_0_t:
mv_t_x=-ref_mv_x*tb/td,mv_t_y=-ref_mv_y*tb/td     (10)
其中,(ref_mv_x和ref_mv_y)为参考块在x轴和y轴上的运动矢量,td为同位参考图像col_pic和同位块的参考图像col_ref之间的POC距离,tb为当前图像curr_pic和参考图像curr_ref之间的POC距离。
假设参考块和当前块都属于同一个整体运动的物体,且这个物体的运动是一个匀速直线运动,而且参考块的运动信息是准确的,将上述确定的第三运动信息中的运动矢量-mv_t的绝对值与当前块的第一运动信息中的运动矢量mv_0的绝对值进行比较,以判断第一运动信息更有可能往哪一个方向偏移。例如,若第三运动信息中的运动矢量-mv_t的绝对值小于当前块的第一运动信息中的运动矢量mv_0的绝对值,则说明参考块按其现有的运动只能达到图17A所示虚线框位置,那么可以推测出第一运动信息中当前块的运动矢量mv_0偏大。再例如,若第三运动信息中的运动矢量-mv_t的绝对值大于当前块的第一运动信息中的运动矢量mv_0的绝对值,则说明参考块按其现有的运动可以达到图17B所示虚线框位置,那么可以推测出第一运动信息中当前块的运动矢量mv_0偏小。
在一种示例中,若当前块采用双向预测时,假设参考块内部的运动都是相同的,解码端根据第一参考块的运动信息可以推测出第一参考块从第一参考图像到当前图像的运动矢量0,记为mv_0_t。以及解码端根据第二参考块的运动信息,推测出第二参考块从第二参考图像到当前图像的运动矢量1,记为mv_1_t。其中,运动矢量0和运动矢量1可以利用时域运动信息的导出方法来推导。
示例性的,通过如下公式(11),确定第一参考块的时域运动信息的运动矢量mv_0_t:
mv_0_t_x=-ref_mv_0_x*tb0/td0,mv_0_t_y=-ref_mv_0_y*tb0/td0     (11)
其中,ref_mv_0_x和ref_mv_0_y为第一参考块在x轴和y轴上的运动矢量。td0为第一同位参考图像col_pic0和同位块的第一参考图像col_ref0之间的POC距离,tb0为当前图像curr_pic和第一参考图像curr_ref0之间的POC距离。
示例性的,通过如下公式(12),确定第二参考块的时域运动信息的运动矢量mv_1_t:
mv_1_t_x=-ref_mv_1_x*tb1/td1,mv_1_t_y=-ref_mv_1_y*tb1/td1      (12)
其中,ref_mv_1_x和ref_mv_1_y为第二参考块在x轴和y轴上的运动矢量。td1为第二同位参考图像col_pic1和同位块的第二参考图像col_ref1之间的POC距离,tb1为当前图像curr_pic和第二参考图像curr_ref1之间的POC距离。
接着,解码端将该mv_0_t的相反向量-mv_0_t确定为第三运动信息中的第一预测方向运动矢量,将mv_1_t的相向量-mv_1_t确定为第三运动信息中的第二预测方向运动矢量。
假设参考块和当前块都属于同一个整体运动的物体,且这个物体的运动是一个匀速直线运动,而且参考块的运动信息是准确的,将上述确定的第三运动信息中的运动信息的绝对值与当前块的第一运动信息中的运动矢量的绝对值进行比较,以确定第一运动信息更有可能往哪一个方向偏移。具体的是,由于本申请实施例的当前块采用双向预测,因此可以确定第一运动信息和第三运动信息均包括第一预测方向运动矢量和第二预测方向运动矢量,这样将每一个方向上的运动信息分别进行比较,首先,将第三运动信息中第一预测方向运动矢量-mv_0_t的绝对值与第一运动信息中第一预测方向运动示例mv_0的绝对值进行比较。例如,若第三运动信息中第一预测方向运动矢量-mv_0_t的绝对值小于第一运动信息中第一预测方向运动矢量mv_0的绝对值,则说明第一参考块按其现有的运动矢量从第一参考图像中的当前位置向当前图像进行移动时,无法到达目前当前块在当前图像中的位置,那么可以推测出第一运动信息中第一预测方向运动矢量mv_0偏大。再例如,若第三运动信息中第一预测方向运动矢量-mv_0_t的绝对值大于第一运动信息中第一预测方向运动矢量mv_0的绝对值,则说明第一参考块按其现有的运动矢量从第一参考块的当前位置向当前图像移动时,其在当前图像中的位置超过当前块在当前图像中的位置,那么可以推测出第一运动信息中第一预测方向运动矢量mv_0偏小。接着,将第三运动信息中第二预测方向运动矢量-mv_1_t的绝对值与第一运动信息中第二预测方向运动矢量mv_1的绝对值进行比较。例如,若第三运动信息中第二预测方向运动矢量-mv_1_t的绝对值小于第一运动信息中第二预测方向运动矢量mv_1的绝对值,则说明第二参考块按其现有的运动矢量从第二参考图像中的当前位置向当前图像进行移动时,无法到达目前当前块在当前图像,那么可以推测出第一运动信息中第二预测方向运动矢量mv_0偏大。再例如,若第三运动信息中第二预测方向运动矢量-mv_1_t的绝对值大于第一运动信息中第二预测方向运动矢量mv_1的绝对值,则说明第二参考块按其现有的运动矢量从第二参考图像中的当前位置向当前图像进行移动后,其在当前图像中的位置超过当前块在当前图像中的位置,那么可以推测出第一运动信息中第二预测方向运动矢量mv_1偏小。
方式二,解码端基于据参考块的运动信息,确定当前块的时域运动信息作为第三运动矢量。
在一种示例中,若当前块采用单向预测时,假设参考块内部的运动都是相同的,解码端根据参考块的运动信息,确定当前块的时域运动信息作为第三运动信息。
示例性的,通过如下公式(13),确定第三运动信息:
mv_t’_x=ref_mv_x*tb/td,mv_t’_y=ref_mv_y*tb/td       (13)
其中,ref_mv_x和ref_mv_y为参考块在x轴和y轴上的运动矢量,mv_t’_x和mv_t’_y为第三运动信息中的第一预测方向运动矢量和第二预测方向运动矢量,td为同位参考图像col_pic和同位块的参考图像col_ref之间的POC距离,tb为当前图像curr_pic和参考图像curr_ref之间的POC距离。
假设参考块和当前块都属于同一个整体运动的物体,且这个物体的运动是一个匀速直线运动,而且参考块的运动信息是准确的,将上述确定的第三运动信息中的运动矢量mv_t’的绝对值与当前块的第一运动信息中的运动矢量mv_0的绝对值进行比较,以判断第一运动信息更有可能往哪一个方向偏移。例如,若第三运动信息中的运动矢量mv_t’的绝对值小于当前块的第一运动信息中的运动矢量mv_0的绝对值,则说明当前块按其现有的运动矢量从当前图像所在的当前位置向参考图像移动时,只能达到图18A所示虚线框位置,那么可以推测出第一运动信息中当前块的运动矢量mv_0偏小。再例如,若第三运动信息中的运动矢量mv_t’的绝对值大于当前块的第一运动信息中的运动矢量mv_0,则说明当前块按其现有的运动矢量从当前图像所在的当前位置向参考图像移动时,可以达到图18B所示虚线框位置,那么可以推测出第一运动信息中当前块的运动矢量mv_0偏大。
在一种示例中,若当前块采用双向预测时,第一运动信息和第三运动信息均包括第一预测方向运动矢量和第二预测方向运动矢量。假设参考块内部的运动都是相同的,解码端根据第一参考块的运动信息,可以推测出当前块从当前图像所在的当前位置向第一参考图像移动时的运动矢量mv_0_t’。以及解码端根据第二参考块的运动信息,推测出当前块从当前图像所在的当前位置向第二参考图像移动时的运动矢量mv_1_t’。其中,mv_0_t’为第三运动信息中的第一预测方向运动矢量,mv_1_t’为第三运动信息中的第二预测方向运动矢量。其中,mv_0_t’和mv_1_t’可以利用时域运动信息的导出方法来推导。
示例性的,通过如下公式(14),确定第三运动信息中的第一预测方向运动矢量mv_0_t’:
mv_0_t’_x=ref_mv_0_x*tb0/td0,mv_0_t’_y=ref_mv_0_y*tb0/td0     (14)
其中,ref_mv_0_x和ref_mv_0_y为第一参考块在x轴和y轴上的运动矢量。td0为第一同位参考图像col_pic0和同位块的第一参考图像col_ref0之间的POC距离,tb0为当前图像curr_pic和第一参考图像curr_ref0之间的POC距离。
示例性的,通过如下公式(15),确第三运动信息中的第二预测方向运动矢量mv_1_t’:
mv_1_t’_x=ref_mv_1_x*tb1/td1,mv_1_t_y=ref_mv_1_y*tb1/td1    (15)
其中,ref_mv_1_x和ref_mv_1_y为第二参考块在x轴和y轴上的运动矢量。td1为第二同位参考图像col_pic1和同位块的第二参考图像col_ref1之间的POC距离,tb1为当前图像curr_pic和第二参考图像curr_ref1之间的POC距离。
假设参考块和当前块都属于同一个整体运动的物体,且这个物体的运动是一个匀速直线运动,而且参考块的运动信息是准确的,将上述确定的第三运动信息中的运动矢量的绝对值与当前块的第一运动信息中的运动矢量的绝对值进行比较,以确定第一运动信息更有可能往哪一个方向偏移。具体的是,首先,将第三运动信息中第一预测方向运动矢量mv_0_t’的绝对值与第一运动信息中第一预测方向运动矢量mv_0的绝对值进行比较。例如,若第三运动信息中第一 预测方向运动矢量mv_0_t’的绝对值小于第一运动信息中第一预测方向运动矢量mv_0的绝对值,则说明当前块按照第一预测方向运动矢量mv_0从当前图像中的当前位置向第一参考图像移动时,无法到达目前第一参考块在第一参考图像中的位置,那么可以推测出第一运动信息中第一预测方向运动矢量mv_0偏大。再例如,若第三运动信息中第一预测方向运动矢量mv_0_t’的绝对值大于第一运动信息中第一预测方向运动矢量mv_0的绝对值,则说明当前块按照第一预测方向运动矢量mv_0从当前图像中的当前位置向第一参考图像移动后在第一参考图像中的位置,超过第一参考块在第一参考图像中的当前位置,那么可以推测出第一运动信息中第一预测方向运动矢量mv_0偏小。接着,将第三运动信息中第二预测方向运动矢量mv_1_t’的绝对值与第一运动信息中第二预测方向运动矢量mv_1的绝对值进行比较。例如,若第三运动信息中第二预测方向运动矢量mv_1_t’的绝对值小于第一运动信息中第二预测方向运动矢量mv_1的绝对值,则说明当前块按照第二预测方向运动矢量mv_0从当前图像中的当前位置向第二参考图像移动,无法到达目前第二参考块在第二参考图像中的当前位置,那么可以推测出第一运动信息中第二预测方向运动矢量mv_0偏大。再例如,若第三运动信息中第二预测方向运动矢量mv_1_t’的绝对值大于第一运动信息中第二预测方向运动矢量mv_1的绝对值,则说明当前块按照第二预测方向运动矢量mv_0从当前图像中的当前位置向第二参考图像移动后在第二参考图像中的位置,超过第二参考块在第二参考图像中的当前位置,那么可以推测出第一运动信息中第二预测方向运动矢量mv_1偏小。
解码端基于上述步骤,确定出第三运动信息,接着,执行上述S102-C的步骤。
在一些实施例中,上述使用第三运动信息改善第一运动信息可能会带来更大的误差,例如物体没有做均速运动。因此本申请实施例中,解码端在基于第三运动信息对第一运动信息进行改善,得到第二运动信息之前,首先确定第一运动信息和第三运动信息之间的差异值。
本申请实施例对解码端确定第一运动信息和第三运动信息之间的差异值的具体方式不做限制。
在一种可能的实现方式中,将第一运动信息中的运动矢量与第三运动信息中的运动矢量差的绝对值,确定为第一运动信息和第三运动信息之间的差异值。
示例性的,可以通过如下公式(16),确定差异值:
diff=abs(mv_0_x-mv_0_t’_x)+abs(mv_0_y-mv_0_t’_y)    (16)
其中,mv_0_x和mv_0_y为第一运动信息中的运动矢量,mv_0_t’_x和mv_0_t’_y为第三运动信息中的运动矢量,diff为第一运动信息和第三运动信息之间的差异值。
若该差异值大于预设阈值thr时,则跳过基于第三运动信息对第一运动信息进行改善,得到第二运动信息的步骤,而是直接DMVR等方式,对第一运动信息进行改善,得到第二运动信息。若该差异值小于或等于预设阈值时,则执行上述S102-C的步骤,基于第三运动信息对第一运动信息进行改善,得到第二运动信息。
本申请实施例对解码端基于第三运动信息对第一运动信息进行改善,得到第二运动信息的具体方式不做限制。
在一些实施例中,解码端将第三运动信息与第一运动信息进行比较,以基于第三运动信息对第一运动信息进行调整,得到改善后的第二运动信息。示例性的,若确定第一运动信息小于第三运动信息,则可以适应性的增大第一运动信息,得到第二运动信息,例如在第一运动信息周围搜索第二运动信息时,可以更倾向于搜索较大的运动矢量。示例性的,若确定第一运动信息大于第三运动信息,则可以适应性的减小第一运动信息,得到第二运动信息,例如在第一运动信息周围搜索第二运动信息时,可以更倾向于搜索较小的运动矢量。
在一些实施例中,上述S102-C包括如下S102-C1和S102-C2的步骤:
S102-C1、基于第三运动信息和第一运动信息,确定第四运动信息;
S102-C2、基于第四运动信息,确定第二运动信息。
在该实现方式中,解码端基于第三运动信息和第一运动信息,确定第四运动信息的实现方式包括但不限于如下几种:
方式1,将第三运动信息和第一运动信息的平均值,确定为确定第四运动信息。
在一种示例中,若当前块采用单向预测,第一运动信息、第二运动信息和第四运动信息均包括单向运动信息时,则解码端可以通过如下公式(17),确定第四运动信息:
mv_0_c_x=(mv_0_x+mv_0_t’_x)/2
mv_0_c_y=(mv_0_y+mv_0_t’_y)/2    (17)
其中,公式(17)中的mv_0_c_x和mv_0_c_y为第四运动信息中的运动矢量,mv_0_x和mv_0_y为第一运动信息中的运动矢量,mv_0_t’_x和mv_0_t’_y为第三运行信息中的运动矢量。
在一种示例中,若当前块采用双向预测,第一运动信息、第二运动信息和第四运动信息包括第一预测方向运动矢量和第二预测方向运动矢量时,则解码端可以通过如下公式(18),确定第四运动信息:
mv_0_c_x=(mv_0_x+mv_0_t’_x)/2,mv_0_c_y=(mv_0_y+mv_0_t’_y)/2
mv_1_c_x=(mv_1_x+mv_1_t’_x)/2,mv_1_c_y=(mv_1_y+mv_1_t’_y)/2    (18)
其中,公式(18)中的mv_0_c_x和mv_0_c_y为第四运动信息中的第一预测方向运动矢量,mv_1_c_x和mv_1_c_y为第四运动信息中的第二预测方向运动矢量,mv_0_x和mv_0_y为第一运动信息中的第一预测方向运动矢量,mv_1_x和mv_1_y为第一运动信息中的第二预测方向运动矢量,mv_0_t’_x和mv_0_t’_y为第三运行信息中的第一预测方向运动矢量,mv_1_t’_x和mv_1_t’_y为第三运行信息中的第二预测方向运动矢量。
方式2,确定第三运动信息和第一运动信息对应的权重,基于权重,确定第三运动信息和第一运动信息的加权平均值,进而将该加权平均值,确定为确定第四运动信息。
本申请实施例对解码端确定第三运动信息和第一运动信息对应的权重的具体方式不做限制。
在一种示例中,第三运动信息对应的权重大于第一运动信息对应的权重。
在一种示例中,第三运动信息对应的权重小于第一运动信息对应的权重。这是因为,第一运动信息是基于编码器给出的相关预测信息确定的一种运动信息的假设(或预测),编码器给出的相关预测信息包括编码器从某种候选列表中 选择了一个编码器认为合适的运动信息,或者是编码器选择的某种它认为合适的预测模式。第三运动信息是解码器推理出一种运动信息的假设(或预测),比如解码器根据参考图像上的运动信息推导出来的当前块的运动信息。在该示例中,可以认为第一运动信息是经过了编码器选择得到的,而第三运动信息是基于参考图像的运动信息导出的,但是参考图像毕竟不是当前图像,因此在该示例中,可以为第一运动信息设置较高的权重,为第三运动信息设置较低的权重。
解码端确定出第三运动信息和第一运动信息对应的权重后,将第三运动信息和第一运动信息进行加权处理,得到第四运动信息。
在一种示例中,若当前块采用单向预测,第一运动信息、第二运动信息和第四运动信息均包括单向运动信息时,则解码端可以通过如下公式(19),确定第四运动信息:
mv_0_c_x=(a0*mv_0_x+b0*mv_0_t’_x)/2
mv_0_c_y=(a0*mv_0_y+b0*mv_0_t’_y)/2    (19)
其中,a0*为第一运动信息对应的权重,b0*为第三运动信息对应的权重,a0大于b0。
在一种示例中,b0=1-a0。
示例性的,a0可以为3/4,5/8等,b0可以为1/4,2/8等。
示例性的,为了不出现小数或分数,则若a为3/4,则上述公式(19)可以写成如下公式(20)的形式:
mv_0_c_x=(3*mv_0_x+mv_0_t’_x)/8
mv_0_c_y=(3*mv_0_y+mv_0_t’_y)/8    (20)
在一些实施例中,上述公式中的除法也可以用右移>>代替。
在一种示例中,若当前块采用双向预测,第一运动信息、第二运动信息和第四运动信息包括第一预测方向运动矢量和第二预测方向运动矢量时,则解码端可以通过如下公式(21),确定第四运动信息:
mv_0_c_x=(a0*mv_0_x+b0*mv_0_t’_x)/2,mv_0_c_y=(a0*mv_0_y+b0*mv_0_t’_y)/2
mv_1_c_x=(a1*mv_1_x+b1*mv_1_t’_x)/2,mv_1_c_y=(a0*mv_1_y+b1*mv_1_t’_y)/2    (21)
其中,公式(21)中的a0*为第一运动信息中的第一预测方向运动信息(或第一预测方向运动矢量)对应的权重,b0*为第一运动信息中的第二预测方向运动信息(或第二预测方向运动矢量)对应的权重。a1*为第三运动信息中的第一预测方向运动信息(或第一预测方向运动矢量)对应的权重,b1*为第三运动信息中的第二预测方向运动信息(或第二预测方向运动矢量)对应的权重。
在一些实施例中,还可以根据具体的情况设置不同的权重,比如根据预测模式的不同设置不同的权重。
解码端基于上述步骤,确定出第四运动信息后,执行上述S102-C2基于该第四运动信息,确定第二运动信息。
本申请实施例中,上述S102-C2中解码端基于第四运动信息,确定第二运动信息的具体实现方式包括但不限于如下几种:
方式一,将第二运动信息的搜索中心,定位在第四运动信息对应的位置上,此时,上述S102-C2包括如下S102-C2-a的步骤:
S102-C2-a、以第四运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中搜索得到第二运动信息。
在本申请实施例中第二运动信息是在第一运动信息的周围搜索得到。例如图19A所示,解码端基于第一运动信息在参考图像中确定出当前块的参考块的定位点,该定位点可以为参考块的左上角的位置或参考块的中心位置等。接着,在该第一运动信息指定的定位点附近进行搜索,例如该定位点为中心进行搜索,或者在该定位点的周围上下左右相同范围进行搜索。示例性的,图19A中一个正方形为一个像素位置,黑色正方形为第一运动信息在参考图像中指定的定位点,白色正方形为第二运动信息的搜索位置,将这些正方形搜索点中代价最小的一个位置对应的运动信息,确定为第二运动信息。
由上述可知,本申请实施例中的第一运动信息可能存在不准确的问题,例如基于参考图像的运动信息,确定出第一运动信息偏大或者偏小,此时以不准确的第一运动信息为搜索中心进行第二运动信息搜索时,使得第二运动信息的搜索不准确。为了解决该技术问题,本申请实施例基于参考图像的运动信息,确定出第三运动信息,进而基于第三运动信息和当前块的第一运动信息修改第二运动信息的搜索中心,具体是,基于第三运动信息和第一运动信息,确定出第四运动信息,进而以第四运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中搜索得到第二运动信息,由于该第四运动信息考虑了参考图像的运动信息和当前块的第一运动信息,将第四运动信息指定的位置作为第二运动信息的搜索中心可以提高搜索中心的准确性,进而提高第二运动信息的搜索准确性,提升解码预测效果。
举例说明,假设由上述分析,第一运动信息大于第三运动信息时,说明第一运行信息偏大,此时在第二运动信息搜索时,更倾向于搜索更小一点的运动信息。如图19B所示,黑色正方形为第一运动信息在参考图像中指定为位置,灰色正方形为第三运动信息在参考图像中指定为位置,将第三运动信息和第一运动信息的平均值(即第四运动信息)在参考图像中对应的位置作为第二运动信息的搜索中心,得到的搜索范围如图19B所示,相比于图19A,搜索范围右移,偏向搜索更小一点的运动信息,进而实现对第二运动信息的准确搜索。再例如,若第一运动信息小于第三运动信息时,说明第一运行信息偏小,此时在第二运动信息搜索时,更倾向于搜索更大一点的运动信息,例如图19C所示,在第二运动信息的搜索范围相比于图19A发生左移,偏向搜索更大一点的运动信息,进而实现对第二运动信息的准确搜索。
在本申请实施例中,当前块采用单向预测或双向预测时,确定第二运动信息的搜索范围的具体过程基本一致。
在一些实施例中,若当前块采用单向预测时,以第四运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中搜索得到第二运动信息的过程,可以参照图19B和图19C所示的方案。例如,以第四运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中该搜索中心点附近的预设搜索范围内进行运动信息搜索,确定搜索到的每一个位置点对应的运动信息的代价,进而将代价最小的位置点对应的运动信息,确定为当前 块的第二运动信息。在单向预测中,每一个位置点对应的运动信息的代价,可以通过该位置点的模板的运动信息和当前块的模板的运动信息的匹配代价表示。
在一些实施例中,若当前块采用双向预测时,上述当前块包括第一参考图像和第二参考图像,第二运动信息和第四运动信息均包括第一预测方向运动信息和第二预测方向运动信息。这样解码端以第四运动信息中的第一预测方向运动信息在第一参考图像中对应的位置,为第二运动信息中的第一预测方向运动信息的搜索中心点,以第四运动信息中的第二预测方向运动信息在第二参考图像中对应的位置,为第二运动信息中的第二预测方向运动信息的搜索中心点,在第一参考图像和第二参考图像的预设搜索范围内进行运动信息搜索,确定搜索到的每一对双边运动信息的双边匹配代价,其中每一对运动信息包括一个第一预测方向运动信息和一个第二预测方向运动信息;进而基于双边匹配代价,从搜索到的多对双边运动信息中,确定第二运动信息。
在该双向预测中,假设两边的预设搜索范围相同,均包括n种可能的搜索位置点,也就是说,每一边均可以搜索到n个可能的MV。
在一种可能的实现方式中,解码端在搜索时,可以将第一参考图像侧对应的n个可能的MV与第二参考图像侧对应的n个可能的MV进行两两组合,得到n2对双边运动信息。
在一种可能的实现方式中,在第一参考图像和第二参考图像中搜索双边运动信息时,在移动时,两个参考图像的MV是按镜像移动的,也就是在各自搜索中心点对应的MV的基础上一边移动MVdiff,另一边移动-MVdiff,此时可以得到n对双边运动信息。
在搜索得到上述多对双边运动信息中每一对双边运动信息时,确定该双边运动信息之间的双边匹配代价。
本申请实施例对确定双边运动信息之间的双边匹配代价的具体方式不做限制。
以多对双边运动信息中的第i对双边运动信息为例,在一种可能的实现方式中,该第i对双边运动信息中包括第一预测方向运动信息(例如第一预测方向运动矢量MV0)和第二预测方向运动信息(例如第二预测方向运动矢量MV1),由于MV0和MV1均为向量,因此可以基于向量距离的方式,确定MV0和MV1之间的距离,进而将该距离,确定为第i对双边运动信息对应的双边匹配代价。
在一种可能的实现方式中,解码端可以基于第i对双边运动信息中的第一预测方向运动信息,在第一参考图像中确定第一预测块,基于第i对双边运动信息中的第二预测方向运动信息,在第二参考图像中确定第二预测块;确定第一预测块和第二预测块的匹配代价;基于第一预测块和第二预测块的匹配代价,确定第i对双边运动信息的双边匹配代价。例如,将第一预测块和第二预测块的SAD代价,确定为第一预测块和第二预测块的匹配代价,进而将该匹配代价确定为第i对双边运动信息的双边匹配代价,或者对该匹配代价乘上或除以预设系数,得到第i对双边运动信息的双边匹配代价。
基于上述步骤,解码端可以确定出搜索到的多对双边运动信息中每一对双边运动信息的双边匹配代价,进而搜索到的多对双边运动信息中双边匹配代价最小的一对双边运动信息,确定为第二运动信息,得到的第二运动信息为双向运动信息。
上述对方式一中,以第四运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中搜索得到第二运动信息的具体过程进行介绍,下面对方式二进行介绍。
方式二,通过第四运动信息,修改第二运动信息搜索过程中的代价,此时上述S102-C2包括如下S102-C2-b1至S102-C2-b4的步骤:
S102-C2-b1、以第一运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价;
S102-C2-b2、基于候选运动信息和第四运动信息,确定候选运动信息对应的代价系数;
S102-C2-b3、基于候选运动信息对应的代价系数对第一代价进行修正,得到候选运动信息的第二代价;
S102-C2-b4、基于搜索到的多个候选运动信息的第二代价,确定第二运动信息。
在该方式二中,可以基于目前的第一运动信息,搜索确定出各候选运动信息的第一代价,接着使用第四运动信息确定的代价系数,对候选运动信息的第一代价进行修正,得到第二代价,进而基于各候选运动信息的第二代价,从各候选运动信息中,选出第二运动信息。
具体的,解码端首先以第一运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像中进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价。例如,参照上述图19A所示,黑色正方形为第一运动信息在参考图像中指定的定位点,将该定位点作为第二运动信息的搜索中心,白色正方形对应的位置的运动信息记为第二运动信息的候选运动信息,确定这些候选运动信息中每一候选运动信息的第一代价。
接着,对于上述搜索到的每一候选运动信息,基于该候选运动信息和第四运动信息,确定该候选运动信息对应的代价系数。例如,将该候选运动信息和第四运动信息的差的绝对值,确定为该候选运动信息对应的代价系数。或者,将该候选运动信息和第四运动信息的差的绝对值与预设值的和值,确定为该候选运动信息对应的代价系数。
这样可以通过上述确定代价系数对该候选运动信息的第一代价进行修正,得到第二代价。例如,将代价系数和该候选运动信息的第一代价的乘积,确定为该候选运动信息的第二代价。
在本申请实施例中,当前块采用单向预测或双向预测时,确定第二运动信息的搜索范围的具体过程基本一致。
在一些实施例中,若当前块采用单向预测时,则第一运动信息和候选运动信息均为单向运动信息,如图19A所示,解码端首先以第一运动信息在参考图像中对应的位置为第二运动信息的搜索中心点,在参考图像的预设搜索范围内中进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价。在单向预测中,每一候选运动信息的第一代价,可以通过该候选运动信息所在位置的模板的运动信息和当前块的模板的运动信息的匹配代价表示。接着,基于每一候选运动信息和第四运动信息,确定每一候选运动信息对应的代价系数,例如将该候选运动信息和第四运动信息的差的绝对值,确定为该候选运动信息对应的代价系数。或者,将该候选运动信息和第四运动信息的差的绝对值与预设值的和 值,确定为该候选运动信息对应的代价系数。然后基于候选运动信息对应的代价系数对第一代价进行修正,得到候选运动信息的第二代价,例如将代价系数和该候选运动信息的第一代价的乘积,确定为该候选运动信息的第二代价。这样可以确定出搜索到的多个候选运动信息的第二代价,进而将多个候选运动信息中第二代价最小的候选运动信息,确定为第二运动信息,该第二运动信息为单向运动信息。
在一些实施例中,若当前块采用双向预测时,上述当前块包括第一参考图像和第二参考图像,第一运动信息、第二运动信息、第四运动信息和所述候选运动信息均包括第一预测方向运动信息和第二预测方向运动信息。此时,上述S102-C2-b1包括如下步骤:
S102-C2-b11、以第一运动信息中的第一预测方向运动信息在第一参考图像中对应的位置,为第二运动信息中的第一预测方向运动信息的搜索中心点,以第一运动信息中的第二预测方向运动信息在第二参考图像中对应的位置,为第二运动信息中的第二预测方向运动信息的搜索中心点,在第一参考图像和第二参考图像的预设搜索范围内进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价。
在该双向预测中,假设两边的预设搜索范围相同,均包括n种可能的搜索位置点,也就是说,每一边均可以搜索到n个可能的MV。
在一种可能的实现方式中,解码端在搜索时,以第一运动信息中的第一预测方向运动信息在第一参考图像中对应的位置,为第二运动信息中的第一预测方向运动信息的搜索中心点,可以在第一参考图像侧搜索到n个可能的MV,以第一运动信息中的第二预测方向运动信息在第二参考图像中对应的位置,为第二运动信息中的第二预测方向运动信息的搜索中心点,可以在第二参考图像侧搜索到n个可能的MV。两侧的n个可能的MV进行两两组合,得到n2对双边运动信息,进而得到n2个候选运动信息,每一个候选运动信息均为双向运动信息。
在一种可能的实现方式中,以第一运动信息中的第一预测方向运动信息在第一参考图像中对应的位置,为第二运动信息中的第一预测方向运动信息的搜索中心点,以第一运动信息中的第二预测方向运动信息在第二参考图像中对应的位置,为第二运动信息中的第二预测方向运动信息的搜索中心点,在第一参考图像和第二参考图像的预设搜索范围内搜索候选运动信息时,两个参考图像的MV是按镜像移动的。也就是在各自搜索中心点对应的MV的基础上一边移动MVdiff,另一边移动-MVdiff,此时可以得到n对双边运动信息,进而得到n个候选运动信息,每一个候选运动信息均为双向运动信息。
在搜索得到上述多个候选运动信息中每一个候选运动信息时,确定该候选运动信息的第一代价,此时第一代价可以为双边匹配代价。本申请实施例中,确定每一个候选运动信息的第一代价的过程相同,以一个候选运动信息为例进行说明。
在一种可能的实现方式中,该候选运动信息中包括第一预测方向运动信息(例如第一预测方向运动矢量MV0)和第二预测方向运动信息(例如第二预测方向运动矢量MV1),由于MV0和MV1均为向量,因此可以基于向量距离的方式,确定MV0和MV1之间的距离,进而将该距离,确定为该候选运动信息的第一代价。
在一种可能的实现方式中,解码端可以基于候选运动信息中的第一预测方向运动信息,在第一参考图像中确定第一预测块,基于候选运动信息中的第二预测方向运动信息,在第二参考图像中确定第二预测块;确定第一预测块和第二预测块的匹配代价;基于第一预测块和第二预测块的匹配代价,确定该候选运动信息的第一代价。例如,将第一预测块和第二预测块的SAD代价,确定为第一预测块和第二预测块的匹配代价,进而将该匹配代价确定为该候选运动信息的第一代价,或者对该匹配代价乘上或除以预设系数,得到该候选运动信息的第一代价。
接着,基于该候选运动信息和第四运动信息,确定该候选运动信息对应的代价系数。
该方式二不改变第二运动信息的搜索范围,而是为更倾向于选择靠近第四运动信息的候选运动信息,进而可以为每个候选运动信息的第一代价乘以一个系数。例如,可以为靠近第四运动信息的候选运动信息设置较小的代价系数,为距离第四运动信息较远的候选运动信息设置较大的代价系数。
在本申请实施例中,若候选运动信息为双向运动信息时,其对应的代价系数可以为一个或两个。
在一些实施例中,若该候选运动信息对应一个代价系数时,则可以确定该候选运动信息中的第一预测方向运动信息和第四运动信息中的第一预测方向运动信息差的绝对值,记为差值1,确定该候选运动信息中的第二预测方向运动信息和第四运动信息中的第二预测方向运动信息差的绝对值,记为差值2。基于差值1和差值2,确定差值3,例如将差值1和差值2之和,或平均值,确定差值3。进而基于该差值3,确定候选运动信息对应一个代价系数,例如,将该差值3,确定为候选运动信息对应的代价系数,或者将该差值3和预设值的和值,确定为候选运动信息对应的代价系数。这样可以确定出的一个代价系数和上述第一代价,确定出该候选运动信息的第二代价,例如将该候选候选运动信息的第一代价与该候选候选运动信息对应的一个代价系数的乘积,确定为该候选运动信息的第二代价。
在一些实施例中,若该候选运动信息对应两个代价系数,即第一代价系数和第二代价系数时,则上述S102-C2-b2包括如下步骤:
S102-C2-b21、基于候选运动信息中的第一预测方向运动信息和第四运动信息中的第一预测方向运动信息,确定候选运动信息中的第一预测方向运动信息对应的第一代价系数;
S102-C2-b21、基于候选运动信息中的第二预测方向运动信息和第四运动信息中的第二预测方向运动信息,确定候选运动信息中的第二预测方向运动信息对应的第二代价系数。
在该实施例中,若候选运动信息对应的代价系数为两个代价系数时,则解码端基于候选运动信息中的第一预测方向运动信息和第四运动信息中的第一预测方向运动信息,确定候选运动信息中的第一预测方向运动信息对应的第一代价系数,以及基于候选运动信息中的第二预测方向运动信息和第四运动信息中的第二预测方向运动信息,确定候选运动信息中的第二预测方向运动信息对应的第二代价系数。
本申请实施例中,解码端确定第一代价系数和确定第二代价系数的过程基本一致。
在一种可能的实现方式中,将候选运动信息中的第一预测方向运动信息和第四运动信息中的第一预测方向运动信息差值的绝对值,确定为候选运动信息中的第一预测方向运动信息对应的第一代价系数。将候选运动信息中的第二预 测方向运动信息和第四运动信息中的第二预测方向运动信息差值的绝对值,确定为候选运动信息中的第二预测方向运动信息对应的第二代价系数。
在一种可能的实现方式中,确定候选运动信息中的第i预测方向运动信息,与第四运动信息中的第i预测方向运动信息之间的差的绝对值,i为一或二;基于差的绝对值,确定第i代价系数,其中第i代价系数与差的绝对值负相关。也就是说,确定候选运动信息中的第一预测方向运动信息,与第四运动信息中的第一预测方向运动信息之间的差的绝对值1,基于该差的绝对值1,确定第一代价系数,该第一代价系数与差的绝对值1负相关,即距离1越大,第一代价系数越小。确定候选运动信息中的第二预测方向运动信息,与第四运动信息中的第二预测方向运动信息之间的差的绝对值2,基于该差的绝对值2,确定第二代价系数,该第二代价系数与差的绝对值2负相关,即差的绝对值2越大,第二代价系数越小。
本申请实施例对上述基于差的绝对值,确定第i代价系数的具体方式不做限制。
在一种示例中,将该差的绝对值,确定为第i代价系数。即将候选运动信息中的第一预测方向运动信息,与第四运动信息中的第一预测方向运动信息之间的差的绝对值1,确定第一代价系数,将候选运动信息中的第二预测方向运动信息,与第四运动信息中的第二预测方向运动信息之间的差的绝对值2,确定第二代价系数。
在另一种示例中,确定差的绝对值与第一预设值中的最小值;基于最小值,确定第i代价系数。即,将候选运动信息中的第一预测方向运动信息与第四运动信息中的第一预测方向运动信息之间的差的绝对值1,与第一预设值进行比较,确定出差的绝对值1与第一预设值中的最小值1,进而基于该最小值2,确定第一代价系数。以及将候选运动信息中的第二预测方向运动信息与第四运动信息中的第二预测方向运动信息之间的差的绝对值2,与第一预设值进行比较,确定出差的绝对值2与第一预设值中的最小值2,进而基于该最小值2,确定第二代价系数。
本申请实施例对上述第一预设值的具体取值不做限制。
示例性的,第一预设值为大于0的数值。
可选的,第一预设值为4。
本申请实施例对基于最小值,确定第i代价系数的具体方式不做限制。
例如,将该最小值,确定为第i代价系数。例如将上述最小值1,确定为第一代价系数,将最小值2,确定为第二代价系数。
再例如,将最小值和第二预设值的和,确定为第i代价系数。例如,将最小值1和第二预设值的和,确定为第一代价系数,将最小值2和第二预设值的和,确定为第二代价系数。
示例性的,解码端基于如下公式(21)明确的第一代价系数和第二代价系数:
coef_0=min(abs(mv_0_x-mv_0_c_x)+abs(mv_0_y-mv_0_c_y),a)+b
coef_1=min(abs(mv_1_x-mv_1_c_x)+abs(mv_1_y-mv_1_c_y),a)+b    (211)
其中,coef_0为候选运动信息的第一代价系数,coef_0为候选运动信息的第二代价系数,mv_0_x和mv_0_y为候选运动信息中第一预测方向运动信息(即第一预测方向运动矢量),mv_1_x和mv_1_y为候选运动信息中第二预测方向运动信息(即第二预测方向运动矢量)。mv_0_c_x和mv_0_c_y为第四运动信息中第一预测方向运动信息(即第一预测方向运动矢量),mv_0_c_x和mv_0_c_y为第四运动信息中第二预测方向运动信息(即第二预测方向运动矢量)。a为第一预设值,b为第二预设值,min()是取最小值的运算。可选的,上述公式(21)中的除法也可以使用右移替代。
本申请实施例对上述第一预设值a和第二预设值b的具体取值不做限制。
可选的,第一预设值a为4。
可选的,第二预设值b为32。
基于上述步骤,解码端确定出候选运动信息的第一代价系数和第二代价系数后,基于该第一代价系数和第二代价系数,对该候选运动信息的第一代价进行修正,得到该候选运动信息的第二代价。
本申请实施例对解码端基于该第一代价系数和第二代价系数,对该候选运动信息的第一代价进行修正,得到该候选运动信息的第二代价的具体方式不做限制。
在一种可能的实现方式中,将第一代价系数和第二代价系数相加后,与该候选运动信息的第一代价相乘,得到该候选运动信息的第二代价。
在一种可能的实现方式中,将该候选运动信息的第一代价与第一代价系数和第二代价系数相乘,得到该候选运动信息第二代价。示例性的,解码端通过如下公式(23),得到该候选运动信息第二代价:
SAD_c=SAD*coef_0*coef_1    (23)
其中,SAD_c为该候选运动信息第二代价,SAD为该候选运动信息的第一代价,coef_0为该候选运动信息的第一代价系数,coef_1为该候选运动信息的第二代价系数。
解码端基于上述步骤,可以得到搜索到的多个候选运动信息中每一个候选运动信息的第二代价,该第二代价是基于参考图像的运动信息修正后的代价,其正确性比第一代价高,进而基于多个候选运动信息中每个候选运动信息的第二代价,从这多个候选运动信息中确定出当前块的第二运动信息,例如将搜索到的多个候选运动信息中第二代价最小的候选运动信息,确定为第二运动信息,实现对第二运动信息的准确确定,进而提高当前块的预测准确性,提升解码端的解码效果。
上述实施例均是将当前块作为一个整块,对整个当前块的第一运动信息进行整块改善的过程进行介绍。需要说明的是,上述实施例的方法也用于基于子块的处理,例如将当前块划分为多个子块,对于每一个子块的第一运动信息单独使用与上述当前块的第一运动信息相同的方式进行改善,得到每一个子块的第二运动信息,进而基于每一个子块的第二运动信息,得到每一个子块的预测值,每一个子块的预测值组成当前块的预测值。
上述实施例对情况1中,若参考图像的运动信息用于直接参与第一运动信息的改善时,则解码端对第一运动信息的改善,得到第二运动信息的具体过程进行介绍。
情况2,本申请实施例的参考图像的运动信息可以用于指导第一运动信息改善过程中,相关块的划分方式。
在一些实施例中,上述S102包括如下S102-D至S102-F的步骤:
S102-D、基于参考图像的运动信息,将当前块划分为至少一个子块;
S102-E、对于至少一个子块中的第i个子块,对第i个子块的第一运动信息进行改善,得到第i个子块的第二运动信息,i为正整数;
S102-F、基于N个子块的第二运动信息,得到当前块的第二运动信息。
示例性的,如图20所示,基于当前块的第一运动信息在参考图像中,确定当前块的参考块,获取参考块的运动信息,参考块中右上角的灰色区域的运动矢量和其他区域差别明显,可以设置一个阈值,两个运动矢量的差超过阈值认为差别明显。由于当前块与参考块的相关性较强,此时当前块的左上角区域的运动矢量与其他区域差别也较明显,若当前块采用整块改善时,改善效果不明显。因此,在本申请实施例中,在对当前块的第一运动信息进行改善时,基于参考图像块的运动信息,将当前块划分为至少一个子块,对于每一个子块,单独改善运动信息,这样不仅降低硬件的实现代价,同时划分成子块进行改善时提供了更好的灵活性,可以为每个子块可以独立改善MV,从一定程度上达到了提高运动信息改善精度的效果。
例如,若当前块中有一个部分的运动信息和其他部分不一样时,将当前块划分为多个子块进行改善时,可以提高运动信息的改善效果。由于当前图像和参考图像中物体的分布相差比较小,因此在本申请实施例中,解码端基于参考图像的运动信息来指示当前块的划分。
本申请实施例对解码端基于参考图像的运动信息,将当前块划分为至少一个子块的具体方式不做限制。
在一些实施例中,解码端首先在参考图像中确定当前块对应的参考块,进而在参考块中随机采样几个点的运动信息,进而将这几个点的运动信息进行比较,将这几个点中运动信息的点划分在一个子块,进而将参考块划分为至少一个子块。接着,解码端将当前块中确定出参考块中各子块对应的子块,进而将当前块划分为至少一个子块。
在一些实施例中,则上述S102-D包括如下S102-D1至S102-D4的步骤:
S102-D1、确定当前块在参考图像中对应的参考块;
S102-D2、获取当前块的M个子块在参考块中对应的M个子块的运动信息,M为大于1的正整数;
S102-D3、对获取的M个子块的运动信息进行分类,得到P种分类结果,P为小于或等于M的正整数;
S102-D4、基于P种分类结果,将当前块划分为至少一个子块。
在该实现方式中,解码端首先基于当前块的第一运动信息,在参考图像中确定当前块的参考块,接着,将当前块预先划分为M个子块,这M个子块的大小可以相同,也可以不同。例如M个子块均为16X16、或8X8、或4X4等,或者当前块的宽度和高度都是2N时,则可以将当前块划分为4个NXN的子块,或者当前块的尺寸是NX2N或2N XN时,可以将当前块划分为2个NXN的子块。接着,确定当前块中这M个子块在参考块对应的M个子块,由于参考块的运动信息已知,因此解码端可以获取当前块的M个子块在参考块中所对应的M个子块的运动信息也已知。当前块与参考块的相关性较强,因此,解码端对获取的参考块中的M个子块的运动信息进行聚类,得到P个分类结果,进而基于这P个分类结果,将当前块划分为至少一个子块。
例如,若上述P=1时,则说明当前块的各区域的运动信息的差异不大,可以使用整块进行第一运动信息的改善,进而不对当前块进行划分,或者说将当前块划分为一个子块,即当前块本身。
再例如,若上述P大于1时,则基于P种分类结果分别对应的子块,将当前块划分为P个子块,其中一个子块对应的一种分类结果。
解码端基于上述步骤,将当前块划分为至少一个子块后,对至少一个子块中的每一个子块进行单独改善,且每一个子块的改善过程基本相同,为了便于描述,在此以第i个子块为例进行说明。
本申请实施例对第i个子块的第一运动信息进行改善,得到第i个子块的第二运动信息的具体方式不做限制。
在一种可能的实现方式中,解码端采用上述DMVR和/或BDOF等的方式,对第i个子块的第一运动信息进行改善,得到第i个子块的第二运动信息。
在一种可能的实现方式中,解码端通过参考图像的运动信息对第i个子块的第一运动信息进行改善,具体的,解码端基于第i个子块的第一运动信息,在参考图像中确定第i个子块对应的参考块;基于第i个子块的参考块的运动信息,确定第i个子块的时域运动信息作为该第i子块对应的第三运动信息;基于第三运动信息对第i个子块的第一运动信息进行改善,得到第i个子块的第二运动信息。该实现方式的具体过程可以参照上述情况1中对当前块的第一运动信息进行改善的具体描述,只需将上述当前块替换为第i个子块即可,在此不再赘述。
解码端可以基于上述步骤,确定出当前块中每一个子块的第二运动信息,进而基于这至少一个子块的第二运动信息,确定当前块的第二运动信息。
例如,将这至少一个子块的第二运动信息,确定为当前块的第二运动信息,此时当前块的第二运动信息包括至少一个子块中每一个子块的第二运动信息。这样在后续基于当前块的第二运动信息,确定当前块的预测值时,可以基于这至少一个子块的第二运动信息,确定出这至少一个子块中每一个子块的预测值,进而这至少一个子块的预测值,组成当前块的预测值。
再例如,将这至少一个子块的第二运动信息的平均值,确定为当前块的第二运动信息,此时当前块的第二运动信息包括一个整块运动信息。
上述实施例对当前块中每一个子块单独进行运动信息改善的过程进行介绍。
在一些实施例中,解码端在对当前块的第一运动信息进行改善时,可以经过多伦改善,此时上述S102包括如下S102-G的步骤:
S102-G、基于当前块的参考图像的运动信息,对第一运动信息进行N轮改善,得到第二运动信息,N为大于1的正整数。
本申请实施例对这N轮改善所使用的具体改善方式不做限制。
在一种可能的实现方式中,这N轮改善所使用的具体改善方式相同。
在一种可能的实现方式中,这N轮改善所使用的具体改善方式均不相同。
在一种可能的实现方式中,这N轮改善所使用的具体改善方式中部分相同,部分不相同。
在一些实施例中,解码端在对当前块的第一运动信息进行N轮改善时,在下一轮改善时,对上一轮的块进行划分。例如,第一轮,解码端基于整块的双向匹配的运动矢量改善方法,对当前块的第一运动信息进行改善。第二轮,将当前块划分为至少一个子块,基于子块的双向匹配的运动矢量改善,对当前块的各子块的第一轮改善后的运动信息进行改善,可选的第二轮中子块大小可以是16x16。第三轮,对第二轮的子块划分为至少一个子块,基于子块的双向光流的运动矢量改善,对第二轮改善后的运动信息进行改善,可选的,这一轮的子块大小可以是8x8。当然在此基础上还可以进一步丰富步骤,如再来一个第四轮改善,例如基于4x4子块的双向光流的运动矢量改善,对第三轮改善后的运动信息进行改善。可选的,还可以采用基于点的双向光流的运动矢量改善等改善方法,进行进一步改善。多轮改善自上而下地分多层来优化运动矢量。在上述多伦改善中,至少一轮中子块的划分是基于参考图像的运动信息执行,具体划分过程可以参照上述S102-D的相关描述,进而提高子块划分的合理性和准确性。
在一些实施例中,上述S102-G包括如下S102-G1至S102-G4的步骤:
S102-G2、对于第j轮对应的至少一个子块中的每一个子块的运动信息进行改善,得到第j轮对应的至少一个子块的改善后的运动信息,若j为1时,第j轮对应的子块为当前块;
S102-G3、基于参考图像的运动信息,对第j轮对应的至少一个子块中的每一个子块进行块划分,得到第j+1轮对应的至少一个子块;
S102-G4、对于第j+1轮对应的至少一个子块中的每一个子块的运动信息进行改善,重复执行N轮,得到第二运动信息。
在该实现方式中,解码端首先对当前块的第一运动信息进行第一轮改善,得到当前块对应的第一轮改善后的运动信息1。接着,基于参考图像的运动信息,对当前块进行块划分,得到第二轮对应的至少一个子块2。对于第二轮对应的至少一个子块2中每一个子块2的运动信息进行改善,此时子块2的运动信息为第一轮改善后的运动信息,进而得到第二轮对应的至少一个子块2的改善后的运动信息。接着,基于参考图像的运动信息,对子块2进行块划分,得到第三轮对应的至少一个子块3。对于第三轮对应的至少一个子块3中每一个子块3的运动信息进行改善,此时子块3的运动信息为第二轮改善后的运动信息,进而得到第三轮对应的至少一个子块3的改善后的运动信息。接着,基于参考图像的运动信息,对子块3进行块划分,得到第四轮对应的至少一个子块4。对于第四轮对应的至少一个子块4中每一个子块4的运动信息进行改善,此时子块4的运动信息为第三轮改善后的运动信息,进而得到第四轮对应的至少一个子块4的改善后的运动信息。重复执行上述步骤,执行N轮改善后,得到当前块的第二运动信息。
本申请实施例中,基于参考图像的运动信息,对当前块或子块进行块划分的方式,与上述S102-D基本相同。例如,解码端首先确定第二块在参考图像中对应的参考块,该第二块为当前块或者为第j轮对应的子块。接着,获取该第二块的M个子块在参考块中对应的M个子块的运动信息,M为大于1的正整数。然后,对获取的M个子块的运动信息进行分类,得到P种分类结果,P为小于或等于M的正整数。最后,基于P种分类结果,将第二块划分为至少一个子块,例如基于P种分类结果分别对应的子块,将第二块划分为P个子块。具体参照上述S102-D的相关描述,在此不再赘述。
本申请实施例中,解码端对子块的运动信息进行改善,得到子块改善后的运动信息的具体方法不做限制。
在一种可能的实现方式中,解码端采用上述DMVR和/或BDOF等的方式,对上述子块的运动信息进行改善。
在一种可能的实现方式中,解码端通过参考图像的运动信息对子块的运动信息进行改善,具体的,解码端基于该子块的运动信息,在参考图像中确定该子块对应的参考块;基于该子块的参考块的运动信息,确定该子块的时域运动信息作为该子块对应的第三运动信息;基于第三运动信息对该子块的运动信息进行改善,得到该子块改善后的运动信息。该实现方式的具体过程可以参照上述情况1中对当前块的第一运动信息进行改善的具体描述,只需将上述当前块替换为该子块即可,在此不再赘述。
上述实施例对当前块的第一运动信息进行多伦改善的过程进行介绍。
在一些实施例中,解码端在基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到当前块的第二运动信息之前,首先确定当前块的是否满足预设的基于整块的运动矢量改善条件。若确定当前块满足该基于整块的运动矢量改善条件时,则基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到当前块的第二运动信息。
在一些实施例中,若当前块不满足该基于整块的运动矢量改善条件,则本申请实施例的方法包括如下步骤:
步骤1、基于参考图像的运动信息,对当前块进行块划分,得到多个第一子块;
步骤2、对于多个第一子块的任一第一子块,确定第一子块是否满足基于整块的运动矢量改善条件;
步骤3、若第一子块不满足基于整块的运动矢量改善条件,则基于参考图像的运动信息,对第一子块进行块划分,得到多个第二子块;
步骤4、对于多个第二子块中的任一第二子块,确定第二子块是否满足基于整块的运动矢量改善条件,重复执行,直到划分后的子块满足基于整块的运动矢量改善条件,或者划分后的子块的大小满足预设大小为止。
在该实现方式中,解码端在确定当前块不满足基于整块的运动矢量改善条件时,则对当前块进行逐层块划分,直到划分后的子块满足基于整块的运动矢量改善条件,或者划分后的子块不满足基于整块的运动矢量改善条件,但划分后的子块的大小满足预设大小时,停止子块的划分。
本申请实施例对上述基于整块的运动矢量改善条件的具体内容不做限制。
在一种可能的实现方式中,基于整块的运动矢量改善条件可以为块的大小满足预设阈值大小,该预设阈值可以为8X8或4X4等。
举例说明,假设上述预设阈值为4X4,若当前块的大小为32X32时,当前块不满足基于整块的运动矢量改善条件,则基于参考图像的运动信息,对当前块进行块划分,得到多个第一子块。假设第一子块的大小为8X8时,第一子块的大小不满足基于整块的运动矢量改善条件。接着,基于参考图像的运动信息,对第一子块进行块划分,得到多个第二子块。假设第二子块的大小为4X4时,则第二子块满足基于整块的运动矢量改善条件,进而基于参考图像的运动信息,对每一个第二子块的第一运动信息进行改善,得到每一个第二子块的第二运动信息。
在一种可能的实现方式中,解码端通过如下步骤a至步骤d,确定当前块或第一子块或第二子块是否满足基于整块的运动矢量改善条件:
步骤a、确定第一块在参考图像中对应的参考块,第一块为当前块或者为第一子块或者为第二子块;
步骤b、获取第一块的M个子块在参考块中对应的M个子块的运动信息,M为大于1的正整数;
步骤c、对获取的M个子块的运动信息进行分类,得到P种分类结果,P为小于或等于M的正整数;
步骤d、基于P种分类结果,确定第一块是否满足基于整块的运动矢量改善条件。
在该实现方式中,解码端判断当前块或第一子块或第二子块是否满足基于整块的运动矢量改善条件的具体方式基本相同,为了便于描述,用第一块代替当前块或第一子块或第二子块。具体的,首先解码端确定该第一块在参考图像中对应的参考块,接着,将第一块划分为M个子块,并获取第一块的M个子块在参考块中对应的M个子块的运动信息。然后,对获取的M个子块的运动信息进行分类,得到P种分类结果。上述步骤a至步骤c的具体实现过程可以参照上述S102-D1至S102-D3的具体描述,只需要将S102-D1至S102-D3中的当前块替换为第一块即可,可以得到第一块对应的P种分类结果。
最后,基于第一块对应的P种分类结果,确定第一块是否满足基于整块的运动矢量改善条件。
例如,若P等于1,则说明该第一块中的M个子块的运动信息相同,无需对第一块进行划分,进而确定该第一块满足基于整块的运动矢量改善条件。
再例如,若P大于1,则说明该第一块中的M个子块的运动信息不完全相同,需要对第一块进行划分,进而确定该第一块不满足基于整块的运动矢量改善条件。
举例说明,通过上述步骤a至步骤d,确定当前块不满足基于整块的运动矢量改善条件时,基于参考图像的运动信息,将当前块划分为多个第一子块,通过上述步骤a至步骤d,判断每一个第一子块是否满足基于整块的运动矢量改善条件。若部分第一子块满足基于整块的运动矢量改善条件时,则基于参考图像的运动信息对该部分第一子块的运动信息进行改善。若部分第一子块不满足基于整块的运动矢量改善条件时,则基于参考图像的运动信息,对这部分第一子块中的每一个第一子块进行划分,得到多个第二子块。进而通过上述步骤a至步骤d,判断这多个第二子块中的每一个第二子块是否满足基于整块的运动矢量改善条件,对于满足条件的第二子块,基于参考图像的运动信息对该第二子块的运动信息进行改善。对于不满足条件的第二子块继续进行划分,重复执行上述步骤,直到所有的子块均满足基于整块的运动矢量改善条件为止。或者,直到划分后的子块的大小满足预设大小,例如8x8或4x4时,则不再进行块划分,而是默认使用基于子块的双向光流的运动矢量改善方式,对该子块的运动信息进行改善。
在该实现方式中,解码端基于参考图像的运动信息,对当前块或第一子块或第二子块进行划分的过程与上述S102-D基本相同。例如,解码端首先确定第二块在参考图像中对应的参考块,该第二块为当前块或者第一子块或第二子块。接着,获取该第二块的M个子块在参考块中对应的M个子块的运动信息,M为大于1的正整数。然后,对获取的M个子块的运动信息进行分类,得到P种分类结果,P为小于或等于M的正整数。最后,基于P种分类结果,将第二块划分为至少一个子块,例如基于P种分类结果分别对应的子块,将第二块划分为P个子块。具体参照上述S102-D的相关描述,在此不再赘述。
上述实施例结合情况1和情况2,对解码端利用参考图像的运动信息改善当前块的第一运动信息的过程,以及利用参考图像的运动信息指导块划分的过程进行介绍。
解码端基于上述步骤,得到当前块的第二运动信息后,执行如下S103的步骤。
S103、基于第二运动信息,确定当前块的预测值。
解码端基于上述步骤,在对当前块的第一运动信息进行改善时,考虑了参考图像的运动信息,实现对当前块的第一运动信息的有效改善,得到准确的第二运动信息,进而基于该准确的第二运动信息,确定当前块的预测值时,可以提高当前块的预测准确性,进而提高视频的解码性能。
在一些实施例中,若当前块采用单向预测时,则第二运动信息包括一个方向的运动矢量,进而基于该运动矢量,在当前块的参考图像中确定出当前块的预测块,进而得到当前块的预测值。
在一些实施例中,若当前块采用双向预测时,则当前块包括第一参考图像和第二参考图像,第二运动信息包括第一预测方向运动矢量和第二预测方向运动矢量。这样解码端基于第二运动信息中的第一预测方向运动矢量,在第一参考图像中确定出预测块1,基于第二运动信息中的第二预测方向运动矢量,在第二参考图像中确定出预测块2,进而基于预测块1和预测块2,得到当前块的预测值。例如,将预测块1和预测块2的平均值或加权平均值,确定为当前块的预测值。
本申请实施例提供的视频解码方法,解码端在对当前块进行解码时,首先确定当前块的第一运动信,接着,基于当前块的参考图像的运动信息,对第一运动信息进行改善,得到第二运动信息。也就是说,在本申请实施例中,在对第一运动信息进行改善时,考虑了参考图像的运动信息,实现对第一运动信息的有效改善,得到准确的第二运动信息,进而基于该准确的第二运动信息,确定当前块的预测值时,可以提高当前块的预测准确性,进而提高视频的解码性能。
应理解,图14至图20仅为本申请的示例,不应理解为对本申请的限制。
以上结合附图详细描述了本申请的优选实施方式,但是,本申请并不限于上述实施方式中的具体细节,在本申请的技术构思范围内,可以对本申请的技术方案进行多种简单变型,这些简单变型均属于本申请的保护范围。例如,在上述具体实施方式中所描述的各个具体技术特征,在不矛盾的情况下,可以通过任何合适的方式进行组合,为了避免 不必要的重复,本申请对各种可能的组合方式不再另行说明。又例如,本申请的各种不同的实施方式之间也可以进行任意组合,只要其不违背本申请的思想,其同样应当视为本申请所公开的内容。
还应理解,在本申请的各种方法实施例中,上述各过程的序号的大小并不意味着执行顺序的先后,各过程的执行顺序应以其功能和内在逻辑确定,而不应对本申请实施例的实施过程构成任何限定。另外,本申请实施例中,术语“和/或”,仅仅是一种描述关联对象的关联关系,表示可以存在三种关系。具体地,A和/或B可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。另外,本申请中字符“/”,一般表示前后关联对象是一种“或”的关系。
上文结合图14至图20,详细描述了本申请的方法实施例,下文结合图21,详细描述本申请的装置实施例。
图21是本申请一实施例提供的视频解码装置的示意性框图,该视频解码装置10应用于上述视频解码器。
如图21所示,视频解码装置10包括:
确定单元11,用于确定当前块的第一运动信息;
改善单元12,用于基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息;
预测单元13,用于基于所述第二运动信息,确定所述当前块的预测值。
在一些实施例中,改善单元12,具体用于基于所述第一运动信息,在所述参考图像中确定所述当前块对应的参考块;基于所述参考块的运动信息,确定所述当前块的时域运动信息作为第三运动信息;基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息。
在一些实施例中,改善单元12,具体用于基于所述第三运动信息和所述第一运动信息,确定第四运动信息;基于所述第四运动信息,确定所述第二运动信息。
在一些实施例中,改善单元12,具体用于将所述第三运动信息和所述第一运动信息的平均值,确定为所述确定第四运动信息。
在一些实施例中,改善单元12,具体用于确定所述第三运动信息和所述第一运动信息对应的权重;基于所述权重,确定所述第三运动信息和所述第一运动信息的加权平均值;将所述加权平均值,确定为所述确定第四运动信息。
在一些实施例中,所述第一运动信息的权重大于所述第三运动信息的权重。
在一些实施例中,改善单元12,具体用于以所述第四运动信息在所述参考图像中对应的位置为所述第二运动信息的搜索中心点,在所述参考图像中搜索得到所述第二运动信息。
在一些实施例中,所述参考图像包括第一参考图像和第二参考图像,所述第二运动信息和所述第四运动信息均包括第一预测方向运动信息和第二预测方向运动信息,改善单元12,具体用于以所述第四运动信息中的第一预测方向运动信息在所述第一参考图像中对应的位置,为所述第二运动信息中的第一预测方向运动信息的搜索中心点,以所述第四运动信息中的第二预测方向运动信息在所述第二参考图像中对应的位置,为所述第二运动信息中的第二预测方向运动信息的搜索中心点,在所述第一参考图像和所述第二参考图像的预设搜索范围内进行运动信息搜索,确定搜索到的每一对双边运动信息的双边匹配代价,所述每一对双边运动信息包括一个第一预测方向运动信息和一个第二预测方向运动信息;基于所述双边匹配代价,从搜索到的多对双边运动信息中,确定所述第二运动信息。
在一些实施例中,改善单元12,具体用于对于搜索到的第i对双边运动信息,基于所述第i对双边运动信息中的第一预测方向运动信息,在所述第一参考图像中确定第一预测块,基于所述第i对双边运动信息中的第二预测方向运动信息,在所述第二参考图像中确定第二预测块,所述i为正整数;确定所述第一预测块和所述第二预测块的匹配代价;基于所述第一预测块和所述第二预测块的匹配代价,确定所述第i对双边运动信息的双边匹配代价。
在一些实施例中,改善单元12,具体用于将所述搜索到的多对双边运动信息中双边匹配代价最小的一对双边运动信息,确定为所述第二运动信息。
在一些实施例中,改善单元12,具体用于以所述第一运动信息在所述参考图像中对应的位置为所述第二运动信息的搜索中心点,在所述参考图像中进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价;基于所述候选运动信息和所述第四运动信息,确定所述候选运动信息对应的代价系数;基于所述候选运动信息对应的代价系数对所述第一代价进行修正,得到所述候选运动信息的第二代价;基于搜索到的多个候选运动信息的第二代价,确定所述第二运动信息。
在一些实施例中,所述参考图像包括第一参考图像和第二参考图像,所述第一运动信息、所述第二运动信息、所述第四运动信息和所述候选运动信息均包括第一预测方向运动信息和第二预测方向运动信息,改善单元12,具体用于以所述第一运动信息中的第一预测方向运动信息在所述第一参考图像中对应的位置,为所述第二运动信息中的第一预测方向运动信息的搜索中心点,以所述第一运动信息中的第二预测方向运动信息在所述第二参考图像中对应的位置,为所述第二运动信息中的第二预测方向运动信息的搜索中心点,在所述第一参考图像和所述第二参考图像的预设搜索范围内进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价。
在一些实施例中,若所述候选运动信息对应的代价系数包括第一代价系数和第二代价系数,改善单元12,具体用于基于所述候选运动信息中的第一预测方向运动信息和所述第四运动信息中的第一预测方向运动信息,确定所述候选运动信息中的第一预测方向运动信息对应的第一代价系数;基于所述候选运动信息中的第二预测方向运动信息和所述第四运动信息中的第二预测方向运动信息,确定所述候选运动信息中的第二预测方向运动信息对应的第二代价系数。
在一些实施例中,改善单元12,具体用于确定所述候选运动信息中的第i预测方向运动信息,与所述第四运动信息中的第i预测方向运动信息之间的差的绝对值,所述i为一或二;基于所述差的绝对值,确定所述第i代价系数,其中所述第i代价系数与所述差的绝对值负相关。
在一些实施例中,改善单元12,具体用于确定所述差的绝对值与第一预设值中的最小值;基于所述最小值,确定所述第i代价系数。
在一些实施例中,改善单元12,具体用于将所述最小值和第二预设值的和,确定为所述第i代价系数。
在一些实施例中,改善单元12,具体用于基于所述第一代价系数和所述第二代价系数,对所述候选运动信息的第一代价进行修正,得到所述候选运动信息的第二代价。
在一些实施例中,改善单元12,具体用于将所述第一代价与所述第一代价系数和所述第二代价系数相乘,得到所述候选运动信息第二代价。
在一些实施例中,改善单元12,具体用于将所述搜索到的多个候选运动信息中第二代价最小的候选运动信息,确定为所述第二运动信息。
在一些实施例中,改善单元12,在基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息之前,还确定所述第一运动信息和所述第三运动信息之间的差异值;若所述差异值小于或等于预设阈值时,则基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息。
在一些实施例中,改善单元12,具体用于基于所述参考图像的运动信息,将所述当前块划分为至少一个子块;对于所述至少一个子块中的第i个子块,对所述第i个子块的第一运动信息进行改善,得到所述第i个子块的第二运动信息,所述i为正整数;基于所述N个子块的第二运动信息,得到所述当前块的第二运动信息。
在一些实施例中,改善单元12,具体用于基于所述第i个子块的第一运动信息,在所述参考图像中确定所述第i个子块对应的参考块;确定所述第i个子块按照所述第i个子块的参考块的运动信息,从当前图像运动到所述参考图像的第三运动信息;基于所述第三运动信息对所述第i个子块的第一运动信息进行改善,得到所述第i个子块的第二运动信息。
在一些实施例中,改善单元12,具体用于基于所述当前块的参考图像的运动信息,对所述第一运动信息进行N轮改善,得到所述第二运动信息,所述N为大于1的正整数。
在一些实施例中,改善单元12,具体用于对于所述第j轮对应的至少一个子块中的每一个子块的运动信息进行改善,得到所述第j轮对应的至少一个子块的改善后的运动信息,若所述j等于1时,则所述第j轮对应的子块为所述当前块;基于所述参考图像的运动信息,对所述第j轮对应的至少一个子块中的每一个子块进行块划分,得到第j+1轮对应的至少一个子块;对于所述第j+1轮对应的至少一个子块中的每一个子块的运动信息进行改善,重复执行N轮,得到所述第二运动信息。
在一些实施例中,改善单元12,在基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息之前,还用于确定所述当前块是否满足预设的基于整块的运动矢量改善条件;若所述当前块满足所述基于整块的运动矢量改善条件,基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息。
在一些实施例中,若所述当前块不满足所述基于整块的运动矢量改善条件,则改善单元12,还用于基于所述参考图像的运动信息,对所述当前块进行块划分,得到多个第一子块;对于所述多个第一子块的任一第一子块,确定所述第一子块是否满足所述基于整块的运动矢量改善条件;若所述第一子块不满足所述基于整块的运动矢量改善条件,则基于所述参考图像的运动信息,对所述第一子块进行块划分,得到多个第二子块;对于所述多个第二子块中的任一第二子块,确定所述第二子块是否满足所述基于整块的运动矢量改善条件,重复执行,直到划分后的子块满足所述基于整块的运动矢量改善条件,或者所述划分后的子块的大小满足预设大小为止。
在一些实施例中,改善单元12,具体用于确定所述第一块在所述参考图像中对应的参考块,所述第一块为所述当前块或者为第一子块或者为第二子块;获取所述第一块的M个子块在所述参考块中对应的M个子块的运动信息,所述M为大于1的正整数;对获取的M个子块的运动信息进行分类,得到P种分类结果,所述P为小于或等于M的正整数;基于所述P种分类结果,确定所述第一块是否满足所述基于整块的运动矢量改善条件。
在一些实施例中,改善单元12,具体用于若所述P等于1,则确定所述第一块满足所述基于整块的运动矢量改善条件;若所述P大于1,则确定所述第一块不满足所述基于整块的运动矢量改善条件。
在一些实施例中,改善单元12,具体用于确定所述第二块在所述参考图像中对应的参考块,所述第二块为所述当前块、第一子块、第二子块、第j轮对应的子块中的任意一个;获取所述第二块的M个子块在所述参考块中对应的M个子块的运动信息,所述M为大于1的正整数;对获取的M个子块的运动信息进行分类,得到P种分类结果,所述P为小于或等于M的正整数;基于所述P种分类结果,将所述第二块划分为至少一个子块。
在一些实施例中,改善单元12,具体用于基于所述P种分类结果分别对应的子块,将所述第二块划分为P个子块。
应理解,装置实施例与方法实施例可以相互对应,类似的描述可以参照方法实施例。为避免重复,此处不再赘述。具体地,图21所示的装置10可以执行本申请实施例的解码端的解码方法,并且装置10中的各个单元的前述和其它操作和/或功能分别为了实现上述解码端的解码方法等各个方法中的相应流程,为了简洁,在此不再赘述。
上文中结合附图从功能单元的角度描述了本申请实施例的装置和系统。应理解,该功能单元可以通过硬件形式实现,也可以通过软件形式的指令实现,还可以通过硬件和软件单元组合实现。具体地,本申请实施例中的方法实施例的各步骤可以通过处理器中的硬件的集成逻辑电路和/或软件形式的指令完成,结合本申请实施例公开的方法的步骤可以直接体现为硬件译码处理器执行完成,或者用译码处理器中的硬件及软件单元组合执行完成。可选地,软件单元可以位于随机存储器,闪存、只读存储器、可编程只读存储器、电可擦写可编程存储器、寄存器等本领域的成熟的存储介质中。该存储介质位于存储器,处理器读取存储器中的信息,结合其硬件完成上述方法实施例中的步骤。
图22是本申请实施例提供的电子设备的示意性框图。
如图22所示,该电子设备30可以为本申请实施例所述的视频解码器,该电子设备30可包括:
存储器33和处理器32,该存储器33用于存储计算机程序34,并将该程序代码34传输给该处理器32。换言之,该处理器32可以从存储器33中调用并运行计算机程序34,以实现本申请实施例中的方法。
例如,该处理器32可用于根据该计算机程序34中的指令执行上述方法200中的步骤。
在本申请的一些实施例中,该处理器32可以包括但不限于:
通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现场可编程门阵列(Field Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等等。
在本申请的一些实施例中,该存储器33包括但不限于:
易失性存储器和/或非易失性存储器。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDR SDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(synch link DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DR RAM)。
在本申请的一些实施例中,该计算机程序34可以被分割成一个或多个单元,该一个或者多个单元被存储在该存储器33中,并由该处理器32执行,以完成本申请提供的方法。该一个或多个单元可以是能够完成特定功能的一系列计算机程序指令段,该指令段用于描述该计算机程序34在该电子设备30中的执行过程。
如图22所示,该电子设备30还可包括:
收发器33,该收发器33可连接至该处理器32或存储器33。
其中,处理器32可以控制该收发器33与其他设备进行通信,具体地,可以向其他设备发送信息或数据,或接收其他设备发送的信息或数据。收发器33可以包括发射机和接收机。收发器33还可以进一步包括天线,天线的数量可以为一个或多个。
应当理解,该电子设备30中的各个组件通过总线系统相连,其中,总线系统除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。
本申请还提供了一种计算机存储介质,其上存储有计算机程序,该计算机程序被计算机执行时使得该计算机能够执行上述方法实施例的方法。或者说,本申请实施例还提供一种包含指令的计算机程序产品,该指令被计算机执行时使得计算机执行上述方法实施例的方法。
本申请还提供了一种码流,该码流是根据上述编码方法生成的。
当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。该计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行该计算机程序指令时,全部或部分地产生按照本申请实施例该的流程或功能。该计算机可以是通用计算机、专用计算机、计算机网络、或者其他可编程装置。该计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,该计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、数字用户线(digital subscriber line,DSL))或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。该计算机可读存储介质可以是计算机能够存取的任何可用介质或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。该可用介质可以是磁性介质(例如,软盘、硬盘、磁带)、光介质(例如数字视频光盘(digital video disc,DVD))、或者半导体介质(例如固态硬盘(solid state disk,SSD))等。
本领域普通技术人员可以意识到,结合本申请中所公开的实施例描述的各示例的单元及算法步骤,能够以电子硬件、或者计算机软件和电子硬件的结合来实现。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,该单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。例如,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
以上内容,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以该权利要求的保护范围为准。

Claims (33)

  1. 一种视频解码方法,其特征在于,包括:
    确定当前块的第一运动信息;
    基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息;
    基于所述第二运动信息,确定所述当前块的预测值。
  2. 根据权利要求1所述的方法,其特征在于,所述基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息,包括:
    基于所述第一运动信息,在所述参考图像中确定所述当前块对应的参考块;
    基于所述参考块的运动信息,确定所述当前块的时域运动信息作为第三运动信息;
    基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息。
  3. 根据权利要求2所述的方法,其特征在于,所述基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息,包括:
    基于所述第三运动信息和所述第一运动信息,确定第四运动信息;
    基于所述第四运动信息,确定所述第二运动信息。
  4. 根据权利要求3所述的方法,其特征在于,所述基于所述第三运动信息和所述第一运动信息,确定第四运动信息,包括:
    将所述第三运动信息和所述第一运动信息的平均值,确定为所述确定第四运动信息。
  5. 根据权利要求3所述的方法,其特征在于,所述基于所述第三运动信息和所述第一运动信息,确定第四运动信息,包括:
    确定所述第三运动信息和所述第一运动信息对应的权重;
    基于所述权重,确定所述第三运动信息和所述第一运动信息的加权平均值;
    将所述加权平均值,确定为所述确定第四运动信息。
  6. 根据权利要求5所述的方法,其特征在于,所述第一运动信息的权重大于所述第三运动信息的权重。
  7. 根据权利要求3所述的方法,其特征在于,所述基于所述第四运动信息,确定所述第二运动信息,包括:
    以所述第四运动信息在所述参考图像中对应的位置为所述第二运动信息的搜索中心点,在所述参考图像中搜索得到所述第二运动信息。
  8. 根据权利要求7所述的方法,其特征在于,所述参考图像包括第一参考图像和第二参考图像,所述第二运动信息和所述第四运动信息均包括第一预测方向运动信息和第二预测方向运动信息,所述以所述第四运动信息在所述参考图像中对应的位置为所述第二运动信息的搜索中心点,在所述参考图像中搜索得到所述第二运动信息,包括:
    以所述第四运动信息中的第一预测方向运动信息在所述第一参考图像中对应的位置,为所述第二运动信息中的第一预测方向运动信息的搜索中心点,以所述第四运动信息中的第二预测方向运动信息在所述第二参考图像中对应的位置,为所述第二运动信息中的第二预测方向运动信息的搜索中心点,在所述第一参考图像和所述第二参考图像的预设搜索范围内进行运动信息搜索,确定搜索到的每一对双边运动信息的双边匹配代价,所述每一对双边运动信息包括一个第一预测方向运动信息和一个第二预测方向运动信息;
    基于所述双边匹配代价,从搜索到的多对双边运动信息中,确定所述第二运动信息。
  9. 根据权利要求8所述的方法,其特征在于,所述确定搜索到的每一对双边运动信息的双边匹配代价,包括:
    对于搜索到的第i对双边运动信息,基于所述第i对双边运动信息中的第一预测方向运动信息,在所述第一参考图像中确定第一预测块,基于所述第i对双边运动信息中的第二预测方向运动信息,在所述第二参考图像中确定第二预测块,所述i为正整数;
    确定所述第一预测块和所述第二预测块的匹配代价;
    基于所述第一预测块和所述第二预测块的匹配代价,确定所述第i对双边运动信息的双边匹配代价。
  10. 根据权利要求9所述的方法,其特征在于,所述基于所述双边匹配代价,从搜索到的多对双边运动信息中,确定所述第二运动信息,包括:
    将所述搜索到的多对双边运动信息中双边匹配代价最小的一对双边运动信息,确定为所述第二运动信息。
  11. 根据权利要求3所述的方法,其特征在于,所述基于所述第四运动信息,确定所述第二运动信息,包括:
    以所述第一运动信息在所述参考图像中对应的位置为所述第二运动信息的搜索中心点,在所述参考图像中进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价;
    基于所述候选运动信息和所述第四运动信息,确定所述候选运动信息对应的代价系数;
    基于所述候选运动信息对应的代价系数对所述第一代价进行修正,得到所述候选运动信息的第二代价;
    基于搜索到的多个候选运动信息的第二代价,确定所述第二运动信息。
  12. 根据权利要求11所述的方法,其特征在于,所述参考图像包括第一参考图像和第二参考图像,所述第一运动信息、所述第二运动信息、所述第四运动信息和所述候选运动信息均包括第一预测方向运动信息和第二预测方向运动信息,所述以所述第一运动信息在所述参考图像中对应的位置为所述第二运动信息的搜索中心点,在所述参考图像中进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价,包括:
    以所述第一运动信息中的第一预测方向运动信息在所述第一参考图像中对应的位置,为所述第二运动信息中的第一预测方向运动信息的搜索中心点,以所述第一运动信息中的第二预测方向运动信息在所述第二参考图像中对应的位置,为所述第二运动信息中的第二预测方向运动信息的搜索中心点,在所述第一参考图像和所述第二参考图像的预设搜索范围内进行运动信息搜索,确定搜索到的每一候选运动信息的第一代价。
  13. 根据权利要求12所述的方法,其特征在于,若所述候选运动信息对应的代价系数包括第一代价系数和第二代价系数,则所述基于所述候选运动信息和所述第四运动信息,确定所述候选运动信息对应的代价系数,包括:
    基于所述候选运动信息中的第一预测方向运动信息和所述第四运动信息中的第一预测方向运动信息,确定所述候选运动信息中的第一预测方向运动信息对应的第一代价系数;
    基于所述候选运动信息中的第二预测方向运动信息和所述第四运动信息中的第二预测方向运动信息,确定所述候选运动信息中的第二预测方向运动信息对应的第二代价系数。
  14. 根据权利要求13所述的方法,其特征在于,基于所述候选运动信息中的第i预测方向运动信息和所述第四运动信息中的第i预测方向运动信息,确定所述确定所述候选运动信息中的第i预测方向运动信息对应的第i代价系数,包括:
    确定所述候选运动信息中的第i预测方向运动信息,与所述第四运动信息中的第i预测方向运动信息之间的差的绝对值,所述i为一或二;
    基于所述差的绝对值,确定所述第i代价系数,其中所述第i代价系数与所述差的绝对值负相关。
  15. 根据权利要求14所述的方法,其特征在于,所述基于所述差的绝对值,确定所述第i代价系数,包括:
    确定所述差的绝对值与第一预设值中的最小值;
    基于所述最小值,确定所述第i代价系数。
  16. 根据权利要求15所述的方法,其特征在于,所述基于所述最小值,确定所述第i代价系数,包括:
    将所述最小值和第二预设值的和,确定为所述第i代价系数。
  17. 根据权利要求13所述的方法,其特征在于,基于所述候选运动信息对应的代价系数对所述第一代价进行修正,得到所述候选运动信息的第二代价,包括:
    基于所述第一代价系数和所述第二代价系数,对所述候选运动信息的第一代价进行修正,得到所述候选运动信息的第二代价。
  18. 根据权利要求17所述的方法,其特征在于,所述基于所述第一代价系数和所述第二代价系数,对所述候选运动信息的第一代价进行修正,得到所述候选运动信息的第二代价,包括:
    将所述第一代价与所述第一代价系数和所述第二代价系数相乘,得到所述候选运动信息第二代价。
  19. 根据权利要求12所述的方法,其特征在于,所述基于搜索到的多个候选运动信息的第二代价,确定所述第二运动信息,包括:
    将所述搜索到的多个候选运动信息中第二代价最小的候选运动信息,确定为所述第二运动信息。
  20. 根据权利要求2-19任一项所述的方法,其特征在于,所述基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息之前,所述方法包括:
    确定所述第一运动信息和所述第三运动信息之间的差异值;
    所述基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息,包括:
    若所述差异值小于或等于预设阈值时,则基于所述第三运动信息对所述第一运动信息进行改善,得到所述第二运动信息。
  21. 根据权利要求1-19任一项所述的方法,其特征在于,所述基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息,包括:
    基于所述参考图像的运动信息,将所述当前块划分为至少一个子块;
    对于所述至少一个子块中的第i个子块,对所述第i个子块的第一运动信息进行改善,得到所述第i个子块的第二运动信息,所述i为正整数;
    基于所述N个子块的第二运动信息,得到所述当前块的第二运动信息。
  22. 根据权利要求21所述的方法,其特征在于,所述对所述第i个子块的第一运动信息进行改善,得到所述第i个子块的第二运动信息,包括:
    基于所述第i个子块的第一运动信息,在所述参考图像中确定所述第i个子块对应的参考块;
    确定所述第i个子块按照所述第i个子块的参考块的运动信息,从当前图像运动到所述参考图像的第三运动信息;
    基于所述第三运动信息对所述第i个子块的第一运动信息进行改善,得到所述第i个子块的第二运动信息。
  23. 根据权利要求1-19任一项所述的方法,其特征在于,所述基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息,包括:
    基于所述当前块的参考图像的运动信息,对所述第一运动信息进行N轮改善,得到所述第二运动信息,所述N为大于1的正整数。
  24. 根据权利要求23所述的方法,其特征在于,所述基于所述当前块的参考图像的运动信息,对所述第一运动信息进行N轮改善,得到所述第二运动信息,包括:
    对于第j轮对应的至少一个子块中的每一个子块的运动信息进行改善,得到所述第j轮对应的至少一个子块的改善后的运动信息,若所述j等于1时,则所述第j轮对应的子块为所述当前块;
    基于所述参考图像的运动信息,对所述第j轮对应的至少一个子块中的每一个子块进行块划分,得到第j+1轮对应的至少一个子块;
    对于所述第j+1轮对应的至少一个子块中的每一个子块的运动信息进行改善,重复执行N轮,得到所述第二运动信息。
  25. 根据权利要求1-19任一项所述的方法,其特征在于,所述基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息之前,所述方法还包括:
    确定所述当前块是否满足预设的基于整块的运动矢量改善条件;
    所述基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息,包括:
    若所述当前块满足所述基于整块的运动矢量改善条件,基于所述当前块的参考图像的运动信息,对所述第一运动 信息进行改善,得到所述当前块的第二运动信息。
  26. 根据权利要求25所述的方法,其特征在于,若所述当前块不满足所述基于整块的运动矢量改善条件,则所述方法还包括:
    基于所述参考图像的运动信息,对所述当前块进行块划分,得到多个第一子块;
    对于所述多个第一子块的任一第一子块,确定所述第一子块是否满足所述基于整块的运动矢量改善条件;
    若所述第一子块不满足所述基于整块的运动矢量改善条件,则基于所述参考图像的运动信息,对所述第一子块进行块划分,得到多个第二子块;
    对于所述多个第二子块中的任一第二子块,确定所述第二子块是否满足所述基于整块的运动矢量改善条件,重复执行,直到划分后的子块满足所述基于整块的运动矢量改善条件,或者所述划分后的子块的大小满足预设大小为止。
  27. 根据权利要求25所述的方法,其特征在于,确定第一块是否满足基于整块的运动矢量改善条件,包括:
    确定所述第一块在所述参考图像中对应的参考块,所述第一块为所述当前块或者为第一子块或者为第二子块;
    获取所述第一块的M个子块在所述参考块中对应的M个子块的运动信息,所述M为大于1的正整数;
    对获取的M个子块的运动信息进行分类,得到P种分类结果,所述P为小于或等于M的正整数;
    基于所述P种分类结果,确定所述第一块是否满足所述基于整块的运动矢量改善条件。
  28. 根据权利要求27所述的方法,其特征在于,所述基于所述P种分类结果,确定所述第一块是否满足所述基于整块的运动矢量改善条件,包括:
    若所述P等于1,则确定所述第一块满足所述基于整块的运动矢量改善条件;
    若所述P大于1,则确定所述第一块不满足所述基于整块的运动矢量改善条件。
  29. 根据权利要求21或24或26所述的方法,其特征在于,基于所述参考图像的运动信息,对第二块进行块划分,包括:
    确定所述第二块在所述参考图像中对应的参考块,所述第二块为所述当前块、第一子块、第二子块、第j轮对应的子块中的任意一个;
    获取所述第二块的M个子块在所述参考块中对应的M个子块的运动信息,所述M为大于1的正整数;
    对获取的M个子块的运动信息进行分类,得到P种分类结果,所述P为小于或等于M的正整数;
    基于所述P种分类结果,将所述第二块划分为至少一个子块。
  30. 根据权利要求29所述的方法,其特征在于,所述基于所述P种分类结果,将所述第二块划分为至少一个子块,包括:
    基于所述P种分类结果分别对应的子块,将所述第二块划分为P个子块。
  31. 一种视频解码装置,其特征在于,包括:
    确定单元,用于确定当前块的第一运动信息;
    改善单元,用于基于所述当前块的参考图像的运动信息,对所述第一运动信息进行改善,得到所述当前块的第二运动信息;
    预测单元,用于基于所述第二运动信息,确定所述当前块的预测值。
  32. 一种电子设备,其特征在于,包括处理器和存储器;
    所示存储器用于存储计算机程序;
    所述处理器用于调用并运行所述存储器中存储的计算机程序,以实现上述权利要求1至30任一项所述的方法。
  33. 一种计算机可读存储介质,其特征在于,用于存储计算机程序;
    所述计算机程序使得计算机执行如上述权利要求1至30任一项所述的方法。
PCT/CN2023/105570 2023-07-03 2023-07-03 视频解码方法、装置、设备、及存储介质 Ceased WO2025007250A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
CN202380100003.XA CN121511598A (zh) 2023-07-03 2023-07-03 视频解码方法、装置、设备、及存储介质
PCT/CN2023/105570 WO2025007250A1 (zh) 2023-07-03 2023-07-03 视频解码方法、装置、设备、及存储介质
MX2025015514A MX2025015514A (es) 2023-07-03 2025-12-18 Metodo y aparato de decodificacion de video, dispositivo y medio de almacenamiento
US19/428,443 US20260113474A1 (en) 2023-07-03 2025-12-22 Video decoding method and apparatus, and device and storage medium

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/105570 WO2025007250A1 (zh) 2023-07-03 2023-07-03 视频解码方法、装置、设备、及存储介质

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US19/428,443 Continuation US20260113474A1 (en) 2023-07-03 2025-12-22 Video decoding method and apparatus, and device and storage medium

Publications (1)

Publication Number Publication Date
WO2025007250A1 true WO2025007250A1 (zh) 2025-01-09

Family

ID=94171087

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/105570 Ceased WO2025007250A1 (zh) 2023-07-03 2023-07-03 视频解码方法、装置、设备、及存储介质

Country Status (4)

Country Link
US (1) US20260113474A1 (zh)
CN (1) CN121511598A (zh)
MX (1) MX2025015514A (zh)
WO (1) WO2025007250A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2018128232A1 (ko) * 2017-01-03 2018-07-12 엘지전자 주식회사 영상 코딩 시스템에서 영상 디코딩 방법 및 장치
WO2019194502A1 (ko) * 2018-04-01 2019-10-10 엘지전자 주식회사 인터 예측 모드 기반 영상 처리 방법 및 이를 위한 장치
CN110876059A (zh) * 2018-09-03 2020-03-10 华为技术有限公司 运动矢量的获取方法、装置、计算机设备及存储介质
CN112565768A (zh) * 2020-12-02 2021-03-26 浙江大华技术股份有限公司 一种帧间预测方法、编解码系统及计算机可读存储介质
CN113747172A (zh) * 2020-05-29 2021-12-03 Oppo广东移动通信有限公司 帧间预测方法、编码器、解码器以及计算机存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2018128232A1 (ko) * 2017-01-03 2018-07-12 엘지전자 주식회사 영상 코딩 시스템에서 영상 디코딩 방법 및 장치
WO2019194502A1 (ko) * 2018-04-01 2019-10-10 엘지전자 주식회사 인터 예측 모드 기반 영상 처리 방법 및 이를 위한 장치
CN110876059A (zh) * 2018-09-03 2020-03-10 华为技术有限公司 运动矢量的获取方法、装置、计算机设备及存储介质
CN113747172A (zh) * 2020-05-29 2021-12-03 Oppo广东移动通信有限公司 帧间预测方法、编码器、解码器以及计算机存储介质
CN112565768A (zh) * 2020-12-02 2021-03-26 浙江大华技术股份有限公司 一种帧间预测方法、编解码系统及计算机可读存储介质

Also Published As

Publication number Publication date
CN121511598A (zh) 2026-02-10
US20260113474A1 (en) 2026-04-23
MX2025015514A (es) 2026-02-03

Similar Documents

Publication Publication Date Title
CN118945317B (zh) 对仿射译码块进行光流预测修正的方法及装置
KR20210072064A (ko) 인터 예측 방법 및 장치
CN117596398A (zh) 用于译码块的几何划分块的帧间预测的装置及方法
CN117956166A (zh) 一种编码器、解码器及去块滤波器的边界强度的对应推导方法
CN111953995A (zh) 一种帧间预测的方法和装置
KR102616713B1 (ko) 이미지 예측 방법, 장치 및 시스템, 디바이스 및 저장 매체
TWI806212B (zh) 視訊編碼器、視訊解碼器及相應方法
CN113196783B (zh) 去块效应滤波自适应的编码器、解码器及对应方法
CN113711601A (zh) 用于推导当前块的插值滤波器索引的方法和装置
JP7640647B2 (ja) 双予測のオプティカルフロー計算および双予測補正におけるブロックレベル境界サンプル勾配計算のための整数グリッド参照サンプルの位置を計算するための方法
CN114913249B (zh) 编码、解码方法和相关设备
US20240372984A1 (en) Prediction methods
CN114079786A (zh) 一种帧间预测的方法和装置
CN118435595A (zh) 帧内预测方法、设备、系统、及存储介质
CN111866502A (zh) 图像预测方法、装置和计算机可读存储介质
WO2024007128A1 (zh) 视频编解码方法、装置、设备、系统、及存储介质
WO2023122968A1 (zh) 帧内预测方法、设备、系统、及存储介质
JP2025539476A (ja) ビデオ処理方法および装置
WO2025007250A1 (zh) 视频解码方法、装置、设备、及存储介质
WO2024108391A1 (zh) 视频编解码方法、装置、设备、系统、及存储介质
CN119137938B (zh) 视频编码方法、装置、设备、系统、及存储介质
WO2025065522A1 (zh) 编解码方法及装置、编解码器、码流、设备、存储介质
WO2024092425A1 (zh) 视频编解码方法、装置、设备、及存储介质
WO2025208348A1 (zh) 编解码方法、编解码器、码流以及存储介质
KR20260036172A (ko) 인코딩 및 디코딩 방법과 장치

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23943984

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: MX/A/2025/015514

Country of ref document: MX

WWE Wipo information: entry into national phase

Ref document number: 202617011327

Country of ref document: IN

WWP Wipo information: published in national office

Ref document number: MX/A/2025/015514

Country of ref document: MX

NENP Non-entry into the national phase

Ref country code: DE

WWP Wipo information: published in national office

Ref document number: 202617011327

Country of ref document: IN