WO2023172002A1 - 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 - Google Patents
영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 Download PDFInfo
- Publication number
- WO2023172002A1 WO2023172002A1 PCT/KR2023/003020 KR2023003020W WO2023172002A1 WO 2023172002 A1 WO2023172002 A1 WO 2023172002A1 KR 2023003020 W KR2023003020 W KR 2023003020W WO 2023172002 A1 WO2023172002 A1 WO 2023172002A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- motion vector
- block
- current
- prediction block
- reference picture
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/109—Selection of coding mode or of prediction mode among a plurality of temporal predictive coding modes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
- H04N19/137—Motion inside a coding unit, e.g. average field, frame or block difference
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/577—Motion compensation with bidirectional frame interpolation, i.e. using B-pictures
Definitions
- the present invention relates to a video encoding/decoding method, device, and recording medium storing bitstreams. Specifically, the present invention relates to an image encoding/decoding method and device using a decoder side motion vector refinement (DMVR) method, and a recording medium storing a bitstream.
- DMVR decoder side motion vector refinement
- motion information such as motion vectors and reference picture information from neighboring blocks of the current block can be used for prediction of the current block.
- DMVR decoder side motion vector refinement
- the purpose of the present invention is to provide a video encoding/decoding method and device with improved encoding/decoding efficiency.
- Another object of the present invention is to provide a recording medium that stores a bitstream generated by the video decoding method or device provided by the present invention.
- An image decoding method includes determining a first basic motion vector of a current block for a first reference picture and a second basic motion vector of the current block for a second reference picture, a first Determining a first basic motion vector of the current block with respect to a reference picture and a second basic motion vector of the current block with respect to the second reference picture, by correcting the first basic motion vector by a first differential motion vector , determining a first corrected motion vector, and determining a second corrected motion vector by correcting the second basic motion vector by a second differential motion vector; Based on the weighted sum of the first prediction block and the second prediction block, determining a first prediction block and a second prediction block of the current block, and determining a final prediction block for the current block. It includes steps to:
- the distance between the current picture including the current block and the first reference picture and the current picture and the second reference picture If the distances of the reference pictures are different, the first correction motion vector and the second correction motion vector may be determined.
- the weighted sum of the first prediction block and the second prediction block is a weight value determined according to the distance between the current picture and the first reference picture and the distance between the current picture and the second reference picture. It can be decided by .
- the first weight applied to the first prediction block is proportional to the distance between the current picture and the second reference picture
- the second weight applied to the second prediction block is proportional to the distance between the current picture and the second reference picture. and may be proportional to the distance between the first reference picture and the first reference picture.
- the ratio of the size of the first differential motion vector and the size of the second differential motion vector is the current picture and It may be proportional to the ratio of the distance between the first reference picture and the distance between the current picture and the second reference picture.
- the sizes of the first differential motion vector and the second differential motion vector may be limited to within a predetermined range.
- the first correction motion vector and the second prediction block are minimized so that distortion between the first prediction block indicated by the first correction motion vector and the second prediction block indicated by the second correction motion vector is minimized.
- a corrective motion vector may be determined.
- the first corrected motion vector and the second corrected motion vector may be determined so that distortion is minimized.
- the distortion may be calculated based on a final template determined by weighting the first template and the second template and the current template.
- the first weight applied to the first prediction block is proportional to the distortion of the second template and the current template
- the second weight applied to the second prediction block is proportional to the distortion of the second template and the current template. and may be proportional to the distortion of the current template.
- the weighted sum of the first prediction block and the second prediction block may be determined by a weight value determined based on the picture type of the current picture including the current block.
- the weighted sum of the first prediction block and the second prediction block may be determined by a weight value determined based on weight value information generated by parsing a bitstream.
- An image encoding method includes determining a first basic motion vector of a current block for a first reference picture and a second basic motion vector of the current block for a second reference picture, determining a first corrected motion vector by correcting a first basic motion vector by a first differential motion vector, and determining a second corrected motion vector by correcting the second basic motion vector by a second differential motion vector, Based on the first correction motion vector and the second correction motion vector, determining a first prediction block and a second prediction block of the current block, and a weighted sum of the first prediction block and the second prediction block Based on this, determining a final prediction block for the current block.
- a non-transitory computer-readable recording medium stores a bitstream generated by the video encoding method.
- the transmission method according to an embodiment of the present invention transmits a bitstream generated by the video encoding method.
- the present invention proposes a method for improving decoder side motion vector refinement (DMVR).
- the frequency of application of the decoding-side motion vector correction method can be increased by changing the application conditions of the decoding-side motion vector correction method.
- coding efficiency can be improved by increasing the accuracy of the motion vector correction method on the decoding side and increasing the accuracy of the final prediction block.
- FIG. 1 is a block diagram showing the configuration of an encoding device to which the present invention is applied according to an embodiment.
- Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment to which the present invention is applied.
- Figure 3 is a diagram schematically showing a video coding system to which the present invention can be applied.
- Figure 4 shows an example of a method for determining one of various inter-screen prediction methods in the inter-screen prediction mode.
- Figure 5 shows an example of motion vector correction on the decoding side when the temporal distance between the current picture and the L0 reference picture is the same as the temporal distance between the current picture and the L1 reference picture.
- Figure 6 shows motion vector correction on the decoding side when the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture are not the same.
- Figure 7 explains motion vector correction on the decoding side based on the template matching method.
- Figure 8 is a flowchart of a decoding-side motion vector correction method according to an embodiment.
- Figure 9 exemplarily shows a content streaming system to which an embodiment according to the present invention can be applied.
- An image decoding method includes determining a first basic motion vector of a current block for a first reference picture and a second basic motion vector of the current block for a second reference picture, a first Determining a first basic motion vector of the current block with respect to a reference picture and a second basic motion vector of the current block with respect to the second reference picture, by correcting the first basic motion vector by a first differential motion vector , determining a first corrected motion vector, and determining a second corrected motion vector by correcting the second basic motion vector by a second differential motion vector; Based on the weighted sum of the first prediction block and the second prediction block, determining a first prediction block and a second prediction block of the current block, and determining a final prediction block for the current block. It includes steps to:
- first and second may be used to describe various components, but the components should not be limited by the terms.
- the above terms are used only for the purpose of distinguishing one component from another.
- a first component may be named a second component, and similarly, the second component may also be named a first component without departing from the scope of the present invention.
- the term and/or includes any of a plurality of related stated items or a combination of a plurality of related stated items.
- each component is listed and included as a separate component for convenience of explanation, and at least two of each component can be combined to form one component, or one component can be divided into a plurality of components to perform a function, and each of these components can perform a function.
- Integrated embodiments and separate embodiments of the constituent parts are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
- the terms used in the present invention are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. Additionally, some of the components of the present invention may not be essential components that perform essential functions in the present invention, but may be merely optional components to improve performance. The present invention can be implemented by including only essential components for implementing the essence of the present invention excluding components used only to improve performance, and a structure including only essential components excluding optional components used only to improve performance. is also included in the scope of rights of the present invention.
- the term “at least one” may mean one of numbers greater than 1, such as 1, 2, 3, and 4. In embodiments, the term “a plurality of” may mean one of two or more numbers, such as 2, 3, and 4.
- video may refer to a single picture that constitutes a video, or may refer to the video itself.
- encoding and/or decoding of a video may mean “encoding and/or decoding of a video,” or “encoding and/or decoding of one of the videos that make up a video.” It may be possible.
- the target image may be an encoding target image that is the target of encoding and/or a decoding target image that is the target of decoding. Additionally, the target image may be an input image input to an encoding device or may be an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
- image may be used with the same meaning and may be used interchangeably.
- target block may be an encoding target block that is the target of encoding and/or a decoding target block that is the target of decoding. Additionally, the target block may be a current block that is currently the target of encoding and/or decoding. For example, “target block” and “current block” may be used with the same meaning and may be used interchangeably.
- a Coding Tree Unit may be composed of two chrominance component (Cb, Cr) coding tree blocks related to one luminance component (Y) coding tree block (CTB). .
- sample may represent the basic unit constituting the block.
- FIG. 1 is a block diagram showing the configuration of an encoding device to which the present invention is applied according to an embodiment.
- the encoding device 100 may be an encoder, a video encoding device, or an image encoding device.
- a video may contain one or more images.
- the encoding device 100 can sequentially encode one or more images.
- the encoding device 100 includes an image segmentation unit 110, an intra prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switch 115, a subtractor 113, A transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 117, a filter unit 180, and a reference picture buffer 190. It can be included.
- the encoding device 100 can generate a bitstream including encoded information through encoding of an input image and output the generated bitstream.
- the generated bitstream can be stored in a computer-readable recording medium or streamed through wired/wireless transmission media.
- the image segmentation unit 110 may divide the input image into various forms to increase the efficiency of video encoding/decoding.
- the input video consists of multiple pictures, and one picture can be hierarchically divided and processed for compression efficiency, parallel processing, etc.
- one picture can be divided into one or multiple tiles or slices and further divided into multiple CTUs (Coding Tree Units).
- one picture may first be divided into a plurality of sub-pictures defined as a group of rectangular slices, and each sub-picture may be divided into the tiles/slices.
- subpictures can be used to support the function of partially independently encoding/decoding and transmitting a picture.
- bricks can be created by dividing tiles horizontally.
- a brick can be used as a basic unit of intra-picture parallel processing.
- one CTU can be recursively divided into a quad tree (QT: Quadtree), and the end node of the division can be defined as a CU (Coding Unit).
- CU can be divided into PU (Prediction Unit), which is a prediction unit, and TU (Transform Unit), which is a transformation unit, and prediction and division can be performed. Meanwhile, CUs can be used as prediction units and/or transformation units themselves.
- each CTU may be recursively partitioned into not only a quad tree (QT) but also a multi-type tree (MTT).
- CTU can begin to be divided into a multi-type tree from the end node of QT, and MTT can be composed of BT (Binary Tree) and TT (Triple Tree).
- MTT can be composed of BT (Binary Tree) and TT (Triple Tree).
- the MTT structure can be divided into vertical binary split mode (SPLIT_BT_VER), horizontal binary split mode (SPLIT_BT_HOR), vertical ternary split mode (SPLIT_TT_VER), and horizontal ternary split mode (SPLIT_TT_HOR).
- the minimum block size (MinQTSize) of the quad tree of the luminance block can be set to 16x16
- the maximum block size (MaxBtSize) of the binary tree can be set to 128x128, and the maximum block size (MaxTtSize) of the triple tree can be set to 64x64.
- the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the triple tree can be set to 4x4, and the maximum depth (MaxMttDepth) of the multi-type tree can be set to 4.
- a dual tree that uses different CTU division structures for the luminance and chrominance components can be applied.
- the luminance and chrominance CTB (Coding Tree Blocks) within the CTU can be divided into a single tree that shares the coding tree structure.
- the encoding device 100 may perform encoding on an input image in intra mode and/or inter mode.
- the encoding device 100 may perform encoding on the input image in a third mode (eg, IBC mode, Palette mode, etc.) other than the intra mode and inter mode.
- a third mode eg, IBC mode, Palette mode, etc.
- the third mode may be classified as intra mode or inter mode for convenience of explanation. In the present invention, the third mode will be classified and described separately only when a detailed explanation is needed.
- intra mode may mean intra-screen prediction mode
- inter mode may mean inter-screen prediction mode.
- the encoding device 100 may generate a prediction block for an input block of an input image. Additionally, after the prediction block is generated, the encoding device 100 may encode the residual block using the residual of the input block and the prediction block.
- the input image may be referred to as the current image that is currently the target of encoding.
- the input block may be referred to as the current block that is currently the target of encoding or the encoding target block.
- the intra prediction unit 120 may use samples of blocks that have already been encoded/decoded around the current block as reference samples.
- the intra prediction unit 120 may perform spatial prediction for the current block using a reference sample and generate prediction samples for the input block through spatial prediction.
- intra prediction may mean prediction within the screen.
- non-directional prediction modes such as DC mode and Planar mode and directional prediction modes (e.g., 65 directions) can be applied.
- the intra prediction method can be expressed as an intra prediction mode or an intra prediction mode.
- the motion prediction unit 121 can search for the area that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched area. . At this time, the search area can be used as the area.
- the reference image may be stored in the reference picture buffer 190.
- it when encoding/decoding of the reference image is processed, it may be stored in the reference picture buffer 190.
- the motion compensation unit 122 may generate a prediction block for the current block by performing motion compensation using a motion vector.
- inter prediction may mean inter-screen prediction or motion compensation.
- the motion prediction unit 121 and the motion compensation unit 122 can generate a prediction block by applying an interpolation filter to some areas in the reference image.
- the motion prediction and motion compensation methods of the prediction unit included in the coding unit based on the coding unit include skip mode, merge mode, and improved motion vector prediction ( It is possible to determine whether it is in Advanced Motion Vector Prediction (AMVP) mode or Intra Block Copy (IBC) mode, and inter-screen prediction or motion compensation can be performed depending on each mode.
- AMVP Advanced Motion Vector Prediction
- IBC Intra Block Copy
- AFFINE mode of sub-PU-based prediction based on the inter-screen prediction method, AFFINE mode of sub-PU-based prediction, Subblock-based Temporal Motion Vector Prediction (SbTMVP) mode, and Merge with MVD (MMVD) mode of PU-based prediction, Geometric Partitioning Mode (GPM) ) mode can also be applied.
- HMVP History based MVP
- PAMVP Packet based MVP
- CIIP Combined Intra/Inter Prediction
- AMVR Adaptive Motion Vector Resolution
- BDOF Bi-Directional Optical-Flow
- BCW Bi-predictive with CU Weights
- BCW Local Illumination Compensation
- TM Template Matching
- OBMC Overlapped Block Motion Compensation
- the subtractor 113 may generate a residual block using the difference between the input block and the prediction block.
- the residual block may also be referred to as a residual signal.
- the residual signal may refer to the difference between the original signal and the predicted signal.
- the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the predicted signal.
- the remaining block may be a residual signal in block units.
- the transform unit 130 may generate a transform coefficient by performing transformation on the remaining block and output the generated transform coefficient.
- the transformation coefficient may be a coefficient value generated by performing transformation on the remaining block.
- the transform unit 130 may skip transforming the remaining blocks.
- Quantized levels can be generated by applying quantization to the transform coefficients or residual signals.
- the quantized level may also be referred to as a transform coefficient.
- the 4x4 luminance residual block generated through intra-screen prediction is transformed using a DST (Discrete Sine Transform)-based basis vector, and the remaining residual blocks are transformed using a DCT (Discrete Cosine Transform)-based basis vector.
- DST Discrete Sine Transform
- DCT Discrete Cosine Transform
- RQT Residual Quad Tree
- the transform block for one block is divided into a quad tree form, and after performing transformation and quantization on each transform block divided through RQT, when all coefficients become 0,
- cbf coded block flag
- MTS Multiple Transform Selection
- RQT Multiple Transform Selection
- SBT Sub-block Transform
- LFNST Low Frequency Non-Separable Transform
- a secondary transform technology that further transforms the residual signal converted to the frequency domain through DCT or DST, can be applied.
- LFNST additionally performs transformation on the 4x4 or 8x8 low-frequency area in the upper left corner, allowing the residual coefficients to be concentrated in the upper left corner.
- the quantization unit 140 may generate a quantized level by quantizing a transform coefficient or a residual signal according to a quantization parameter (QP), and output the generated quantized level. At this time, the quantization unit 140 may quantize the transform coefficient using a quantization matrix.
- QP quantization parameter
- a quantizer using QP values of 0 to 51 can be used.
- 0 to 63 QP can be used.
- a DQ (Dependent Quantization) method that uses two quantizers instead of one quantizer can be applied. DQ performs quantization using two quantizers (e.g., Q0, Q1), but even without signaling information about the use of a specific quantizer, the quantizer to be used for the next transformation coefficient is determined based on the current state through a state transition model. It can be applied to be selected.
- the entropy encoding unit 150 can generate a bitstream by performing entropy encoding according to a probability distribution on the values calculated by the quantization unit 140 or the coding parameter values calculated during the encoding process. and bitstream can be output.
- the entropy encoding unit 150 may perform entropy encoding on information about image samples and information for decoding the image. For example, information for decoding an image may include syntax elements, etc.
- the entropy encoding unit 150 may use encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) for entropy encoding. For example, the entropy encoding unit 150 may perform entropy encoding using a Variable Length Coding/Code (VLC) table.
- VLC Variable Length Coding/Code
- the entropy encoding unit 150 derives a binarization method of the target symbol and a probability model of the target symbol/bin, and then uses the derived binarization method, probability model, and context model. Arithmetic coding can also be performed using .
- the table probability update method may be changed to a table update method using a simple formula. Additionally, two different probability models can be used to obtain more accurate symbol probability values.
- the entropy encoder 150 can change a two-dimensional block form coefficient into a one-dimensional vector form through a transform coefficient scanning method to encode the transform coefficient level (quantized level).
- Coding parameters include information (flags, indexes, etc.) encoded in the encoding device 100 and signaled to the decoding device 200, such as syntax elements, as well as information derived from the encoding or decoding process. It may include and may mean information needed when encoding or decoding an image.
- signaling a flag or index may mean that the encoder entropy encodes the flag or index and includes it in the bitstream, and the decoder may include the flag or index from the bitstream. This may mean entropy decoding.
- the encoded current image can be used as a reference image for other images to be processed later. Accordingly, the encoding device 100 can restore or decode the current encoded image, and store the restored or decoded image as a reference image in the reference picture buffer 190.
- the quantized level may be dequantized in the dequantization unit 160. It may be inverse transformed in the inverse transform unit 170.
- the inverse-quantized and/or inverse-transformed coefficients may be combined with the prediction block through the adder 117.
- a reconstructed block may be generated by combining the inverse-quantized and/or inverse-transformed coefficients with the prediction block.
- the inverse-quantized and/or inverse-transformed coefficient refers to a coefficient on which at least one of inverse-quantization and inverse-transformation has been performed, and may refer to a restored residual block.
- the inverse quantization unit 160 and the inverse transform unit 170 may be performed as reverse processes of the quantization unit 140 and the transform unit 130.
- the restored block may pass through the filter unit 180.
- the filter unit 180 includes a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter (BIF), and an LMCS (Luma). Mapping with Chroma Scaling) can be applied to restored samples, restored blocks, or restored images as all or part of the filtering techniques.
- the filter unit 180 may also be referred to as an in-loop filter. At this time, in-loop filter is also used as a name excluding LMCS.
- the deblocking filter can remove block distortion occurring at the boundaries between blocks. To determine whether to perform a deblocking filter, it is possible to determine whether to apply a deblocking filter to the current block based on the samples included in a few columns or rows included in the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering strength.
- Sample adaptive offset can correct the offset of the deblocked image with the original image on a sample basis. You can use a method of dividing the samples included in the image into a certain number of regions, then determining the region to perform offset and applying the offset to that region, or a method of applying the offset by considering the edge information of each sample.
- Bilateral filter can also correct the offset from the original image on a sample basis for the deblocked image.
- the adaptive loop filter can perform filtering based on a comparison value between the restored image and the original image. After dividing the samples included in the video into predetermined groups, filtering can be performed differentially for each group by determining the filter to be applied to that group. Information related to whether to apply an adaptive loop filter may be signaled for each coding unit (CU), and the shape and filter coefficients of the adaptive loop filter to be applied may vary for each block.
- CU coding unit
- LMCS Luma Mapping with Chroma Scaling
- LM luma-mapping
- CS chroma scaling
- This refers to a technology that scales the residual value of the color difference component according to the luminance value.
- LMCS can be used as an HDR correction technology that reflects the characteristics of HDR (High Dynamic Range) images.
- the reconstructed block or reconstructed image that has passed through the filter unit 180 may be stored in the reference picture buffer 190.
- the restored block that has passed through the filter unit 180 may be part of a reference image.
- the reference image may be a reconstructed image composed of reconstructed blocks that have passed through the filter unit 180.
- the stored reference image can then be used for inter-screen prediction or motion compensation.
- Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment to which the present invention is applied.
- the decoding device 200 may be a decoder, a video decoding device, or an image decoding device.
- the decoding device 200 includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, a motion compensation unit 250, and an adder 201. , it may include a switch 203, a filter unit 260, and a reference picture buffer 270.
- the decoding device 200 may receive the bitstream output from the encoding device 100.
- the decoding device 200 may receive a bitstream stored in a computer-readable recording medium or receive a bitstream streamed through a wired/wireless transmission medium.
- the decoding device 200 may perform decoding on a bitstream in intra mode or inter mode. Additionally, the decoding device 200 can generate a restored image or a decoded image through decoding, and output the restored image or a decoded image.
- the decoding device 200 can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block.
- the decoding device 200 may generate a restored block to be decoded by adding the restored residual block and the prediction block.
- the block to be decrypted may be referred to as the current block.
- the entropy decoding unit 210 may generate symbols by performing entropy decoding according to a probability distribution for the bitstream.
- the generated symbols may include symbols in the form of quantized levels.
- the entropy decoding method may be the reverse process of the entropy encoding method described above.
- the entropy decoder 210 can change one-dimensional vector form coefficients into two-dimensional block form through a transform coefficient scanning method in order to decode the transform coefficient level (quantized level).
- the intra prediction unit 240 may generate a prediction block by performing spatial prediction on the current block using sample values of already decoded blocks surrounding the decoding target block.
- the intra prediction unit 240 applied to the decoding device may use the same technology as the intra prediction unit 120 applied to the above-described encoding device.
- the motion compensation unit 250 may generate a prediction block by performing motion compensation on the current block using a motion vector and a reference image stored in the reference picture buffer 270.
- the motion compensator 250 may generate a prediction block by applying an interpolation filter to a partial area in the reference image.
- To perform motion compensation based on the coding unit, it can be determined whether the motion compensation method of the prediction unit included in the coding unit is skip mode, merge mode, AMVP mode, or current picture reference mode, and each mode Motion compensation can be performed according to .
- the motion compensation unit 250 applied to the decoding device may use the same technology as the motion compensation unit 122 applied to the above-described encoding device.
- the adder 201 may generate a restored block by adding the restored residual block and the prediction block.
- the filter unit 260 may apply at least one of inverse-LMCS, deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or reconstructed image.
- the filter unit 260 applied to the decoding device may apply the same filtering technology as the filtering technology applied to the filter unit 180 applied to the above-described encoding device.
- the filter unit 260 may output a restored image.
- the reconstructed block or reconstructed image may be stored in the reference picture buffer 270 and used for inter prediction.
- the restored block that has passed through the filter unit 260 may be part of the reference image.
- the reference image may be a reconstructed image composed of reconstructed blocks that have passed through the filter unit 260.
- the stored reference image can then be used for inter-screen prediction or motion compensation.
- Figure 3 is a diagram schematically showing a video coding system to which the present invention can be applied.
- a video coding system may include an encoding device 10 and a decoding device 20.
- the encoding device 10 may transmit encoded video and/or image information or data in file or streaming form to the decoding device 20 through a digital storage medium or network.
- the encoding device 10 may include a video source generator 11, an encoder 12, and a transmitter 13.
- the decoding device 20 may include a receiving unit 21, a decoding unit 22, and a rendering unit 23.
- the encoder 12 may be called a video/image encoder
- the decoder 22 may be called a video/image decoder.
- the transmission unit 13 may be included in the encoding unit 12.
- the receiving unit 21 may be included in the decoding unit 22.
- the rendering unit 23 may include a display unit, and the display unit may be composed of a separate device or external component.
- the video source generator 11 may acquire video/image through a video/image capture, synthesis, or creation process.
- the video source generator 11 may include a video/image capture device and/or a video/image generation device.
- a video/image capture device may include, for example, one or more cameras, a video/image archive containing previously captured video/images, etc.
- Video/image generating devices may include, for example, computers, tablets, and smartphones, and are capable of generating video/images (electronically). For example, a virtual video/image may be created through a computer, etc., and in this case, the video/image capture process may be replaced by the process of generating related data.
- the encoder 12 can encode the input video/image.
- the encoder 12 can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency.
- the encoder 12 may output encoded data (encoded video/image information) in the form of a bitstream.
- the detailed configuration of the encoding unit 12 may be the same as that of the encoding device 100 of FIG. 1 described above.
- the transmission unit 13 may transmit encoded video/image information or data output in the form of a bitstream to the reception unit 21 of the decoding device 20 through a digital storage medium or network in the form of a file or streaming.
- Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
- the transmission unit 13 may include elements for creating a media file through a predetermined file format and may include elements for transmission through a broadcasting/communication network.
- the receiving unit 21 may extract/receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
- the decoder 22 can decode the video/image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoder 12.
- the detailed configuration of the decoding unit 22 may be the same as that of the decoding device 200 of FIG. 2 described above.
- the rendering unit 23 may render the decrypted video/image.
- the rendered video/image may be displayed through the display unit.
- FIG. 4 illustrates an example 400 of a method for determining one of various inter-prediction methods in the inter-prediction mode.
- the inter-prediction mode of the current block is merge mode or AMVP (Advanced Motion Vector Prediction). It is determined whether it is a mode or not.
- AMVP Advanced Motion Vector Prediction
- Merge mode is an inter-screen prediction mode that obtains motion information such as the motion vector and reference picture of the current block from adjacent blocks of the current block.
- AMVP mode a predicted motion vector is obtained from neighboring blocks of the current block, and other motion information, such as a differential motion vector and reference picture information excluding the predicted motion vector, is obtained by parsing the bitstream. Therefore, AMVP mode is different from merge mode in which all motion information of neighboring blocks is used to predict the current block.
- step 420 it is determined whether the inter-screen prediction mode of the current block is the subblock merge mode. If the inter-screen prediction mode of the current block is the sub-block merge mode, in step 422, the current block is divided into a plurality of sub-blocks according to the sub-block merge mode, and each sub-block moves according to the Affine transform. Can be predicted based on vectors.
- step 430 it is determined whether the inter-screen prediction mode of the current block is the regular merge mode. If the inter-prediction mode of the current block is not a general merge mode, in step 432, it is determined whether the inter-prediction mode of the current block is a combined intra inter prediction (CIIP) mode.
- CIIP combined intra inter prediction
- the inter-screen prediction mode of the current block is determined to be a geometric partition mode.
- Geometric partition mode is an inter-screen prediction mode that divides the current block into two partitions based on a predetermined boundary and determines the final prediction block of the current block by combining two prediction blocks for the two partitions.
- the inter-prediction mode of the current block is determined as the inter-prediction mode within the combined screen. According to the combined intra-screen inter-prediction mode, the final prediction block of the current block can be determined by combining a prediction block based on intra-prediction of the current block and a prediction block based on inter-screen prediction of the current block.
- step 440 it is determined whether the inter-screen prediction mode of the current block is Merge mode with Motion Vector Difference (MMVD).
- MMVD Motion Vector Difference
- the final motion vector of the current block is determined by adding the differential motion vector to the motion vector obtained from the neighboring block.
- the direction of the differential motion vector may be limited to one of +x, -x, +y, and -y.
- the size of the differential motion vector may be limited to being selected from a limited number of predetermined size candidates.
- the current block is predicted according to the differential motion vector merge mode. And if the current block is bi-directionally predicted, the prediction block of the current block may be adjusted according to the Bi-Directional Optical Flow (BDOF) mode in step 446.
- BDOF Bi-Directional Optical Flow
- step 444 the general merge mode is applied to the current block. And if the current block is bi-predicted, the motion vector of the current block may be corrected according to the decoding side motion vector correction mode in step 450. And in step 452, like step 446, the prediction block of the current block may be adjusted by the bidirectional optical flow mode.
- Decoding-side motion vector correction is a method of correcting a motion vector through a two-way matching-based motion vector search process without parsing additional information when decoding a two-way motion vector derived from a general merge mode. According to motion vector correction on the decoding side, motion vector accuracy can be improved in general merge mode. And, accordingly, the encoding efficiency of the general merge mode can be improved.
- the decoding-side motion vector correction mode is explained as being applied only in the general merge mode, but depending on the embodiment, the decoding side may also be used in the sub-block merge mode, geometric partition mode, combined intra-screen inter-screen prediction mode, and differential motion vector merge mode.
- a motion vector correction mode may be applied.
- the inter-screen prediction mode of the current block may be determined in steps 410, 420, 430, 432, and 440 of FIG. 4. For example, a merge flag in step 410, a subblock merge flag in step 420, a general merge flag in step 430, a combined intra-picture inter-picture prediction flag in step 432, and a differential motion vector merge flag in step 440 are respectively generated from the bitstream. Can be parsed.
- the bidirectional motion vector predicted in the general merge mode is corrected through motion search based on bilateral matching (BM).
- BM bilateral matching
- neighboring blocks of a reference block may be searched from the initial motion vector to find the optimal reference block.
- a motion vector that minimizes the degree of distortion between two reference blocks located in two reference pictures is searched from the initial motion vector based on bilateral matching (BM).
- Decoding-side motion vector correction is not applied to all bidirectional motion vectors predicted in the general merge mode, but can be applied only to cases where one or more predetermined conditions are all satisfied. The above predetermined conditions are explained below.
- decoding-side motion vector correction when a coding unit (CU) block is in merge mode, decoding-side motion vector correction may be applied. Additionally, when the coding unit block is not in the sub-block merge mode to which inter-screen prediction according to affine transform is applied, motion vector correction on the decoding side may be applied. Additionally, when merge mode with motion vector difference (MMVD) is not applied to the coding unit block, decoding-side motion vector correction may be applied.
- MMVD motion vector difference
- decoding-side motion vector correction when the coding unit block is in bi-prediction mode, decoding-side motion vector correction may be applied. Additionally, according to one embodiment, when two reference pictures referenced by a coding unit block are located in opposite temporal directions from the current picture, decoding-side motion vector correction may be applied. For example, when the first reference picture among two reference pictures temporally precedes the current picture and the second reference picture temporally lags the current picture, decoding-side motion vector correction may be applied.
- decoding-side motion vector correction when the temporal distance between two reference pictures and the current picture is the same, decoding-side motion vector correction can be applied.
- the temporal distance may mean the size of the POC (Picture Order Count) difference between the reference picture and the current picture.
- the distance between pictures represents the size of the temporal distance and POC difference.
- decoding-side motion vector correction may be applied even when the temporal distance between two reference pictures and the current picture is different.
- more motion vector correction on the decoding side can be applied in the general merge mode. Accordingly, by relaxing the performance conditions, more motion vector correction on the decoding side can be performed, thereby improving coding efficiency.
- motion vector correction on the decoding side may be applied.
- whether to apply motion vector correction on the decoding side may be determined depending on the size of the coding unit block. For example, when the size of a coding unit block is larger than a predetermined size, decoding-side motion vector correction may be applied.
- the predetermined size can be expressed as the number of luminance samples included in the coding unit block. And the predetermined size may be a power of 2, such as 64, 128, 256, 512, or 1024.
- whether to apply motion vector correction on the decoding side may be determined depending on the width and height of the coding unit block. For example, when the height and/or width of the coding unit block is greater than a predetermined value, decoding-side motion vector correction may be applied.
- the predetermined value may be a power of 2, such as 4, 8, 16, or 32.
- decoding-side motion vector correction may be applied.
- the final prediction block of the coding unit block is determined as the weighted average of two prediction blocks obtained from bidirectional prediction.
- the bidirectional coding unit weight value is used to determine the weighted average of the two prediction blocks.
- decoding-side motion vector correction may be set to be applied even when the bidirectional coding unit weight values applied to the two prediction blocks are different.
- decoding-side motion vector correction when combined intra inter prediction (CIIP) is not applied, decoding-side motion vector correction may be applied.
- CIIP intra inter prediction
- combined intra-screen inter-prediction it is a prediction method that determines the final prediction block by performing a weighted average of the first prediction block derived from intra-prediction and the second prediction block derived from inter-screen prediction for one block.
- the conditions for performing decoding-side motion vector correction may include at least one of the plurality of conditions described above.
- the frequency of decoding-side motion vector correction may decrease. Conversely, as the conditions for performing decoding-side motion vector correction decrease, the frequency of decoding-side motion vector correction may increase. Therefore, depending on the conditions of motion vector correction on the decoding side, the frequency of motion vector correction and the resulting encoding efficiency of merge mode can be determined.
- Figure 5 shows an example of motion vector correction on the decoding side when the temporal distance between the current picture and the L0 reference picture is the same as the temporal distance between the current picture and the L1 reference picture.
- the current block 500 and the L0 reference picture 520 have a POC distance of N.
- the current block 500 and the L1 reference picture 540 also have a POC distance of N. Accordingly, the distance between the current block 500 and the L0 reference picture 520 and the distance between the current block 500 and the L1 reference picture 540 are the same.
- the current block 500 can obtain initial motion vectors MV0 (522) and MV1 (542) in the L0 and L1 directions based on the motion information of neighboring blocks.
- MV0 (522) and MV1 (542) point to the first reference blocks (524 and 544), respectively.
- MV0 522 and MV1 542 are the initial motion vectors in the L0 and L1 directions respectively derived from the general merge mode.
- MV0' (528) is a motion vector obtained by correcting MV0 (522), the initial motion vector in the L0 direction, by MV diff (526).
- MV1' (548) is a motion vector obtained by correcting MV1 (542), the initial motion vector in the L1 direction, by - MV diff (546).
- MV diff (526) and -MV diff (546) are the L0 differential motion vector and L1 differential motion vector, respectively.
- the L0 differential motion vector and the L1 differential motion vector have the same size, but the directions are set to be opposite.
- MV0 '528' and MV1' (548) may be determined as the final motion vector of the current block 502.
- various distortion measurement methods can be used, such as the sum of absolute difference (SAD) or the sum of squared error (SSE) between two reference blocks.
- a prediction block In bi-prediction with Coding Unit-level weight (BCW), adaptive weighting values are applied to two prediction blocks derived from motion vectors in the L0 direction and L1 direction, so that the prediction block can be created.
- a prediction block In bidirectional prediction using coding unit weight values, a prediction block can be calculated using Equation 1.
- P0 and P1 represent a prediction block motion-compensated from the L0 reference picture and a prediction block motion-compensated from the L1 reference picture, respectively.
- w refers to the weight value applied to each reference block, and the value of w is determined differently depending on whether the current picture is a low delay picture or not. If the current picture is a low-latency picture, five weight values (w ⁇ ⁇ -2, 3, 4, 5, 10 ⁇ ) can be used. If the current picture is not a low-latency picture, three weight values (w ⁇ ⁇ 3, 4, 5 ⁇ ) can be used.
- weight value information can be derived from the merge candidate block. If the current block is not in merge mode, weight value information may be parsed after motion vector difference information is parsed.
- Equation 2 represents a method of generating the final prediction block using the decoding-side motion vector correction method.
- P L0 , P L1 , and Pred Final are a prediction block indicated by the MV0' motion vector in the L0 reference picture, a prediction block indicated by the MV1' motion vector in the L1 reference picture, and a prediction block generated using the two prediction blocks, respectively. Indicates the final prediction block. As shown in Figure 5, since the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture are the same, the same weight value is assigned to the prediction blocks (P L0 and P L1 ) generated from each reference picture. A final prediction block may be generated.
- the method of generating a prediction block according to the decoding-side motion vector correction method according to an embodiment can be used only when the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture are the same.
- a prediction block may be generated by averaging the two prediction blocks generated from the two reference pictures.
- the final prediction block can be generated by assigning different weight values to the prediction block P L0 of L0 and the prediction block P L1 of L1.
- the final prediction block is generated in bi-prediction with CU-level weight (BCW) using the coding unit weight value. Equation 1 for can be used.
- weight value information may be derived from the merge candidate block. Otherwise, weight value information may be parsed from the bitstream when not in merge mode.
- a prediction block may be generated according to the determined weight value. For example, five weight values (w ⁇ ⁇ -2, 3, 4, 5, 10 ⁇ ) used in low-latency pictures can be used to generate a prediction block. Alternatively, three weight values (w ⁇ ⁇ 3, 4, 5 ⁇ ), which are used in cases where the picture is not a low-latency picture, can be used to generate a prediction block. Alternatively, an arbitrary predetermined weight value can be used. At this time, the number of weighting values used may be arbitrarily determined.
- weight value information determined by the encoder is transmitted and parsed.
- the method of transmitting/parsing the weight value information is a fixed length code (FLC), unary code, or truncated binary code, taking into account the weight value used.
- FLC fixed length code
- unary code unary code
- truncated binary code taking into account the weight value used.
- the bidirectional coding unit weight value may be determined based on the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- the decoding-side motion vector correction method is applied to the current block. Correction methods may be applied.
- Figure 6 shows motion vector correction on the decoding side when the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture are not the same.
- the current picture 600 is a picture at time t
- the L0 reference picture 620 and L1 reference picture 640 are pictures at time t-M (t-M > 0) and time t+N (t+N > 0), respectively. is (t-M ⁇ t+N).
- M and N are different (M ⁇ N) arbitrary positive integer values. Therefore, as shown in FIG. 6, the temporal distance between the L0 reference picture 620 and the current picture 600 and the temporal distance between the L1 reference picture 640 and the current picture 600 are not the same.
- the L0 direction motion vector MV0 (622) and the L1 direction motion vector MV1 (642) derived from the general merge mode are initial motion vectors.
- MV0 (622) and MV1 (642) point to the first reference blocks (624 and 644), respectively.
- the motion vector MV0' (628) is determined by correcting MV0 (622) by MV diff_L0 (626).
- the motion vector MV1' (648) is determined by correcting MV1 (642) by MV diff_L1 (646).
- MV diff_L1 (646) can be determined from MV diff_L0 (626) by considering the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- the motion vectors MV0' (628) and MV1' (648) can be derived as final motion vectors in the L0 direction and L1 direction, respectively, when the distortion between the block P L0 (630) and the block P L1 (650) is minimal.
- Equation 3 shows a method of calculating MV diff_L1 corresponding to MV diff_L0 by considering the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- M and N mean the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture, respectively.
- M and N are different random positive integer values.
- MV diff_L1_x , MV diff_L1_y , MV diff_L0_x , and MV diff_L0_y are the x-direction motion information of MV diff_L1 , the y-direction motion information of MV diff_L1 , the x-direction motion information of MV diff_L0 , and the y-direction motion of MV diff_L0 , respectively. Indicates information.
- MV diff_L1 corresponding to MV diff_L0 is calculated symmetrically considering the temporal distance between the current picture and the L0 reference picture and the ratio (M:N) of the temporal distance between the current picture and the L1 reference picture.
- the motion information of MV diff_L1 was calculated by reflecting the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture in the motion information of MV diff_L0 .
- the motion information of MV diff_L0 may be calculated by reflecting the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- MV diff_L1_x and MV diff_L1_y determined in Equation 3 are added to MV diff_L0_x and MV diff_L0_y As determined by multiplying N/M, the values of MV diff_L1_x and MV diff_L1_y may be non-integer values. Therefore, according to one embodiment, MV diff_L1_x and MV diff_L1_y may be adjusted to values in integer units. At this time, MV diff_L1_x and MV diff_L1_y can be integerized according to a rounding or truncation process. Alternatively, MV diff_L1_x and MV diff_L1_y may be adjusted to a predetermined precision other than integer units according to a rounding or truncation process. The predetermined precision may be 1/2, 1/4, etc.
- the absolute values of MV diff_L1_x , MV diff_L1_y , MV diff_L0_x , and MV diff_L0_y may be limited to a predetermined range.
- the predetermined range may be determined in units of 1, 2, 4, or 8 integer samples, etc.
- the predetermined ranges of MV diff_L1_x and MV diff_L1_y may be determined differently from the predetermined ranges of MV diff_L0_x and MV diff_L0_y by considering the ratio of M and N.
- the absolute values of MV diff_L1_x , MV diff_L1_y , MV diff_L0_x , and MV diff_L0_y can be set to be equal to or less than one of 1, 2, 4, or 8.
- motion vector accuracy can be improved by correcting the motion vector by considering the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- a bidirectional coding unit weight value may be determined by considering the temporal distance between the current picture and the L1 reference picture. And, a final prediction block may be generated according to the determined weight value.
- the corrected motion vectors MV0' and MV1' are as shown in Equation 2.
- the final prediction block for the current block may be generated by using the same weight value for the indicated prediction blocks P L0 and P L1 .
- the L0 and L1 direction motion vectors MV0 and MV1 may be determined as the initial motion vectors. And using the motion vector MV0' in which MV0, the motion vector in the L0 direction, is corrected by MV diff_L0 , and the motion vector MV1' in which the motion vector in the L1 direction, MV1, is corrected by MV diff_L1 , prediction blocks P L0 and L1 in the L0 direction are generated.
- the prediction block P L1 of can be determined.
- the final prediction block is calculated using different weight values according to the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- Equation 4 generates each prediction block using the motion vector corrected in the decoder, and considers the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture. Indicates how to generate the final prediction block by assigning weight values.
- M and N mean the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture, respectively.
- M and N are different (M ⁇ N) arbitrary positive integer values.
- the weight value of P L0 (N/(M+N)) and the weight value of P L1 (M/(M+N)) are the temporal distance between the current picture and the L0 reference picture and the current picture and L1 It is calculated by considering the ratio of the temporal distance of the reference picture. According to Equation 4, a larger weight value (N/(M+N)>M/(M+N)) is assigned to the prediction block (P L0 ) of the reference picture that is closer to the current picture.
- a smaller weight value (N/(M+N)>M/(M+N)) is assigned to the prediction block (P L1 ) of the current picture and a farther reference picture.
- the weight value used in Equation 4 is derived by considering the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture, so there is no need to explicitly obtain weight value information from the bitstream.
- weight value information can be parsed from the bitstream even when the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture are not the same.
- the weight value is determined considering the characteristics of the pictures, or weight value information is obtained from the bitstream. Thus, the weight value can be determined.
- a predetermined arbitrary weight value may be used to generate a prediction block. That is, the weighting value used in low-delay pictures may be applied, or the weighting value used in cases where it is not a low-delay picture may be applied. Alternatively, an arbitrary predetermined weight value may be applied. At this time, the number of weight values used may be arbitrarily determined.
- Weight value information determined by the encoder can be parsed from the bitstream.
- the method of transmitting/parsing the weight value information is any efficient method such as a fixed length code (FLC), unary code, or truncated binary code, considering the weight value used. method may be used.
- FLC fixed length code
- unary code unary code
- truncated binary code considering the weight value used. method may be used.
- decoding-side motion vector correction and final prediction block generation may be performed using a template matching method.
- the template matching (TM) method can be applied instead of the bilateral matching (BM) method. Regardless of whether the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture are the same, the template matching method can be applied to the motion vector correction method on the decoding side.
- Figure 7 explains a decoding-side motion vector correction method based on the template matching method.
- the current picture 700 is a picture at time t
- the L0 reference picture 720 and L1 reference picture 740 are pictures at time t-M (t-M > 0) and time t+N (t+N > 0), respectively.
- M and N are different (M ⁇ N) arbitrary positive integer values. Therefore, as shown in Figure 7, the temporal distance between the L0 reference picture 720 and the current picture 700 and the temporal distance between the L1 reference picture 740 and the current picture 700 are not the same as M and N, respectively.
- the final motion vector MV0 is obtained by comparing the current template 704, L0 reference template 732, and L1 reference template 752 based on the template matching method.
- ' (728) and MV1' (748) may be determined.
- a final prediction block may be generated.
- the motion vector MV0 (722) in the L0 direction and the motion vector MV1 (742) in the L1 direction derived from the general merge mode are determined as the initial motion vectors.
- a motion vector that minimizes the degree of distortion of the current template 704, the L0 reference template 732, and the L1 reference template 752 can be searched based on the two initial motion vectors.
- the motion vectors with the least distortion can be determined as the final motion vectors MV0' (728) and MV1' (748).
- the MV diff_L1 can be calculated from MV diff_L0 by considering the ratio of the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- MV diff_L1 can be calculated according to Equation 3 described above. Contrary to Equation 3, MV diff_L0 may be calculated based on the ratio of MV diff_L1 and the temporal distance between the current picture and the L0 reference picture and the temporal distance between the current picture and the L1 reference picture.
- a motion vector that minimizes inter-template distortion is searched for in two reference pictures centered on the initial motion vectors MV0 and MV1.
- the search is performed based on MV diff_L0 and MV diff_L1 , which represent the corrected motion vector of the initial motion vector.
- the MV diff_L0 and MV diff_L1 have opposite signs, and the ratio of the sizes of MV diff_L0 and MV diff_L1 can be determined according to the ratio of the temporal distance between the L0 reference picture and the current picture and the temporal distance between the L1 reference picture and the current picture. there is.
- the corrected motion vectors MV0' and MV1' are determined.
- the reference template (L0 reference template) of the block indicated by MV0' and the reference template (L1 reference template) of the block indicated by MV1' are used to create a final template.
- (final template) is created.
- the degree of distortion between the final generated template and the current template is determined.
- the degree of distortion is calculated for MV0' and MV1' candidates for a certain range of MV diff_L0 and MV diff_L1 , respectively.
- MV0' and MV1', which have the minimum distortion are determined as the final motion vectors.
- any distortion measurement method such as sum of absolute difference (SAD) or sum of squared error (SSE) can be used. Equation 5 and Equation 6 show how to calculate the final template from the L0 reference template and the L1 reference template.
- Template L0 , Template L1 , and Template final represent the L0 reference template, L1 reference template, and final template, respectively.
- the final template is calculated as the average of the templates of the two reference pictures, regardless of the ratio of the temporal distance between the L0 reference picture and the current picture and the temporal distance between the L1 reference picture and the current picture.
- the final template is calculated from the L0 reference template and the L1 reference template according to the ratio of the temporal distance between the L0 reference picture and the current picture and the temporal distance between the L1 reference picture and the current picture.
- the final template can be calculated using any of Equation 5 and Equation 6.
- the template matching-based decoding-side motion vector correction method can be applied to correction of the bidirectional motion vector predicted in general merge mode.
- the motion vectors MV0' (728) and MV1 (728) are obtained by correcting MV0 (722) by MV diff_L0 (726)
- a motion vector MV1' (748) is obtained by correcting 742) by MV diff_L1 (746).
- the L0 reference template 732 of the block indicated by the motion vector MV0' 728 in the L0 reference picture 720, and the L1 reference template 752 of the block indicated by the motion vector MV1' 748 in the L1 reference picture 740. ) is determined. If the distortion between the L0 reference template 732 and the current template 704 is defined as D L0 and the distortion between the L1 reference template 752 and the current template 704 is defined as D L1 , the final prediction block is calculated using Equation 7. It can be.
- the weight value (D L1 /(D L0 +D L1 )) of the prediction block P L0 in the L0 direction is calculated using the distortion D L1 between the L1 reference template and the current template.
- the weight value (D L0 /(D L0 +D L1 )) of the prediction block P L1 in the L1 direction is calculated using the distortion D L0 between the L0 reference template and the current template.
- the difference in these weight values is due to the fact that the smaller the distortion value, the greater the similarity between the current signal and the reference signal, and conversely, the larger the distortion value, the smaller the similarity between the current signal and the reference signal.
- the distortion of the current template and the reference template can be determined based on any distortion measurement method, such as sum of absolute difference (SAD) or sum of squared error (SSE).
- Embodiments of the decoding-side motion vector correction method described above can be combined within a range that does not conflict with each other.
- the distance ratio between two reference pictures and the current picture can be considered.
- the decoding-side motion vector correction method may be performed without considering the distance ratio between the two reference pictures and the current picture.
- a two-way matching method or a template matching method may be used.
- weight value may be determined by transmitting/parsing weight value information from the bitstream.
- the weight value may be determined based on the template matching-based distortion value.
- any of the various methods mentioned above may be used for motion vector correction, and any of the various methods mentioned above may be used to determine the weight value. That is, the motion vector correction method and the weight value determination method can be selected and used independently of each other.
- any combination of the decoding-side motion vector correction method for selecting a bidirectional prediction block in the decoder and the method for determining weight values to generate the final prediction block can be implemented.
- Figure 8 is a flowchart of a decoding-side motion vector correction method according to an embodiment.
- a first basic motion vector of the current block for a first reference picture and a second basic motion vector of the current block for a second reference picture are determined.
- the first reference picture and the second reference picture are located in different temporal directions from the current block.
- a first corrected motion vector is determined by correcting the first basic motion vector by a first differential motion vector
- a second corrected motion vector is determined by correcting the second basic motion vector by the second differential motion vector. is decided.
- the first correction motion vector and the second correction motion A vector can be determined. Or, regardless of the relationship between the distance between the current picture including the current block and the first reference picture and the distance between the current picture and the second reference picture, the first corrected motion vector and the second corrected motion vector are determined. You can.
- the ratio of the size of the first differential motion vector and the size of the second differential motion vector is the distance between the current picture and the first reference picture and the distance between the current picture and the second reference picture. It can be set to be proportional to the ratio.
- the sizes of the first differential motion vector and the second differential motion vector are limited to within a predetermined range.
- the first corrected motion vector and the second corrected motion vector may be determined.
- the first template of the first prediction block indicated by the first correction motion vector, the second template of the second prediction block indicated by the second correction motion vector, The first correction motion vector and the second correction motion vector may be determined so that distortion between the template and the current template of the current block is minimized.
- the distortion may be calculated based on the current template and a final template determined by weighting the first template and the second template.
- the first weight applied to the first template is proportional to the distance between the current picture and the second reference picture
- the first weight value applied to the first template is proportional to the distance between the current picture and the second reference picture.
- the second weight value applied to the template may be set to be proportional to the distance between the current picture and the first reference picture.
- a weighted average value of the first template and the second template may be determined based on the set first and second weight values.
- step 806 a first prediction block and a second prediction block of the current block are determined based on the first correction motion vector and the second correction motion vector.
- a final prediction block for the current block is determined based on the weighted sum of the first prediction block and the second prediction block.
- the weighted sum of the first prediction block and the second prediction block is a weight value determined according to the distance between the current picture and the first reference picture and the distance between the current picture and the second reference picture. It can be decided by .
- the first weight applied to the first prediction block is proportional to the distance between the current picture and the second reference picture
- the second weight applied to the second prediction block is proportional to the distance between the current picture and the second reference picture. and may be set to be proportional to the distance between the first reference picture and the first reference picture.
- the first weight applied to the first prediction block is proportional to the distortion of the second template and the current template
- the second weight applied to the second prediction block is proportional to the distortion of the second template and the current template. and may be determined to be proportional to the distortion of the current template.
- the weighted sum of the first prediction block and the second prediction block may be determined by a weight value determined based on the picture type of the current picture including the current block.
- the weighted sum of the first prediction block and the second prediction block may be determined by a weight value determined based on weight value information generated by parsing a bitstream.
- the current block may be restored based on the final prediction block of the current block restored according to the decoding-side motion vector correction method of FIG. 8.
- the current block may be encoded based on the final prediction block of the current block restored according to the decoding-side motion vector correction method of FIG. 8.
- a computer-readable recording medium storing a bitstream generated by a video encoding method applied to the decoding-side motion vector correction method of FIG. 8 may be provided.
- the bitstream generated by the video encoding method applied to the decoding-side motion vector correction method of FIG. 8 can be stored in a computer-recordable recording medium. Additionally, the bitstream generated by the video encoding method applied to the decoding-side motion vector correction method of FIG. 8 can be transmitted from the video encoding device to the video decoding device.
- a bitstream of video data stored in a computer-recordable recording medium can be decoded by a video decoding method applied to the decoding-side motion vector correction method of FIG. 8. Additionally, the bitstream transmitted from the video encoding device to the video decoding device can be decoded by the video decoding method applied to the decoding-side motion vector correction method of FIG. 8.
- Figure 9 is a diagram illustrating a content streaming system to which an embodiment according to the present invention can be applied.
- a content streaming system to which an embodiment of the present invention is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
- the encoding server compresses content input from multimedia input devices such as smartphones, cameras, CCTV, etc. into digital data, generates a bitstream, and transmits it to the streaming server.
- multimedia input devices such as smartphones, cameras, CCTV, etc. directly generate bitstreams
- the encoding server may be omitted.
- the bitstream may be generated by an image encoding method and/or an image encoding device to which an embodiment of the present invention is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
- the streaming server transmits multimedia data to the user device based on a user request through a web server, and the web server can serve as a medium to inform the user of what services are available.
- the web server delivers it to a streaming server, and the streaming server can transmit multimedia data to the user.
- the content streaming system may include a separate control server, and in this case, the control server may control commands/responses between each device in the content streaming system.
- the streaming server may receive content from a media repository and/or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a certain period of time.
- Examples of the user devices include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, slate PCs, Tablet PC, ultrabook, wearable device (e.g. smartwatch, smart glass, head mounted display), digital TV, desktop There may be computers, digital signage, etc.
- PDAs personal digital assistants
- PMPs portable multimedia players
- navigation slate PCs
- Tablet PC ultrabook
- wearable device e.g. smartwatch, smart glass, head mounted display
- digital TV desktop There may be computers, digital signage, etc.
- Each server in the content streaming system may be operated as a distributed server, and in this case, data received from each server may be distributedly processed.
- an image can be encoded/decoded using at least one or a combination of at least one of the above embodiments.
- the order in which the above embodiments are applied may be different in the encoding device and the decoding device. Alternatively, the order in which the above embodiments are applied may be the same in the encoding device and the decoding device.
- the above embodiments can be performed for each luminance and chrominance signal.
- the above embodiments for luminance and chrominance signals can be performed in the same way.
- the above embodiments may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium.
- the computer-readable recording medium may include program instructions, data files, data structures, etc., singly or in combination.
- Program instructions recorded on the computer-readable recording medium may be specially designed and configured for the present invention, or may be known and usable by those skilled in the computer software field.
- the bitstream generated by the encoding method according to the above embodiment may be stored in a non-transitory computer-readable recording medium. Additionally, the bitstream stored in the non-transitory computer-readable recording medium can be decoded using the decoding method according to the above embodiment.
- examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, and magneto-optical media such as floptical disks. -optical media), and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, flash memory, etc.
- Examples of program instructions include not only machine language code such as that created by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
- the hardware device may be configured to operate as one or more software modules to perform processing according to the invention and vice versa.
- the present invention can be used in devices that encode/decode images and recording media that store bitstreams.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Description
Claims (16)
- 비디오 복호화 방법에 있어서,제1 참조 픽처에 대한 현재 블록의 제1 기초 움직임 벡터 및 제2 참조 픽처에 대한 상기 현재 블록의 제2 기초 움직임 벡터를 결정하는 단계;상기 제1 기초 움직임 벡터를 제1 차분 움직임 벡터만큼 보정함으로써, 제1 보정 움직임 벡터를 결정하고, 상기 제2 기초 움직임 벡터를 제2 차분 움직임 벡터만큼 보정함으로써, 제2 보정 움직임 벡터를 결정하는 단계;상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터에 기초하여, 상기 현재 블록의 제1 예측 블록 및 제2 예측 블록을 결정하는 단계; 및상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합에 기초하여, 상기 현재 블록에 대한 최종 예측 블록을 결정하는 단계를 포함하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터를 결정하는 단계에 있어서,상기 현재 블록이 포함된 현재 픽처와 상기 제1 참조 픽처의 거리와 상기 현재 픽처와 상기 제2 참조 픽처의 거리가 다른 경우, 상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터가 결정되는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합은,상기 현재 픽처와 상기 제1 참조 픽처의 거리와 상기 현재 픽처와 상기 제2 참조 픽처의 거리에 따라 결정된 가중 값에 의하여 결정되는 것을 특징으로 하는 비디오 복호화 방법.
- 제3 항에 있어서,상기 제1 예측 블록에 적용되는 제1 가중 값은 상기 현재 픽처와 상기 제2 참조 픽처의 거리에 비례하고, 상기 제2 예측 블록에 적용되는 제2 가중 값은 상기 현재 픽처와 상기 제1 참조 픽처의 거리에 비례하는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터를 결정하는 단계에 있어서,상기 제1 차분 움직임 벡터의 크기와 상기 제2 차분 움직임 벡터의 크기의 비는 상기 현재 픽처와 상기 제1 참조 픽처의 거리와 상기 현재 픽처와 상기 제2 참조 픽처의 거리의 비에 비례하는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 차분 움직임 벡터 및 상기 제2 차분 움직임 벡터의 크기는 소정의 범위 내로 한정되는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 보정 움직임 벡터가 가리키는 상기 제1 예측 블록과 상기 제2 보정 움직임 벡터가 가리키는 상기 제2 예측 블록 간의 왜곡이 최소가 되도록, 상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터가 결정되는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 보정 움직임 벡터가 가리키는 상기 제1 예측 블록의 제1 템플릿, 상기 제2 보정 움직임 벡터가 가리키는 상기 제2 예측 블록의 제2 템플릿, 및 상기 현재 블록의 현재 템플릿 간의 왜곡이 최소가 되도록, 상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터가 결정되는 것을 특징으로 하는 비디오 복호화 방법.
- 제8 항에 있어서,상기 왜곡은,상기 제1 템플릿과 상기 제2 템플릿을 가중 평균함으로써 결정된 최종 템플릿과 상기 현재 템플릿에 기초하여 계산되는 것을 특징으로 하는 비디오 복호화 방법.
- 제9 항에 있어서,상기 제1 템플릿과 상기 제2 템플릿을 가중 평균함에 있어서,상기 제1 템플릿에 적용되는 제1 가중 값은 상기 현재 픽처와 상기 제2 참조 픽처의 거리에 비례하고, 상기 제2 템플릿에 적용되는 제2 가중 값은 상기 현재 픽처와 상기 제1 참조 픽처의 거리에 비례하는 것을 특징으로 하는 비디오 복호화 방법.
- 제8 항에 있어서,상기 제1 예측 블록에 적용되는 제1 가중 값은 상기 제2 템플릿과 상기 현재 템플릿의 왜곡에 비례하고, 상기 제2 예측 블록에 적용되는 제2 가중 값은 상기 제1 템플릿과 상기 현재 템플릿의 왜곡에 비례하는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합은,상기 현재 블록이 포함된 현재 픽처의 픽처 타입에 기초하여 결정된 가중 값에 의하여 결정되는 것을 특징으로 하는 비디오 복호화 방법.
- 제1 항에 있어서,상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합은,비트스트림을 파싱함으로써 생성된 가중 값 정보에 기초하여 결정된 가중 값에 의하여 결정되는 것을 특징으로 하는 비디오 복호화 방법.
- 비디오 부호화 방법에 있어서,제1 참조 픽처에 대한 현재 블록의 제1 기초 움직임 벡터 및 제2 참조 픽처에 대한 상기 현재 블록의 제2 기초 움직임 벡터를 결정하는 단계;상기 제1 기초 움직임 벡터를 제1 차분 움직임 벡터만큼 보정함으로써, 제1 보정 움직임 벡터를 결정하고, 상기 제2 기초 움직임 벡터를 제2 차분 움직임 벡터만큼 보정함으로써, 제2 보정 움직임 벡터를 결정하는 단계;상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터에 기초하여, 상기 현재 블록의 제1 예측 블록 및 제2 예측 블록을 결정하는 단계; 및상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합에 기초하여, 상기 현재 블록에 대한 최종 예측 블록을 결정하는 단계를 포함하는 비디오 부호화 방법.
- 비디오 부호화 방법에 의하여 생성된 비트스트림을 저장한 컴퓨터로 판독가능한 기록매체에 있어서,상기 비디오 부호화 방법은,제1 참조 픽처에 대한 현재 블록의 제1 기초 움직임 벡터 및 제2 참조 픽처에 대한 상기 현재 블록의 제2 기초 움직임 벡터를 결정하는 단계;상기 제1 기초 움직임 벡터를 제1 차분 움직임 벡터만큼 보정함으로써, 제1 보정 움직임 벡터를 결정하고, 상기 제2 기초 움직임 벡터를 제2 차분 움직임 벡터만큼 보정함으로써, 제2 보정 움직임 벡터를 결정하는 단계;상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터에 기초하여, 상기 현재 블록의 제1 예측 블록 및 제2 예측 블록을 결정하는 단계; 및상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합에 기초하여, 상기 현재 블록에 대한 최종 예측 블록을 결정하는 단계를 포함하는 기록매체.
- 비디오 부호화 방법에 의하여 생성된 비트스트림을 전송하는 비트스트림 전송 방법에 있어서,상기 비디오 부호화 방법에 기초하여 영상을 부호화하는 단계; 및상기 부호화된 영상이 포함된 비트스트림을 전송하는 단계를 포함하고,상기 비디오 부호화 방법은,제1 참조 픽처에 대한 현재 블록의 제1 기초 움직임 벡터 및 제2 참조 픽처에 대한 상기 현재 블록의 제2 기초 움직임 벡터를 결정하는 단계;상기 제1 기초 움직임 벡터를 제1 차분 움직임 벡터만큼 보정함으로써, 제1 보정 움직임 벡터를 결정하고, 상기 제2 기초 움직임 벡터를 제2 차분 움직임 벡터만큼 보정함으로써, 제2 보정 움직임 벡터를 결정하는 단계;상기 제1 보정 움직임 벡터 및 상기 제2 보정 움직임 벡터에 기초하여, 상기 현재 블록의 제1 예측 블록 및 제2 예측 블록을 결정하는 단계; 및상기 제1 예측 블록 및 상기 제2 예측 블록의 가중합에 기초하여, 상기 현재 블록에 대한 최종 예측 블록을 결정하는 단계를 포함하는 비트스트림 전송 방법.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/843,196 US20250184496A1 (en) | 2022-03-11 | 2023-03-06 | Method and apparatus for encoding/decoding image and recording medium for storing bitstream |
| CN202380026795.0A CN118947122A (zh) | 2022-03-11 | 2023-03-06 | 图像编码/解码的方法、设备及存储比特流的记录介质 |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2022-0030592 | 2022-03-11 | ||
| KR20220030592 | 2022-03-11 | ||
| KR10-2023-0028852 | 2023-03-06 | ||
| KR1020230028852A KR20230133775A (ko) | 2022-03-11 | 2023-03-06 | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023172002A1 true WO2023172002A1 (ko) | 2023-09-14 |
Family
ID=87935353
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2023/003020 Ceased WO2023172002A1 (ko) | 2022-03-11 | 2023-03-06 | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 |
Country Status (2)
| Country | Link |
|---|---|
| US (1) | US20250184496A1 (ko) |
| WO (1) | WO2023172002A1 (ko) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20190095348A (ko) * | 2017-01-04 | 2019-08-14 | 삼성전자주식회사 | 비디오 복호화 방법 및 그 장치 및 비디오 부호화 방법 및 그 장치 |
| KR20200027487A (ko) * | 2011-06-21 | 2020-03-12 | 한국전자통신연구원 | 인터 예측 방법 및 그 장치 |
| KR20210153548A (ko) * | 2020-06-10 | 2021-12-17 | 주식회사 케이티 | 비디오 신호 부호화/복호화 방법 및 장치, 그리고 비트스트림을 저장한 기록 매체 |
| KR20220003027A (ko) * | 2019-06-21 | 2022-01-07 | 항조우 힉비젼 디지털 테크놀로지 컴퍼니 리미티드 | 인코딩 및 디코딩 방법, 장치 및 그 기기 |
Family Cites Families (39)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3340620B1 (en) * | 2015-08-23 | 2024-10-02 | LG Electronics Inc. | Inter prediction mode-based image processing method and apparatus therefor |
| WO2018226015A1 (ko) * | 2017-06-09 | 2018-12-13 | 한국전자통신연구원 | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 |
| KR20200098520A (ko) * | 2018-01-08 | 2020-08-20 | 삼성전자주식회사 | 움직임 정보의 부호화 및 복호화 방법, 및 움직임 정보의 부호화 및 복호화 장치 |
| CN116962717A (zh) * | 2018-03-14 | 2023-10-27 | Lx 半导体科技有限公司 | 图像编码/解码方法、存储介质和发送方法 |
| CN119324983A (zh) * | 2018-03-30 | 2025-01-17 | 英迪股份有限公司 | 图像编码/解码方法以及存储介质 |
| US20220038734A1 (en) * | 2018-09-20 | 2022-02-03 | Lg Electronics Inc. | Method and device for processing image signal |
| CN112868238B (zh) * | 2018-10-23 | 2023-04-21 | 北京字节跳动网络技术有限公司 | 局部照明补偿和帧间预测编解码之间的并置 |
| CN112913247B (zh) * | 2018-10-23 | 2023-04-28 | 北京字节跳动网络技术有限公司 | 使用局部照明补偿的视频处理 |
| CN113039780B (zh) * | 2018-11-17 | 2023-07-28 | 北京字节跳动网络技术有限公司 | 视频处理中用运动矢量差的Merge |
| WO2020103852A1 (en) * | 2018-11-20 | 2020-05-28 | Beijing Bytedance Network Technology Co., Ltd. | Difference calculation based on patial position |
| CN113196773B (zh) * | 2018-12-21 | 2024-03-08 | 北京字节跳动网络技术有限公司 | 具有运动矢量差的Merge模式中的运动矢量精度 |
| WO2020141816A1 (ko) * | 2018-12-31 | 2020-07-09 | 한국전자통신연구원 | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 |
| WO2020141911A1 (ko) * | 2019-01-02 | 2020-07-09 | 엘지전자 주식회사 | 화면간 예측을 사용하여 비디오 신호를 처리하기 위한 방법 및 장치 |
| CN113302918B (zh) * | 2019-01-15 | 2025-01-17 | 北京字节跳动网络技术有限公司 | 视频编解码中的加权预测 |
| WO2020147804A1 (en) * | 2019-01-17 | 2020-07-23 | Beijing Bytedance Network Technology Co., Ltd. | Use of virtual candidate prediction and weighted prediction in video processing |
| CN112514384B (zh) * | 2019-01-28 | 2024-12-24 | 苹果公司 | 视频信号编码/解码方法及其装置 |
| WO2020175915A1 (ko) * | 2019-02-26 | 2020-09-03 | 주식회사 엑스리스 | 영상 신호 부호화/복호화 방법 및 이를 위한 장치 |
| KR102733344B1 (ko) * | 2019-03-05 | 2024-11-21 | 엘지전자 주식회사 | 인터 예측을 위한 비디오 신호의 처리 방법 및 장치 |
| KR102635518B1 (ko) * | 2019-03-06 | 2024-02-07 | 베이징 바이트댄스 네트워크 테크놀로지 컴퍼니, 리미티드 | 변환된 단예측 후보의 사용 |
| WO2020180155A1 (ko) * | 2019-03-07 | 2020-09-10 | 엘지전자 주식회사 | 비디오 신호를 처리하기 위한 방법 및 장치 |
| US12010305B2 (en) * | 2019-03-11 | 2024-06-11 | Apple Inc. | Method for encoding/decoding image signal, and device therefor |
| WO2020185022A1 (ko) * | 2019-03-12 | 2020-09-17 | 주식회사 엑스리스 | 영상 신호 부호화/복호화 방법 및 이를 위한 장치 |
| CN119967162A (zh) * | 2019-05-02 | 2025-05-09 | 皇家飞利浦有限公司 | 图像信号编码/解码方法及其装置 |
| CN114009047B (zh) * | 2019-06-23 | 2024-11-08 | Lg电子株式会社 | 视频/图像编译系统中用于合并数据语法的信令方法和装置 |
| SI3989584T1 (sl) * | 2019-06-23 | 2024-10-30 | Lg Electronics Inc. | Postopek in naprava za odstranjevanje redundančne sintakse iz sintakse združenih podatkov |
| EP3989582B1 (en) * | 2019-06-23 | 2024-08-07 | LG Electronics Inc. | Method and device for syntax signaling in video/image coding system |
| US11259016B2 (en) * | 2019-06-30 | 2022-02-22 | Tencent America LLC | Method and apparatus for video coding |
| WO2021025168A1 (en) * | 2019-08-08 | 2021-02-11 | Panasonic Intellectual Property Corporation Of America | System and method for video coding |
| WO2021025167A1 (en) * | 2019-08-08 | 2021-02-11 | Panasonic Intellectual Property Corporation Of America | System and method for video coding |
| US11197030B2 (en) * | 2019-08-08 | 2021-12-07 | Panasonic Intellectual Property Corporation Of America | System and method for video coding |
| CN120302049A (zh) * | 2019-08-08 | 2025-07-11 | 松下电器(美国)知识产权公司 | 用于视频编码的系统和方法 |
| CN120017829A (zh) * | 2019-08-08 | 2025-05-16 | 松下电器(美国)知识产权公司 | 用于视频编码的系统和方法 |
| WO2021045130A1 (en) * | 2019-09-03 | 2021-03-11 | Panasonic Intellectual Property Corporation Of America | System and method for video coding |
| WO2021049593A1 (en) * | 2019-09-11 | 2021-03-18 | Panasonic Intellectual Property Corporation Of America | System and method for video coding |
| CN114450963B (zh) * | 2019-09-18 | 2025-06-20 | 松下电器(美国)知识产权公司 | 用于视频编码的系统和方法 |
| KR20260044264A (ko) * | 2020-02-25 | 2026-04-01 | 엘지전자 주식회사 | 인터 예측에 기반한 영상 인코딩/디코딩 방법 및 장치, 그리고 비트스트림을 저장한 기록 매체 |
| US12219166B2 (en) * | 2021-03-12 | 2025-02-04 | Lemon Inc. | Motion candidate derivation with search order and coding mode |
| US11671616B2 (en) * | 2021-03-12 | 2023-06-06 | Lemon Inc. | Motion candidate derivation |
| US11936899B2 (en) * | 2021-03-12 | 2024-03-19 | Lemon Inc. | Methods and systems for motion candidate derivation |
-
2023
- 2023-03-06 US US18/843,196 patent/US20250184496A1/en active Pending
- 2023-03-06 WO PCT/KR2023/003020 patent/WO2023172002A1/ko not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20200027487A (ko) * | 2011-06-21 | 2020-03-12 | 한국전자통신연구원 | 인터 예측 방법 및 그 장치 |
| KR20190095348A (ko) * | 2017-01-04 | 2019-08-14 | 삼성전자주식회사 | 비디오 복호화 방법 및 그 장치 및 비디오 부호화 방법 및 그 장치 |
| KR20220003027A (ko) * | 2019-06-21 | 2022-01-07 | 항조우 힉비젼 디지털 테크놀로지 컴퍼니 리미티드 | 인코딩 및 디코딩 방법, 장치 및 그 기기 |
| KR20210153548A (ko) * | 2020-06-10 | 2021-12-17 | 주식회사 케이티 | 비디오 신호 부호화/복호화 방법 및 장치, 그리고 비트스트림을 저장한 기록 매체 |
Non-Patent Citations (1)
| Title |
|---|
| Z. LV (VIVO), C. ZHOU (VIVO), J. ZHANG (VIVO): "Non-EE2: Template Matching-based OBMC Design", 25. JVET MEETING; 20220112 - 20220121; TELECONFERENCE; (THE JOINT VIDEO EXPLORATION TEAM OF ISO/IEC JTC1/SC29/WG11 AND ITU-T SG.16 ), 13 January 2022 (2022-01-13), XP030300298 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US20250184496A1 (en) | 2025-06-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2023200206A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2023239147A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024053963A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2023200214A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2023172002A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024043666A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024144118A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2023171988A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2026071463A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024210648A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2025048441A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024147600A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2025018679A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024080849A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024005456A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2026054290A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2026019034A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024210624A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024191219A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024258110A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2025029091A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024191221A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2025048492A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2024191196A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 | |
| WO2026054352A1 (ko) | 영상 부호화/복호화 방법, 장치 및 비트스트림을 저장한 기록 매체 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23767103 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202417064970 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18843196 Country of ref document: US |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202380026795.0 Country of ref document: CN |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23767103 Country of ref document: EP Kind code of ref document: A1 |
|
| WWP | Wipo information: published in national office |
Ref document number: 18843196 Country of ref document: US |


