WO2020132333A1 - On block level bi-prediction with weighted averaging - Google Patents
On block level bi-prediction with weighted averaging Download PDFInfo
- Publication number
- WO2020132333A1 WO2020132333A1 PCT/US2019/067619 US2019067619W WO2020132333A1 WO 2020132333 A1 WO2020132333 A1 WO 2020132333A1 US 2019067619 W US2019067619 W US 2019067619W WO 2020132333 A1 WO2020132333 A1 WO 2020132333A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- prediction
- weight
- coded block
- syntax element
- value
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/157—Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
- H04N19/159—Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/577—Motion compensation with bidirectional frame interpolation, i.e. using B-pictures
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/109—Selection of coding mode or of prediction mode among a plurality of temporal predictive coding modes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/184—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being bits, e.g. of the compressed video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/42—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
Definitions
- the present disclosure generally relates to video processing, and more
- BWA weighted averaging
- Video coding systems are often used to compress digital video signals, for instance to reduce storage space consumed or to reduce transmissi on bandwidth consumption associated with such signals.
- temporal motion prediction is an effective method to increase the coding efficiency and provides high compression.
- the temporal motion prediction may be a single prediction using one reference picture or a bi-prediction using two reference pictures. In some conditions, such as when fading occurs, bi-prediction may not yield the most accurate prediction. To compensate for this, weighted prediction may be used to weigh the two prediction signals differently.
- the different coding tools are not always compatible. For example, it may not be suitable to apply the above-mentioned temporal prediction, bi-prediction, or weighted prediction to the same coding block (e.g., coding unit), or m the same slice or same picture. Therefore, it is desirable to make the different coding tools interact with each other properly.
- Embodiments of the present disclosure relate to methods of coding and signaling weights of weighted-averaging based bi-prediction, at coding-unit (CU) level.
- a computer- implemented video signaling method includes signaling, by a processor to a video encoder, a bitstream including weight information used for prediction of a coding unit (CU).
- the weight information indicates: if weighted prediction is enabled for a bi-prediction mode of the CU, disabling weighted averaging for the bi-prediction mode.
- a computer-implemented video coding method includes constructing, by a processor, a merge candidate list for a coding unit, the merge candidate list including motion information of a non-affine inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-affine inter-coded block.
- the video coding method also includes coding, by the processor, based on the motion information.
- a computer-implemented video signaling method includes determining, by a processor, a value of a bi-prediction weight used for a coding unit (CU) of a video frame.
- the video signaling method also includes determining, by the processor, whether the bi-prediction weight is an equal weight.
- the video signaling method further includes in response to the determination, signaling, by the processor to a video decoder: a bitstream including a first syntax element indicating the equal weight when the bi-prediction weight is an equal weight, or after determining that the bi-prediction weight is an unequal weight, a bitstream including a second syntax element indicating a value of the bi-prediction weight corresponding to the unequal weight.
- a computer-implemented signaling method performed by a decoder includes receiving, by the decoder from a video encoder, a bitstream including weight information used for prediction of a coding unit (CU). The signaling method also includes determining, based on the weight information, that weighted averaging for the bi-prediction mode is disabled if weighted prediction is enabled for a bi-prediction mode of the CU.
- a computer-implemented video coding method performed by a decoder includes receiving, by the decoder, a merge candidate list for a coding unit from an encoder, the merge candidate list including motion information of a non-adjacent inter-coded block of the coding unit.
- the video coding method also includes determining a bi-prediction weight associated with the non-adjacent inter-coded block based on the motion information.
- a computer-implemented signaling method performed by a decoder includes receiving, by the decoder, from a video encoder: a bitstream including a first syntax element corresponding to a bi-prediction weight used for a coding unit (CU) of a video frame, or a bitstream including a second syntax element corresponding to the bi-prediction weight.
- the signaling method also includes in response to receiving the first syntax element, determining, by the processor, the bi-prediction weight is an equal weight.
- the signaling method further includes in response to receiving the first syntax element, determining, by the processor, the bi-prediction weight is an unequal weight, and determining, by the processor based on the second syntax element, a value of the unequal weight.
- aspects of the disclosed embodiments may include non-transitory, tangible computer-readable media that store software instructions that, when executed by one or more processors, are configured for and capable of performing and executing one or more of the methods, operations, and the like consistent with the disclosed embodiments. Also, aspects of the disclosed embodiments may be performed by one or more processors that are configured as special-purpose processor(s) based on software instructions that are programmed with logic and instructions that perform, when executed, one or more operations consistent with the disclosed embodiments.
- FIG. 1 is a schematic diagram illustrating an exemplary video encoding and decoding system, consistent with embodiments of the present disclosure.
- FIG. 2 is a schematic diagram illustrating an exemplary video encoder that may be a part of the exemplary system of FIG. 1, consistent with embodiments of the present disclosure.
- FIG, 3 is a schematic diagram illustrating an exemplary' video decoder that may be a part of the exemplary system of FIG, 1 , consistent with embodiments of the present disclosure.
- FIG. 4 is a table of syntax elements used for weighted prediction (WP), consistent with embodiments of the present disclosure.
- FIG. 5 is a schematic diagram illustrating bi-prediction, consistent with embodiments of the present disclosure.
- FIG. 6 is a table of syntax elements used for bi-prediction with weighted averaging (BWA), consistent with embodiments of the present disclosure.
- FIG. 7 is a schematic diagram illustrating spatial neighbors used in merge candidate list construction, consistent with embodiments of the present disclosure.
- FIG. 8 is a table of syntax elements used for signaling enablement or disablement of WP at picture level, consistent with embodiments of the present disclosure.
- FIG. 9 is a table of syntax elements used for signaling enablement or disablement of WP at slice level, consistent with embodiments of the present disclosure.
- FIG. 10 is a table of syntax elements used for maintaining exclusivity ofWP and BWA at CU level, consistent with embodiments of the present disclosure.
- FIG. 11 is a table of syntax elements used for maintaining exclusivity of WP and BWA at CU level, consistent with embodiments of the present disclosure.
- FIG. 12 is a flowchart of a BWA weight signaling process used for LD picture, consistent with embodiments of the present disclosure.
- FIG. 13 is a flowchart of a BWA weight signaling process used for non-LD picture, consistent with embodiments of the present disclosure.
- FIG. 14 is a block diagram of a video processing apparatus, consistent with embodiments of the present disclosure. DESCRIPTION OF THE EMBODIMENTS
- FIG. 1 is a block diagram illustrating an example video encoding and decoding system 100 that may utilize techniques in compliance w th various video coding standards, such as HEVC/H.265 and WC/H.266.
- system 100 includes a source device 120 that provides encoded video data to be decoded at a later time by a destination device 140.
- each of source device 120 and destination device 140 may include any of a wide range of devices, including a desktop computer, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a mobile phone, a television, a camera, a wearable device (e.g., a smart watch or a wearable camera), a display device, a digital media player, a video gaming console, a video streaming device, or the like.
- Source device 120 and destination device 140 may be equipped for wireless or wired communication
- source device 120 may include a video source 122, a video encoder 124, and an output interface 126.
- Destination device 140 may include an input interface 142, a video decoder 144, and a display device 146.
- a source device and a destination device may include other components or arrangements.
- source device 120 may receive video data from an external video source (not shown), such as an external camera.
- destination device 140 may interface with an external display device, rather than including an integrated display device.
- Source device 120 and destination device 140 are merely examples of such coding devices in which source device 120 generates coded video data for transmission to destination device 140.
- source device 120 and destination device 140 may operate in a substantially symmetrical manner such that each of source device 120 and destination device 140 includes video encoding and decoding components.
- system 100 may support one-way or two-way video transmission between source device 120 and destination device 140, e.g., for video streaming, video playback, video broadcasting, or video telephony.
- Video source 122 of source device 120 may include a video capture device, such as a video camera, a video archive containing previously captured video, or a video feed interface to receive video from a video content provider.
- video source 122 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video.
- the captured, pre-captured, or computer-generated video may be encoded by video encoder 124.
- the encoded video information may then be output by output interface 126 onto a communication medium 160.
- Output interface 126 may include any type of medium or device capable of transmitting the encoded video data from source device 120 to destination device 140.
- output interface 126 may include a transmitter or a transceiver configured to transmit encoded video data from source device 120 directly to destination device 140 m real-time.
- the encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
- Communication medium 160 may include transient media, such as a wireless broadcast or wired network transmission.
- communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., a cable).
- RF radio frequency
- Communication medium 160 may form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet.
- communication medium 160 may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from source device 120 to destination device 140.
- a network server (not shown) may receive encoded video data from source device 120 and provide the encoded video data to destination device 140, e.g., via network transmission.
- Communication medium 160 may also be in the form of a storage media (e.g., non-transitory storage media), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.
- a computing device of a medium production facility such as a disc stamping facility, may receive encoded video data from source device 120 and produce a disc containing the encoded video data.
- Input interface 142 of destination device 140 receives information from communication medium 160.
- the received information may include syntax information including syntax elements that describe characteristics or processing of blocks and other coded units.
- the syntax information is defined by video encoder 124 and used by video decoder 144.
- Display device 146 displays the decoded video data to a user and may include any of a variety of display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
- CTR cathode ray tube
- LCD liquid crystal display
- plasma display an organic light emitting diode
- OLED organic light emitting diode
- the encoded video generated by source device 120 may be stored on a file server or a storage device.
- Input interface 142 may access stored video data from the file server or storage device via streaming or download.
- the file server or storage device may be any type of computing device capable of storing encoded video data and transmitting that encoded video data to destination device 140. Examples of a file server include a web server that supports a website, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive.
- FTP file transfer protocol
- NAS network attached storage
- the transmission of encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
- Video encoder 124 and video decoder 144 each may be implemented as any of a variety of suitable encoder circuitry 7 , such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof.
- DSPs digital signal processors
- ASICs application specific integrated circuits
- FPGAs field programmable gate arrays
- a device may store instructions for the software in a suitable, non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of tins disclosure.
- Each of video encoder 124 and video decoder 144 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder/decoder (CODEC) in a respective device.
- CDEC combined encoder/decoder
- Video encoder 124 and video decoder 144 may operate according to any video coding standard, such as the Versatile Video Coding (VVC/H.266) standard, the High Efficiency Video Coding (HEVC/H.265) standard, the ITU-T H.264 (also known as MPEG-4) standard, etc, although not shown in FIG. 1, in some embodiments, video encoder 124 and video decoder 144 may each be integrated with an audio encoder and decoder, and may include appropriate MUX-DEMUX units, or other hardware and software, to handle encoding of both audio and video in a common data stream or separate data streams.
- VVC/H.266 Versatile Video Coding
- HEVC/H.265 High Efficiency Video Coding
- ITU-T H.264 also known as MPEG-4 standard
- FIG. 2 is a schematic diagram illustrating an exemplary video encoder 200, consistent with the disclosed embodiments.
- video encoder 200 may be used as video encoder 124 in system 100 (FIG. 1).
- Video encoder 200 may perform mtra- or
- Intra-coding may rely on spatial prediction to reduce or remove spatial redundancy in video within a given video frame.
- Inter-coding may rely on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames of a video sequence.
- Intra modes may refer to a number of spatial based compression modes and inter modes (such as uni-prediction or bi -prediction) may refer to a number of temporal-based compression modes.
- input video signal 202 may be processed block by block.
- the video block unit may be a 16x 16 pixel block (e.g., a macroblock (MB)).
- extended block sizes e.g., a coding unit (Cl . ⁇ may be used to compress video signals of resolution, e.g., 1080p and beyond.
- a CU may include up to 64x64 luma samples and corresponding chroma samples.
- the size of a CU may be further increased to include 128x128 luma samples and corresponding chroma samples.
- a CU may be partitioned into prediction units (PUs), for which separate prediction methods may be applied.
- Each input video block (e.g., MB, CU, PIT, etc.) may be processed by using spatial prediction unit 260 or temporal prediction unit 262.
- Spatial prediction unit 260 perforins spatial prediction (e.g., intra prediction) to the current CU using information on the same picture/slice containing the current CU. Spatial prediction may use pixels from the already coded neighboring blocks in the same video picture/slice to predict the current video block. Spatial prediction may reduce spatial redundancy inherent in the video signal. Temporal prediction (e.g., inter prediction or motion compensated prediction) may use samples from the already coded video pictures to predict the current video block. Temporal prediction may reduce temporal redundancy inherent in the video signal.
- spatial prediction e.g., intra prediction
- Spatial prediction may use pixels from the already coded neighboring blocks in the same video picture/slice to predict the current video block. Spatial prediction may reduce spatial redundancy inherent in the video signal.
- Temporal prediction e.g., inter prediction or motion compensated prediction
- Temporal prediction may reduce temporal redundancy inherent in the video signal.
- Temporal prediction unit 262 performs temporal prediction (e.g., inter prediction) to the current CXI using information from picture(s)/slice(s) different from the picture/slice containing the current CU.
- Temporal prediction for a video block may be signaled by one or more motion vectors.
- the motion vectors may indicate the amount and the direction of motion between the current block and one or more of its prediction b!ock(s) in the reference frames. If multiple reference pictures are supported, one or more reference picture indices may be sent for a video block.
- the one or more reference indices may be used to identify from which reference picture(s) in the reference picture store or Decoded Picture Buffer (DPB) 264, the temporal prediction signal may come.
- DPB Decoded Picture Buffer
- the mode decision and encoder control unit 280 in the encoder may choose the prediction mode, for example based on a rate-distortion optimization method.
- the prediction block may be subtracted from the current video block at adder 216.
- the prediction residual may be transformed by transformation unit 204 and quantized by quantization unit 206.
- the quantized residual coefficients may be inverse quantized at inverse quantization unit 210 and inverse transformed at inverse transform unit 212 to form the reconstructed residual.
- the reconstructed block may be added to the prediction block at adder 226 to form the reconstructed video block.
- the in-loop filtering such as deblocking filter and adaptive loop filters 266, may be applied on the reconstructed video block before it is put in the reference picture store 264 and used to code future video blocks.
- coding mode e.g., inter or intra
- prediction mode information e.g., motion information
- quantized residual coefficients may be sent to the entropy coding unit 208 to be compressed and packed to form the bitstream 220.
- the systems and methods, and instrumentalities described herein may be implemented, at least partially, within the temporal prediction unit 262.
- FIG. 3 is a schematic diagram illustrating a video decoder 300, consistent with the disclosed embodiments.
- video decoder 300 may be used as video decoder 144 in system 100 (FIG. 1).
- a video bitstream 302 may be unpacked or entropy decoded at entropy decoding unit 308.
- the coding mode or prediction information may be sent to the spatial prediction unit 360 (e.g., if intra coded) or the temporal prediction unit 362(e.g., if inter coded) to form the prediction block.
- the prediction information may comprise prediction block sizes, one or more motion vectors (e.g., which may indicate direction and amount of motion), or one or more reference indices (e.g., which may indicate from which reference picture the prediction signal is to be obtained).
- Motion compensated prediction may be applied by the temporal prediction unit 362 to form the temporal prediction block.
- the residual transform coefficients may be sent to inverse quantization unit 310 and inverse transform unit 312 to reconstruct the residual block.
- the prediction block and the residual block may be added together at 326.
- the reconstructed block may go through in-loop filtering (via loop filer 366) before it is stored in reference picture store 364.
- the reconstructed video in the reference picture store 364 may be used to drive a display device or used to predict future video blocks.
- Decoded video 320 may be displayed on a display.
- the above-described video encoder and video decoder may use various video coding/decoding tools to process, compress, and decompress video data.
- Weighted prediction is used to provide significantly better temporal prediction when there is fading in the video sequence. Fading refers to the phenomenon when the average illumination levels of the pictures m the video content exhibit noticeable change in the temporal domain, such as fade to white or fade to black. Fading is often used by content creators to create the desired special effect and to express their artistic views. Fading causes the average illumination level of the reference picture and that of the current picture to be significantly different, making it more difficult to obtain an accurate prediction signal from the temporal neighboring pictures. As part of an effort to solve this problem, WP may provide a powerful tool to adjust the illumination level of the prediction signal obtained from the reference picture and match it to that of the current picture, thus significantly improving the temporal prediction accuracy.
- parameters used for WP are signaled for each of the reference pictures used to code the current picture.
- the WP parameters include a pair of weight and offset, (>v, o), which may be signaled for each color component of the reference picture.
- FIG. 4 depicts a table 400 of syntax elements used for WP, according to the disclosed embodiments. Referring to Table 400, the pred _w eight jahleQ syntax is signaled as part of the slice header.
- the flags luma _vr eight Jx Jlagfi/ and chroma _yr eight Jx jlagfi] are signaled to indicate whether weighted prediction is applied to the luma and chroma component of the ?-th reference picture, respectively.
- Equation (1) uses luma as an example to illustrate signaling of WP parameters. Specifically, if the flag luma weight lx flag[i] is 1 for the i- th reference picture m the reference picture list Lx, the WP parameters (wfij, ofij) are signaled for the luma component. Then, when applying temporal prediction using a given reference picture, the following Equation (1) applies:
- WPfx, y) WPfx, y) w P(x, y) + o Equation (1), where: WPfx, y) is the weighted prediction signal at sample location (x, y); (w, o) is the WP parameter pair associated with the reference picture; and P(x, yj re/fx mvx. y mvy) is the prediction before WP is applied, (mvx, mvy ) being the motion vector associated with the reference picture, and reffx, y) being the reference signal at location (x, y). If the motion vector (mvx, mvy) has fractional sample precision, then interpolation may be applied, such as using the 8-tap luma interpolation filter in HEVC.
- bi-prediction may be used to improve temporal prediction accuracy, so as to improve the compression performance of a video encoder. It is used in various video coding standards, such as H.264/A VC, HEVC, and VVC.
- FIG. 5 is a schematic diagram illustrating an exemplary bi-prediction. Referring to FIG.
- CU 503 is part of the current picture.
- CU 501 is from reference picture 511
- CU 502 is from reference picture 512.
- reference pictures 511 and 512 may be selected from two different reference picture lists L0 and LI, respectively.
- Two motion vectors, (mvx 0 , mvy 0 ) and (mvx ⁇ mvy ⁇ , may be generated with reference to CU 501 and CU 502, respectively. These two motion vectors form two prediction signals that may be averaged to obtain the bi-predicted signal, i.e., a prediction corresponding to CU 503.
- reference pictures 511 and 512 may come from the same or different picture sources.
- FIG. 5 depicts that reference pictures 51 1 and 512 are two difference physical reference pictures corresponding to different points in time, in some embodiments reference pictures 511 and 512 may be the same physical reference picture because the same physical reference picture is allowed to appear one or more times in either or both of the reference picture lists L0 and LI.
- FIG. 5 depicts that reference pictures 511 and 512 are from the past and the future in the temporal domain, respectively, in some embodiments, reference pictures 511 and 512 are allowed to be both from the past or both from the future, in relationship to the current picture.
- bi-prediction may be performed based on the following equation:
- ( mvx 0 , mvy 0 ) is a motion vector associated with a reference picture (e.g., reference picture 511) selected from reference picture lists L0:
- (mvx l mvy ) is a motion vector associated with a reference picture (e.g., reference picture 512) selected from reference picture lists Li;
- ref Q (x, y) is a reference signal at location (x, y) m reference picture 511; and refi(x, y) is a reference signal at location (x, y) in reference picture 512
- weighted prediction may be applied to bi-prediction.
- an equal weight of 0.5 is given to each prediction signal, such that the prediction signals are averaged based on the following equation: Equation (3),
- BWA is used to apply unequal weights with weighted averaging to bi-prediction, which may improve coding efficiency.
- BWA may be applied adaptively at the block level.
- a weight index gbijdx is signaled if certain conditions are met.
- FIG. 6 depicts a table 600 of syntax elements used for BWA, according to the disclosed embodiments. Referring to Table 600, a CU containing, for example, at least 256 luma samples may be bi-predicted using the syntax at 601 and 602. Based on the value of the gbijdx , a weight w is determined, and is applied to the reference signals, according to the following:
- the value of the BWA weight w may be selected from five possible values, e.g., w e
- a low-delay (LD) picture is defined as a picture whose reference pictures all precede itself in display order.
- non-low-delay (non-LD) pictures only 3 BWA
- weights w 6 ⁇ - ⁇ are used.
- the value of weight index gbi idx is in the range of
- the value of the BWA weight for the current CU is selected by the encoder, for example, by rate-distortion optimization.
- rate-distortion optimization One method is to try all allowed weight values w and select the one that has the lowest rate distortion cost.
- exhaustive search of optimal combination of weights and motion vectors may significantly increase encoding time. Therefore, fast encoding methods may be applied to reduce encoding time without degrading coding efficiency.
- the BWA weight w may be determined and signaled in one of two ways: 1) for a non-merge CU, the weight index is signaled after the motion vector difference, as shown in Table 600 (FIG. 6); and 2) for a merge CU, the weight index gbi idx is inferred from neighboring blocks based on the merge candidate index.
- Table 600 FIG. 6
- merge mode is explained in detail below.
- the merge candidates of a CU may come from neighboring blocks of the current CU, or the collocated block in the temporal collocated picture of the current CU.
- FIG. 7 is a schematic diagram illustrating spatial neighbors used in merge candidate list construction, according to an exemplary embodiment.
- FIG. 7 depicts the positions of an example of five spatial candidates of motion information.
- the five spatial candidates may be checked and may be added into the list, for example according to the order A1 , B! , B0, A0 and A2. If the block located at a spatial position is intra-coded or outside the boundary of the current slice, it may be considered as unavailable. Redundant entries, for example where candidates have the same motion information, may be excluded from the merge candidate list.
- the merge mode has been supported since the HE VC standard.
- the merge mode is an effective way of reducing motion signaling overhead. Instead of signaling the motion information (prediction mode, motion vectors, reference indices, etc.) of the current CU explicitly, motion information from the neighboring blocks of the current CU is used to construct a merge candidate list. Both spatial and temporal neighboring blocks can be used to construct the merge candidate list. After the merge candidate list is constructed, an index is signaled to indicate which one of the merge candidates is used to code the current CU. The motion information from that merge candidate is then used to predict the current CU.
- the motion information that it inherits from its merge candidate may include not only the motion vectors and reference indices, but also the weight index gbi idx of that merge candidate.
- weighted averaging of the two prediction signals are performed for the current CU according to its neighbor block’s weight index gbi idx.
- the weight index gbi idx is only inherited from the merge candidate if the merge candidate is a spatial neighbor, and is not inherited if the candidate is a temporal neighbor.
- the merge mode in HE, VC constructs merge candidate list using spatial neighboring blocks and temporal neighboring block.
- all spatial neighboring blocks are adjacent (i.e. connected) to the current CU.
- all spatial neighboring blocks are adjacent (i.e. connected) to the current CU.
- non-adjacent neighbors may be used in the merge mode to further increase the coding efficiency of the merge mode.
- a merge mode using non-adjacent neighbors is called extended merge mode.
- HMVP HMVP
- VVC extended merge mode
- a table of HMVP candidates is maintained and updated continuously during the video encoding/decoding process.
- the HDVLVP table may include up to six entries.
- the HMVP candidates are inserted in the middle of the merge candidate list of the spatial neighbors and may be selected using the merge candidate index as other merge candidates to code the current CU.
- a first-in-first-out (FIFO) rule is applied to remove and add entries to the table
- the table is updated by adding the associated motion information as a new HMVP candidate to the last entry of the table and removing the oldest HMVP candidate in the table.
- the table is emptied when a new slice is encountered. In some embodiments, the table may be emptied more frequently, for example, when a new coding tree unit (CTU) is encountered, or when a new row of CTU is encountered.
- CTU coding tree unit
- Equation (4) BWA applies weights m a normalized manner. That is, the weights applied to L0 prediction and LI prediction are (1-vr) and w, respectively. Because the weights add up to 1 , BWA defines how the two prediction signals are combined but does not change the total energy of the bi-prediction signal.
- WP does not have the normalization constraint. That is, w 0 and w do not need to add up to I. Further, WP can add the constant offsets o 0 and o 1 according to Equation (3). Moreover,
- BWA and WP are suitable for different kinds of video content.
- WP is effective in fading video sequences (or other video content with global illumination change in the temporal domain)
- it does not improve coding efficiency for normal sequences when the illumination level does not change in the temporal domain.
- BWA is a block-level adaptive tool that adaptively selects how to combine the two prediction signals.
- BWA is effective on normal sequences without illumination change, it is far less effective on fading sequences than the WP method.
- the BWA tool and the WP tool may be both supported in a video coding standard but work in a mutually exclusive manner. Therefore, a mechanism is needed to disable one tool in the presence of the other.
- the BWA tool may be combined with the merge mode by allowing the weight index gbi idx from the selected merge candidate to be inherited, if the selected merge candidate is a spatial neighbor adjacent to the current CU.
- HMVP weight index
- FIG. 8 is a table 800 of syntax elements used for signaling enablement or disablement of WP at picture level, consistent with embodiments of the present disclosure. As shown at 801 in Table 800, weighted pred fiag and weighted bipred jlag are sent in the PPS to indicate whether WP is enabled for uni-prediction and bi-prediction respectively depending on the slice type of the slices that refer to this PPS.
- FIG. 8 is a table 800 of syntax elements used for signaling enablement or disablement of WP at picture level, consistent with embodiments of the present disclosure.
- weighted pred fiag and weighted bipred jlag are sent in the PPS to indicate whether WP is enabled for uni-prediction and bi-prediction respectively depending on the slice type of the slices that refer to this PPS.
- FIG. 9 is a table 900 of syntax elements used for signaling enablement or disablement of WP at slice level, consistent with embodiments of the present disclosure. As shown at 901 in Table 900, at the slice/picture level, if the PPS that the slice refers to (which is determined by matching the slice j pic j parameter sei d of the slice header with the
- FIG. 4 is sent to a decoder to indicate the WP parameters for each of the reference pictures of the current picture.
- an additional condition may be added in the CU-level weight index gbijdx signaling.
- the additional condition signals that: weighted averaging is disabled for the bi -prediction mode of the current CU, if WP is enabled for the picture containing the current CU.
- FIG. 10 is a table 1000 of syntax elements used for maintaining exclusivity' of WP and BWA at CU level, consistent with embodiments of the present disclosure. Referring to Table 1000, condition 1001 may be added to indicate that: if the PPS that the current slice refers to allows WP for bi-prediction, then BWA is completely disabled for all the CUs in the current slice. This ensures that WP and BW A are exclusive.
- the above method may completely disable BWA for all of the CUs in the current slice, regardless of whether the current CU uses reference pictures for which WP is enabled or not. This may reduce the coding efficiency.
- whether WP is enabled for its reference pictures can be determined by the values of luma weight 10 flag ref idx 10 ],
- FIG. 11 is a table 1100 of syntax elements used for maintaining exclusivity of WP and BWA at CU level, consistent with embodiments of the present disclosure.
- condition 1 101 is added to control the exclusivity of WP and BWA at CU level, regardless of whether weight index gbijdx is signaled or not.
- weight index gbijdx is not signaled, it is inferred to be the default value (i.e., 1 or 2 depending on whether 3 or 5 BWA weights are allowed) that represents the equal-weight case.
- weight index gbi idx values that correspond to unequal weights can only be sent if WP is not enabled for both the luma and chroma components for both the and LI reference pictures.
- this signaling is redundant, because the Context Adaptive Binary Arithmetic Coding (CAB AC) engine in the entropy coding stage can adapt to the statistics of the weight index gbi idx values, the actual bit cost of this redundant signaling may be negligible. Further, this simplifies the parsing process.
- CAB AC Context Adaptive Binary Arithmetic Coding
- a decoder (e.g., encoder 300 in FIG. 3) receives a bitstream including the above- described syntax for maintaining the exclusivity of WP and BWA, the decoder may parse the bitstream and determine, based on the syntax, whether BWA is disabled or not.
- the embodiments of the disclosure can provide a solution for symmetric signaling BWA at CU level.
- the CU level weight used in BWA is signaled as a weight index, gbi idx , with the value of gbi idx being in the range of [0, 4] for low-delay (LD) pictures and in the range of [0, 2] for non-LD pictures.
- LD low-delay
- the same BWA weight value is represented by different gbi idx values in LD and non-LD pictures.
- the signaling of weight index gbi idx may be modified into a first flag indicating if the BWA weight is equal weight, followed by either an index or a flag for non-equal weights.
- FIG. 12 and FIG. 13 illustrate flow-charts of exemplary BWA weight signaling processes used for LD picture and non-LD picture, respectively. For LD pictures that allow 5 BWA weight values, the signaling flow in FIG. 12 is used, and for non-LD pictures that allow' 3 BWA weight values, the signaling flow in FIG. 13 is used.
- the first flag gbi ew flag indicates whether equal weight is applied m BWA.
- FIGs. 12 and 13 only illustrate one example of possible mapping relationship between BWA weight values and the index/flag values. It is contemplated that other mappings between the weight values and index/flag values may be used.
- Another benefit of splitting the weight index gbi idx into two syntax elements, gbi ew flag and gbi new val idx (or gbi new val flag) is that separate CAB AC contexts may be used to code these values. Further, for the LD pictures when a 2-bit value gbi new val idx is used, separate CAB AC contexts may be used to code the first bit and second bit.
- the decoder may parse the signaling and determine, based on the signaling, whether the BWA uses an equal weight. If the BWA is determined to be an unequal weight, the decoder may further determine, based on the signaling, a value of the unequal weight.
- Some embodiments of the present disclosure provide a solution to combine BWA and HMVP. If the motion information stored in the HMVP table only includes motion vectors, reference indices, and prediction modes (e.g. uni-prediction vs. bi-prediction) for the merge candidates, the merge candidates cannot be used with BWA because no BWA weight is stored or updated in the HMVP table. Therefore, according to some disclosed embodiments, B WA weights are included as part of the motion information stored in the HMVP table. When the HMVP table is updated, the BWA weights are also updated together with other motion information, such as motion vectors, reference indices, and prediction modes.
- prediction modes e.g. uni-prediction vs. bi-prediction
- a partial pruning may be applied to avoid having too many identical candidates in the merge candidate list Identical candidates are defined as candidates whose motion information is the same as at least one of the existing merge candidate in the merge candidate list. An identical candidate takes up a space in the merge candidate list but does not provide any additional motion information. Partial pruning detects some of these cases and may prevent some of these identical candidates to be added into the merge candidate list. By including BWA weights in the HMVP table, the pruning process also considers BWA weights in deciding whether two merge candidates are identical.
- the new candidate may be considered to be not identical, and may not he pruned.
- a decoder (e.g., encoder 300 m FIG. 3) receives a bitstream including the above-described HMVP table
- the decoder may parse the bitstream and determine the BWA weights of the merge candidates included in the HMVP table.
- FIG. 14 is a block diagram of a video processing apparatus 1400, consistent with embodiments of the present disclosure.
- apparatus 1400 may embody a video encoder (e.g., video encoder 200 in FIG. 2) or video decoder (e.g., video decoder 300 in FIG. 3) described above.
- apparatus 1400 may be configured to perform the above-described methods for coding and signaling the BWA weights.
- apparatus 1400 may include a processing component 1402, a memory 1404, and an input/output (I/O) interface 1406.
- Apparatus 1400 may also include one or more of a power component and a multimedia component (not shown), or any other suitable hardware or software components.
- Processing component 1402 may control overall operations of apparatus 1400.
- processing component 1402 may include one or more processors that execute instructions to perform the above-described methods for coding and signaling the BWA weights.
- processing component 1402 may include one or more modules that facilitate the interaction between processing component 1402 and other components.
- processing component 1402 may include an I/O module to facilitate the interaction between the I/O interface and processing component 1402
- Memory 1404 is configured to store various types of data or instructions to support the operation of apparatus 1400.
- Memory 1404 may include a non- transitory
- non-transitory media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, cloud storage, a FLASH -EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and networked versions of the same.
- I/O interface 1406 provides an interface between processing component 1402 and peripheral interface modules, such as a camera or a display.
- I/O interface 1406 may employ communication protocols/methods such as audio, analog, digital, serial bus, universal serial bus (USB), infrared, PS/2, BNC, coaxial, RF antennas, Bluetooth, etc.
- I/O interface 1406 may also be configured to facilitate communication, wired or wirelessly, between apparatus 1400 and other devices, such as devices connected to the Internet. Apparatus can access a wireless network based on one or more communication standards, such as WiFi, LTE, 2G, 3G, 4G, 5G, etc.
- the term“or” encompasses all possible combinations, except where infeasible.
- a database may include A or B, then, unless specifically stated otherwise or infeasible, the database may include A, or B, or A and B.
- the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Video encoding and decoding techniques for bi-prediction with weighted averaging are disclosed. According to certain embodiments, a computer- implemented video signaling method includes signaling, by a processor to a video decoder, a bitstream including weight information used for prediction of a coding unit (CU). The weight information indicates: if weighted prediction is enabled for a bi-prediction mode of the CU, disabling weighted averaging for the bi-prediction mode.
Description
ON BLOCK LEVEL BI-PREDICTION WITH WEIGHTED AVERAGING
RELATED APPLICATION
[001 ] This application claims priority to U.S Patent Application No. 16/228,741, filed on December 20, 2018, which is incorporated herein by reference in its entirety.
TECHNICAL FIELD
[002] The present disclosure generally relates to video processing, and more
particularly, to video coding and decoding using bi-prediction with weighted averaging (BWA) at block for coding unit) level.
BACKGROUND
[003] Video coding systems are often used to compress digital video signals, for instance to reduce storage space consumed or to reduce transmissi on bandwidth consumption associated with such signals.
[004] A video coding system may use various tools or techniques to solve different problems. For example, temporal motion prediction is an effective method to increase the coding efficiency and provides high compression. The temporal motion prediction may be a single prediction using one reference picture or a bi-prediction using two reference pictures. In some conditions, such as when fading occurs, bi-prediction may not yield the most accurate prediction. To compensate for this, weighted prediction may be used to weigh the two prediction signals differently.
[005] However, the different coding tools are not always compatible. For example, it may not be suitable to apply the above-mentioned temporal prediction, bi-prediction, or weighted prediction to the same coding block (e.g., coding unit), or m the same slice or same picture.
Therefore, it is desirable to make the different coding tools interact with each other properly.
SUMMARY
[006] Embodiments of the present disclosure relate to methods of coding and signaling weights of weighted-averaging based bi-prediction, at coding-unit (CU) level. In some embodiments, a computer- implemented video signaling method is provided. The video signaling method includes signaling, by a processor to a video encoder, a bitstream including weight information used for prediction of a coding unit (CU). The weight information indicates: if weighted prediction is enabled for a bi-prediction mode of the CU, disabling weighted averaging for the bi-prediction mode.
[007] In some embodiments, a computer-implemented video coding method is provided. The video coding method includes constructing, by a processor, a merge candidate list for a coding unit, the merge candidate list including motion information of a non-affine inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-affine inter-coded block. The video coding method also includes coding, by the processor, based on the motion information.
[008] In some embodiments, a computer-implemented video signaling method is provided. The video signaling method includes determining, by a processor, a value of a bi-prediction weight used for a coding unit (CU) of a video frame. The video signaling method also includes determining, by the processor, whether the bi-prediction weight is an equal weight. The video signaling method further includes in response to the determination, signaling, by the processor to a video decoder: a bitstream including a first syntax element indicating the equal weight when the bi-prediction weight is an equal weight, or after determining that the bi-prediction weight is an unequal weight, a bitstream including a second syntax element
indicating a value of the bi-prediction weight corresponding to the unequal weight.
[009] In some embodiments, a computer-implemented signaling method performed by a decoder is provided. The signaling method includes receiving, by the decoder from a video encoder, a bitstream including weight information used for prediction of a coding unit (CU). The signaling method also includes determining, based on the weight information, that weighted averaging for the bi-prediction mode is disabled if weighted prediction is enabled for a bi-prediction mode of the CU.
[010] In some embodiments, a computer-implemented video coding method performed by a decoder is provided. The video coding method includes receiving, by the decoder, a merge candidate list for a coding unit from an encoder, the merge candidate list including motion information of a non-adjacent inter-coded block of the coding unit. The video coding method also includes determining a bi-prediction weight associated with the non-adjacent inter-coded block based on the motion information.
[011] In some embodiments, a computer-implemented signaling method performed by a decoder is provided. The signaling method includes receiving, by the decoder, from a video encoder: a bitstream including a first syntax element corresponding to a bi-prediction weight used for a coding unit (CU) of a video frame, or a bitstream including a second syntax element corresponding to the bi-prediction weight. The signaling method also includes in response to receiving the first syntax element, determining, by the processor, the bi-prediction weight is an equal weight. The signaling method further includes in response to receiving the first syntax element, determining, by the processor, the bi-prediction weight is an unequal weight, and determining, by the processor based on the second syntax element, a value of the unequal weight.
[012] Aspects of the disclosed embodiments may include non-transitory, tangible
computer-readable media that store software instructions that, when executed by one or more processors, are configured for and capable of performing and executing one or more of the methods, operations, and the like consistent with the disclosed embodiments. Also, aspects of the disclosed embodiments may be performed by one or more processors that are configured as special-purpose processor(s) based on software instructions that are programmed with logic and instructions that perform, when executed, one or more operations consistent with the disclosed embodiments.
[013] Additional objects and advantages of the disclosed embodiments will be set forth in part in the following description, and in part wall be apparent from the description, or may be learned by practice of the embodiments. The objects and advantages of the disclosed
embodiments may be realized and attained by the elements and combinations set forth in the claims.
[014] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.
BRIEF DESCRIPTION OF THE DRAWINGS
[015] FIG. 1 is a schematic diagram illustrating an exemplary video encoding and decoding system, consistent with embodiments of the present disclosure.
[016] FIG. 2 is a schematic diagram illustrating an exemplary video encoder that may be a part of the exemplary system of FIG. 1, consistent with embodiments of the present disclosure.
[017] FIG, 3 is a schematic diagram illustrating an exemplary' video decoder that may be a part of the exemplary system of FIG, 1 , consistent with embodiments of the present
disclosure.
[018] FIG. 4 is a table of syntax elements used for weighted prediction (WP), consistent with embodiments of the present disclosure.
[019] FIG. 5 is a schematic diagram illustrating bi-prediction, consistent with embodiments of the present disclosure.
[020] FIG. 6 is a table of syntax elements used for bi-prediction with weighted averaging (BWA), consistent with embodiments of the present disclosure.
[021] FIG. 7 is a schematic diagram illustrating spatial neighbors used in merge candidate list construction, consistent with embodiments of the present disclosure.
[022] FIG. 8 is a table of syntax elements used for signaling enablement or disablement of WP at picture level, consistent with embodiments of the present disclosure.
[023] FIG. 9 is a table of syntax elements used for signaling enablement or disablement of WP at slice level, consistent with embodiments of the present disclosure.
[024] FIG. 10 is a table of syntax elements used for maintaining exclusivity ofWP and BWA at CU level, consistent with embodiments of the present disclosure.
[025] FIG. 11 is a table of syntax elements used for maintaining exclusivity of WP and BWA at CU level, consistent with embodiments of the present disclosure.
[026] FIG. 12 is a flowchart of a BWA weight signaling process used for LD picture, consistent with embodiments of the present disclosure.
[027] FIG. 13 is a flowchart of a BWA weight signaling process used for non-LD picture, consistent with embodiments of the present disclosure.
[028] FIG. 14 is a block diagram of a video processing apparatus, consistent with embodiments of the present disclosure.
DESCRIPTION OF THE EMBODIMENTS
[029] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers in different drawings represent the same or similar elements unless otherwise represented. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with the invention. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the invention as recited in the appended claims.
[030] FIG. 1 is a block diagram illustrating an example video encoding and decoding system 100 that may utilize techniques in compliance w th various video coding standards, such as HEVC/H.265 and WC/H.266. As shown in FIG, 1, system 100 includes a source device 120 that provides encoded video data to be decoded at a later time by a destination device 140. Consistent with the disclosed embodiments, each of source device 120 and destination device 140 may include any of a wide range of devices, including a desktop computer, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a mobile phone, a television, a camera, a wearable device (e.g., a smart watch or a wearable camera), a display device, a digital media player, a video gaming console, a video streaming device, or the like. Source device 120 and destination device 140 may be equipped for wireless or wired communication
[031 ] Referring to FIG. 1, source device 120 may include a video source 122, a video encoder 124, and an output interface 126. Destination device 140 may include an input interface 142, a video decoder 144, and a display device 146. In other examples, a source device and a destination device may include other components or arrangements. For example, source
device 120 may receive video data from an external video source (not shown), such as an external camera. Likewise, destination device 140 may interface with an external display device, rather than including an integrated display device.
[032] Although in the following description the disclosed techniques are explained as being performed by a video encoding device, the techniques may also be performed by a video encoder/decoder, typically referred to as a“CODEC.” Moreover, the techniques of this disclosure may also be performed by a video preprocessor. Source device 120 and destination device 140 are merely examples of such coding devices in which source device 120 generates coded video data for transmission to destination device 140. In some examples, source device 120 and destination device 140 may operate in a substantially symmetrical manner such that each of source device 120 and destination device 140 includes video encoding and decoding components. Hence, system 100 may support one-way or two-way video transmission between source device 120 and destination device 140, e.g., for video streaming, video playback, video broadcasting, or video telephony.
[033] Video source 122 of source device 120 may include a video capture device, such as a video camera, a video archive containing previously captured video, or a video feed interface to receive video from a video content provider. As a further alternative, video source 122 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. The captured, pre-captured, or computer-generated video may be encoded by video encoder 124. The encoded video information may then be output by output interface 126 onto a communication medium 160.
[034] Output interface 126 may include any type of medium or device capable of transmitting the encoded video data from source device 120 to destination device 140. For
example, output interface 126 may include a transmitter or a transceiver configured to transmit encoded video data from source device 120 directly to destination device 140 m real-time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
[035] Communication medium 160 may include transient media, such as a wireless broadcast or wired network transmission. For example, communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., a cable).
Communication medium 160 may form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. In some embodiments, communication medium 160 may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from source device 120 to destination device 140. For example, a network server (not shown) may receive encoded video data from source device 120 and provide the encoded video data to destination device 140, e.g., via network transmission.
[036] Communication medium 160 may also be in the form of a storage media (e.g., non-transitory storage media), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In some embodiments, a computing device of a medium production facility, such as a disc stamping facility, may receive encoded video data from source device 120 and produce a disc containing the encoded video data.
[037] Input interface 142 of destination device 140 receives information from communication medium 160. The received information may include syntax information including syntax elements that describe characteristics or processing of blocks and other coded
units. The syntax information is defined by video encoder 124 and used by video decoder 144. Display device 146 displays the decoded video data to a user and may include any of a variety of display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[038] In a further example, the encoded video generated by source device 120 may be stored on a file server or a storage device. Input interface 142 may access stored video data from the file server or storage device via streaming or download. The file server or storage device may be any type of computing device capable of storing encoded video data and transmitting that encoded video data to destination device 140. Examples of a file server include a web server that supports a website, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. The transmission of encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[039] Video encoder 124 and video decoder 144 each may be implemented as any of a variety of suitable encoder circuitry7, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. When the techniques are implemented partially in software, a device may store instructions for the software in a suitable, non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of tins disclosure. Each of video encoder 124 and video decoder 144 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder/decoder (CODEC) in a respective device.
[040] Video encoder 124 and video decoder 144 may operate according to any video
coding standard, such as the Versatile Video Coding (VVC/H.266) standard, the High Efficiency Video Coding (HEVC/H.265) standard, the ITU-T H.264 (also known as MPEG-4) standard, etc, Although not shown in FIG. 1, in some embodiments, video encoder 124 and video decoder 144 may each be integrated with an audio encoder and decoder, and may include appropriate MUX-DEMUX units, or other hardware and software, to handle encoding of both audio and video in a common data stream or separate data streams.
[041] FIG. 2 is a schematic diagram illustrating an exemplary video encoder 200, consistent with the disclosed embodiments. For example, video encoder 200 may be used as video encoder 124 in system 100 (FIG. 1). Video encoder 200 may perform mtra- or
inter-coding of blocks within video frames, including video blocks, or partitions or sub-partitions of video blocks. Intra-coding may rely on spatial prediction to reduce or remove spatial redundancy in video within a given video frame. Inter-coding may rely on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames of a video sequence.
Intra modes may refer to a number of spatial based compression modes and inter modes (such as uni-prediction or bi -prediction) may refer to a number of temporal-based compression modes.
[042] Referring to FIG. 2, input video signal 202 may be processed block by block. For example, the video block unit may be a 16x 16 pixel block (e.g., a macroblock (MB)). In HEVC, extended block sizes (e.g., a coding unit (Cl . }} may be used to compress video signals of resolution, e.g., 1080p and beyond. In HEVC, a CU may include up to 64x64 luma samples and corresponding chroma samples. In WC, the size of a CU may be further increased to include 128x128 luma samples and corresponding chroma samples. A CU may be partitioned into prediction units (PUs), for which separate prediction methods may be applied. Each input video block (e.g., MB, CU, PIT, etc.) may be processed by using spatial prediction unit 260 or temporal
prediction unit 262.
[043] Spatial prediction unit 260 perforins spatial prediction (e.g., intra prediction) to the current CU using information on the same picture/slice containing the current CU. Spatial prediction may use pixels from the already coded neighboring blocks in the same video picture/slice to predict the current video block. Spatial prediction may reduce spatial redundancy inherent in the video signal. Temporal prediction (e.g., inter prediction or motion compensated prediction) may use samples from the already coded video pictures to predict the current video block. Temporal prediction may reduce temporal redundancy inherent in the video signal.
[044] Temporal prediction unit 262 performs temporal prediction (e.g., inter prediction) to the current CXI using information from picture(s)/slice(s) different from the picture/slice containing the current CU. Temporal prediction for a video block may be signaled by one or more motion vectors. The motion vectors may indicate the amount and the direction of motion between the current block and one or more of its prediction b!ock(s) in the reference frames. If multiple reference pictures are supported, one or more reference picture indices may be sent for a video block. The one or more reference indices may be used to identify from which reference picture(s) in the reference picture store or Decoded Picture Buffer (DPB) 264, the temporal prediction signal may come. After spatial or temporal prediction, the mode decision and encoder control unit 280 in the encoder may choose the prediction mode, for example based on a rate-distortion optimization method. The prediction block may be subtracted from the current video block at adder 216. The prediction residual may be transformed by transformation unit 204 and quantized by quantization unit 206. The quantized residual coefficients may be inverse quantized at inverse quantization unit 210 and inverse transformed at inverse transform unit 212 to form the reconstructed residual. The reconstructed block may be added to the
prediction block at adder 226 to form the reconstructed video block. The in-loop filtering, such as deblocking filter and adaptive loop filters 266, may be applied on the reconstructed video block before it is put in the reference picture store 264 and used to code future video blocks. To form the output video bitstream 220, coding mode (e.g., inter or intra), prediction mode information, motion information, and quantized residual coefficients may be sent to the entropy coding unit 208 to be compressed and packed to form the bitstream 220. The systems and methods, and instrumentalities described herein may be implemented, at least partially, within the temporal prediction unit 262.
[045] FIG. 3 is a schematic diagram illustrating a video decoder 300, consistent with the disclosed embodiments. For example, video decoder 300 may be used as video decoder 144 in system 100 (FIG. 1). Referring to FIG. 3, a video bitstream 302 may be unpacked or entropy decoded at entropy decoding unit 308. The coding mode or prediction information may be sent to the spatial prediction unit 360 (e.g., if intra coded) or the temporal prediction unit 362(e.g., if inter coded) to form the prediction block. If inter coded, the prediction information may comprise prediction block sizes, one or more motion vectors (e.g., which may indicate direction and amount of motion), or one or more reference indices (e.g., which may indicate from which reference picture the prediction signal is to be obtained).
[046] Motion compensated prediction may be applied by the temporal prediction unit 362 to form the temporal prediction block. The residual transform coefficients may be sent to inverse quantization unit 310 and inverse transform unit 312 to reconstruct the residual block. The prediction block and the residual block may be added together at 326. The reconstructed block may go through in-loop filtering (via loop filer 366) before it is stored in reference picture store 364. The reconstructed video in the reference picture store 364 may be used to drive a
display device or used to predict future video blocks. Decoded video 320 may be displayed on a display.
[047] Consistent with the disclosed embodiments, the above-described video encoder and video decoder may use various video coding/decoding tools to process, compress, and decompress video data. Three tools— weighted prediction (WP), bi-prediction with weighted averaging (BWA), and history-based motion vector prediction (HMVP)— are described below.
[048] Weighted prediction (WP) is used to provide significantly better temporal prediction when there is fading in the video sequence. Fading refers to the phenomenon when the average illumination levels of the pictures m the video content exhibit noticeable change in the temporal domain, such as fade to white or fade to black. Fading is often used by content creators to create the desired special effect and to express their artistic views. Fading causes the average illumination level of the reference picture and that of the current picture to be significantly different, making it more difficult to obtain an accurate prediction signal from the temporal neighboring pictures. As part of an effort to solve this problem, WP may provide a powerful tool to adjust the illumination level of the prediction signal obtained from the reference picture and match it to that of the current picture, thus significantly improving the temporal prediction accuracy.
[049] Consistent with the disclosed embodiments, parameters used for WP (or“WP parameters”) are signaled for each of the reference pictures used to code the current picture. For each reference picture, the WP parameters include a pair of weight and offset, (>v, o), which may be signaled for each color component of the reference picture. FIG. 4 depicts a table 400 of syntax elements used for WP, according to the disclosed embodiments. Referring to Table 400, the pred _w eight jahleQ syntax is signaled as part of the slice header. For the z-th reference
picture in the reference picture list Lx (x can be 0 or 1), the flags luma _vr eight Jx Jlagfi/ and chroma _yr eight Jx jlagfi] are signaled to indicate whether weighted prediction is applied to the luma and chroma component of the ?-th reference picture, respectively.
[050] Without loss of generality, the following description uses luma as an example to illustrate signaling of WP parameters. Specifically, if the flag luma weight lx flag[i] is 1 for the i- th reference picture m the reference picture list Lx, the WP parameters (wfij, ofij) are signaled for the luma component. Then, when applying temporal prediction using a given reference picture, the following Equation (1) applies:
WPfx, y) w P(x, y) + o Equation (1), where: WPfx, y) is the weighted prediction signal at sample location (x, y); (w, o) is the WP parameter pair associated with the reference picture; and P(x, yj re/fx mvx. y mvy) is the prediction before WP is applied, (mvx, mvy ) being the motion vector associated with the reference picture, and reffx, y) being the reference signal at location (x, y). If the motion vector (mvx, mvy) has fractional sample precision, then interpolation may be applied, such as using the 8-tap luma interpolation filter in HEVC.
[051 ] For the bi-prediction with weighted averaging (BWA) tool, bi-prediction may be used to improve temporal prediction accuracy, so as to improve the compression performance of a video encoder. It is used in various video coding standards, such as H.264/A VC, HEVC, and VVC. FIG. 5 is a schematic diagram illustrating an exemplary bi-prediction. Referring to FIG.
5, CU 503 is part of the current picture. CU 501 is from reference picture 511, and CU 502 is from reference picture 512. In some embodiments, reference pictures 511 and 512 may be selected from two different reference picture lists L0 and LI, respectively. Two motion vectors,
(mvx0, mvy0) and (mvx^ mvy^, may be generated with reference to CU 501 and CU 502, respectively. These two motion vectors form two prediction signals that may be averaged to obtain the bi-predicted signal, i.e., a prediction corresponding to CU 503.
[052] In the disclosed embodiments, reference pictures 511 and 512 may come from the same or different picture sources. In particular, although FIG. 5 depicts that reference pictures 51 1 and 512 are two difference physical reference pictures corresponding to different points in time, in some embodiments reference pictures 511 and 512 may be the same physical reference picture because the same physical reference picture is allowed to appear one or more times in either or both of the reference picture lists L0 and LI. Further, although FIG. 5 depicts that reference pictures 511 and 512 are from the past and the future in the temporal domain, respectively, in some embodiments, reference pictures 511 and 512 are allowed to be both from the past or both from the future, in relationship to the current picture.
[053] Specifically, referring to FIG. 5, bi-prediction may be performed based on the following equation:
where: ( mvx0 , mvy0 ) is a motion vector associated with a reference picture (e.g., reference picture 511) selected from reference picture lists L0: (mvxl mvy ) is a motion vector associated with a reference picture (e.g., reference picture 512) selected from reference picture lists Li; and refQ(x, y) is a reference signal at location (x, y) m reference picture 511; and refi(x, y) is a reference signal at location (x, y) in reference picture 512
[054] Still referring to FIG, 5, weighted prediction may be applied to bi-prediction. In some embodiments, an equal weight of 0.5 is given to each prediction signal, such that the prediction signals are averaged based on the following equation:
Equation (3),
512, respectively'.
[055] In some embodiments, BWA is used to apply unequal weights with weighted averaging to bi-prediction, which may improve coding efficiency. BWA may be applied adaptively at the block level. For each CU, a weight index gbijdx is signaled if certain conditions are met. FIG. 6 depicts a table 600 of syntax elements used for BWA, according to the disclosed embodiments. Referring to Table 600, a CU containing, for example, at least 256 luma samples may be bi-predicted using the syntax at 601 and 602. Based on the value of the gbijdx , a weight w is determined, and is applied to the reference signals, according to the following:
P(x, y) (1 - w) P0{x, y) + w P^x, y )
= (1— w) re f(i(x— mvx0, y— mvy0) + iv ref^x— mvx^, y— mvy^)
Equation (4).
[056] In some embodiments, the value of the BWA weight w may be selected from five possible values, e.g., w e
A low-delay (LD) picture is defined as a picture whose reference pictures all precede itself in display order. For LD pictures, all of the above five values may be used for the BWA weight. That is, m the signaling of the BW A weights, the value of weight index gbi idx is in the range of [0, 4] with the center value ( gbi idx = 2) corresponding to the value of equal weight w = ~
. For non-low-delay (non-LD) pictures, only 3 BWA
3 5
[057] If explicit signaling of weight index gbi idx is used, the value of the BWA weight for the current CU is selected by the encoder, for example, by rate-distortion optimization. One method is to try all allowed weight values w and select the one that has the lowest rate distortion cost. However, exhaustive search of optimal combination of weights and motion vectors may significantly increase encoding time. Therefore, fast encoding methods may be applied to reduce encoding time without degrading coding efficiency.
[058] For each bi-predieted CU, the BWA weight w may be determined and signaled in one of two ways: 1) for a non-merge CU, the weight index is signaled after the motion vector difference, as shown in Table 600 (FIG. 6); and 2) for a merge CU, the weight index gbi idx is inferred from neighboring blocks based on the merge candidate index. The merge mode is explained in detail below.
[059] The merge candidates of a CU may come from neighboring blocks of the current CU, or the collocated block in the temporal collocated picture of the current CU. FIG. 7 is a schematic diagram illustrating spatial neighbors used in merge candidate list construction, according to an exemplary embodiment. FIG. 7 depicts the positions of an example of five spatial candidates of motion information. To construct the list of merge candidates, the five spatial candidates may be checked and may be added into the list, for example according to the order A1 , B! , B0, A0 and A2. If the block located at a spatial position is intra-coded or outside the boundary of the current slice, it may be considered as unavailable. Redundant entries, for example where candidates have the same motion information, may be excluded from the merge candidate list.
[060] The merge mode has been supported since the HE VC standard. The merge mode
is an effective way of reducing motion signaling overhead. Instead of signaling the motion information (prediction mode, motion vectors, reference indices, etc.) of the current CU explicitly, motion information from the neighboring blocks of the current CU is used to construct a merge candidate list. Both spatial and temporal neighboring blocks can be used to construct the merge candidate list. After the merge candidate list is constructed, an index is signaled to indicate which one of the merge candidates is used to code the current CU. The motion information from that merge candidate is then used to predict the current CU.
[061] When BWA is enabled, if the current CU is in a merge mode, then the motion information that it inherits from its merge candidate may include not only the motion vectors and reference indices, but also the weight index gbi idx of that merge candidate. In other words, when performing motion compensated prediction, weighted averaging of the two prediction signals are performed for the current CU according to its neighbor block’s weight index gbi idx. In some embodiments, the weight index gbi idx is only inherited from the merge candidate if the merge candidate is a spatial neighbor, and is not inherited if the candidate is a temporal neighbor.
[062] The merge mode in HE, VC constructs merge candidate list using spatial neighboring blocks and temporal neighboring block. In the example shown in FIG. 7, all spatial neighboring blocks are adjacent (i.e. connected) to the current CU. However, in some
embodiments, non-adjacent neighbors may be used in the merge mode to further increase the coding efficiency of the merge mode. A merge mode using non-adjacent neighbors is called extended merge mode. In some embodiments, the History-based Motion Vector Prediction
(HMVP) method in VVC may be used for inter-coding in an extended merge mode, to improve compression performance with minimum implementation cost. In HMVP, a table of HMVP candidates is maintained and updated continuously during the video encoding/decoding process.
The HDVLVP table may include up to six entries. The HMVP candidates are inserted in the middle of the merge candidate list of the spatial neighbors and may be selected using the merge candidate index as other merge candidates to code the current CU.
[063] A first-in-first-out (FIFO) rule is applied to remove and add entries to the table After decoding a non-affine inter-coded block, the table is updated by adding the associated motion information as a new HMVP candidate to the last entry of the table and removing the oldest HMVP candidate in the table. The table is emptied when a new slice is encountered. In some embodiments, the table may be emptied more frequently, for example, when a new coding tree unit (CTU) is encountered, or when a new row of CTU is encountered.
[064] The above description about WP, BWA, and HMVP demonstrate a need to harmonize these tools in video coding and signaling. For example, the BWA tool and the WP tool both introduce weighting factors to the inter prediction process to improve the motion compensated prediction accuracy. However, the BWA tool’s functionality is different from the WP tool. According to Equation (4), BWA applies weights m a normalized manner. That is, the weights applied to L0 prediction and LI prediction are (1-vr) and w, respectively. Because the weights add up to 1 , BWA defines how the two prediction signals are combined but does not change the total energy of the bi-prediction signal. On the other hand, according to Equation (3), WP does not have the normalization constraint. That is, w0 and w do not need to add up to I. Further, WP can add the constant offsets o0 and o1 according to Equation (3). Moreover,
BWA and WP are suitable for different kinds of video content. Whereas WP is effective in fading video sequences (or other video content with global illumination change in the temporal domain), it does not improve coding efficiency for normal sequences when the illumination level does not change in the temporal domain. In contrast, BWA is a block-level adaptive tool that
adaptively selects how to combine the two prediction signals. Though BWA is effective on normal sequences without illumination change, it is far less effective on fading sequences than the WP method. For these reasons, in some embodiments, the BWA tool and the WP tool may be both supported in a video coding standard but work in a mutually exclusive manner. Therefore, a mechanism is needed to disable one tool in the presence of the other.
[065] Moreover, as discussed above, the BWA tool may be combined with the merge mode by allowing the weight index gbi idx from the selected merge candidate to be inherited, if the selected merge candidate is a spatial neighbor adjacent to the current CU. To harness the benefit of HMVP, methods are needed to combine BWA with HMVP to use non-adjacent neighbors in an extended merge mode.
[066] At least some of the disclosed embodiments provide a solution to maintain the exclusivity of WP and BWA. Whether WP is enabled for a picture or not is indicated using a combination of syntax in the Picture Parameter Set (PPS) and the slice header. FIG, 8 is a table 800 of syntax elements used for signaling enablement or disablement of WP at picture level, consistent with embodiments of the present disclosure. As shown at 801 in Table 800, weighted pred fiag and weighted bipred jlag are sent in the PPS to indicate whether WP is enabled for uni-prediction and bi-prediction respectively depending on the slice type of the slices that refer to this PPS. FIG. 9 is a table 900 of syntax elements used for signaling enablement or disablement of WP at slice level, consistent with embodiments of the present disclosure. As shown at 901 in Table 900, at the slice/picture level, if the PPS that the slice refers to (which is determined by matching the slice jpic jparameter sei d of the slice header with the
pps pic parameter _set d of the PPS) enables WP, then pred _w eight table () in Table 400
(FIG. 4) is sent to a decoder to indicate the WP parameters for each of the reference pictures of
the current picture.
[067] Based on such signaling, in some embodiment of this disclosure, an additional condition may be added in the CU-level weight index gbijdx signaling. The additional condition signals that: weighted averaging is disabled for the bi -prediction mode of the current CU, if WP is enabled for the picture containing the current CU. FIG. 10 is a table 1000 of syntax elements used for maintaining exclusivity' of WP and BWA at CU level, consistent with embodiments of the present disclosure. Referring to Table 1000, condition 1001 may be added to indicate that: if the PPS that the current slice refers to allows WP for bi-prediction, then BWA is completely disabled for all the CUs in the current slice. This ensures that WP and BW A are exclusive.
[068] However, the above method may completely disable BWA for all of the CUs in the current slice, regardless of whether the current CU uses reference pictures for which WP is enabled or not. This may reduce the coding efficiency. At the CU level, whether WP is enabled for its reference pictures can be determined by the values of luma weight 10 flag ref idx 10 ],
/ chroma weight 10 flagj ref idx 10 j. luma weight 11 JIagf ref idx 11 /, and
luma/chroma weight 11 flag[ ref idx 11 /, where ref idx 10 and ref idx 11 are the reference picture indices of the current CU in L0 and LI, respectively luma/chroma weight 10 flag and luma/chroma weight // flag are signaled /w weight la ble() for both the L 0 and LI reference pictures for the current slice, as shown m Table 400 (FIG. 4). FIG. 11 is a table 1100 of syntax elements used for maintaining exclusivity of WP and BWA at CU level, consistent with embodiments of the present disclosure. Referring to Table 1100, condition 1 101 is added to control the exclusivity of WP and BWA at CU level, regardless of whether weight index gbijdx is signaled or not. When weight index gbijdx is not signaled, it is inferred to be the default value (i.e., 1 or 2 depending on whether 3 or 5 BWA weights are allowed) that represents the
equal-weight case.
[069] The methods illustrated in Table 1000 and Table 1 100 both add condition(s) to the signaling of weight index gbijdx at the CU level, which could complicate the parsing process at the decoder. Therefore, in a third embodiment, the weight index gbi jdx signaling conditions are kept the same as those m Table 600 (FIG. 6). And it becomes a bitstream conformance constraint for the encoder to always send the default value of weight index gbi idx for the current CU if WP is enabled for either the luma or chroma component of either the
or LI reference picture. That is, weight index gbi idx values that correspond to unequal weights can only be sent if WP is not enabled for both the luma and chroma components for both the
and LI reference pictures. Though this signaling is redundant, because the Context Adaptive Binary Arithmetic Coding (CAB AC) engine in the entropy coding stage can adapt to the statistics of the weight index gbi idx values, the actual bit cost of this redundant signaling may be negligible. Further, this simplifies the parsing process.
[070] After a decoder (e.g., encoder 300 in FIG. 3) receives a bitstream including the above- described syntax for maintaining the exclusivity of WP and BWA, the decoder may parse the bitstream and determine, based on the syntax, whether BWA is disabled or not.
[071] At least some of the embodiments of the disclosure can provide a solution for symmetric signaling BWA at CU level. As discussed above, in some embodiments, the CU level weight used in BWA is signaled as a weight index, gbi idx , with the value of gbi idx being in the range of [0, 4] for low-delay (LD) pictures and in the range of [0, 2] for non-LD pictures. However, this creates an inconsistency between the LD pictures and non-LD pictures, as shown below:
Here, the same BWA weight value is represented by different gbi idx values in LD and non-LD pictures.
[072] In order to improve signaling consistency, according to some disclosed embodiments, the signaling of weight index gbi idx may be modified into a first flag indicating if the BWA weight is equal weight, followed by either an index or a flag for non-equal weights. FIG. 12 and FIG. 13 illustrate flow-charts of exemplary BWA weight signaling processes used for LD picture and non-LD picture, respectively. For LD pictures that allow 5 BWA weight values, the signaling flow in FIG. 12 is used, and for non-LD pictures that allow' 3 BWA weight values, the signaling flow in FIG. 13 is used. The first flag gbi ew flag indicates whether equal weight is applied m BWA. If gbi ew flag is 1, then no further signaling is needed, because equal weight is applied (w = ½); otherwise, a flag (1 bit for 2 values) or an index (2 bits for 4 values) is signaled to indicate which of the unequal weights is applied. FIGs. 12 and 13 only illustrate one example of possible mapping relationship between BWA weight values and the index/flag values. It is contemplated that other mappings between the weight values and index/flag values may be used. Another benefit of splitting the weight index gbi idx into two syntax elements, gbi ew flag and gbi new val idx (or gbi new val flag) is that separate CAB AC contexts may be used to code these values. Further, for the LD pictures when a 2-bit value gbi new val idx is used, separate CAB AC contexts may be used to code the first bit and second bit.
[073] After a decoder (e.g., encoder 300 in FIG. 3) receives the above-described signaling of BWA at CU level, the decoder may parse the signaling and determine, based on the
signaling, whether the BWA uses an equal weight. If the BWA is determined to be an unequal weight, the decoder may further determine, based on the signaling, a value of the unequal weight.
[074] Some embodiments of the present disclosure provide a solution to combine BWA and HMVP. If the motion information stored in the HMVP table only includes motion vectors, reference indices, and prediction modes (e.g. uni-prediction vs. bi-prediction) for the merge candidates, the merge candidates cannot be used with BWA because no BWA weight is stored or updated in the HMVP table. Therefore, according to some disclosed embodiments, B WA weights are included as part of the motion information stored in the HMVP table. When the HMVP table is updated, the BWA weights are also updated together with other motion information, such as motion vectors, reference indices, and prediction modes.
[075] Moreover, a partial pruning may be applied to avoid having too many identical candidates in the merge candidate list Identical candidates are defined as candidates whose motion information is the same as at least one of the existing merge candidate in the merge candidate list. An identical candidate takes up a space in the merge candidate list but does not provide any additional motion information. Partial pruning detects some of these cases and may prevent some of these identical candidates to be added into the merge candidate list. By including BWA weights in the HMVP table, the pruning process also considers BWA weights in deciding whether two merge candidates are identical. Specifically, if a new candidate has identical motion vectors, reference indices, and prediction modes as another candidate in the merge candidate list, but has a different BWA weight from the other candidate, the new candidate may be considered to be not identical, and may not he pruned.
[076] After a decoder (e.g., encoder 300 m FIG. 3) receives a bitstream including the
above-described HMVP table, the decoder may parse the bitstream and determine the BWA weights of the merge candidates included in the HMVP table.
[077] FIG. 14 is a block diagram of a video processing apparatus 1400, consistent with embodiments of the present disclosure. For example, apparatus 1400 may embody a video encoder (e.g., video encoder 200 in FIG. 2) or video decoder (e.g., video decoder 300 in FIG. 3) described above. In the disclosed embodiments, apparatus 1400 may be configured to perform the above-described methods for coding and signaling the BWA weights. Referring to FIG. 14, apparatus 1400 may include a processing component 1402, a memory 1404, and an input/output (I/O) interface 1406. Apparatus 1400 may also include one or more of a power component and a multimedia component (not shown), or any other suitable hardware or software components.
[078] Processing component 1402 may control overall operations of apparatus 1400. For example, processing component 1402 may include one or more processors that execute instructions to perform the above-described methods for coding and signaling the BWA weights. Moreover, processing component 1402 may include one or more modules that facilitate the interaction between processing component 1402 and other components. For instance, processing component 1402 may include an I/O module to facilitate the interaction between the I/O interface and processing component 1402
[079] Memory 1404 is configured to store various types of data or instructions to support the operation of apparatus 1400. Memory 1404 may include a non- transitory
computer-readable storage medium including instructions for applications or methods operated on apparatus 1400, executable by the one or more processors of apparatus 1400. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical
data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, cloud storage, a FLASH -EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and networked versions of the same.
[080] I/O interface 1406 provides an interface between processing component 1402 and peripheral interface modules, such as a camera or a display. I/O interface 1406 may employ communication protocols/methods such as audio, analog, digital, serial bus, universal serial bus (USB), infrared, PS/2, BNC, coaxial, RF antennas, Bluetooth, etc. I/O interface 1406 may also be configured to facilitate communication, wired or wirelessly, between apparatus 1400 and other devices, such as devices connected to the Internet. Apparatus can access a wireless network based on one or more communication standards, such as WiFi, LTE, 2G, 3G, 4G, 5G, etc.
[081] As used herein, unless specifically stated otherwise, the term“or” encompasses all possible combinations, except where infeasible. For example, if it is stated that a database may include A or B, then, unless specifically stated otherwise or infeasible, the database may include A, or B, or A and B. As a second example, if it is stated that a database may include A, B, or C, then, unless specifically stated otherwise or infeasible, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[082] It will be appreciated that the present invention is not limited to the exact construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made without departing from the scope thereof. It is intended that the scope of the invention should only be limited by the appended claims.
Claims
1. A computer-implemented signaling method, comprising:
signaling, by a processor to a video decoder, a bitstream including weight information used for prediction of a coding unit (CU), the weight information indicating:
if weighted prediction is enabled for a bi-prediction mode of the CU, disabling weighted averaging for the bi-prediction mode
2. The method of claim 1 , wherein the weight information indicates:
if weighted prediction is enabled for bi-prediction of a picture including the CU, disabling weighted averaging for the bi-prediction mode
3. The method of claim 2, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for bi-prediction of the picture including the CU.
4. The method of claim 1, wherein the weight information indicates:
if weighted prediction is enabled for at least one of a luma component or a chroma component of a reference picture of the CU, disabling weighted averaging for the bi-prediction mode.
5. The method of claim 4, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for at least one of the luma component or the chroma component of the reference picture.
6. The method of claim 1, wherein the weight information includes a value of a
bi-prediction weight associated with the CU, the method further comprising:
if weighted prediction is enabled for at least one of a luma component or a chroma component of a reference picture of the CU, setting the value of the bi-prediction weight to be a default value.
7. The method of claim 6, wherein the default value corresponds to an equal weight.
8. A computer-implemented video coding method, comprising:
constructing, by a processor, a merge candidate list for a coding unit, the merge candidate list including motion information of a non-adjacent inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-adjacent inter-coded block; and
coding, by the processor, based on the motion information.
9. The method of claim 8, further comprising:
determining, by the processor, a new non-adjacent inter-coded block of the coding unit: and
updating, by the processor, the merge candidate list by inserting a bi-prediction weight associated with the new non-adjacent inter-coded block into the merge candidate list.
10. The method of claim 8, wherein:
determining, by the processor, motion information of a new non-adjacent inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-adjacent inter-coded block:
comparing, by the processor, the motion information of the newr non-adjacent inter-coded block to motion information of each inter-coded block included in the merge candidate list;
in response to the comparison that no inter-coded block included in the merge candidate list has motion information identical to the motion information of the new non-adjacent inter-coded block, adding, by the processor, the new non-adjacent inter-coded block to the merge candidate list; or
in response to the comparison that the motion information of the new non-adjacent inter-coded block is identical to motion information of at least one
inter-coded block included in the merge candidate list, determining that the new non-adjacent inter-coded block is redundant for the merge candidate list.
11. The method of claim 8, wherein the non-adjacent inter-coded block is from a Histoiy based Motion Vector Prediction (HMVP) table.
12. The method of claim 8, wherein the non-adjacent inter-coded block is in a video frame including the coding unit and is a spatial non-adjacent neighbor of the coding unit.
13. The method of claim 8, wherein the motion information of the non-adjacent inter-coded block further includes:
a reference index associated with the non-adjacent inter-coded block, a motion vector associated with the non-affine inter-coded block, and at least one of an indicator of a uni-prediction mode or an indicator of a bi-prediction mode.
14. The method of claim 8, wherein coding based on the motion information comprises:
selecting, from the merge candidate list, the non-adjacent inter-coded block for coding the coding unit; and
signaling, to a decoder, an index indicating the non-adjacent inter-coded block is selected for coding the coding unit.
15. A computer-implemented signaling method, comprising:
determining, by a processor, a value of a bi-prediction weight used for a coding unit (CU) of a v ideo frame;
determining, by the processor, whether the bi-prediction weight is an equal weight; and
in response to the determination, signaling, by the processor to a video decoder: a bitstream including a first syntax element indicating the equal weight when the bi-prediction weight is an equal weight, or
after determining that the bi-prediction weight is an unequal weight, a bitstream including a second syntax element indicating a value of the bi-prediction weight corresponding to the unequal weight.
16. The method of claim 15, wherein the first syntax element is a flag having one bit.
17. The method of claim 15, wherein signaling the second syntax element further comprises:
determining a number of unequal weights usable by the CU; and
in response to the determination that there are more than two unequal weights usable by the CU, signaling, by the processor to the video decoder, the second syntax element as an index having at least two bits; or
in response to the determination that there are one or two unequal weights usable by the CU, signaling, by the processor to the video decoder, the second syntax element as a flag having one bit.
18. The method of claim 17, further comprising:
coding, by the processor, each bit of the second syntax element using a different Context Adaptive Binary' Arithmetic Coding (CAB AC) context.
19. The method of claim 15, further comprising:
determining whether the CU is part of a low?-delay picture;
determining that:
the CU uses more than two unequal weights m response to the determination that the CU is part of the low-delay picture, or
the CU uses one or two unequal weights in response to the determination that the CU is not a part of the low-delay picture.
20. The method of claim 19, wherein:
when the CU is part of a low-delay picture and the value of the bi-prediction weight is 3/8, determining that the value of the second syntax element is 0;
when the CU is part of a lowr-delay picture and the value of the bi-prediction weight is 5/8, determining that the value of the second syntax element is 1 ;
when the CU is not part of a low-delay picture and the value of the bi-prediction weight is 3/8, determining that the value of the second syntax element is 0; and
when the CU is not part of a low-delay picture and the value of the bi-prediction weight is 5/8, determining that the value of the second syntax element is 1.
21. The method of claim 15, further comprising:
assigning, by the processor, different numbers of bits to the second syntax element, for a low-delay picture and a non-lowr-delay picture, respectively.
22. The method of claim 21, wherein a value of the second syntax element corresponds to a same value of bi-prediction weight, for the lowr-delay picture and the non-low?-delay picture.
23. The method of claim 15, wherein each of the first and second syntax elements
corresponds to one or more pre-assigned bits in a CU-level weight index.
24. The method of claim 15, further comprising:
coding, by the processor, a value of the first syntax element and a value of the second syntax element using different Context Adaptive Binary Arithmetic Coding (CABAC) contexts.
25. A device comprising:
a memory storing instructions; and
a processor configured to execute the instructions to cause the device to:
signal, to a video encoder, a bitstream including weight information used for prediction of a coding unit (CU), the weight information indicating:
if weighted prediction is enabled for a bi-prediction mode of the CU, disabling weighted averaging for the bi-prediction mode.
26. The device of claim 25, wherein the weight information indicates:
if weighted prediction is enabled for bi-prediction of a picture including the CU, disabling weighted averaging for the bi-prediction mode.
27. The device of claim 26, wherein the bitstream includes a flag indicating whether
weighted prediction is enabled for bi-prediction of the picture including the CU.
28. The device of claim 25, wherein the weight information indicates:
if weighted prediction is enabled for at least one of a luma component or a chroma component of a reference picture of the CU, disabling weighted averaging for the bi-prediction mode.
29. The device of claim 28, wherein the bitstream includes a flag indicating whether
weighted prediction is enabled for at least one of the luma component or the chroma component of the reference picture.
30. The device of claim 25, wherein the weight information includes a value of a
bi-prediction weight associated with the CU, and the processor is further configured to execute the instructions to:
if weighted prediction is enabled for at least one of a luma component or a chroma component of a reference picture of the CU, set the value of the bi-prediction weight to be a default value.
31. The device of claim 30, wherein the default value corresponds to an equal weight.
32. A device comprising:
a memory storing instructions; and
a processor configured to execute the instructions to cause the device to:
construct a merge candidate list for a coding unit, the merge candidate list including motion information of a non-adjacent inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-adjacent inter-coded block; and
code based on the motion information.
33. The device of claim 32, wherein the processor is further configured to execute the
instructions to:
determine a newr non-adjacent inter-coded block of the coding unit; and update the merge candidate list by inserting a bi-prediction weight associated with the new non-adjacent inter-coded block into the merge candidate list.
34. The device of claim 32, wherein the processor is further configured to execute the
instructions to:
determine motion information of a new non-adjacent inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-adjacent inter-coded block;
compare the motion information of the new non-adjacent mter-coded block to motion information of each mter-coded block included in the merge candidate list;
in response to the comparison that no inter-coded block included in the merge candidate list has motion information identical to the motion information of the new non-adjacent inter-coded block, add the ne non-adjacent inter-coded block to the merge candidate list; and
m response to the comparison that the motion information of the new non-adjacent inter-coded block is identical to motion information of at least one inter-coded block included in the merge candidate list, determine that the newr non-adjacent inter-coded block is redundant for the merge candidate list.
35. The device of claim 32, wherein the non-adjacent inter-coded block is from a History based Motion Vector Prediction (HMVP) table.
36. The device of claim 32, wherein the non-adjacent inter-coded block is in a video frame including the coding unit and is a spatial non-adjacent neighbor of the coding unit.
37. The device of claim 32, wherein the motion information of the non-adjacent inter-coded block further includes:
a reference index associated with the non-adjacent inter-coded block.
a motion vector associated with the non-affine inter-coded block, and at least one of an indicator of a uni-prediction mode or an indicator of a bi-prediction mode.
38. The device of claim 32, wherein the processor is further configured to execute the
instructions to:
select, from the merge candidate list, the non-adjacent rater-coded block for coding the coding unit; and
signal, to a decoder, an index indicating the non-adjacent inter-coded block is selected for coding the coding unit.
39. A device compri sing :
a memory storing instructions; and
a processor configured to execute the instructions to cause the device to:
determine a value of a bi-prediction weight used for a coding unit (CU) of a video frame;
determine whether the bi-prediction weight is an equal weight; and in response to the determination, signal to a video decoder:
a bitstream including a first syntax element indicating the equal weight when the bi-prediction weight is an equal weight, or
after determining that the bi-prediction weight is an unequal weight, a bitstream including a second syntax element indicating a value of the bi-prediction weight corresponding to the unequal weight.
40. The device of claim 39, wherein the first syntax element is a flag having one bit.
41. The device of claim 39, wherein the processor is further configured to execute the
instructions to:
determine a number of unequal weights usable by the CU; and
in response to the determination that there are more than two unequal weights usable by the CU, signal, to the video decoder, the second syntax element as an index having at least two bits; and
in response to the determination that there are one or two unequal weights usable by the CU, signal, to the video decoder, the second syntax element as a flag having one bit.
42. The device of claim 41, wherein the processor is further configured to execute the
instructions to:
code each bit of the second syntax element using a different Context Adaptive
Binary Arithmetic Coding (CABAC) context.
43. The device of claim 39, wherein the processor is further configured to execute the
instructions to:
determine whether the CU is part of a low-delay picture;
determine that:
the CU uses more than two unequal weights in response to the determination that the CU is part of the low-delay picture, and
the CU uses one or two unequal weights in response to the determination that the CU is not a part of the low-delay picture.
44. The device of claim 43, wherein the processor is further configured to execute the
instructions to:
when the CU is part of a lowr-delay picture and the value of the bi-prediction weight is 3/8, determine that the valise of the second syntax element is 0;
when the CU is part of a low-delay picture and the value of the bi-prediction weight is 5/8, determine that the valise of the second syntax element is 1 ;
when the CU is not part of a low-delay picture and the value of the bi-prediction weight is 3/8, determine that the valise of the second syntax element is 0; and
when the CU is not part of a low-delay picture and the value of the bi-prediction weight is 5/8, determine that the value of the second syntax element is 1.
45. The device of claim 39, wherein the processor is further configured to execute the
instructions to:
assigning different numbers of bits to the second syntax element, for a low-delay picture and a non-low?-delay picture, respectively.
46. The device of claim 45, wherein a value of the second syntax element corresponds to a same value of bi-prediction weight, for the low-delay picture and the non-low-delay picture.
47. The device of claim 39, wherein each of the first and second syntax elements corresponds to one or more pre-assigned bits in a CU-level weight index.
48. The device of claim 39, wherein the processor is further configured to execute the
instructions to:
code a value of the first syntax element and a value of the second syntax element using different Context Adaptive Binary Arithmetic Coding (CAB AC) contexts.
49. A non-transitory computer-readable medium storing a set of instructions that is executable by one or more processors of a device to cause the device to perform a method comprising:
signaling, to a video decoder, a bitstream including weight information used for prediction of a coding unit (CU), the weight information indicating:
if weighted prediction is enabled for a bi-prediction mode of the CU, disabling weighted averaging for the bi-prediction mode.
50. The medium of claim 49, wherein the weight information indicates:
if weighted prediction is enabled for bi-prediction of a picture including the CU, disabling weighted averaging for the bi-prediction mode.
51. The medium of claim 50, wherein the bitstream includes a flag indicating whether
weighted prediction is enabled for bi-prediction of the picture including the CU.
52. The medium of claim 49, wherein the weight information indicates:
if weighted prediction is enabled for at least one of a luma component or a chroma component of a reference picture of the CU, disabling weighted averaging for the bi-prediction mode.
53. The medium of claim 52, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for at least one of the luma component or the chroma component of the reference picture.
54. The medium of claim 49, wherein the weight information includes a value of a
bi-prediction weight associated with the CU, and the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
if weighted prediction is enabled for at least one of a luma component or a chroma component of a reference picture of the CU, setting the value of the bi-prediction weight to be a default value.
55. The medium of claim 54, wherein the default value corresponds to an equal weight.
56. A non-transitory computer-readable medium storing a set of instructions that is
executable by one or more processors of a device to cause the device to perform a method comprising:
constructing a merge candidate list for a coding unit, the merge candidate list including motion information of a non- adjacent inter-coded block of the coding unit, the motion information including a bi-prediction weight associated with the non-adjacent inter-coded block; and
coding based on the motion information.
57. The medium of claim 56, further comprising:
determining a new non-adjacent inter-coded block of the coding unit; and updating the merge candidate list by inserting a bi-prediction weight associated with the new non-adjacent inter-coded block into the merge candidate list.
58. The medium of claim 56, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
determining motion information of a new non-adjacent inter-coded block of the coding unit, the motion information including a bi-prediction wreight associated with the non-adjacent inter-coded block;
comparing the motion information of the newr non-adjacent inter-coded block to motion information of each inter-coded block included in the merge candidate list;
in response to the comparison that no inter-coded block included in the merge candidate list has motion information identical to the motion information of the new non-adjacent inter-coded block, adding the new non-adjacent mter-coded block to the merge candidate list; and
in response to the comparison that the motion information of the new non-adjacent mter-coded block is identical to motion information of at least one
inter-coded block included in the merge candidate list, determining that the new non-adjacent inter-coded block is redundant for the merge candidate list.
59. The medium of claim 56, wherein the non-adjacent inter-coded block is from a History based Motion Vector Prediction (HMVP) table.
60. The medium of claim 56, wherein the non-adjacent inter-coded block is in a video frame including the coding unit and is a spatial non-adjacent neighbor of the coding unit.
61. The medium of claim 56, wherein the motion information of the non-adjacent inter-coded block further includes:
a reference index associated with the non-adjacent inter-coded block, a motion vector associated with the non-affine inter-coded block, and at least one of an indicator of a uni-prediction mode or an indicator of a bi-prediction mode.
62. The medium of claim 56, wherein coding based on the motion information comprises:
selecting, from the merge candidate list, the non-adjacent inter-coded block for coding the coding unit; and
signaling, to a decoder, an index indicating the non-adjacent inter-coded block is selected for coding the coding unit.
63. A non-transitory computer-readable medium storing a set of instructions that is executable by one or more processors of a device to cause the device to perform a method comprising:
determining a value of a bi-prediction weight used for a coding unit (CU) of a video frame;
determining whether the bi-prediction weight is an equal weight; and in response to the determination, signaling, to a video decoder:
a bitstream including a first syntax element indicating the equal weight when the bi-prediction weight is an equal weight, or
after determining that the bi-prediction weight is an unequal weight, a bitstream including a second syntax element indicating a value of the bi-prediction weight corresponding to the unequal weight.
64. The medium of claim 63, wherein the first syntax element is a flag having one bit.
65. The medium of claim 63, wherein signaling the second syntax element further comprises:
determining a number of unequal weights usable by the CU; and
in response to the determination that there are more than two unequal weights usable by the CU, signaling, to the video decoder, the second syntax element as an index having at least two bits; and
in response to the determination that there are one or two unequal weights usable by the CU, signaling, to the video decoder, the second syntax element as a flag having one bit.
66. The medium of claim 65, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
coding each bit of the second syntax element using a different Context Adaptive Binary Arithmetic Coding (CAB AC) context.
67. The medium of claim 63, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
determining whether the CU is part of a low-delay picture;
determining that:
the CU uses more than two unequal weights in response to the
determination that the CU is part of the low-delay picture, and
the CU uses one or two unequal weights in response to the determination that the CU is not a part of the lo w-delay picture.
68. The medium of claim 67, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
when the CU is part of a low-delay picture and the value of the bi-prediction weight is 3/8, determining that the value of the second syntax element is 0;
when the CU is part of a low-delay picture and the value of the bi-prediction weight is 5/8, determining that the value of the second syntax element is 1 ;
when the CU is not part of a low-delay picture and the value of the bi-prediction weight is 3/8, determining that the value of the second syntax element is 0; and
when the CU is not part of a low-delay picture and the value of the bi-prediction weight is 5/8, determining that the value of the second syntax element is 1.
69. The medium of claim 63, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
assigning different numbers of bits to the second syntax element, for a low-delay picture and a non-low-delay picture, respectively.
70. The medium of claim 69, wherein a value of the second syntax element corresponds to a same value of bi-prediction weight, for the low-delay picture and the non-low-delay picture.
71. The medium of claim 63, wherein each of the first and second syntax elements corresponds to one or more pre-assigned bits in a CU-levei weight index.
72. The medium of claim 63, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:
coding a value of the first syntax element and a value of the second syntax element using different Context Adaptive Binary Arithmetic Coding (CABAC) contexts.
73. A computer- implemented signaling method performed by a decoder, the method
comprising:
receiving, by the decoder from a video encoder, a bitstream including weight information used for prediction of a coding unit (CU);
determining, based on the weight information, that weighted averaging for the bi-prediction mode is disabled if weighted prediction is enabled for a bi-prediction mode of the CU.
74. A computer-implemented video coding method performed by a decoder, the method comprising:
receiving, by the decoder, a merge candidate list for a coding unit from an encoder, the merge candidate list including motion information of a non-adjacent inter-coded block of the coding unit; and
determining a bi-prediction weight associated with the non-adjacent inter-coded block based on the motion information.
75. A computer- implemented signaling method performed by a decoder, the method
comprising:
receiving, by the decoder, from a video encoder:
a bitstream including a first syntax element corresponding to a
bi-prediction weight used for a coding unit (CU) of a video frame, or
a bitstream including a second syntax element corresponding to the bi-prediction weight;
in response to receiving the first syntax element, determining, by the processor, the bi-prediction weight is an equal weight; and
in response to receiving the first syntax element, determining, by the processor, the bi-prediction weight is an unequal weight, and determining, by the processor based on the second syntax element, a value of the unequal weight.
Priority Applications (6)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310311779.8A CN116347097B (en) | 2018-12-20 | 2019-12-19 | Video decoding method, video encoder, and computer-readable storage medium |
| CN202511566501.0A CN121262381A (en) | 2018-12-20 | 2019-12-19 | Signal transmission method, video encoder and computer readable storage medium |
| CN202310315808.8A CN116347098B (en) | 2018-12-20 | 2019-12-19 | Signal transmission method, video encoder and computer-readable storage medium |
| CN202310310153.5A CN116347096A (en) | 2018-12-20 | 2019-12-19 | Video decoding method, video encoding method, and computer-readable storage medium |
| CN201980085317.0A CN113228672B (en) | 2018-12-20 | 2019-12-19 | Block-Level Bidirectional Prediction with Weighted Averaging |
| CN202511566474.7A CN121239864A (en) | 2018-12-20 | 2019-12-19 | Signal transmission method, video encoder and computer readable storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/228,741 | 2018-12-20 | ||
| US16/228,741 US10855992B2 (en) | 2018-12-20 | 2018-12-20 | On block level bi-prediction with weighted averaging |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020132333A1 true WO2020132333A1 (en) | 2020-06-25 |
Family
ID=71097931
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2019/067619 Ceased WO2020132333A1 (en) | 2018-12-20 | 2019-12-19 | On block level bi-prediction with weighted averaging |
Country Status (4)
| Country | Link |
|---|---|
| US (5) | US10855992B2 (en) |
| CN (6) | CN121239864A (en) |
| TW (2) | TWI907873B (en) |
| WO (1) | WO2020132333A1 (en) |
Families Citing this family (22)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112913241B (en) | 2018-10-22 | 2024-03-26 | 北京字节跳动网络技术有限公司 | Restriction of decoder-side motion vector derivation |
| WO2020084461A1 (en) | 2018-10-22 | 2020-04-30 | Beijing Bytedance Network Technology Co., Ltd. | Restrictions on decoder side motion vector derivation based on coding information |
| CN111083489B (en) | 2018-10-22 | 2024-05-14 | 北京字节跳动网络技术有限公司 | Multiple iteration motion vector refinement |
| CN111083484B (en) | 2018-10-22 | 2024-06-28 | 北京字节跳动网络技术有限公司 | Sub-block based prediction |
| CN112868238B (en) | 2018-10-23 | 2023-04-21 | 北京字节跳动网络技术有限公司 | Juxtaposition between local illumination compensation and inter-prediction codec |
| CN112913247B (en) | 2018-10-23 | 2023-04-28 | 北京字节跳动网络技术有限公司 | Video processing using local illumination compensation |
| CN111436227B (en) | 2018-11-12 | 2024-03-29 | 北京字节跳动网络技术有限公司 | Use of combined inter-intra prediction in video processing |
| WO2020103877A1 (en) | 2018-11-20 | 2020-05-28 | Beijing Bytedance Network Technology Co., Ltd. | Coding and decoding of video coding modes |
| WO2020103852A1 (en) | 2018-11-20 | 2020-05-28 | Beijing Bytedance Network Technology Co., Ltd. | Difference calculation based on patial position |
| US11876957B2 (en) * | 2018-12-18 | 2024-01-16 | Lg Electronics Inc. | Method and apparatus for processing video data |
| KR20210118068A (en) * | 2018-12-29 | 2021-09-29 | 브이아이디 스케일, 인크. | History-based motion vector prediction |
| CN112042191B (en) * | 2019-01-01 | 2024-03-19 | Lg电子株式会社 | Method and apparatus for predicting and processing video signals based on history-based motion vectors |
| CN113302918B (en) * | 2019-01-15 | 2025-01-17 | 北京字节跳动网络技术有限公司 | Weighted Prediction in Video Codecs |
| WO2020147804A1 (en) | 2019-01-17 | 2020-07-23 | Beijing Bytedance Network Technology Co., Ltd. | Use of virtual candidate prediction and weighted prediction in video processing |
| KR102635518B1 (en) | 2019-03-06 | 2024-02-07 | 베이징 바이트댄스 네트워크 테크놀로지 컴퍼니, 리미티드 | Use of converted single prediction candidates |
| US12149730B2 (en) * | 2019-03-11 | 2024-11-19 | Telefonaktiebolaget Lm Ericsson (Publ) | Motion refinement and weighted prediction |
| EP4243417A3 (en) * | 2019-03-11 | 2023-11-15 | Alibaba Group Holding Limited | Method, device, and system for determining prediction weight for merge mode |
| US12088839B2 (en) * | 2019-10-06 | 2024-09-10 | Hyundai Motor Company | Method and apparatus for encoding and decoding video using inter-prediction |
| JP2023011955A (en) * | 2019-12-03 | 2023-01-25 | シャープ株式会社 | Dynamic image coding device and dynamic image decoding device |
| JP7475908B2 (en) * | 2020-03-17 | 2024-04-30 | シャープ株式会社 | Prediction image generating device, video decoding device, and video encoding device |
| US12081736B2 (en) * | 2021-04-26 | 2024-09-03 | Tencent America LLC | Bi-prediction without signaling cu-level weights |
| WO2023025178A1 (en) * | 2021-08-24 | 2023-03-02 | Beijing Bytedance Network Technology Co., Ltd. | Method, apparatus, and medium for video processing |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140072041A1 (en) * | 2012-09-07 | 2014-03-13 | Qualcomm Incorporated | Weighted prediction mode for scalable video coding |
| US20150103898A1 (en) * | 2012-04-09 | 2015-04-16 | Vid Scale, Inc. | Weighted prediction parameter signaling for video coding |
| WO2017197146A1 (en) * | 2016-05-13 | 2017-11-16 | Vid Scale, Inc. | Systems and methods for generalized multi-hypothesis prediction for video coding |
| US20180295385A1 (en) * | 2015-06-10 | 2018-10-11 | Samsung Electronics Co., Ltd. | Method and apparatus for encoding or decoding image using syntax signaling for adaptive weight prediction |
Family Cites Families (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| GB2469679B (en) * | 2009-04-23 | 2012-05-02 | Imagination Tech Ltd | Object tracking using momentum and acceleration vectors in a motion estimation system |
| US11445172B2 (en) * | 2012-01-31 | 2022-09-13 | Vid Scale, Inc. | Reference picture set (RPS) signaling for scalable high efficiency video coding (HEVC) |
| US9503720B2 (en) * | 2012-03-16 | 2016-11-22 | Qualcomm Incorporated | Motion vector coding and bi-prediction in HEVC and its extensions |
| US9479778B2 (en) * | 2012-08-13 | 2016-10-25 | Qualcomm Incorporated | Device and method for coding video information using base layer motion vector candidate |
| US20140198846A1 (en) * | 2013-01-16 | 2014-07-17 | Qualcomm Incorporated | Device and method for scalable coding of video information |
| US10999574B2 (en) * | 2013-11-05 | 2021-05-04 | Arris Enterprises Llc | Simplified processing of weighted prediction syntax and semantics using a bit depth variable for high precision data |
| JPWO2015194669A1 (en) * | 2014-06-19 | 2017-04-20 | シャープ株式会社 | Image decoding apparatus, image encoding apparatus, and predicted image generation apparatus |
| CN105493505B (en) * | 2014-06-19 | 2019-08-06 | 微软技术许可有限责任公司 | Unified Intra Block Copy and Inter Prediction Modes |
| US9854237B2 (en) * | 2014-10-14 | 2017-12-26 | Qualcomm Incorporated | AMVP and merge candidate list derivation for intra BC and inter prediction unification |
| US10306229B2 (en) * | 2015-01-26 | 2019-05-28 | Qualcomm Incorporated | Enhanced multiple transforms for prediction residual |
| US10404992B2 (en) * | 2015-07-27 | 2019-09-03 | Qualcomm Incorporated | Methods and systems of restricting bi-prediction in video coding |
| WO2017205703A1 (en) * | 2016-05-25 | 2017-11-30 | Arris Enterprises Llc | Improved weighted angular prediction coding for intra coding |
| CN109479149B (en) * | 2016-07-05 | 2022-04-15 | 株式会社Kt | Video signal processing method and apparatus |
| US10750203B2 (en) * | 2016-12-22 | 2020-08-18 | Mediatek Inc. | Method and apparatus of adaptive bi-prediction for video coding |
| WO2018128232A1 (en) * | 2017-01-03 | 2018-07-12 | 엘지전자 주식회사 | Image decoding method and apparatus in image coding system |
| US10715827B2 (en) * | 2017-01-06 | 2020-07-14 | Mediatek Inc. | Multi-hypotheses merge mode |
| US20180332298A1 (en) * | 2017-05-10 | 2018-11-15 | Futurewei Technologies, Inc. | Bidirectional Prediction In Video Compression |
-
2018
- 2018-12-20 US US16/228,741 patent/US10855992B2/en active Active
-
2019
- 2019-08-28 TW TW112145065A patent/TWI907873B/en active
- 2019-08-28 TW TW108130764A patent/TWI847998B/en active
- 2019-12-19 CN CN202511566474.7A patent/CN121239864A/en active Pending
- 2019-12-19 WO PCT/US2019/067619 patent/WO2020132333A1/en not_active Ceased
- 2019-12-19 CN CN201980085317.0A patent/CN113228672B/en active Active
- 2019-12-19 CN CN202310310153.5A patent/CN116347096A/en active Pending
- 2019-12-19 CN CN202511566501.0A patent/CN121262381A/en active Pending
- 2019-12-19 CN CN202310315808.8A patent/CN116347098B/en active Active
- 2019-12-19 CN CN202310311779.8A patent/CN116347097B/en active Active
-
2020
- 2020-11-12 US US17/095,839 patent/US11425397B2/en active Active
-
2022
- 2022-08-22 US US17/821,267 patent/US12101487B2/en active Active
-
2023
- 2023-11-20 US US18/514,291 patent/US12401798B2/en active Active
-
2025
- 2025-03-07 US US19/073,993 patent/US20250211760A1/en active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20150103898A1 (en) * | 2012-04-09 | 2015-04-16 | Vid Scale, Inc. | Weighted prediction parameter signaling for video coding |
| US20140072041A1 (en) * | 2012-09-07 | 2014-03-13 | Qualcomm Incorporated | Weighted prediction mode for scalable video coding |
| US20180295385A1 (en) * | 2015-06-10 | 2018-10-11 | Samsung Electronics Co., Ltd. | Method and apparatus for encoding or decoding image using syntax signaling for adaptive weight prediction |
| WO2017197146A1 (en) * | 2016-05-13 | 2017-11-16 | Vid Scale, Inc. | Systems and methods for generalized multi-hypothesis prediction for video coding |
Also Published As
| Publication number | Publication date |
|---|---|
| TWI847998B (en) | 2024-07-11 |
| TWI907873B (en) | 2025-12-11 |
| US20210067787A1 (en) | 2021-03-04 |
| US20250211760A1 (en) | 2025-06-26 |
| CN116347097B (en) | 2025-08-12 |
| US12401798B2 (en) | 2025-08-26 |
| TW202415074A (en) | 2024-04-01 |
| US10855992B2 (en) | 2020-12-01 |
| US11425397B2 (en) | 2022-08-23 |
| CN121262381A (en) | 2026-01-02 |
| US12101487B2 (en) | 2024-09-24 |
| US20240098276A1 (en) | 2024-03-21 |
| CN116347096A (en) | 2023-06-27 |
| CN116347098B (en) | 2026-02-10 |
| CN116347098A (en) | 2023-06-27 |
| US20220400268A1 (en) | 2022-12-15 |
| CN116347097A (en) | 2023-06-27 |
| CN113228672B (en) | 2023-02-17 |
| CN113228672A (en) | 2021-08-06 |
| TW202025778A (en) | 2020-07-01 |
| CN121239864A (en) | 2025-12-30 |
| US20200204807A1 (en) | 2020-06-25 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12401798B2 (en) | On block level bi-prediction with weighted averaging | |
| US11218694B2 (en) | Adaptive multiple transform coding | |
| CN112840658B (en) | Inter-frame prediction method and device | |
| CN109996081B (en) | Image prediction method, device and codec | |
| KR102310752B1 (en) | Slice level intra block copy and other video coding improvements | |
| JP2022548650A (en) | History-based motion vector prediction | |
| JP2022523350A (en) | Methods, devices and systems for determining predictive weighting for merge modes | |
| EP3520418A1 (en) | Memory and bandwidth reduction of stored data in image/video coding | |
| KR20230150284A (en) | Efficient video encoder architecture | |
| JP2022511873A (en) | Video picture decoding and encoding methods and equipment | |
| KR20210104904A (en) | Video encoders, video decoders, and corresponding methods | |
| CN113615191A (en) | Method and device for determining image display sequence and video coding and decoding equipment | |
| HK40090447A (en) | Video decoding method, video encoding method and computer-readable storage medium | |
| HK40090446A (en) | Video decoding method, video encoder and computer-readable storage medium | |
| HK40090446B (en) | Video decoding method, video encoder and computer-readable storage medium | |
| HK40090448A (en) | Signal transmission method, video encoder and computer-readable storage medium | |
| TW202610298A (en) | On block level bi-prediction with weighted averaging | |
| TW202610297A (en) | On block level bi-prediction with weighted averaging |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19898870 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19898870 Country of ref document: EP Kind code of ref document: A1 |
