WO2025149263A1 - Encoding and decoding methods using virtual intra prediction modes and corresponding apparatuses - Google Patents
Encoding and decoding methods using virtual intra prediction modes and corresponding apparatusesInfo
- Publication number
- WO2025149263A1 WO2025149263A1 PCT/EP2024/085294 EP2024085294W WO2025149263A1 WO 2025149263 A1 WO2025149263 A1 WO 2025149263A1 EP 2024085294 W EP2024085294 W EP 2024085294W WO 2025149263 A1 WO2025149263 A1 WO 2025149263A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- intra prediction
- current block
- prediction mode
- predictor
- block
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/129—Scanning of coding units, e.g. zig-zag scan of transform coefficients or flexible macroblock ordering [FMO]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/593—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques
Definitions
- At least one of the present embodiments generally relates to a method and an apparatus for encoding and decoding a block (e.g., a picture block) using virtual intra prediction modes.
- image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content.
- intra or inter prediction is used to exploit the intra or inter picture correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded.
- the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction.
- a virtual intra prediction mode may be derived for an IBC (Intra Block Copy) coded block.
- the VIMP for the IBC coded block may be obtained by applying a decoder-side intra mode derivation (DIMD)-like process to a prediction block or a residual block, e.g. a block identified by a block vector associated with the current block, of the IBC coded block.
- DIMD decoder-side intra mode derivation
- VIPMs of non-intra coded neighboring blocks of a current block could be used as the derived intra prediction mode candidate for deriving (e.g., filling) a generic MPM (Most Probable Mode) list for the current block.
- a VIPM of the current block could be used as the derived intra prediction mode for generating the intra predicted samples which may be used by prediction process, e.g., by GPM (Geometric Partitioning Mode)-Intra, IBC-GPM, CIIP (Combined Inter and Intra Prediction), IBC-CIIP.
- GPM Gaussian Partitioning Mode
- IBC-GPM IBC-GPM
- CIIP Combined Inter and Intra Prediction
- VIPMs of non-intra coded blocks neighboring could be used as the derived intra prediction modes for generating the intra predicted samples which may be used by prediction process, e.g., by GPM-Intra, SGPM (Spatial geometric partitioning mode), IBC-GPM, CIIP, IBC-CIIP.
- the scanning order of coefficient groups and/or coefficients may be selected based on VIPM for non-intra coded block.
- FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented
- FIG. 2 illustrates a block diagram of an embodiment of a video encoder
- FIG. 3 illustrates a block diagram of an embodiment of a video decoder
- FIG.4 illustrates the principles of a decoder-side intra mode derivation (DIMD)-like process to applied to an inter prediction block ;
- FIG. 5 depicts locations of neighboring blocks surrounding a current block
- FIG. 6 illustrates partitioning of a current block (e.g., a CU) split into partitions by a geometrically located straight line;
- a current block e.g., a CU
- FIG. 7 depicts possible intra prediction mode (IPM) candidates for a current block split into an inter region and an intra region ;
- FIG. 8 illustrates various scanning of coefficients within coefficient groups (CGs) ;
- FIG. 9 depicts a Look-up table mapping a virtual intra prediction mode (VIPM) mode to a LFNST index
- FIG. 10 depicts a flowchart of a method for deriving a VIPM mode for an IBC Intra Block Copy) coded block according to an example
- FIG. 11 depicts a flowchart of a method for decoding a current block, e.g., an intra coded block ;
- FIG. 12 depicts the above neighboring inter-coded block of a current block ;
- FIG. 13 depicts a flowchart of a method for decoding a current block according to an example.
- FIG. 14 depicts a flowchart of a method for decoding a current block according to an example.
- FIGs. 1, 2 and 3 below provide some embodiments, but other embodiments are contemplated and the discussion of FIGs. 1, 2 and 3 does not limit the breadth of the implementations.
- At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded.
- These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium (e.g. a non-transitory computer readable storage medium) having stored thereon instructions for encoding or decoding video data according to any of the methods described, and/or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
- the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” or “coded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably and the terms “image,” “picture” and “frame” may be used interchangeably.
- the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side.
- each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
- satisfying, failing to satisfy a condition and configuring condition parameter(s) are described throughout embodiments described herein as relative to a threshold (e.g., greater, or lower than), a (e.g., threshold) value, configuring the (e.g., threshold) value, etc.).
- a condition may be described as being above a (e.g., threshold) value
- failing to satisfy a condition e.g., performance criteria
- a condition e.g., performance criteria
- Embodiments described herein are not limited to thresholdbased conditions. Any kind of other condition and parameter(s) (such as e.g., belonging or not belonging to a range of values) may be applicable to embodiments described herein.
- VVC VVC
- HEVC High Efficiency Video Coding
- present aspects are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
- FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented.
- System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers.
- Elements of system 100 singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and/or discrete components.
- the processing and encoder/decoder elements of system 100 are distributed across multiple ICs and/or discrete components.
- System 100 includes a storage device 140, which may include non-volatile memory and/or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and/or optical disk drive.
- the storage device 140 may include an internal storage device, an attached storage device, and/or a network accessible storage device, as non-limiting examples.
- System 100 includes an encoder/decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder/decoder module 130 may include its own processor and memory.
- the encoder/decoder module 130 represents module(s) that may be included in a device to perform the encoding and/or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder/decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
- Program code to be loaded onto processor 110 or encoder/decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110.
- one or more of processor 110, memory 120, storage device 140, and encoder/decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
- memory inside of the processor 110 and/or the encoder/decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding.
- a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder/decoder module 130) is used for one or more of these functions.
- the external memory may be the memory 120 and/or the storage device 140, for example, a dynamic volatile memory and/or a non-volatile flash memory.
- an external non-volatile flash memory is used to store the operating system of a television.
- a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2, (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO/IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).
- MPEG refers to the Moving Picture Experts Group
- MPEG-2 is also referred to as ISO/IEC 13818
- 13818-1 is also known as H.222
- 13818-2 is also known as H.262
- HEVC High Efficiency Video Coding
- VVC Very Video Coding
- the input to the elements of system 100 may be provided through various input devices as indicated in block 105.
- Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and/or (iv) a High Definition Multimedia Interface (HDMI) input terminal.
- RF radio frequency
- COMP Component
- USB Universal Serial Bus
- HDMI High Definition Multimedia Interface
- the input devices of block 105 have associated respective input processing elements as known in the art.
- the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets.
- the RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers.
- the RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband.
- the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band.
- Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter.
- the RF portion includes an antenna.
- USB and/or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and/or HDMI connections.
- various aspects of input processing for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary.
- aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary.
- the demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder/decoder 130 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
- connection arrangement 115 for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
- the system 100 includes communication interface 150 that enables communication with other devices via communication channel 190.
- the communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190.
- the communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and/or a wireless medium.
- Wi-Fi Wireless Fidelity
- IEEE 802.11 IEEE refers to the Institute of Electrical and Electronics Engineers
- the Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications.
- the communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications.
- Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105.
- Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.
- various embodiments provide data in a non-streaming manner.
- various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
- control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention.
- the output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150.
- the display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television.
- the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
- the display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box.
- the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
- the embodiments can be carried out by computer software implemented by the processor 110 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits.
- the memory 120 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples.
- the processor 110 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
- the video sequence may go through pre-encoding processing (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one of the color components).
- Metadata can be associated with the preprocessing and attached to the bitstream.
- a picture is encoded by the encoder elements as described below.
- the picture to be encoded is partitioned (202) and processed in units of, for example, CTU (Coding Tree Units), CUs (Coding Units).
- Each unit is encoded using, for example, either an intra or inter mode.
- a unit may be encoded using intra block copy (IBC).
- IBC intra block copy
- a unit When a unit is encoded in an intra mode, it performs intra prediction (260), e.g., using an intraprediction tool such as Decoder Side Intra Mode Derivation (DIMD).
- DIMD Decoder Side Intra Mode Derivation
- motion estimation 275
- compensation 270
- IBC is a predictive coding mode that explores the similarity, e.g.
- the encoder decides (205) which one of the intra mode, inter mode or IBC mode to use for encoding the unit, e.g., based on rate-distortion optimization, and indicates the intra/inter/IBC decision by, for example, a prediction mode flag.
- Prediction residuals are calculated, for example, by subtracting (210) the predicted block (a.k.a. prediction block) from the original image block.
- the encoder decodes an encoded block to provide a reference for further predictions.
- the quantized transform coefficients are de-quantized (240) (a.k.a. scaled) and inverse transformed (250) to decode prediction residuals.
- In-loop filters (265) are applied to the reconstructed picture to perform, for example, deblocking/SAO (Sample Adaptive Offset)/ALF (Adaptive Loop Filter) filtering to reduce encoding artifacts.
- the filtered image is stored in a reference picture buffer (280).
- In-loop filters (265) are thus used to enhance reconstructed images before storing them in the reference picture buffer (280).
- Inloop filters form a whole family.
- deblocking filters aim at reducing blocking artifacts occurring along block boundaries.
- Deblocking filters are usually designed to improve subjective quality, that is, the noticeability of such coding errors by the human psycho visual system.
- deblocking filters are predetermined based on coding information (such as prediction modes, motion vectors, transform coefficients) and on local variations across block boundaries.
- ALF adaptive loop filters
- ALF adaptive loop filters
- Adaptive loop filters are usually applied at CTU-level, while deblocking filters are applied along block borders.
- FIG. 3 illustrates a block diagram of an example video decoder 300.
- a io bitstream is decoded by the decoder elements as described below.
- Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG. 2.
- the encoder 200 also generally performs video decoding as part of encoding video data.
- the input of the decoder includes a video bitstream, which can be generated by video encoder 200.
- the bitstream is first entropy decoded (330) to obtain quantization levels c q (a.k.a. transform coefficient levels or quantization levels on the decoder side), prediction modes, motion vectors, and other coded information.
- the picture partition information indicates how the picture is partitioned.
- the decoder may therefore divide (335) the picture according to the decoded picture partitioning information.
- the quantization levels c q are dequantized (340) into reconstructed transform coefficients c r . De-quantization is also named scaling.
- the reconstructed transform coefficients c r are inverse transformed (350) to obtain the prediction residuals.
- an image block is reconstructed.
- the predicted block can be obtained (370) from intra prediction (360), motion-compensated prediction (i.e., inter prediction) (375) or from IBC prediction.
- In-loop filters (365) are applied to the reconstructed image.
- the filtered image is stored at a reference picture buffer (380). Note that, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side is identical to the contents of the reference picture buffer 280 on the encoder 200 side for the same picture.
- a plurality of transform sets (e.g., a total of four transform sets) and one or more (e.g., two) non separable transform matrices (kernels) per transform set may be used (e.g., predefined).
- the transform set to be used may be determined based on intra prediction modes.
- the separable DCT-II plus LFNST transform combinations may be replaced with NSPT for the block shapes 4x4, 4x8, 8x4 and 8x8, 4x16, 16x4, 8x16 and 16x8.
- a VIPM (e.g., each VIPM) represents a kind of feature of the inter-coded block, for example, a dominant direction.
- the VIPM may be derived by applying a decoder-side intra mode derivation (DIMD)-like process to an inter prediction block (as shown in FIG. 4.
- DIMD decoder-side intra mode derivation
- a first step horizontal and vertical gradients of a pixel (e.g., of each pixel) inside the prediction block are calculated, e.g., with Sobel operators.
- the obtained gradients are used to derive a direction.
- the target direction is given by 0 and the derived direction is the direction of the intra prediction mode that is the closest to this target direction.
- a block vector is obtained (e.g., decoded) for an IBC coded block called the current block.
- the block vector (BV) indicates a displacement from the current block to a reference block, which is already reconstructed inside the current picture.
- the samples inside the reference block may be copied as the prediction samples for the current IBC block.
- a direction of an intra prediction mode is obtained (for each pixel inside the IBC prediction block) based on the gradients computed for that pixel, e.g., direction of an intra prediction mode that is the closest to that of a target direction defined by an angle 0.
- , and tan(0) I ⁇ VBR I/ I ⁇ HOR I otherwise. If
- , the reference axis is the horizontal axis. Otherwise, the reference axis is the vertical axis.
- a neighboring block of a current (e.g., intra-coded) block is an inter-coded block or an IBC coded block.
- a VIPM mode is determined for the neighboring block, e.g., by applying a DIMD- like process to the prediction or reconstruction of this neighboring block.
- the method illustrated by FIG. 10 may be used to determine the corresponding VIPM mode.
- the IPM list used for GPM-Intra may include VIPM modes of adjacent spatial neighboring, non-adjacent spatial neighboring, history-based, and/or temporal non-intra coded blocks.
- the IPM candidates derived from spatial neighboring blocks e.g., from the five spatial neighboring blocks of FIG. 5
- a VIPM could be derived by applying a DIMD-like process to the prediction or reconstruction of this above neighboring block
- the corresponding VIPM mode may be inserted into the IPM list to be used for decoding/ encoding the current block, e.g., if it does not exist (e.g., is not already present) in the list.
- the above neighboring block is an IBC coded block.
- FIG. 13 depicts a flowchart of a method for decoding a current block according to an example.
- a second predictor is obtained for a second part of the current block or for the whole block based on the VIPM mode.
- the second predictor for the second part of the current block or for the whole current block may be obtained directly from the virtual intra prediction mode. In this case, the signaling overhead is reduced since no index identifying a mode in a list needs to be signaled.
- the virtual intra prediction mode may be inserted as an additional candidate in an intra prediction mode candidate list and the second predictor may be obtained from a particular intra prediction mode identified in the list by an index (e.g., a signaled or decoded index).
- the inter prediction signal of the partial or whole block may be derived with the selected merge candidate using its MV.
- the horizontal and vertical gradients of pixels inside the partial or whole inter prediction block may be calculated.
- the obtained gradients may be used to fill a histogram of oriented gradients (HOG).
- HOG histogram of oriented gradients
- the intra prediction mode whose index corresponds to the HOG bin of largest magnitude may be selected as the VIPM, which could be used to derive the intra prediction signal of the remaining part of the inter prediction block or the whole inter prediction block.
- the IBC prediction signal of the partial or whole block may be derived with the selected merge candidate using its block vector (BV).
- the VIPM may be derived by applying the same DIMD-like process to the partial or whole IBC prediction block, which could be used to derive the intra prediction signal of the remaining part of the IBC prediction block or the whole IBC prediction block.
- a VIPM may replace an IPM list or a TIMD derived intra prediction mode. For example, there’s no need to build the IPM candidate list and signal a selected index, a VIPM could be directly employed for GPM-Intra, IBC-GPM and IBC-CIIP, which results in reducing the signaling overhead and restricting the complexity.
- a VIPM may be included as a possible candidate in an IPM list.
- a VIPM could be registered first in the IPM candidate list used for GPM-Intra, and IBC-CIIP and IBC-GPM.
- FIG. 14 depicts a flowchart of a method for decoding a current block according to an example.
- the scanning order of coefficient groups and/or coefficients may be selected for a non-intra coded block (i.e., an inter-coded block or an IBC coded block) based on its VIPM.
- a VIPM mode is obtained for the current block based on a prediction block or on a residual block of the current block, e.g., by applying a DIMD-like process to the prediction block or to the residual block.
- horizontal and vertical gradients are computed (e.g., with Sobel operators) for each pixel inside the prediction block of the current block.
- the prediction block is obtained from one or more motion vectors in case of an inter-coded block and from one or more block vectors in case of an IBC coded block.
- steps SI 004 to SI 008 apply to derive the VIPM mode for the current block.
- the coefficient groups inside an inter block may be scanned and coded according to a scan pattern selected among three pre-defined scan orders: diagonal, horizontal, vertical.
- the scanning order of an inter block may be determined by the VIPM using a fixed LUT.
- the LUT of luma blocks could be proposed as shown in Table 3: vertical VIPMs 42-58 use the horizontal scan, horizontal VIP Ms 10-26 use the vertical scan, other VIPMs use the diagonal scan.
- the present aspects are not limited to ECM, VVC or HEVC, and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
- encoding refers only to entropy encoding
- encoding refers only to differential encoding
- encoding refers to a combination of differential encoding and entropy encoding.
- DASH MPD Media Presentation Description
- a Descriptor is associated with a Representation or collection of Representations to provide additional characteristic to the content Representation.
- RTP header extensions for example as used during RTP streaming.
- ISO Base Media File Format for example as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length also known as 'atoms' in some specifications.
- HLS HTTP live Streaming
- a manifest can be associated, for example, to a version or collection of versions of a content to provide characteristics of the version or collection of versions.
- Some embodiments may refer to rate distortion optimization.
- the rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion.
- the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding.
- Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one.
- the implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program).
- An apparatus can be implemented in, for example, appropriate hardware, software, and firmware.
- the methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants ("PDAs”), and other devices that facilitate communication of information between end-users.
- PDAs portable/personal digital assistants
- references to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment.
- the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
- Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
- Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
- this application may refer to “receiving” various pieces of information.
- Receiving is, as with “accessing”, intended to be a broad term.
- Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory).
- “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
- such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C).
- This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
- the word “signal” refers to, among other things, indicating something to a corresponding decoder.
- the encoder signals a particular one of a coding mode.
- the same parameter is used at both the encoder side and the decoder side.
- an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter.
- signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments.
- implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted.
- the information can include, for example, instructions for performing a method, or data produced by one of the described implementations.
- a signal can be formatted to carry the bitstream of a described embodiment.
- Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal.
- the formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream.
- the information that the signal carries can be, for example, analog or digital information.
- the signal can be transmitted over a variety of different wired or wireless links, as is known.
- the signal can be stored on a processor-readable medium.
- features described herein may be implemented in a bitstream or signal that includes information generated as described herein. The information may allow a decoder to decode a bitstream, the encoder, bitstream, and/or decoder according to any of the embodiments described.
- features described herein may be implemented by creating and/or transmitting and/or receiving and/or decoding a bitstream or signal.
- features described herein may be implemented a method, process, apparatus, medium storing instructions, medium storing data, or signal.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
A decoding method is disclosed. A first predictor is obtained (S1300) for a first part of a current block or for the whole current block. A virtual intra prediction mode is determined (S1302) for the current block based on the first predictor. A second predictor is obtained (S1304) for a second part of the current block or for the whole current block based on the virtual intra prediction mode. The current block is decoded (S1306) based on the first and second predictors.
Description
ENCODING AND DECODING METHODS USING VIRTUAL INTRA PREDICTION MODES AND CORRESPONDING APPARATUSES
CROSS REFERENCE TO RELATED APPLICATIONS
This application claims the benefit of European Application No. 24305028.3, filed on January 08, 2024 which is incorporated herein by reference in its entirety.
TECHNICAL FIELD
At least one of the present embodiments generally relates to a method and an apparatus for encoding and decoding a block (e.g., a picture block) using virtual intra prediction modes.
BACKGROUND
To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter picture correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction.
SUMMARY
In one implementation, a virtual intra prediction mode (VIPM) may be derived for an IBC (Intra Block Copy) coded block. In an example, the VIMP for the IBC coded block may be obtained by applying a decoder-side intra mode derivation (DIMD)-like process to a prediction block or a residual block, e.g. a block identified by a block vector associated with the current block, of the IBC coded block.
In an example, VIPMs of non-intra coded neighboring blocks of a current block could be used as the derived intra prediction mode candidate for deriving (e.g., filling) a generic MPM (Most Probable Mode) list for the current block.
In an example, a VIPM of the current block could be used as the derived intra prediction mode
for generating the intra predicted samples which may be used by prediction process, e.g., by GPM (Geometric Partitioning Mode)-Intra, IBC-GPM, CIIP (Combined Inter and Intra Prediction), IBC-CIIP.
VIPMs of non-intra coded blocks neighboring could be used as the derived intra prediction modes for generating the intra predicted samples which may be used by prediction process, e.g., by GPM-Intra, SGPM (Spatial geometric partitioning mode), IBC-GPM, CIIP, IBC-CIIP.
In an example, the scanning order of coefficient groups and/or coefficients may be selected based on VIPM for non-intra coded block.
BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented;
FIG. 2 illustrates a block diagram of an embodiment of a video encoder;
FIG. 3 illustrates a block diagram of an embodiment of a video decoder;
FIG.4 illustrates the principles of a decoder-side intra mode derivation (DIMD)-like process to applied to an inter prediction block ;
FIG. 5 depicts locations of neighboring blocks surrounding a current block;
FIG. 6 illustrates partitioning of a current block (e.g., a CU) split into partitions by a geometrically located straight line;
FIG. 7 depicts possible intra prediction mode (IPM) candidates for a current block split into an inter region and an intra region ;
FIG. 8 illustrates various scanning of coefficients within coefficient groups (CGs) ;
FIG. 9 depicts a Look-up table mapping a virtual intra prediction mode (VIPM) mode to a LFNST index;
FIG. 10 depicts a flowchart of a method for deriving a VIPM mode for an IBC Intra Block Copy) coded block according to an example;
FIG. 11 depicts a flowchart of a method for decoding a current block, e.g., an intra coded block ;
FIG. 12 depicts the above neighboring inter-coded block of a current block ;
FIG. 13 depicts a flowchart of a method for decoding a current block according to an example; and
FIG. 14 depicts a flowchart of a method for decoding a current block according to an example.
DETAILED DESCRIPTION
This application describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
The aspects described and contemplated in this application can be implemented in many different forms. FIGs. 1, 2 and 3 below provide some embodiments, but other embodiments are contemplated and the discussion of FIGs. 1, 2 and 3 does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium (e.g. a non-transitory computer readable storage medium) having stored thereon instructions for encoding or decoding video data according to any of the methods described, and/or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” or “coded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably and the terms “image,” “picture” and “frame” may be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side.
Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used
in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
For the sake of clarity, satisfying, failing to satisfy a condition and configuring condition parameter(s) are described throughout embodiments described herein as relative to a threshold (e.g., greater, or lower than), a (e.g., threshold) value, configuring the (e.g., threshold) value, etc.). For example, satisfying a condition may be described as being above a (e.g., threshold) value, and failing to satisfy a condition (e.g., performance criteria) may be described as being below a (e.g., threshold) value. Embodiments described herein are not limited to thresholdbased conditions. Any kind of other condition and parameter(s) (such as e.g., belonging or not belonging to a range of values) may be applicable to embodiments described herein.
The present aspects are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and/or discrete components. For example, in at least one embodiment, the processing and encoder/decoder elements of system 100 are distributed across multiple ICs and/or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.
The system 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device, and/or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and/or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and/or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and/or a network accessible storage device, as non-limiting examples.
System 100 includes an encoder/decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder/decoder module 130 may include its own processor and memory. The encoder/decoder module 130 represents module(s) that may be included in a device to perform the encoding and/or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder/decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
Program code to be loaded onto processor 110 or encoder/decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder/decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
In some embodiments, memory inside of the processor 110 and/or the encoder/decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder/decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and/or the storage device 140, for example, a dynamic volatile memory
and/or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2, (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO/IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).
The input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and/or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. 1, include composite video.
In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and/or add other elements performing similar or different functions. Adding elements may include inserting elements in between
existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
Additionally, the USB and/or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and/or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder/decoder 130 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
Various elements of system 100 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and/or a wireless medium.
Data is streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105. As indicated above, various embodiments provide
data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
The system 100 may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The display 165 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and/or a foldable display. The display 165 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 165 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and/or a lighting system. Various embodiments use one or more peripheral devices 185 that provide a function based on the output of the system 100. For example, a disk player performs the function of playing the output of the system 100.
In various embodiments, control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
The embodiments can be carried out by computer software implemented by the processor 110 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 120
can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 110 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
FIG. 2 illustrates an example video encoder 200, such as a VVC (Versatile Video Coding) encoder. FIG. 2 may also illustrate an encoder in which improvements are made to the VVC standard or an encoder employing technologies similar to VVC.
Before being encoded, the video sequence may go through pre-encoding processing (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one of the color components). Metadata can be associated with the preprocessing and attached to the bitstream.
In the encoder 200, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of, for example, CTU (Coding Tree Units), CUs (Coding Units). Each unit is encoded using, for example, either an intra or inter mode. In another example, a unit may be encoded using intra block copy (IBC). When a unit is encoded in an intra mode, it performs intra prediction (260), e.g., using an intraprediction tool such as Decoder Side Intra Mode Derivation (DIMD). In an inter mode, motion estimation (275) and compensation (270) are performed. IBC is a predictive coding mode that explores the similarity, e.g. repeating patterns, within the same picture. In this mode, a block (e.g., a CU) is predicted by an already reconstructed reference block within the same picture (reference samples are derived from inside the reconstructed part of current picture). The offset from the current block to its reference block is referred as a block vector (BV) or displacement vector, i.e., a vector that indicates the displacement from the current block to a reference block. If IBC is used by one particular CU, the CU is represented with a BV and possibly a residual signal of that CU, which is the approach similar to an inter mode in interframe motion estimation. The encoder decides (205) which one of the intra mode, inter mode or IBC mode to use for encoding the unit, e.g., based on rate-distortion optimization, and
indicates the intra/inter/IBC decision by, for example, a prediction mode flag. Prediction residuals are calculated, for example, by subtracting (210) the predicted block (a.k.a. prediction block) from the original image block.
The prediction residuals are then transformed (225) into transform coefficients c (a.k.a prediction residual transform coefficients) which are quantized (230) into quantization indexes cq (a.k.a transform coefficient levels or quantized transform coefficients on the encoder side). The quantization levels (a.k.a quantization indexes) cq. as well as motion vectors and other syntax elements such as the picture partitioning information, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes.
The encoder decodes an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized (240) (a.k.a. scaled) and inverse transformed (250) to decode prediction residuals. Combining (255) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (265) are applied to the reconstructed picture to perform, for example, deblocking/SAO (Sample Adaptive Offset)/ALF (Adaptive Loop Filter) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280). In-loop filters (265) are thus used to enhance reconstructed images before storing them in the reference picture buffer (280). Inloop filters form a whole family. Among them, deblocking filters (DBF) aim at reducing blocking artifacts occurring along block boundaries. Deblocking filters are usually designed to improve subjective quality, that is, the noticeability of such coding errors by the human psycho visual system. In usual video coding standards such as HEVC and VVC, deblocking filters are predetermined based on coding information (such as prediction modes, motion vectors, transform coefficients) and on local variations across block boundaries. On the other hand, adaptive loop filters (ALF) are learnt at encoder side in order to minimize a mean squared error with respect to source images, then the learned filter weights are encoded into the bitstream. Adaptive loop filters are usually applied at CTU-level, while deblocking filters are applied along block borders.
FIG. 3 illustrates a block diagram of an example video decoder 300. In the decoder 300, a io
bitstream is decoded by the decoder elements as described below. Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG. 2. The encoder 200 also generally performs video decoding as part of encoding video data.
In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain quantization levels cq (a.k.a. transform coefficient levels or quantization levels on the decoder side), prediction modes, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (335) the picture according to the decoded picture partitioning information. The quantization levels cq are dequantized (340) into reconstructed transform coefficients cr. De-quantization is also named scaling. The reconstructed transform coefficients cr are inverse transformed (350) to obtain the prediction residuals. Combining (355) the prediction residuals and the predicted block (a.k.a. prediction block), an image block is reconstructed. The predicted block can be obtained (370) from intra prediction (360), motion-compensated prediction (i.e., inter prediction) (375) or from IBC prediction. In-loop filters (365) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380). Note that, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side is identical to the contents of the reference picture buffer 280 on the encoder 200 side for the same picture.
The decoded picture can further go through post-decoding processing (385), for example, an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
For intra prediction of a given block, a most probable mode (MPM) list-based signaling scheme may be used to efficiently signal the index of the intra prediction mode selected to predict this block with little signaling overhead. The MPM list-based signaling scheme may implement a sophisticated entropy coding of the index of the selected intra prediction mode at low complexity. A generic MPM list may be decomposed into a list of (e.g., 6) primary MPMs (PMPM) and a list of (e.g., 16) secondary MPMs (SMPM).
A generic MPM list with several entries (e.g., 22 entries) may be built by sequentially adding candidate intra prediction mode indices, from the index of the intra prediction mode being most
likely the index of the selected one to the index of the intra prediction mode being least likely the index of the selected one. Some candidates of the MPM list may be obtained from the indices of the intra modes of the above (A), left (L), below-left (BL), above-right (AR), and above-left (AL) neighboring blocks in order. The locations of neighboring blocks surrounding the current block are shown in FIG. 5.
When a neighboring block is intra coded, the index of its selected intra prediction mode can directly be inserted in the generic MPM list if it does not already exist in this list (e.g., if it is not already present in the list). Otherwise, e.g., when this neighboring block is inter-coded or intra block copy (IBC) coded, the generic MPM list may be filled with the index of a “propagated” intra prediction mode. The “propagated” intra prediction mode refers to the selected intra prediction mode of this neighboring block’s reference block, if the reference block is intra coded. If the reference block is inter-coded, its “propagated” intra prediction mode may be picked instead.
A geometric partitioning mode (GPM) may be used, e.g., featuring 64 partitions in total for inter prediction. When the GPM is used, a block (e.g., a CU) is split into two partitions by a geometrically located straight line as illustrated on FIG. 6). On FIG. 6, several straight lines are represented, some have the same orientation but different distance offset. The location of the splitting line may be mathematically derived from an angle and a distance offset of a specific partition. A partition in the block (e.g., each partition in the block) is inter-predicted using its own motion parameters. In an example, only uni-prediction is allowed for each partition, that is, each partition has one motion vector (MV) and one reference index. After predicting each of the partitions, the sample values along the splitting edge may be adjusted, e.g., using a blending process with adaptive weights.
Combining an inter prediction with an intra prediction may be used by adding intra modes to GPM. In GPM with inter and intra prediction (GPM-Intra), the final prediction samples may be generated by weighting inter predicted samples and intra predicted samples for GPM- separated regions (e.g., for each GPM-separated region). More precisely, one GPM-region may be intra predicted and another one may be inter predicted as illustrated on FIG. 7. After predicting each GPM-separated region, the sample values along the splitting edge (a.k.a the GPM block boundary) may be adjusted, e.g., using a blending process with adaptive weights. The inter predicted samples may be derived by the same scheme as the GPM (e.g., the scheme in VVC and ECM) whereas the intra predicted samples may be derived based on an intra
prediction mode (IPM) candidate list and an index signaled by the encoder (e.g., encoded in a bitstream), the index identifying a particular intra prediction mode in the IPM list.
The size of the IPM candidate list may be set to a pre-defined (e.g., fixed) value, e.g., to 3. The IPM candidates (e.g., the initial available IPM candidates) may be the parallel angular mode against the GPM block boundary (Parallel mode) as shown on FIG. 7(a), the perpendicular angular mode against the GPM block boundary (Perpendicular mode) as shown on FIG. 7(b), and the Planar mode as shown on FIG. 7(c). The IPM candidate list may comprise additional candidates. For example, Decoder-Side Intra Mode Derivation (DIMD) and template-based intra mode derivation (TIMD) as defined in ECM may also be used as IPM candidates of GPM- Intra to further improve the coding performance.
DIMD is a process that may be used by both the encoder and the decoder. According to DIMD, indices of two intra prediction modes (e.g., two intra prediction modes that most likely yield the predictions of the current block (e.g., luminance CB) of highest qualities according to DIMD) are derived (e.g., selected). The derivation may comprise creation (e.g., filling) of a Histogram of Oriented Gradients (HOG) of a context (e.g., an L-shape template) of decoded reference samples surrounding the current block. The indices of the two derived intra prediction modes may be the indices of the two HOG bins of largest magnitudes.
TIMD is a process that may be used by both the encoder and the decoder. According to TIMD, indices of the two intra prediction modes (e.g., two intra prediction modes that most likely yield the predictions of the current block (e.g., luminance CB) of highest qualities according to TIMD) are derived (e.g., selected). The derivation may comprise testing a plurality of intra prediction mode. For each tested intra prediction mode, a template of the current block (e.g., luminance CB) may be predicted from a set of decoded reference samples surrounding the template via this tested mode and the prediction SATD may be computed. The indices of the two derived intra prediction modes may be the indices of the two tested modes incurring the two smallest prediction SATDs.
In an example, the IPM candidate list may thus be obtained by adding (e.g., registering) first the Parallel mode. Then, IPM candidates of TIMD, DIMD, and neighboring blocks, and the Planar mode, may be added (e.g., registered), e.g., if there is not the same IPM candidate in the list (i.e., if not already present in the list). The available IPM may thus be added (e.g., registered) in the IPM list in following order:
1) Parallel mode;
2) one derived mode from DIMD;
3) one derived mode from TIMD;
4) five candidate mode derived from the neighboring blocks; and
5) the Planar mode.
As for the neighboring mode derivation, several positions (e.g., at most five positions) of available neighboring blocks may be used (as shown in FIG. 5). In an example, they may be restricted by the angle of GPM block boundary as shown in Table 1. Table 1 represents the position of available neighboring blocks for IPM candidate derivation based on the angle of GPM block boundary. A and L denotes the above and left side of the prediction block.
Table 1
A Spatial geometric partitioning mode (SGPM) may be defined as an intra mode that resembles the inter coding tool of GPM. According to SGPM, two prediction parts (e.g., predictions) are generated from an intra prediction process. For each partition mode, an IPM list may be derived for each part. The IPM list size may be set to a pre-defined (e.g., fixed) value (e.g., 3). In an example, the available IPM candidates may be:
1) two derived modes from TIMD with horizontal and vertical orientations;
2) Parallel mode;
3) one derived mode from DIMD;
4) five candidates derived from the neighboring blocks; and
5) the Planar mode.
The neighboring mode derivation may be restricted by the angle of SGPM block boundary.
Intra block copy with geometry partitioning mode (IBC-GPM) is a coding tool which divides a block (e.g., a CU) into two sub-partitions geometrically. The prediction signals of the two sub-partitions may generated using IBC and possibly intra prediction. An IPM candidate list may be derived (e.g., constructed) using the same method as the method used for GPM-Intra for intra prediction. The IPM candidate list size may be set to a pre-defined (e.g., fixed) value (e.g., 3). In an example, there are several geometry partitioning modes (e.g., 48 geometry
partitioning modes in total), which may be divided into two geometry partitioning mode sets.
When IBC-GPM is used, an IBC-GPM geometry partitioning mode set flag may be signaled to indicate whether the first or the second geometry partitioning mode set is selected, followed by a geometry partitioning mode index. A first IBC-GPM intra flag may be signaled (e.g., is always signaled) to indicate whether the first partition is intra predicted or not. When intra prediction is used for the first sub-partition, an intra prediction mode index may be signaled. When IBC is used for the first sub-partition, a merge index may be signaled. The second IBC- GPM intra flag only needs to be signaled when the first flag is false (i.e. , the first partition is IBC predicted) to indicate whether the intra prediction is applied to the second partition. Additionally, when both flags are false (e.g., indicating that both partitions are generated using IBC), the maximum codeword of the second partition may be reduced by 1, e.g., due to the fact that the merge indices of two partitions cannot be identical.
The combined inter/intra prediction (CIIP) combines an inter prediction signal with an intra prediction signal. The inter prediction signal in the CIIP mode Plnter may be derived using the same inter prediction process applied to regular merge mode; and the intra prediction signal Ptntra maY be using a TIMD derived intra prediction mode. The TIMD derivation method may be used to derive the intra prediction mode in CIIP. Specifically, an intra prediction mode for which the sum of absolute transformed differences (SATD) between the template of the current block and the prediction of the template of the current block via this mode is the smallest is first selected (e.g., among a large set of intra prediction modes, e.g. corresponding to 130 angles) and then mapped to one of the 67 regular intra prediction modes. The intra and inter prediction signals may finally be combined using weighted averaging to generate a final prediction. The weights may be calculated depending on the derived intra prediction mode.
Combined intra block copy and intra prediction (IBC-CIIP) is a coding tool for an IBC coded block which uses IBC and intra predictions to obtain two prediction signals. A final prediction P may be obtained (e.g. generated) by a weighted sum of the two prediction signals as follows:
P = (wibc * Pibc + ((1 « shift) - wlbc) * Plntra + (1 « (shift - 1))) » shift where Pibc and Pintra denote the IBC prediction signal and intra prediction signal respectively. In an example, (wibc, shift) are set equal to (13, 4) and (1, 1) for IBC merge mode and IBC AMVP mode respectively.
An IPM candidate list may be used to generate the intra prediction signal. The IPM candidate
list size may be set to a pre-defined value (e.g., 2). An IPM index may be signaled (e.g., in the bitstream) to indicate which IPM is used. The available IPM candidates may be:
1) derived mode from TIMD;
2) the Planar mode; and
3) the Horizontal mode (HOR IDX).
Quantized transform coefficients of a coding block may be coded using non-overlapped coefficient groups (CGs). In an example, a CG (e.g., each CG) contains the coefficients of a 4x4 block of a coding block. The CGs contained in an 8x8 block are illustrated on FIG. 8.
The CGs inside a coding block, and the 16 transform coefficients within a CG, may be scanned and coded (e.g., in HEVC) according to a scan pattern selected among three pre-defined scan orders: diagonal, horizontal, vertical. For inter coded blocks, the diagonal scanning on the left of FIG. 8 is always used; while for 4x4 and 8x8 intra block, the scanning order depends on the intra prediction mode active for that block, which is known as mode dependent coefficient scanning (MDCS), as depicted in FIG. 8.
The scanning order of an intra block may be determined by the intra prediction mode using a fixed look-up table (LUT). The LUT of luma blocks (e.g., in HEVC) is given in Table 2. Table 2 is a LUT for intra luma coefficient scan index selection based on intra prediction mode. D, H and V indicate diagonal, horizontal and vertical scanning. Vertical modes 6-14 use the horizontal scan, horizontal modes 22-30 use the vertical scan, other modes use the diagonal scan.
Table 2
As disclosed with respect to FIG. 2 and 3, a block may be encoded (decoded respectively) in inter mode. Such an inter-coded block may use virtual intra prediction mode (VIPM), e.g., to utilize low frequency non-separable transform (LFNST)Znon-separable primary transform (NSPT). Low-frequency non-separable transform (LFNST) can be applied to the low- frequency components of a primary transform to better exploit the directionality characteristics. It may be applied between the forward primary transform and quantization at the encoder side and between the inverse quantization scaling and inverse primary transform at the decoder side.
In an example, a plurality of transform sets (e.g., a total of four transform sets) and one or more (e.g., two) non separable transform matrices (kernels) per transform set may be used (e.g., predefined). The transform set to be used may be determined based on intra prediction modes (a.k.a intra-picture prediction mode). For a (e.g. each) transform set, the selected non separable secondary transform candidate may be further specified by an explicitly LFNST index that is signaled for the CU. Non-separable primary transform (NSPT) may be applied as the primary transform for intra coding. Similar to LFNST, a plurality of transform sets (e.g., a total of four transform sets) and one or more (e.g., two) non separable transform matrices (kernels) per transform set may be used (e.g., predefined). The transform set to be used may be determined based on intra prediction modes. The separable DCT-II plus LFNST transform combinations may be replaced with NSPT for the block shapes 4x4, 4x8, 8x4 and 8x8, 4x16, 16x4, 8x16 and 16x8.
A VIPM (e.g., each VIPM) represents a kind of feature of the inter-coded block, for example, a dominant direction. The VIPM may be derived by applying a decoder-side intra mode derivation (DIMD)-like process to an inter prediction block (as shown in FIG. 4. In a first step, horizontal and vertical gradients of a pixel (e.g., of each pixel) inside the prediction block are calculated, e.g., with Sobel operators. Then, in a second step, the obtained gradients are used to derive a direction. With reference to FIG. 4, the target direction is given by 0 and the derived direction is the direction of the intra prediction mode that is the closest to this target direction. In an example, the amplitudes of the gradient of each pixel are accumulated for the corresponding direction, e.g., by filling a histogram of oriented gradients (HOG) wherein each bin of the HOG is associated with the index of a different directional intra prediction mode. More precisely, the horizontal and vertical gradients computed at a given pixel of the interprediction block enable to compute an index of a directional intra prediction mode whose prediction gives rise to this gradient, and, in the HOG, the bin of this computed index is incremented by the absolute gradients. Finally, in a third step, the intra prediction mode whose index corresponds to the HOG bin of largest magnitude (e.g., strongest accumulation) is selected as the VIPM.
From the VIPM, the transform set of LFNST/NSPT may thus be derived for an inter-coded block. More precisely, the index of the selected VIPM maps to the transform set index of low frequency non-separable transform/non-separable primary transform (LFNST/NSPT). FIG. 9 shows an example of mapping VIPM mode to LFNST set index. For inter coding, the kernel mapping method may be the same as the method used for intra LFNST/NSPT.
As disclosed above, virtual intra prediction mode (VIPM) may be employed for an inter coded block. In the previous example, the derived VIPM is only used to select low frequency non- separable transform/non-separable primary transform (LFNST/NSPT) transform.
In contrast, a decoding method (an encoding method respectively) is disclosed below whereby an IBC coded block (i.e., different from an inter coded block) uses VIPM. Besides, the decoding method (encoding method respectively) may use the VIPM as the derived intra prediction mode for use with several tools. This makes it possible to improve the encoding performance.
FIG. 10 depicts a flowchart of a method for deriving a VIPM mode for an IBC coded block in a current picture according to an example. The method may be used by an encoder and/or a decoder. An intra prediction mode is derived according to an IBC prediction block (or to a corresponding IBC residual block). This derived intra prediction mode may, for example, be used for generating corresponding intra prediction samples of the current IBC coded block.
At S1000, a block vector is obtained (e.g., decoded) for an IBC coded block called the current block. For example, the block vector (BV) indicates a displacement from the current block to a reference block, which is already reconstructed inside the current picture. The samples inside the reference block may be copied as the prediction samples for the current IBC block.
At SI 002, horizontal and vertical gradients are computed (e.g., with Sobel operators) for each pixel inside the IBC prediction block. The IBC prediction block is the block comprising the prediction samples for the current IBC block and is thus obtained from the BV.
At SI 004, a direction of an intra prediction mode is obtained (for each pixel inside the IBC prediction block) based on the gradients computed for that pixel, e.g., direction of an intra prediction mode that is the closest to that of a target direction defined by an angle 0. The angle 6 between a reference axis and the direction perpendicular to the gradient G of components GVER and GHOR is such that tan(0) = | GH0R | / 1 GVER | if IG^ I > |tiH0R | , and tan(0) = I ^VBR I/ I ^HOR I otherwise. If |CK£R| > |tiH0R | , the reference axis is the horizontal axis. Otherwise, the reference axis is the vertical axis.
At SI 006, the amplitudes of the horizontal and vertical gradients for the obtained direction are accumulated. In an example, a histogram of oriented gradients is filled based on the computed horizontal and vertical gradients, e.g., the bin associated with the index i of the directional intra
prediction mode (a.k.a. target intra prediction mode) whose direction is the closest to the target direction is incremented by the amplitudes of both gradients. Said otherwise, the bin associated with the index i of the intra prediction mode corresponding to the direction obtained at SI 104 is incremented by | GH0R 11+ 1 GVER | . The bin associated with the index i of the target intra prediction mode, H0G[i]=H0G[i]+ \GH0R | +
|. Note that, for the current decoded pixel of interest, if GH0R = GVER = 0, no bin in the HOG is incremented.
At SI 008, an intra prediction mode corresponding to the highest accumulation is selected. For example, the intra prediction mode whose index corresponds to the HOG bin of largest magnitude is selected. The selected intra prediction mode is the VIPM associated with the IBC coded block.
FIG. 11 depicts a flowchart of a method for decoding a current block, e.g., an intra coded block.
As explained previously, a generic MPM list could be filled with the derived intra prediction mode of non-intra coded blocks (i.e., inter and/or IBC blocks) using intra mode propagation method. However, these propagated intra prediction modes might not be relevant for the current block, because the neighboring block’s reference block could be quite far away (spatially and/or temporally) from the current block and its direction feature could be non-representative of the texture of the current block. To improve the accuracy of the MPM list generation, VIPM of non-intra coded blocks may be used to fill a generic MPM list. In an example, the VIPM of non-intra coded blocks may replace (e.g., may be used instead of) the propagated intra prediction modes in the generic MPM list. In a variant, both the propagated intra mode and the VIPMs of non-intra coded blocks can be used to fill the MPM list.
At SI 100, it is determined that a neighboring block of a current (e.g., intra-coded) block is an inter-coded block or an IBC coded block.
At SI 102, a VIPM mode is determined for the neighboring block, e.g., by applying a DIMD- like process to the prediction or reconstruction of this neighboring block. In the case where the neighboring block is an IBC coded block, the method illustrated by FIG. 10 may be used to determine the corresponding VIPM mode.
At SI 104, this VIPM mode may be added to (e.g., inserted in) an MPM list associated with the current block, e.g., if this mode is not already in the list.
At SI 106, the current block is decoded based on the MPM list.
In one example, when the neighboring block of the current block is inter coded or IBC coded, the VIPM may be derived based on (e.g., according to) the reconstructed inter/IBC neighboring block, and then this derived VIPM could be used as a possible candidate to be inserted in the generic MPM list, e.g., if it is not already present in the list. In this example, the horizontal and vertical gradients used to derive the VIPM may be computed directly on the reconstructed neighboring block. In one variant, the VIPM may be derived based on a prediction block (or a corresponding residual block) used for reconstructing the inter/IBC neighboring block.
The VIPM mode of the neighboring block may be inserted in place of (e.g., instead of) propagated intra prediction modes. In a variant, both the propagated intra modes and the VIPMs of non-intra coded neighboring blocks can be used to fill the MPM list.
VIPM modes of non-intra coded blocks (e.g., of non-intra coded neighboring blocks) may further be used to generate the intra predicted samples in GPM-Intra, SGPM, IBC-GPM, CIIP and IBC-CIIP.
In an example, the IPM list used for GPM-Intra may include VIPM modes of adjacent spatial neighboring, non-adjacent spatial neighboring, history-based, and/or temporal non-intra coded blocks. Specifically, for the IPM candidates derived from spatial neighboring blocks (e.g., from the five spatial neighboring blocks of FIG. 5), if the above neighboring block is inter coded as shown in FIG. 12, then a VIPM could be derived by applying a DIMD-like process to the prediction or reconstruction of this above neighboring block, the corresponding VIPM mode may be inserted into the IPM list to be used for decoding/ encoding the current block, e.g., if it does not exist (e.g., is not already present) in the list. The same may apply in the case where the above neighboring block is an IBC coded block.
FIG. 13 depicts a flowchart of a method for decoding a current block according to an example.
As explained previously, the intra predicted samples in GPM-Intra, IBC-GPM and IBC-CIIP may be derived by an intra prediction mode candidate from an IPM list, e.g., by an intra prediction mode candidate identified in the IPM list, e.g., using a signaled index. The intra predicted samples in CIIP may be generated by a TIMD derived intra prediction mode. To further improve the compression efficiency (that is, to reduce the bitrate while maintaining the quality, or equivalently to improve the quality while maintaining the bitrate), VIPM mode may be used to derive the intra predicted samples in GPM-Intra, IBC-GPM, CIIP and IBC-CIIP.
At S1300, a first predictor (a.k.a. a first prediction part) is obtained for a first part of a current block, e.g., for a block partition, a block sub-partition, a block region, or for the whole current block. The prediction block first predictor may be obtained from one or more motion vectors (e.g., by motion compensation) in case of an inter-coded block and from one or more block vectors in case of an IBC coded block.
At S1302, a VIPM mode is obtained for the current block based on (e.g., derived from) the first predictor (or from the corresponding residual part). In an example, horizontal and vertical gradients are computed (e.g., with Sobel operators) for each pixel inside the first predictor. Then, steps SI 004 to SI 008 apply to derive the VIPM mode. As an example, for each pixel inside the first predictor : vertical and horizontal gradients are computed, a direction of an intra prediction mode is derived based on the vertical and horizontal gradients and amplitudes of the vertical and horizontal gradients are accumulated for the direction. An intra prediction mode corresponding to strongest accumulation is selected as the virtual intra prediction mode for the current block. HOG may be used.
At SI 304, a second predictor is obtained for a second part of the current block or for the whole block based on the VIPM mode. In an example, the second predictor for the second part of the current block or for the whole current block may be obtained directly from the virtual intra prediction mode. In this case, the signaling overhead is reduced since no index identifying a mode in a list needs to be signaled. In another example, the virtual intra prediction mode may be inserted as an additional candidate in an intra prediction mode candidate list and the second predictor may be obtained from a particular intra prediction mode identified in the list by an index (e.g., a signaled or decoded index).
At SI 306, the current block is decoded from the first and second predictors.
For GPM-Intra and CIIP, the inter prediction signal of the partial or whole block may be derived with the selected merge candidate using its MV. The horizontal and vertical gradients of pixels inside the partial or whole inter prediction block may be calculated. Then, the obtained gradients may be used to fill a histogram of oriented gradients (HOG). Finally, the intra prediction mode whose index corresponds to the HOG bin of largest magnitude may be selected as the VIPM, which could be used to derive the intra prediction signal of the remaining part of the inter prediction block or the whole inter prediction block.
For IBC-GPM and IBC-CIIP, the IBC prediction signal of the partial or whole block may be
derived with the selected merge candidate using its block vector (BV). The VIPM may be derived by applying the same DIMD-like process to the partial or whole IBC prediction block, which could be used to derive the intra prediction signal of the remaining part of the IBC prediction block or the whole IBC prediction block.
In an example, a VIPM may replace an IPM list or a TIMD derived intra prediction mode. For example, there’s no need to build the IPM candidate list and signal a selected index, a VIPM could be directly employed for GPM-Intra, IBC-GPM and IBC-CIIP, which results in reducing the signaling overhead and restricting the complexity.
In another variant, a VIPM may be included as a possible candidate in an IPM list. For example, a VIPM could be registered first in the IPM candidate list used for GPM-Intra, and IBC-CIIP and IBC-GPM.
FIG. 14 depicts a flowchart of a method for decoding a current block according to an example.
As explained previously, mode dependent coefficient scanning (MDCS) may be applied to some intra-coded blocks, wherein the scanning order of coefficient groups and/or coefficients is selected based on the intra prediction mode (e.g., activated) for that block. However, only diagonal scanning order is allowed for the inter blocks due to the absence of intra prediction mode information.
To improve compression efficiency, the scanning order of coefficient groups and/or coefficients may be selected for a non-intra coded block (i.e., an inter-coded block or an IBC coded block) based on its VIPM.
At SI 400, a VIPM mode is obtained for the current block based on a prediction block or on a residual block of the current block, e.g., by applying a DIMD-like process to the prediction block or to the residual block. In an example, horizontal and vertical gradients are computed (e.g., with Sobel operators) for each pixel inside the prediction block of the current block. The prediction block is obtained from one or more motion vectors in case of an inter-coded block and from one or more block vectors in case of an IBC coded block. Then, steps SI 004 to SI 008 apply to derive the VIPM mode for the current block. As an example, for each pixel inside the prediction block : vertical and horizontal gradients are computed, a direction of an intra prediction mode is derived based on the vertical and horizontal gradients and amplitudes of the vertical and horizontal gradients are accumulated for the direction. An intra prediction mode
corresponding to strongest accumulation is selected as the virtual intra prediction mode for the current block. HOG may be used.
At SI 402, a scanning pattern is selected based on the VIPM mode, e.g., using a Look-Up Table associating scanning pattern (a.k.a. scanning order) with VIPM modes. An example of such is LUT is given by Table 3.
At S1404, the current block is decoded using the selected scanning pattern.
For example, the coefficient groups inside an inter block, and the coefficients within a coefficient group, may be scanned and coded according to a scan pattern selected among three pre-defined scan orders: diagonal, horizontal, vertical. The scanning order of an inter block may be determined by the VIPM using a fixed LUT. The LUT of luma blocks could be proposed as shown in Table 3: vertical VIPMs 42-58 use the horizontal scan, horizontal VIP Ms 10-26 use the vertical scan, other VIPMs use the diagonal scan.
Table 3
The examples disclosed with respect to FIGs 10-14 that relate to the decoding of a block apply in the same way to the encoding of a block since the encoding method comprises a decoding loop. In addition, the examples disclosed with respect to FIGs 10-14 may be combined. As an example, the VIPM derived for the current block (FIG. 13) may be added to the IPM candidate list in addition to the VIPMs derived (as illustrated by FIG. 12) for the neighboring blocks of the current block. In addition, a VIPM of a current block may be used to derive intra predicted samples and further to select a scanning pattern.
Moreover, the present aspects are not limited to ECM, VVC or HEVC, and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.
Note that syntax elements as used herein, such as terms in equations and algorithms, signal (e.g., flag, mode) labels/names, etc., such as geometry partitioning mode set flag, IBC-GPM intra flag, index identifying a mode in a list (e.g., in an MPM list or in a IPM list) and so on, are descriptive terms. As such, they do not preclude the use of other syntax element names.
Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application, for example, decode a block using a VIPM mode obtained for the current block or for a neighboring block of the current block.
As further examples, in one embodiment “decoding” refers only to entropy decoding, in another embodiment “decoding” refers only to differential decoding, and in another embodiment “decoding” refers to a combination of entropy decoding and differential decoding, and in another embodiment “decoding” refers to the whole reconstructing picture process including entropy decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application, for example, encoding a block using a VIPM mode obtained for the current block or for a neighboring block of the current block.
As further examples, in one embodiment “encoding” refers only to entropy encoding, in another embodiment “encoding” refers only to differential encoding, and in another embodiment “encoding” refers to a combination of differential encoding and entropy encoding.
Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
This disclosure has described various pieces of information, such as for example syntax, that can be transmitted or stored, for example. This information can be packaged or arranged in a variety of manners, including for example manners common in video standards such as putting the information into an SPS, a PPS, a NAL unit, a header (for example, a NAL unit header, or a slice header), or an SEI message. Other manners are also available, including for example manners common for system level or application level standards such as putting the information into one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, for example as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. DASH MPD (Media Presentation Description) Descriptors, for example as used in DASH and transmitted over HTTP, a Descriptor is associated with a Representation or collection of Representations to provide additional characteristic to the content Representation. c. RTP header extensions, for example as used during RTP streaming. d. ISO Base Media File Format, for example as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length also known as 'atoms' in some specifications. e. HLS (HTTP live Streaming) manifest transmitted over HTTP. A manifest can be associated, for example, to a version or collection of versions of a content to provide characteristics of the version or collection of versions.
When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method/process.
Some embodiments may refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. The rate distortion optimization is usually
formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. Mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion.
The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
Additionally, this application may refer to “determining” various pieces of information.
Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
It is to be appreciated that the use of any of the following
“and/or”, and “at least one of’, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and/or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a particular one of a coding mode. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit
signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Many examples are described herein. Features of examples may be provided alone or in any combination, across various claim categories and types. Further, examples may include one or more of the features, devices, or aspects described herein, alone or in any combination, across various claim categories and types. For example, features described herein may be implemented in a bitstream or signal that includes information generated as described herein. The information may allow a decoder to decode a bitstream, the encoder, bitstream, and/or decoder according to any of the embodiments described. For example, features described herein may be implemented by creating and/or transmitting and/or receiving and/or decoding a bitstream or signal. For example, features described herein may be implemented a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a TV, set-top box, cell phone, tablet, or other electronic device that performs decoding. The TV, set-top box, cell phone, tablet, or other
electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image from residual reconstruction of the video bitstream). The TV, set-top box, cell phone, tablet, or other electronic device may receive a signal including an encoded image and perform decoding. A number of embodiments has been described above. Features of these embodiments can be provided alone or in any combination, across various claim categories and types.
Claims
1. A decoding method comprising: obtaining a first predictor for a first part of a current block or for the whole current block ; determining a virtual intra prediction mode for the current block based on the first predictor ; obtaining a second predictor for a second part of the current block or for the whole current block based on the virtual intra prediction mode; and decoding the current block based on the first and second predictors.
2. The method of claim 1 , wherein determining a virtual intra prediction mode for the current block based on the first predictor comprises: for each pixel inside the first predictor : computing vertical and horizontal gradients ; deriving a direction of an intra prediction mode based on the vertical and horizontal gradients ; accumulating amplitudes of the vertical and horizontal gradients for the direction ; and selecting an intra prediction mode corresponding to strongest accumulation as the virtual intra prediction mode for the current block.
3. The method of claim 1 or 2, wherein obtaining the second predictor for the second part of the current block or for the whole current block based on the virtual intra prediction mode comprises obtaining the second predictor directly from the virtual intra prediction mode.
4. The method of claim 1 or 2, wherein obtaining the second predictor for the second part of the current block or for the whole current block based on the virtual intra prediction mode comprises inserting the virtual intra prediction mode as an additional candidate in an intra prediction mode candidate list and obtaining the second predictor from an intra prediction mode identified by an index in the intra prediction mode candidate list.
5. The method of any one of claims 1 to 4, wherein obtaining the first predictor for the first part of the current block or for the whole current block comprises obtaining the first predictor based on at least one motion vector or a block vector associated with the current block.
6. The method of any one of claims 1 to 5, wherein decoding the current block based on the first and second predictors comprises selecting a scan pattern based on the virtual intra prediction mode.
7. An encoding method comprising: obtaining a first predictor for a first part of a current block or for the whole current block ; determining a virtual intra prediction mode for the current block based on the first predictor ; obtaining a second predictor for a second part of the current block or for the whole current block based on the virtual intra prediction mode; and encoding the current block based on the first and second predictors.
8. The method of claim 7, wherein determining a virtual intra prediction mode for the current block based on the first predictor comprises: for each pixel inside the first predictor : computing vertical and horizontal gradients ; deriving a direction of an intra prediction mode based on the vertical and horizontal gradients ; accumulating amplitudes of the vertical and horizontal gradients for the direction ; and selecting an intra prediction mode corresponding to strongest accumulation as the virtual intra prediction mode for the current block.
9. The method of claim 7 or 8, wherein obtaining the second predictor for the second part of the current block or for the whole current block based on the virtual intra prediction mode comprises obtaining the second predictor directly from the virtual intra prediction mode.
10. The method of claim 7 or 8, wherein obtaining the second predictor for the second part of the current block or for the whole current block based on the virtual intra prediction mode comprises inserting the virtual intra prediction mode as an additional candidate in an intra prediction mode candidate list and obtaining the second predictor from an intra prediction mode identified by an index in the intra prediction mode candidate list.
11. The method of any one of claims 7 to 10, wherein obtaining the first predictor for the first part of the current block or for the whole current block comprises obtaining the first predictor based on at least one motion vector or a block vector associated with the current block.
12. The method of any one of claims 7 to 11, wherein decoding the current block based on the first and second predictors comprises selecting a scan pattern based on the virtual intra prediction mode.
13. A decoding apparatus comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform the method of any one of claims 1-6.
14. An encoding apparatus comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform the method of any one of claims 7-12.
15. A computer program comprising program code instructions for implementing the method according to any one of claims 1-6 when executed by a processor.
16. A computer readable storage medium having stored thereon instructions for implementing the method of any one of claims 7-12.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24305028 | 2024-01-08 | ||
| EP24305028.3 | 2024-01-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025149263A1 true WO2025149263A1 (en) | 2025-07-17 |
Family
ID=89661113
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2024/085294 Pending WO2025149263A1 (en) | 2024-01-08 | 2024-12-09 | Encoding and decoding methods using virtual intra prediction modes and corresponding apparatuses |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025149263A1 (en) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023194556A1 (en) * | 2022-04-08 | 2023-10-12 | Interdigital Ce Patent Holdings, Sas | Implicit intra mode for combined inter merge/intra prediction and geometric partitioning mode intra/inter prediction |
-
2024
- 2024-12-09 WO PCT/EP2024/085294 patent/WO2025149263A1/en active Pending
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023194556A1 (en) * | 2022-04-08 | 2023-10-12 | Interdigital Ce Patent Holdings, Sas | Implicit intra mode for combined inter merge/intra prediction and geometric partitioning mode intra/inter prediction |
Non-Patent Citations (6)
| Title |
|---|
| CHEN J ET AL: "Algorithm description for Versatile Video Coding and Test Model 4 (VTM 4)", no. m46628, 16 February 2019 (2019-02-16), XP030215566, Retrieved from the Internet <URL:http://phenix.int-evry.fr/mpeg/doc_end_user/documents/125_Marrakech/wg11/m46628-JVET-M1002-v1-JVET-M1002-v1.zip JVET-M1002-v1.docx> [retrieved on 20190216] * |
| COBAN M ET AL: "Algorithm description of Enhanced Compression Model 11 (ECM 11)", no. m65800 ; JVET-AF2025, 28 December 2023 (2023-12-28), XP030313758, Retrieved from the Internet <URL:https://dms.mpeg.expert/doc_end_user/documents/144_Hannover/wg11/m65800-JVET-AF2025-v1-JVET-AF2025.zip JVET-AF2025-v1.docx> [retrieved on 20231228] * |
| J-Y HUO ET AL: "EE2-4.1: Modification of LFNST for MIP coded block", no. JVET-AB0067 ; m60796, 14 October 2022 (2022-10-14), XP030304501, Retrieved from the Internet <URL:https://jvet-experts.org/doc_end_user/documents/28_Mainz/wg11/JVET-AB0067-v2.zip JVET-AB0067-v2.docx> [retrieved on 20221014] * |
| WANG (OPPO) F ET AL: "Non-EE2: On LFNST/NSPT for inter coding", no. JVET-AF0082, 6 October 2023 (2023-10-06), XP030312109, Retrieved from the Internet <URL:https://jvet-experts.org/doc_end_user/documents/32_Hannover/wg11/JVET-AF0082-v1.zip JVET-AF0082-v1.docx> [retrieved on 20231006] * |
| YUNFEI ZHENG (QUALCOMM) ET AL: "CE11: Mode Dependent Coefficient Scanning", no. JCTVC-D393, 26 January 2011 (2011-01-26), XP030226951, Retrieved from the Internet <URL:http://phenix.int-evry.fr/jct/doc_end_user/documents/4_Daegu/wg11/JCTVC-D393-v4.zip JCTVC-D393_r4.doc> [retrieved on 20110126] * |
| ZHANG KAI ET AL: "Intra-Prediction Mode Propagation for Video Coding", IEEE JOURNAL ON EMERGING AND SELECTED TOPICS IN CIRCUITS AND SYSTEMS, IEEE, PISCATAWAY, NJ, USA, vol. 9, no. 1, 1 March 2019 (2019-03-01), pages 110 - 121, XP011714050, ISSN: 2156-3357, [retrieved on 20190308], DOI: 10.1109/JETCAS.2019.2896792 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20220159277A1 (en) | Method and apparatus for video encoding and decoding with subblock based local illumination compensation | |
| US12407817B2 (en) | Intra prediction with geometric partition | |
| EP4289141A1 (en) | Spatial local illumination compensation | |
| US20250392748A1 (en) | Methods and apparatuses for encoding and decoding an image or a video using combined intra modes | |
| US20250047867A1 (en) | Extension of template based intra mode derivation (timd) with isp mode | |
| US20250365419A1 (en) | Methods and apparatuses for encoding/decoding a video | |
| WO2023194105A1 (en) | Intra mode derivation for inter-predicted coding units | |
| WO2021130025A1 (en) | Estimating weighted-prediction parameters | |
| US20230262268A1 (en) | Chroma format dependent quantization matrices for video encoding and decoding | |
| EP4625985A1 (en) | Hybrid explicit/implicit lfnst/nspt | |
| EP4661395A1 (en) | Encoding and decoding methods using multiple transform set selection and corresponding apparatuses | |
| US20260136023A1 (en) | Methods and apparatuses for padding reference samples | |
| EP4668739A1 (en) | Encoding and decoding methods using geometric partition modes and corresponding apparatuses | |
| US20250106428A1 (en) | Methods and apparatuses for encoding/decoding a video | |
| EP4730792A1 (en) | Adaptive selection of transform for spatial geometric partition | |
| WO2026002580A1 (en) | Encoding and decoding methods using multiple motion compensation filters and corresponding apparatuses | |
| WO2025146297A1 (en) | Encoding and decoding methods using intra prediction with sub-partitions and corresponding apparatuses | |
| WO2024083500A1 (en) | Methods and apparatuses for padding reference samples | |
| WO2025114149A1 (en) | Template-based reordering of gpm candidates | |
| WO2025114195A1 (en) | Template-based intra mode derivation (timd) & decoder side intra mode derivation (dimd) | |
| WO2026008385A1 (en) | Dimd and obic histogram adaptation to mip modes | |
| WO2025067883A1 (en) | Deriving a coding mode from one or more template-based costs | |
| WO2025201862A1 (en) | Video coding: coding parameter restrictions | |
| WO2024099962A1 (en) | ENCODING AND DECODING METHODS OF INTRA PREDICTION MODES USING DYNAMIC LISTS OF MOST PROBABLE MODEs AND CORRESPONDING APPARATUSES | |
| WO2026002532A1 (en) | Template based intra mode derivation improvements |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24820765 Country of ref document: EP Kind code of ref document: A1 |