WO2025201110A1 - Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system - Google Patents

Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system

Info

Publication number
WO2025201110A1
WO2025201110A1 PCT/CN2025/083081 CN2025083081W WO2025201110A1 WO 2025201110 A1 WO2025201110 A1 WO 2025201110A1 CN 2025083081 W CN2025083081 W CN 2025083081W WO 2025201110 A1 WO2025201110 A1 WO 2025201110A1
Authority
WO
WIPO (PCT)
Prior art keywords
lfnst
transform
sets
residual data
current block
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/083081
Other languages
French (fr)
Inventor
Chen-Yen LAI
Chih-Hsuan Lo
Chih-Wei Hsu
Ching-Yeh Chen
Tzu-Der Chuang
Yi-Wen Chen
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
MediaTek Inc
Original Assignee
MediaTek Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by MediaTek Inc filed Critical MediaTek Inc
Publication of WO2025201110A1 publication Critical patent/WO2025201110A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/12Selection from among a plurality of transforms or standards, e.g. selection between discrete cosine transform [DCT] and sub-band transform or selection between H.263 and H.264
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/12Selection from among a plurality of transforms or standards, e.g. selection between discrete cosine transform [DCT] and sub-band transform or selection between H.263 and H.264
    • H04N19/122Selection of transform size, e.g. 8x8 or 2x4x8 DCT; Selection of sub-band transforms of varying structure or type
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/132Sampling, masking or truncation of coding units, e.g. adaptive resampling, frame skipping, frame interpolation or high-frequency transform coefficient masking
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/146Data rate or code amount at the encoder output
    • H04N19/147Data rate or code amount at the encoder output according to rate distortion criteria
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/157Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/157Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
    • H04N19/159Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/18Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being a set of transform coefficients
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/70Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards

Definitions

  • the present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63/570,842, filed on March 28, 2024, U.S. Provisional Patent Application No. 63/637,412, filed on April 23, 2024 and U.S. Provisional Patent Application No. 63/710,651, filed on October 23, 2024.
  • the U.S. Provisional Patent Applications are hereby incorporated by reference in their entireties.
  • the present invention relates to video coding system.
  • the present invention relates to designing multiple sets of LFNST or designing different sets of LFNST or NSPT for different prediction modes.
  • VVC Versatile video coding
  • JVET Joint Video Experts Team
  • MPEG ISO/IEC Moving Picture Experts Group
  • ISO/IEC 23090-3 2021
  • Information technology -Coded representation of immersive media -Part 3 Versatile video coding, published Feb. 2021.
  • VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
  • HEVC High Efficiency Video Coding
  • incoming video data undergoes a series of processing in the encoding system.
  • the reconstructed video data from REC 128 may be subject to various impairments due to a series of processing.
  • in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality.
  • deblocking filter (DF) may be used.
  • SAO Sample Adaptive Offset
  • ALF Adaptive Loop Filter
  • the loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream.
  • DF deblocking filter
  • SAO Sample Adaptive Offset
  • ALF Adaptive Loop Filter
  • Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134.
  • the system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
  • HEVC High Efficiency Video Coding
  • the decoder can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126.
  • the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) .
  • the Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140.
  • the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
  • LFNST is applied between forward primary transform and quantization (at encoder) and between de-quantization and inverse primary transform (at decoder side) as shown in Fig. 2.
  • Forward Primary Transform 210 Forward Low-Frequency Non-Separable Transform LFNST 220 is applied to top-left region 222 of the Forward Primary Transform output, for example, 16 coefficients for 4x4 forward LFNST and/or 64 coefficients for 8x8 forward LFNST.
  • 4x4 non-separable transform or 8x8 non-separable transform is applied according to block size.
  • 4x4 LFNST is applied for small blocks (i.e., min (width, height) ⁇ 8) and 8x8 LFNST is applied for larger blocks (i.e., min (width, height) > 4) .
  • the transform coefficients are quantized by Quantization 230.
  • the quantized transform coefficients are de-quantized using De-Quantization 240 to obtain the de-quantized transform coefficients.
  • Inverse LFNST 250 is applied to the top-left region 252 (8 coefficients for 4x4 inverse LFNST or 16 coefficients for 8x8 inverse LFNST) .
  • inverse Primary Transform 260 is applied to recover the input signal.
  • the non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16x16 transform matrix.
  • T is a 16x16 transform matrix.
  • the 16x1 coefficient vector is subsequently re-organized as 4x4 block using the scanning order for that block (horizontal, vertical or diagonal) .
  • the coefficients with smaller index will be placed with the smaller scanning index in the 4x4 coefficient block.
  • LFNST is based on direct matrix multiplication approach to apply non-separable transform so that it is implemented in a single pass without multiple iterations.
  • the non-separable transform matrix dimension needs to be reduced to minimize computational complexity and memory space to store the transform coefficients.
  • reduced non-separable transform (or RST) method is used in LFNST.
  • the main idea of the reduced non-separable transform is to map an N (N is commonly equal to 64 for 8x8 NSST) dimensional vector to an R dimensional vector in a different space, where N/R (R ⁇ N) is the reduction factor.
  • NxN matrix instead of NxN matrix, RST matrix becomes an R ⁇ N matrix as follows: where the R rows of the transform are R bases of the N dimensional space.
  • the inverse transform matrix for RT is the transpose of its forward transform.
  • 8x8 LFNST a reduction factor of 4 is applied, and 64x64 direct matrix, which is conventional 8x8 non-separable transform matrix size, is reduced to16x48 direct matrix.
  • the 48 ⁇ 16 inverse RST matrix is used at the decoder side to generate core (primary) transform coefficients in 8 ⁇ 8 top-left regions.
  • 16x48 matrices are applied instead of 16x64 with the same transform set configuration, each of which takes 48 input data from three 4x4 blocks in a top-left 8x8 block excluding right-bottom 4x4 block.
  • LFNST index coding depends on the position of the last significant coefficient.
  • the LFNST index is context coded, but does not depend on intra prediction mode, and only the first bin is context coded.
  • LFNST is applied for intra CU in both intra and inter slices, and for both luma and chroma. If a dual tree is enabled, LFNST indices for luma and chroma are signalled separately. For inter slice (the dual tree is disabled) , a single LFNST index is signalled and used for both luma and chroma.
  • an LFNST index search could increase data buffering by four times for a certain number of decode pipeline stages. Therefore, the maximum size that LFNST is allowed is restricted to 64x64. Note that LFNST is enabled with DCT2 only. The LFNST index signalling is placed before MTS index signalling.
  • the forward LFNST is applied to top-left low frequency region, which is called Region-Of-Interest (ROI) .
  • ROI Region-Of-Interest
  • the ROI for LFNST16 is depicted in Fig. 3A. It consists of six 4x4 sub-blocks, which are consecutive in scan order. Since the number of input samples is 96, transform matrix for forward LFNST16 can be Rx96. R is chosen to be 32 in this case, 32 coefficients (two 4x4 sub-blocks) are generated from forward LFNST16 accordingly, which are placed following the coefficient scan order.
  • the ROI for LFNST8 is shown in Fig. 3B.
  • the forward LFNST8 matrix can be Rx64 and R is chosen to be 32.
  • the generated coefficients are located in the same manner as with LFNST16.
  • weights which are defined for a block shape and intra mode, is introduced, those weights are multiplied by the neighbouring reference template to derive the prediction samples replacing conventional intra prediction.
  • the weights are applied to the reference samples of the L shaped causal neighbouring template as shown in the Fig. 4.
  • the reference samples in the causal neighbourhood are denoted as r, and F (x, y) is the matrix of weights.
  • the prediction is used for block size with both width and height up to 32 (except for 4x32, 32x4, 8x32 and 32x8) .
  • the template size is 2 for blocks with both width and height up to 16 and it is only used for mode 0, 1, and (2+2*k) .
  • template size is set to 1; is used for mode 0, 1, and (2+4*k) ; prediction is only performed for 16x16 positions, and the rest of the samples are generated by bilinear interpolation.
  • block shape and mode-based symmetry is used for Reference length is set to W and H for modes greater than 18 and less than 50 and set to 2*W and 2*H otherwise.
  • JVET-AI0050 it is proposed to enable LFNST/NSPT for SBT-coded blocks.
  • a CU-level LFNST/NSPT index is signalled for an SBT-coded block to indicate the usage of LFNST/NSPT.
  • LFNST/NSPT is applied to an SBT-coded CU, the TU with non-zero residual will perform LFNST/NSPT, and the prediction signal within the TU region is used to derive the IPM for transform.
  • All NSPTs consist of 35 sets and 3 candidates (similar to the current LFNST) .
  • the kernels of NSPTs have the following shapes: ⁇ NSPT4x4: 16x16 ⁇ NSPT4x8/NSPT8x4: 32x20 ⁇ NSPT8x8: 64x32 ⁇ NSPT4x16/NSPT16x4: 64x24 ⁇ NSPT8x16/NSPT16x8: 128x40 ⁇ NSPT4x32/NSPT32x4: 128x20 ⁇ NSPT8x32/NSPT32x8: 256x24.
  • JVET-AI0163 proposes MTSS method to allow CUs coded with intra modes to select one LFNST/NSPT transform set out of two candidate sets.
  • the current block is coded with intra modes in combination with LFNST/NSPT, one more bin is employed to indicate whether the first or the second candidate transform set is selected.
  • LFNST/NSPT can be applied to inter-coded blocks according to JVET-AI0050, where the transform kernels for intra LFNST/NSPT are reused by inter blocks.
  • a Histogram of Gradients (HoG) is built in a similar way as DIMD (Decoder-side Intra Mode Derivation) , and the first DIMD intra prediction mode (IPM) corresponding to the highest amplitude is used to determine the LFNST/NSPT kernel set.
  • DIMD Decoder-side Intra Mode Derivation
  • LFNST/NSPT for SBT-coded blocks.
  • a CU-level LFNST/NSPT index is signalled for an SBT-coded block to indicate the usage of LFNST/NSPT.
  • LFNST/NSPT is applied to an SBT-coded CU, the TU with non-zero residual will perform LFNST/NSPT, and the prediction signal within the TU region is used to derive the IPM for transform.
  • JVET-M0425 In the multi-hypothesis inter prediction mode (JVET-M0425) , one or more additional motion-compensated prediction signals are signalled, in addition to the conventional bi-prediction signal.
  • the resulting overall prediction signal is obtained by sample-wise weighted superposition.
  • the weighting factor ⁇ is specified by the new syntax element add_hyp_weight_idx, according to the following mapping: add_hyp_weight_idx ⁇ 0 1/4 1 -1/8.
  • the motion parameters of each additional prediction hypothesis can be signalled either explicitly by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly by specifying a merge index.
  • a separate multi-hypothesis merge flag distinguishes between these two signalling modes.
  • MHP is only applied if non-equal weight in BCW is selected in bi-prediction mode.
  • a geometric partition index indicating the partition mode of the geometric partition i.e., angle and offset
  • two merge indices one for each partition
  • the number of maximum GPM candidate size is signalled explicitly in SPS and specifies syntax binarization for GPM merge indices.
  • a method and apparatus for video decoding for using LFNST are disclosed. According to this method, input data associated with a current block is received, wherein the input data comprises transformed residual data associated with the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined. One or more target sets of LFNST is selected from said two or more sets of LFNST. Inverse transform comprising said one or more target sets of LFNST is applied to the transformed residual data to derive reconstructed residual data. The reconstructed residual data is provided.
  • LFNST Low-Frequency Non-Separable Transform
  • a first set of LFNST is derived based on a DIMD (Decoder-side Intra Mode Derivation) -related scheme and one or more additional sets of LFNST are derived by adding one or more predefined offsets to the first set of LFNST.
  • DIMD Decoder-side Intra Mode Derivation
  • signalling of syntax related to selection of said two or more sets of LFNST is constrained by sum of absolute transform coefficients of the current block.
  • said selecting said one or more target sets of LFNST from said two or more sets of LFNST is conditioned by one or more neighbouring blocks.
  • said two or more sets of LFNST are determined by taking into account of boundary matching (BM) cost calculated between current prediction samples and neighbouring reconstruction samples in a current frame, or reference block samples and the reconstruction samples adjacent to a reference block in a reference frame.
  • BM boundary matching
  • a method and apparatus of designing different multiple sets of LFNST are also disclosed.
  • input data associated with a current block is received, wherein the input data comprises transformed residual data associated with the current block, wherein the current block is coded in a target prediction mode.
  • One or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) are derived based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique.
  • a target set of LFNST or NSPT is selected from said one or more sets of LFNST or NSPT for the target prediction mode.
  • Inverse transform comprising the target set of LFNST is applied to the transformed residual data to derive reconstructed residual data.
  • the reconstructed residual data is provided.
  • the sub-sample technique is only applied when the current block is larger than a pre-defined threshold.
  • Fig. 1A illustrates an exemplary adaptive Inter/Intra video encoding system incorporating loop processing.
  • Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
  • Fig. 2 illustrates an example of Low-Frequency Non-Separable Transform (LFNST) process.
  • Fig. 3A illustrates an example of Region-Of-Interest (ROI) for LFNST16.
  • ROI Region-Of-Interest
  • Fig. 3B illustrates an example of Region-Of-Interest (ROI) for LFNST8.
  • ROI Region-Of-Interest
  • Fig. 6 illustrates an example of boundary samples used for boundary matching cost calculation.
  • Fig. 7A illustrates an example of boundary matching cost calculated in 4x1 or 1x4 blocks of the prediction and neighbouring reconstruction samples in the current frame.
  • Fig. 7B illustrates an example of boundary matching cost calculated in 4x1 or 1x4 blocks of the prediction and neighbouring reconstruction samples in the reference frame.
  • Fig. 8 illustrates a flowchart of an exemplary video decoding system that derives multiple sets of LFNST and selecting one target set of LFNST to code a block according to an embodiment of the present invention.
  • Fig. 9 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 8.
  • Fig. 10 illustrates a flowchart of an exemplary video decoding system that signals syntax to select a target set of LFNST with constraint by a sum of absolute transform coefficients of the current block according to an embodiment of the present invention.
  • Fig. 11 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 10.
  • Fig. 12 illustrates a flowchart of an exemplary video decoding system that derives one or more sets of LFNST or NSPT for different prediction modes according to an embodiment of the present invention.
  • Fig. 13 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 12.
  • LFNST can be tested with more than one transform set.
  • the transform sets used for LFNST are determined based on DIMD related method. In that, a histogram is generated, and the top N intra modes with higher amplitude values can be used to derive LFNST transform sets for LFNST. N can be any integer larger than 0.
  • a flag is signalled in the bitstream to the decoder to indicate whether the second transform set is used for LFNST. If the flag is false, the first transform set is used. Otherwise, the second transform set is used.
  • the first transform set can be derived based on DIMD related method.
  • the second transform set can be derived by adding a predefined offset to the first transform set.
  • the first transform set is set 4.
  • the second transform set will be set 5 (4 + 1) , set 3 (4 -1) , or set 21 (4 + 17) .
  • the first transform set can be derived based on DIMD related method.
  • the intraMode with the highest amplitude value will be used to derive LFNST transform set.
  • the second transform set is derived by using the intraMode with the second highest amplitude value.
  • the first transform set can be derived based on DIMD related method.
  • the intraMode1 with the highest amplitude value will be used to derive LFNST transform set.
  • the second transform set is derived by using intraMode2 with the second highest amplitude value.
  • the distance between intraMode1 and intraMode2 shall be larger than TH.
  • TH can be any integer.
  • two flags are signalled in the bitstream.
  • the first flag is used to indicate whether the main transform set is used.
  • the second flag is used to indicate which additional transform set is used. Only if the first flag is false, the second flag needs to be signalled.
  • the derived transform sets of intraMode1, and intraMode2 are constrained to be different.
  • the derived transform sets of intraMode1, intraMode2 and intraMode3 are constrained to be different.
  • the first bin is signalled to the decoder in the bitstream to indicate whether the additional transform set is used for LFNST.
  • the first bin can be coded by a context coded bin.
  • some truncated unary coded bins are used to indicate the selected LFNST transform set.
  • the truncated unary coded bins will only be sent if the bin used to indicate whether the additional transform set is used for LFNST to decoder is true. Or they will only be sent if the bin used to indicate whether the main transform set is used for LFNST to decoder is false.
  • the signalling of syntax related to LFNST transform set can be constrained by the absolute sum of transform coefficients of the current CU.
  • the signalling of syntax related to LFNST transform set can be constrained by the number of non-zero transform coefficients of current CU. For example, if the number of non-zero transform coefficients of current CU is smaller than a predefined threshold, the syntax related to LFNST transform set will not be signalled.
  • the predefined thresholds for LFNST transform set signalling for intra LFNST and inter LFNST can be different.
  • the predefined thresholds for LFNST transform set signalling for intra LFNST and inter LFNST can be the same.
  • the signalling of syntax related to LFNST transform set can be constrained by CU size, prediction mode, slice-type or QP value.
  • one on/off control flag is signalled at CU level, slice level, picture level, and/or sequence level to indicate the proposed method in the above is enabled or not.
  • more than one predefined threshold can be used for selecting LFNST/NSPT transform set. For example, if the number of non-zero quantized coefficients of current block is larger than threshold 1, three different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream. If the number of non-zero quantized coefficients of current block is larger than threshold 2, five different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream.
  • the syntax design for signalling best LFNST/NSPT transform set can be different for different CUs, and it is based on the number of non-zero quantized coefficients to choose the syntax design.
  • the alternative tested transform sets are the neighbouring transform sets of the main transform set. For example, if the main transform set is set 4. Two alternative tested transform sets are set 3 (4 –1) and set 5 (4 + 1) .
  • the total number of transform sets tested in encoder is 3 for each CU.
  • the alternative tested transform sets are the neighbouring transform sets of the main transform set. For example, if the main transform set is set 4. Two alternative tested transform sets are set 3 (4 –1) and set 5 (4 + 1) .
  • the alternative tested transform sets are derived by DIMD histogram. The first two intramodes with highest amplitudes will be used to derive the alternative tested transform sets.
  • the alternative transform set derivation for LFNST/NSPT can be different from different prediction modes.
  • the maximum number of transform sets signalled to the decoder is dependent by prediction mode.
  • the maximum number of transform sets signalled to the decoder is N for fusion related predicted CUs.
  • the maximum number of transform sets signalled to decoder is M for non-fusion related predicted CUs.
  • N and M can be any integer larger than 0.
  • N is larger than M.
  • N is smaller than M.
  • the alternative transform set derivation for LFNST/NSPT can be different from different prediction modes.
  • the alternative tested transform sets are derived by planar or DC mode.
  • the maximum number of transform set signalled to the decoder is N for CU area smaller than TH.
  • the maximum number of transform set signalled to the decoder is M for other CUs.
  • N and M can be any integer larger than 0.
  • TH can be any pre-defined integer value.
  • the alternative transform set derivation for matrix based intra prediction mode mentioned in Section I. 3 can be aligned with the replaced intra mode or determined by DIMD method or derived by planar or dc mode.
  • the main transform set of replaced intra mode is set 4
  • the main transform set of the matrix-based intra prediction mode is also set 4 and the two alternative tested transform sets are set 3 (4 –1) and set 5 (4 + 1) .
  • the main transform set of the matrix-based intra prediction mode is set 4 (e.g., can be determined by replaced intra mode or DIMD histogram) and the two alternative tested transform sets are derived by DIMD histogram.
  • the main transform set of the matrix-based intra prediction mode is set 4 (e.g., can be determined by replaced intra mode or DIMD histogram) and the two alternative tested transform sets are derived by planar or dc mode.
  • the syntax related to LFNST/NSPT kernel index will only be signalled when some LFNST transform sets are used.
  • the pre-defined kernel index can be different from luma LFNST and chroma LFNST, but not limit to.
  • the pre-defined kernel index can be different from fusion related prediction mode and non-fusion related prediction mode, but not limit to.
  • the pre-defined kernel index can be different for different CU size, but not limit to.
  • first kernel will be used.
  • second kernel will be used.
  • the first kernel will be used. Otherwise, second kernel will be used.
  • EP bins will be used to code the kernel index.
  • non-alternative transform set context coded bins will be used to code the kernel index.
  • the histogram used to derive the transform set of LFNST and NSPT is derived by collecting the statistic information of neighbouring reconstruction samples.
  • the statistic information includes x-axis and y-axis gradients.
  • the histogram used to derive the transform set of LFNST and NSPT is derived by collecting the statistic information current predictors.
  • the statistic information includes x-axis and y-axis gradients.
  • the derivation of histogram used for inter LFNST/NSPT and intra LFNST/NSPT can be different, but not limit to.
  • the histogram used for inter LFNST/NSPT is derived by using the statistic information of current predictors.
  • the histogram used for intra LFNST/NSPT is derived by using neighbouring reconstruction samples.
  • the derivation of histogram used for LFNST/NSPT can be determined by statistical analysis. For example, if the histogram derived by neighbouring reconstruction samples is very different from the histogram derived by current predictors, the histogram derived by current predictors will be used to derive LFNST/NSPT transform set. For example, if the histogram derived by neighbouring reconstruction samples is very different from the histogram derived by current predictors, the histogram derived by neighbouring reconstruction samples will be used to derive LFNST/NSPT transform set.
  • one on/off control flag is signalled at CU level, slice level, picture level, and/or sequence level to indicate the proposed method in the above is enabled or not.
  • neighbouring reconstruction samples or the current predictor will be used to derive DIMD histogram for LFNST/NSPT transform set selection for fusion modes and non-fusion mode. Otherwise, for intra blocks, the neighbouring reconstruction samples will always be used to derive DIMD histogram for LFNST/NSPT transform set selection. For inter blocks, the current predictor will always be used to derive DIMD histogram for LFNST/NSPT transform set selection.
  • the selection of using neighbouring reconstruction samples or the current predictor to derive DIMD histogram for LFNST/NSPT transform set derivation is signalled in the bitstream. It can be signalled at CU level, slice level, picture level, and/or sequence level.
  • the current predictors mentioned above are the final predictors after blending if the prediction mode of current CU is a fusion mode. In this case, more than one predictor will be blended to generate a blended final predictor.
  • sub-sample technology can be applied to derive the DIMD histogram for LFNST/NSPT transform set derivation.
  • current predictors i.e., final predictors after blending
  • Step number is set to be 2.Every two samples in each row and each column are collected to do the statistical analysis. In that, the step number can be designed based on current CU size.
  • the step number for row and column can be different, but not limit to.
  • two DIMD histograms can be derived by using predictor L0 and predictor L1.
  • Two DIMD histograms are referenced to derive LFNST/NSPT transform set.
  • the histogram derived based on predictor L0 is used to derived the second transform set.
  • the histogram derived based on predictor L1 is used to derive the third transform set.
  • the intramode of previously coded intra CU can be used to derive the LFNST/NSPT transform set of current CU.
  • the distance between the previously coded intra CU and the current CU can be constrained, but not limit to. For example, only if the distance between the top-left position of previously coded intra CU and the current CU is smaller than a threshold, the intramode of previously coded intra CU can be referenced by the current block.
  • a pre-defined intramode is used (i.e., planar mode) if the previously coded intra CU cannot be referenced.
  • the gradients at the boundary of current inter prediction samples and neighbouring reconstruction samples) , QP value of current block, QP value of reference blocks, transform block size, GPM partition directions, BCW weights, absolute sum of quantized coefficients, number of non-zero quantized coefficients, MV amplitudes, sum of gradient amplitudes of current predictor, and/or gradient directions of current predictor can be used to determine the LFSNT/NSPT kernel set.
  • the LFNST/NSPT kernel set can be determined by considering the boundary matching (BM) cost calculated between current prediction samples and neighbouring reconstruction samples on the current frame, or reference block samples and the reconstruction samples adjacent to reference block on the reference frame.
  • the boundary matching cost is calculated as follows: where R x, -n indicates the neighbouring samples n column left to the prediction samples P x, 0 , and R y, -n indicates the neighboring samples n row above prediction samples P y, 0 As shown in Fig. 6.
  • the boundary matching cost can be calculated in sample-based, block-based, or CU-based. In one example, the boundary matching cost is calculated in 4x1 or 1x4 block-based by the prediction and neighbouring reconstruction samples on the current frame as shown in Fig. 7A.
  • the boundary matching cost is calculated in 4x1 or 1x4 block-based by the reference block samples and the reconstruction samples adjacent to reference block on reference frame as shown in Fig. 7B.
  • the LFNST/NSPT kernel set can be determined by considering the QP value of current block or the QP values of reference blocks.
  • QP thresholds are pre-defined or determined adaptively according to base QP, picture resolution, transform block size, or other coding information.
  • the first LFNST/NSPT kernel set is used.
  • the second LFNST/NSPT kernel set is used.
  • the (i+1) -th LFNST/NSPT kernel set is used.
  • the minimum or maximum QP value of the QP values of reference blocks or the QP value of the reference block with higher prediction blending weight or the QP value selected by the predefined method is set as the selected QP value.
  • the first LFNST/NSPT kernel set is used for the blocks with the selected QP value smaller than or equal to the first threshold.
  • the second LFNST/NSPT kernel set is used for the blocks with the selected QP value larger than first threshold and smaller than or equal to the second threshold.
  • the (i+1) -th LFNST/NSPT kernel set is used.
  • the LFNST/NSPT kernel set is determined by the prediction mode of current CU.
  • the GPM partition modes can be used to determine the kernel sets for GPM-coded CUs.
  • the CUs coded by subblock modes e.g. affine, SBTMVP, DMVR modes
  • the CUs coded by multi-hypothesis coding modes e.g. bi-predictive inter mode, CIIP, GPM, MHP, DIMD, TIMD, etc.
  • each inter/intra mode can use a specific kernel set.
  • the second LFNST/NSPT kernel set is used.
  • the (i+1) -th LFNST/NSPT kernel set is used.
  • the gradients of current predictor are calculated in sample-based or block-based scheme.
  • a histogram is derived by collecting the calculated gradients and the gradient direction with highest peak can be used to determine the LFNST/NSPT kernel set.
  • the inter LFNST/NSPT kernel set can be inherited from the reference blocks or neighbouring blocks. That is, the inter LFNST/NSPT kernel set of current block is set as one of the kernel sets of reference blocks or neighbouring blocks.
  • an LFNST/NSPT kernel set list is constructed from neighbouring blocks.
  • An index is signalled to determine the kernel set used by current CU or a predefined selection method (e.g. generating the templates at the CU boundary by different kernel sets and using BM costs to find the kernel sets with higher accuracy) is applied to select the kernel set used by current CU.
  • Fig. 9 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 8.
  • residual data associated with a current block is received in step 910, wherein the residual data is generated by applying prediction to the current block.
  • Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 920.
  • One or more target sets of LFNST are selected from said two or more sets of LFNST in step 930.
  • Transform comprising said one or more target sets of LFNST is applied to the residual data to generate transformed data in step 940.
  • the transformed data is provided in step 950.
  • Fig. 10 illustrates a flowchart of an exemplary video decoding system that signals syntax to select a target set of LFNST with constraint by a sum of absolute transform coefficients of the current block according to an embodiment of the present invention.
  • input data associated with a current block is received in step 1010, wherein the input data comprises transformed residual data associated with the current block.
  • Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 1020.
  • One or more syntax elements related to selecting a target set of LFNST are parsed from said two or more sets of LFNST in step 1030, wherein said parsing said one or more syntax elements is constrained by a sum of absolute transform coefficients of the current block.
  • Inverse transform comprising the target set of LFNST is applied to the transformed residual data to derive reconstructed residual data in step 1040.
  • the reconstructed residual data is provided in step 1050.
  • Fig. 11 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 10.
  • residual data associated with a current block is received in step 1110, wherein the residual data is generated by applying prediction to the current block.
  • Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 1120.
  • One or more syntax elements related to selecting a target set of LFNST from said two or more sets of LFNST are signalled in step 1130, wherein said signalling said one or more syntax elements is constrained by sum of absolute transform coefficients of the current block.
  • Transform comprising the target set of LFNST is applied to the residual data to generate transformed data in step 1140.
  • the transformed data is provided in step 1150.
  • Fig. 13 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 12.
  • data associated with a current block is received in step 1310, wherein the residual data is generated by applying prediction to the current block, wherein the current block is coded in a target prediction mode.
  • One or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) are derived based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique in step 1320.
  • DIMD Decoder-side Intra Mode Derivation

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Discrete Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

A method and apparatus for video coding using Low-Frequency Non-Separable Transform (LFNST) or NSPT are disclosed. According to one method for the decoder side, two or more sets of LFNST are determined. A target set of LFNST is selected from two or more sets of LFNST. Inverse transform comprising the target set of LFNST is applied to the transformed residual data to derive reconstructed residual data. In another method for the decoder side, different multiple sets of LFNST or Non-Separable Primary Transform (NSPT) for different prediction modes are derived. Target multiple sets of LFNST or NSPT are selected from said different multiple sets of LFNST or NSPT for the target prediction mode. Inverse transform comprising the target multiple sets of LFNST to the transformed residual data is applied to derive reconstructed residual data.

Description

METHODS AND APPARATUS FOR LOW-FREQUENCY NON-SEPARABLE TRANSFORM WITH MULTIPLE TRANSFORM SETS IN A VIDEO CODING SYSTEM
CROSS REFERENCE TO RELATED APPLICATIONS
The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63/570,842, filed on March 28, 2024, U.S. Provisional Patent Application No. 63/637,412, filed on April 23, 2024 and U.S. Provisional Patent Application No. 63/710,651, filed on October 23, 2024. The U.S. Provisional Patent Applications are hereby incorporated by reference in their entireties.
FIELD OF THE INVENTION
The present invention relates to video coding system. In particular, the present invention relates to designing multiple sets of LFNST or designing different sets of LFNST or NSPT for different prediction modes.
BACKGROUND
Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO/IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
Fig. 1A illustrates an exemplary adaptive Inter/Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
I. 1 Low-Frequency Non-Separable Transform (LFNST)
In VVC, LFNST is applied between forward primary transform and quantization (at encoder) and between de-quantization and inverse primary transform (at decoder side) as shown in Fig. 2. As shown in Fig. 2, after Forward Primary Transform 210, Forward Low-Frequency Non-Separable Transform LFNST 220 is applied to top-left region 222 of the Forward Primary Transform output, for example, 16 coefficients for 4x4 forward LFNST and/or 64 coefficients for 8x8 forward LFNST. In LFNST, 4x4 non-separable transform or 8x8 non-separable transform is applied according to block size. For example, 4x4 LFNST is applied for small blocks (i.e., min (width, height) < 8) and 8x8 LFNST is applied for larger blocks (i.e., min (width, height) > 4) . After LFNST, the transform coefficients are quantized by Quantization 230. To reconstruct the input signal, the quantized transform coefficients are de-quantized using De-Quantization 240 to obtain the de-quantized transform coefficients. Inverse LFNST 250 is applied to the top-left region 252 (8 coefficients for 4x4 inverse LFNST or 16 coefficients for 8x8 inverse LFNST) . After invers LFNST, inverse Primary Transform 260 is applied to recover the input signal.
Application of a non-separable transform, which is being used in LFNST, is described as follows using input as an example. To apply 4x4 LFNST, the 4x4 input block X,

is first represented as a vector
The non-separable transform is calculated aswhereindicates the transform coefficient vector, and T is a 16x16 transform matrix. The 16x1 coefficient vectoris subsequently re-organized as 4x4 block using the scanning order for that block (horizontal, vertical or diagonal) . The coefficients with smaller index will be placed with the smaller scanning index in the 4x4 coefficient block.
I. 1.1 Reduced non-separable transform
LFNST is based on direct matrix multiplication approach to apply non-separable transform so that it is implemented in a single pass without multiple iterations. However, the non-separable transform matrix dimension needs to be reduced to minimize computational complexity and memory space to store the transform coefficients. Hence, reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N (N is commonly equal to 64 for 8x8 NSST) dimensional vector to an R dimensional vector in a different space, where N/R (R < N) is the reduction factor. Hence, instead of NxN matrix, RST matrix becomes an R×N matrix as follows:

where the R rows of the transform are R bases of the N dimensional space.
The inverse transform matrix for RT is the transpose of its forward transform. For 8x8 LFNST, a reduction factor of 4 is applied, and 64x64 direct matrix, which is conventional 8x8 non-separable transform matrix size, is reduced to16x48 direct matrix. Hence, the 48×16 inverse RST matrix is used at the decoder side to generate core (primary) transform coefficients in 8×8 top-left regions. When16x48 matrices are applied instead of 16x64 with the same transform set configuration, each of which takes 48 input data from three 4x4 blocks in a top-left 8x8 block excluding right-bottom 4x4 block. With the help of the reduced dimension, memory usage for storing all LFNST matrices is reduced from 10KB to 8KB with reasonable performance drop. In order to reduce the complexity, LFNST is restricted to be applicable only if all coefficients outside the first coefficient sub-group are non-significant. Hence, all primary-only transform coefficients have to be zero when LFNST is applied. This allows conditioning the LFNST index signalling on the last-significant position, and hence avoids the extra coefficient scanning in the current LFNST design, which is needed for checking for significant coefficients at specific positions only.
The worst-case handling of LFNST (in terms of multiplications per pixel) restricts the non-separable transforms for 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In those cases, the last-significant scan position has to be less than 8 when LFNST is applied, for other sizes less than 16. For blocks with a shape of 4xN and Nx4 and N > 8, the proposed restriction implies that the LFNST is now applied only once, and to the top-left 4x4 region only. As all primary-only coefficients are zero when LFNST is applied, the number of operations needed for the primary transforms is reduced in such cases. From encoder perspective, the quantization of coefficients is remarkably simplified when LFNST transforms are tested. A rate-distortion optimized quantization has to be done at maximum for the first 16 coefficients (in scan order) , the remaining coefficients are enforced to be zero.
I. 1.2 LFNST transform selection
There are 4 transform sets and 2 non-separable transform matrices (kernels) in total per transform set are used in LFNST. The mapping from the intra prediction mode to the transform set is pre-defined as shown in Table 1. If one of three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM or INTRA_L_CCLM) is used for the current block (i.e., 81 <= predModeIntra <= 83) , transform set 0 is selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate is further specified by the explicitly signalled LFNST index. The index is signalled in a bit-stream once per Intra CU after transform coefficients.
Table 1. Transform selection table
I. 1.3 LFNST index signalling and interaction with other tools
Since LFNST is restricted to be applicable only if all coefficients outside the first coefficient sub-group are non-significant, LFNST index coding depends on the position of the last significant coefficient. In addition, the LFNST index is context coded, but does not depend on intra prediction mode, and only the first bin is context coded. Furthermore, LFNST is applied for intra CU in both intra and inter slices, and for both luma and chroma. If a dual tree is enabled, LFNST indices for luma and chroma are signalled separately. For inter slice (the dual tree is disabled) , a single LFNST index is signalled and used for both luma and chroma.
Considering that a large CU greater than 64x64 is implicitly split (TU tiling) due to the existing maximum transform size restriction (64x64) , an LFNST index search could increase data buffering by four times for a certain number of decode pipeline stages. Therefore, the maximum size that LFNST is allowed is restricted to 64x64. Note that LFNST is enabled with DCT2 only. The LFNST index signalling is placed before MTS index signalling.
The use of scaling matrices for perceptual quantization is not evident that the scaling matrices that are specified for the primary matrices may be useful for LFNST coefficients. Hence, the use of the scaling matrices for LFNST coefficients is not allowed. For single-tree partition mode, chroma LFNST is not applied.
I. 2 Secondary Transformation: LFNST Extension with Large Kernel
The LFNST design in VVC is extended as follows:
● The number of LFNST sets (S) and candidates (C) are extended to S=35 and C=3, and the LFNST set 
(lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula:
○ For predModeIntra < 2, lfnstTrSetIdx is equal to 2
○ lfnstTrSetIdx = predModeIntra, for predModeIntra in [0, 34]
○ lfnstTrSetIdx = 68 –predModeIntra, for predModeIntra in [35, 66]
● Three different kernels, LFNST4, LFNST8, and LFNST16, are defined to indicate LFNST kernel sets, 
which are applied to 4xN/Nx4 (N≥4) , 8xN/Nx8 (N≥8) , and MxN (M, N≥16) , respectively.
The kernel dimensions are specified by:
(LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96) .
The forward LFNST is applied to top-left low frequency region, which is called Region-Of-Interest (ROI) . When LFNST is applied, primary-transformed coefficients that exist in the region other than ROI are zeroed out, which is not changed from the VVC standard.
The ROI for LFNST16 is depicted in Fig. 3A. It consists of six 4x4 sub-blocks, which are consecutive in scan order. Since the number of input samples is 96, transform matrix for forward LFNST16 can be Rx96. R is chosen to be 32 in this case, 32 coefficients (two 4x4 sub-blocks) are generated from forward LFNST16 accordingly, which are placed following the coefficient scan order.
The ROI for LFNST8 is shown in Fig. 3B. The forward LFNST8 matrix can be Rx64 and R is chosen to be 32. The generated coefficients are located in the same manner as with LFNST16.
The mapping from intra prediction modes to these sets is shown in Table 2.
Table 2. Mapping of intra prediction modes to LFNST set index
I. 3 Matrix Based Intra Prediction Replacing Conventional Intra Modes
A matrix of weights, which are defined for a block shape and intra mode, is introduced, those weights are multiplied by the neighbouring reference template to derive the prediction samples replacing conventional intra prediction. The weights are applied to the reference samples of the L shaped causal neighbouring template as shown in the Fig. 4.
The reference samples in the causal neighbourhood are denoted as r, and F (x, y) is the matrix of weights. Then the prediction P (x, y) can be derived as:
P (x, y) = ∑k F (x, y, k) *r (k) ,
where k denotes the index of the reference sample in the template.
The prediction is used for block size with both width and height up to 32 (except for 4x32, 32x4, 8x32 and 32x8) . The template size is 2 for blocks with both width and height up to 16 and it is only used for mode 0, 1, and (2+2*k) . For other blocks, template size is set to 1; is used for mode 0, 1, and (2+4*k) ; prediction is only performed for 16x16 positions, and the rest of the samples are generated by bilinear interpolation. For all block sizes, block shape and mode-based symmetry is used. Reference length is set to W and H for modes greater than 18 and less than 50 and set to 2*W and 2*H otherwise.
I. 4 LFNST/NSPT for SBT-Coded Blocks
In JVET-AI0050, it is proposed to enable LFNST/NSPT for SBT-coded blocks. A CU-level LFNST/NSPT index is signalled for an SBT-coded block to indicate the usage of LFNST/NSPT. When LFNST/NSPT is applied to an SBT-coded CU, the TU with non-zero residual will perform LFNST/NSPT, and the prediction signal within the TU region is used to derive the IPM for transform.
I. 5 Non-Separable Primary Transform (NSPT) for Intra Coding
The separable DCT-II plus LFNST transform combinations are replaced with NSPT for the block shapes 4x4, 4x8, 8x4 and 8x8, 4x16, 16x4, 8x16 and 16x8.
The affected block sizes are summarized in Fig. 5.
All NSPTs consist of 35 sets and 3 candidates (similar to the current LFNST) . The kernels of NSPTs have the following shapes:
● NSPT4x4: 16x16
● NSPT4x8/NSPT8x4: 32x20
● NSPT8x8: 64x32
● NSPT4x16/NSPT16x4: 64x24
● NSPT8x16/NSPT16x8: 128x40
● NSPT4x32/NSPT32x4: 128x20
● NSPT8x32/NSPT32x8: 256x24.
Therefore, 12, 32, 40 and 88 coefficients are zeroed-out using NSPT4x8/NSPT8x4, NSPT8x8, NSPT4x16/NSPT16x4 and NSPT8x16/NSPT16x8 respectively. For NSPT4x32/NSPT32x4 and NSPT8x32/NSPT32x8, remaining 108 and 232 positions in each transform block are zeroed-out, respectively.
I. 6 Multiple Transform Set Selection (MTSS) for Intra LFNST/NSPT
In JVET-AI0163, it proposes MTSS method to allow CUs coded with intra modes to select one LFNST/NSPT transform set out of two candidate sets. When the current block is coded with intra modes in combination with LFNST/NSPT, one more bin is employed to indicate whether the first or the second candidate transform set is selected.
I. 7 Improvements on Inter LFNST/NSPT
LFNST/NSPT can be applied to inter-coded blocks according to JVET-AI0050, where the transform kernels for intra LFNST/NSPT are reused by inter blocks. Based on the prediction signal of the current block, a Histogram of Gradients (HoG) is built in a similar way as DIMD (Decoder-side Intra Mode Derivation) , and the first DIMD intra prediction mode (IPM) corresponding to the highest amplitude is used to determine the LFNST/NSPT kernel set.
For a GPM-coded block, an alternative IPM is derived for inter LFNST/NSPT. In addition to the first DIMD IPM, a second DIMD IPM corresponding to the second highest HOG amplitude is used as an additional IPM candidate. A CU-level flag indicating the IPM index is signalled.
It is proposed to enable LFNST/NSPT for SBT-coded blocks. A CU-level LFNST/NSPT index is signalled for an SBT-coded block to indicate the usage of LFNST/NSPT. When LFNST/NSPT is applied to an SBT-coded CU, the TU with non-zero residual will perform LFNST/NSPT, and the prediction signal within the TU region is used to derive the IPM for transform.
I. 8 Multi-Hypothesis Prediction (MHP)
In the multi-hypothesis inter prediction mode (JVET-M0425) , one or more additional motion-compensated prediction signals are signalled, in addition to the conventional bi-prediction signal. The resulting overall prediction signal is obtained by sample-wise weighted superposition. With the bi-prediction signal p_bi and the first additional inter prediction signal/hypothesis h_3, the resulting prediction signal p_3 is obtained as follows:
p_3= (1-α) p_bi+αh_3.
The weighting factor α is specified by the new syntax element add_hyp_weight_idx, according to the following mapping:
add_hyp_weight_idx       α
0                   1/4
1                   -1/8.
Similar to the above case, more than one additional prediction signal can be used. The resulting overall prediction signal is accumulated iteratively with each additional prediction signal:
p_ (n+1) = (1-α_ (n+1) ) p_n+α_ (n+1) h_ (n+1) .
The resulting overall prediction signal is obtained as the last p_n (i.e., the p_n having the largest index n) . Within this EE, up to two additional prediction signals can be used (i.e., n is limited to 2) .
The motion parameters of each additional prediction hypothesis can be signalled either explicitly by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly by specifying a merge index. A separate multi-hypothesis merge flag distinguishes between these two signalling modes.
For inter AMVP mode, MHP is only applied if non-equal weight in BCW is selected in bi-prediction mode.
Combination of MHP and BDOF is possible. However, the BDOF is only applied to the bi-prediction signal part of the prediction signal (i.e., the ordinary first two hypotheses) .
I. 9 Introduction to Geometric Partitioning Mode (GPM)
In VVC, a geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is signalled using a CU-level flag as one kind of merge mode, with other merge modes including the regular merge mode, the MMVD mode, the CIIP mode and the subblock merge mode. In total 64 partitions are supported by geometric partitioning mode for each possible CU size w×h=2m×2n with m, n ∈ {3…6} excluding 8x64 and 64x8.
When this mode is used, a CU is split into two parts by a geometrically located straight line. The location of the splitting line is mathematically derived from the angle and offset parameters of a specific partition. Each part of a geometric partition in the CU is inter-predicted using its own motion; only uni-prediction is allowed for each partition, that is, each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure that same as the conventional bi-prediction, only two motion compensated prediction are needed for each CU.
If geometric partitioning mode is used for the current CU, then a geometric partition index indicating the partition mode of the geometric partition (i.e., angle and offset) , and two merge indices (one for each partition) are further signalled. The number of maximum GPM candidate size is signalled explicitly in SPS and specifies syntax binarization for GPM merge indices. After predicting each of part of the geometric partition, the sample values along the geometric partition edge are adjusted using a blending processing with adaptive weights. This is the prediction signal for the whole CU, and transform and quantization process will be applied to the whole CU as in other prediction modes. Finally, the motion field of a CU predicted using the geometric partition modes is stored.
In the present invention, schemes for improving coding perform for systems using a set of LFNST or NSPT are disclosed.
BRIEF SUMMARY OF THE INVENTION
A method and apparatus for video decoding for using LFNST are disclosed. According to this method, input data associated with a current block is received, wherein the input data comprises transformed residual data associated with the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined. One or more target sets of LFNST is selected from said two or more sets of LFNST. Inverse transform comprising said one or more target sets of LFNST is applied to the transformed residual data to derive reconstructed residual data. The reconstructed residual data is provided.
In one embodiment, at least one of said two or more sets of LFNST is derived based on a DIMD (Decoder-side Intra Mode Derivation) -related scheme.
In one embodiment, a first set of LFNST is derived based on a DIMD (Decoder-side Intra Mode Derivation) -related scheme and one or more additional sets of LFNST are derived by adding one or more predefined offsets to the first set of LFNST.
In one embodiment, an index is signalled in a bitstream to indicate the target set of LFNST selected from said two or more sets of LFNST.
In one embodiment, said two or more sets of LFNST consist of a first set of LFNST and a second set of LFNST, a flag is signalled in a bitstream to indicate whether the second set of LFNST is used for the current block. In one embodiment, if the flag is false, the second set of LFNST is used for the current block; otherwise the first set of LFNST is used for the current block.
In one embodiment, signalling of syntax related to selection of said two or more sets of LFNST is constrained by sum of absolute transform coefficients of the current block.
In one embodiment, said selecting said one or more target sets of LFNST from said two or more sets of LFNST is conditioned by one or more neighbouring blocks. In one embodiment, said two or more sets of LFNST are determined by taking into account of boundary matching (BM) cost calculated between current prediction samples and neighbouring reconstruction samples in a current frame, or reference block samples and the reconstruction samples adjacent to a reference block in a reference frame.
A method and apparatus for video decoding for signalling syntax to select a target set of LFNST with constraint by a sum of absolute transform coefficients of the current block are disclosed. According to this method, input data associated with a current block is received, wherein the input data comprises transformed residual data associated with the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined. One or more syntax elements related to selecting a target set of LFNST are parsed from said two or more sets of LFNST, wherein said parsing said one or more syntax elements is constrained by a sum of absolute transform coefficients of the current block. Inverse transform comprising the target set of LFNST is applied to the transformed residual data to derive reconstructed residual data. The reconstructed residual data is provided.
Corresponding methods for the encoder side are also disclosed.
A method and apparatus of designing different multiple sets of LFNST are also disclosed. At the decoder side, input data associated with a current block is received, wherein the input data comprises transformed residual data associated with the current block, wherein the current block is coded in a target prediction mode. One or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) are derived based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique. A target set of LFNST or NSPT is selected from said one or more sets of LFNST or NSPT for the target prediction mode. Inverse transform comprising the target set of LFNST is applied to the transformed residual data to derive reconstructed residual data. The reconstructed residual data is provided.
In one embodiment, the sub-sample technique is only applied when the current block is larger than a pre-defined threshold.
BRIEF DESCRIPTION OF THE DRAWINGS
Fig. 1A illustrates an exemplary adaptive Inter/Intra video encoding system incorporating loop processing.
Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
Fig. 2 illustrates an example of Low-Frequency Non-Separable Transform (LFNST) process.
Fig. 3A illustrates an example of Region-Of-Interest (ROI) for LFNST16.
Fig. 3B illustrates an example of Region-Of-Interest (ROI) for LFNST8.
Fig. 4 illustrates an example of L-shaped neighbourhood for a given predicted block.
Fig. 5 illustrates an example of NSPT and LFNST used for various block sizes.
Fig. 6 illustrates an example of boundary samples used for boundary matching cost calculation.
Fig. 7A illustrates an example of boundary matching cost calculated in 4x1 or 1x4 blocks of the prediction and neighbouring reconstruction samples in the current frame.
Fig. 7B illustrates an example of boundary matching cost calculated in 4x1 or 1x4 blocks of the prediction and neighbouring reconstruction samples in the reference frame.
Fig. 8 illustrates a flowchart of an exemplary video decoding system that derives multiple sets of LFNST and selecting one target set of LFNST to code a block according to an embodiment of the present invention.
Fig. 9 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 8.
Fig. 10 illustrates a flowchart of an exemplary video decoding system that signals syntax to select a target set of LFNST with constraint by a sum of absolute transform coefficients of the current block according to an embodiment of the present invention.
Fig. 11 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 10.
Fig. 12 illustrates a flowchart of an exemplary video decoding system that derives one or more sets of LFNST or NSPT for different prediction modes according to an embodiment of the present invention.
Fig. 13 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 12.
DETAILED DESCRIPTION OF THE INVENTION
It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
II: LFNST with Multiple Transform Sets
In one embodiment, LFNST can be tested with more than one transform set.
In one embodiment, the transform sets used for LFNST are determined based on DIMD related method. In that, a histogram is generated, and the top N intra modes with higher amplitude values can be used to derive LFNST transform sets for LFNST. N can be any integer larger than 0.
In one embodiment, an index is signalled in the bitstream to the decoder to indicate the selected transform set for LFNST. For example, the selected transform set is a transform set from top N candidates with highest amplitude values in the derived histogram. For another example, the selected transform set is a transform set from a derived transform set list.
In one embodiment, a flag is signalled in the bitstream to the decoder to indicate whether the second transform set is used for LFNST. If the flag is false, the first transform set is used. Otherwise, the second transform set is used.
For example, the first transform set can be derived based on DIMD related method. The second transform set can be derived by adding a predefined offset to the first transform set. For example, the first transform set is set 4. The second transform set will be set 5 (4 + 1) , set 3 (4 -1) , or set 21 (4 + 17) .
For example, the first transform set can be derived based on DIMD related method. The intraMode with the highest amplitude value will be used to derive LFNST transform set. The second transform set is derived by using the intraMode with the second highest amplitude value.
For example, the first transform set can be derived based on DIMD related method. The intraMode1 with the highest amplitude value will be used to derive LFNST transform set. The second transform set is derived by using intraMode2 with the second highest amplitude value. In this case, the distance between intraMode1 and intraMode2 shall be larger than TH. TH can be any integer.
In one embodiment, two flags are signalled in the bitstream. The first flag is used to indicate whether the main transform set is used. The second flag is used to indicate which additional transform set is used. Only if the first flag is false, the second flag needs to be signalled.
For example, the first transform set can be derived based on DIMD related method. The second transform set and the third transform set are derived by adding a predefined offset to the first transform set. For example, the first transform set is set 4. The second transform set will be set 5 (4 + 1) . The third transform set will be set 3 (4 -1) .
For example, the first transform set can be derived based on DIMD related method. The intraMode1 with the highest amplitude value will be used to derive LFNST transform set. The second transform set is derived by using intraMode2 with the second highest amplitude value. In this case, the distance between intraMode1 and intraMode2 shall be larger than TH. TH can be any integer. The third transform set is derived by using intraMode3 with the third highest amplitude value. In this case, the distance between intraMode2 and intraMode3 shall be larger than TH. TH can be any integer.
For another example, the derived transform sets of intraMode1, and intraMode2 are constrained to be different.
For another example, the derived transform sets of intraMode1, intraMode2 and intraMode3 are constrained to be different.
In one embodiment, the first bin is signalled to the decoder in the bitstream to indicate whether the additional transform set is used for LFNST.
For example, the first bin can be coded by a context coded bin.
For example, the context variables used for intra LFNST and inter LFNST can be different, but not limited to.
For example, the context variables used for intra LFNST in dual tree case and non-dual tree case can be different, but not limited to.
In one embodiment, some truncated unary coded bins are used to indicate the selected LFNST transform set. In this case, the truncated unary coded bins will only be sent if the bin used to indicate whether the additional transform set is used for LFNST to decoder is true. Or they will only be sent if the bin used to indicate whether the main transform set is used for LFNST to decoder is false.
In one embodiment, the signalling of syntax related to LFNST transform set can be constrained by the absolute sum of transform coefficients of the current CU.
For example, if the absolute sum of quantized coefficients of the current CU is smaller than a predefined threshold, the syntax related to LFNST transform set will not be signalled.
In one embodiment, the signalling of syntax related to LFNST transform set can be constrained by the number of non-zero transform coefficients of current CU. For example, if the number of non-zero transform coefficients of current CU is smaller than a predefined threshold, the syntax related to LFNST transform set will not be signalled.
For example, the predefined thresholds for LFNST transform set signalling for intra LFNST and inter LFNST can be different.
For example, the predefined thresholds for LFNST transform set signalling for intra LFNST and inter LFNST can be the same.
In one embodiment, the signalling of syntax related to LFNST transform set can be constrained by CU size, prediction mode, slice-type or QP value.
In one embodiment, the signalling of syntax related to LFNST transform set can be constrained by any combined conditions mentioned above.
In another embodiment, one on/off control flag is signalled at CU level, slice level, picture level, and/or sequence level to indicate the proposed method in the above is enabled or not.
In another embodiment, the proposed method in the above is enabled or disabled, according to one or a combination of the selected reference picture indices, temporal distance between the reference picture and the current picture, quantization parameter, the coded information of current CU, prediction mode, motion vectors, motion vector resolution, residual of current CU, and reference samples.
In one embodiment, the above-mentioned schemes for LFSNT can also be applied on NSPT and MTS for selection of the transform set.
In one embodiment, more than one predefined threshold can be used for selecting LFNST/NSPT transform set. For example, if the absolute sum of the quantized coefficients of current block is larger than threshold 1, three different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream. If the absolute sum of the quantized coefficients of current block is larger than threshold 2, five different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream. The syntax design for signalling best LFNST/NSPT transform set can be different for different CUs, and it is based on the absolute sum of the quantized coefficients to choose the syntax design.
In one embodiment, more than one predefined threshold can be used for selecting LFNST/NSPT transform set. For example, if the number of non-zero quantized coefficients of current block is larger than threshold 1, three different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream. If the number of non-zero quantized coefficients of current block is larger than threshold 2, five different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream. The syntax design for signalling best LFNST/NSPT transform set can be different for different CUs, and it is based on the number of non-zero quantized coefficients to choose the syntax design.
In one embodiment, TU size, and prediction mode can also be used to select different signalling methods for transform set related syntax. For example, if current block is intra CU, three different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream. If the current block is inter CU, five different transform sets will be tested in the encoder, and the best one will be signalled in the bitstream.
In one embodiment, the predefined thresholds mentioned above can be different for luma and chroma LFNST/NSPT to select the transform set.
In one embodiment, the syntax related to transform set of LFNST/NSPT will only be signalled when LFNST kernel N is used. N can be 0, 1, 2, or 3.
In one embodiment, the syntax related to transform set of LFNST/NSPT will only be signalled when some LFNST kernels are used. For example, only when LFNST kernel 0 and kernel 1 are applied, the additional syntax for LFNST/NSPT transform set will be signalled.
In one embodiment, the LFNST/NSPT transform set of current CU can be conditioned by neighbouring CUs. For example, when the intra/inter mode, the DIMD histogram, or other intra/inter prediction information of current CU is similar to the respective intra/inter mode, the DIMD histogram or other intra/inter prediction information of one of the neighbouring CUs. The transform set of current CU can be determined by the neighbouring CU. For one example, the transform set of current CU is set as the transform set of neighbouring CU with similar intra/inter mode condition and transform set syntax of current can be skipped. For another example, the candidate of transform set of current CU is set as the candidate of neighbouring CU with similar intra/inter mode condition, the transform set syntax of current CU is signalled to indicate the selected transform set.
In one example, the search region of the neighbouring CU for transform set referencing can be constrained in the current CTU or N CTU rows or any specific region.
In one embodiment, the alternative transform set derivation for LFNST/NSPT can be different from different prediction modes.
For example, the total number of transform sets tested in encoder is 3 for each CU.
For intra predicted CUs, the alternative tested transform sets are the neighbouring transform sets of the main transform set. For example, if the main transform set is set 4. Two alternative tested transform sets are set 3 (4 –1) and set 5 (4 + 1) .
For inter predicted CU, the alternative tested transform sets are derived by DIMD histogram. The first two intramode with highest amplitudes will be used to derive the alternative tested transform sets.
For another example, the total number of transform sets tested in encoder is 3 for each CU.
For fusion related predicted CUs (i.e., SGPM, DIMD, TIMD, ITMP, GPM…) , the alternative tested transform sets are the neighbouring transform sets of the main transform set. For example, if the main transform set is set 4. Two alternative tested transform sets are set 3 (4 –1) and set 5 (4 + 1) .
For other non-fusion related predicted CUs, the alternative tested transform sets are derived by DIMD histogram. The first two intramodes with highest amplitudes will be used to derive the alternative tested transform sets.
In one embodiment, the alternative transform set derivation for LFNST/NSPT can be different from different prediction modes. The maximum number of transform sets signalled to the decoder is dependent by prediction mode.
For example, the maximum number of transform sets signalled to the decoder is N for fusion related predicted CUs. The maximum number of transform sets signalled to decoder is M for non-fusion related predicted CUs. N and M can be any integer larger than 0.
For example, N is larger than M. For another example, N is smaller than M.
In one embodiment, the alternative transform set derivation for LFNST/NSPT can be different from different prediction modes.
For another example, the total number of transform sets tested in encoder is 3 for each CU.
For some special mode predicted CUs e.g. MIP, EIP) , the alternative tested transform sets are derived by planar or DC mode.
In one embodiment, the alternative transform set derivation for LFNST/NSPT can be different from different CU sizes.
For example, the maximum number of transform set signalled to the decoder is N for CU area smaller than TH. The maximum number of transform set signalled to the decoder is M for other CUs. N and M can be any integer larger than 0. TH can be any pre-defined integer value.
For example, N is larger than M. For another example, N is smaller than M.
In another embodiment, the alternative transform set derivation for matrix based intra prediction mode mentioned in Section I. 3 can be aligned with the replaced intra mode or determined by DIMD method or derived by planar or dc mode. For example, if the main transform set of replaced intra mode is set 4, the main transform set of the matrix-based intra prediction mode is also set 4 and the two alternative tested transform sets are set 3 (4 –1) and set 5 (4 + 1) . For another example, the main transform set of the matrix-based intra prediction mode is set 4 (e.g., can be determined by replaced intra mode or DIMD histogram) and the two alternative tested transform sets are derived by DIMD histogram. For another example, the main transform set of the matrix-based intra prediction mode is set 4 (e.g., can be determined by replaced intra mode or DIMD histogram) and the two alternative tested transform sets are derived by planar or dc mode.
In addition, the LFNST/NSPT kernel index related syntax can be signalled after transform set related syntax.
In one embodiment, the syntax related to LFNST/NSPT kernel index will only be signalled when some LFNST transform sets are used.
For example, only when LFNST/NSPT transform set related syntax indicates that the alternative transform set is not used on LFNST/NSPT, the syntax related to LFNST/NSPT kernel index needs to be signalled. In other words, when the alternative transform set of LFNST/NSPT is used, LFNST/NSPT kernel index related syntax doesn’t need to be signalled. The LFNST/NSPT kernel index can be assigned with a pre-defined value (i.e., 1st kernel is always used, or the corresponding kernel will be derived by a pre-defined formula (e.g. intraMode %3) ) .
For another example, only when LFNST/NSPT transform set related syntax indicates that the alternative transform set is not used on LFNST/NSPT or the alternative transform sets are some special transform sets (e.g. transform set number %2 == 0) , the syntax related to LFNST/NSPT kernel index needs to be signalled.
The pre-defined kernel index can be determined based on alternative transform set index, but not limit to.
The pre-defined kernel index can be different from inter LFNST/NSPT and intra LFNST/NSPT, but not limit to.
The pre-defined kernel index can be different from luma LFNST and chroma LFNST, but not limit to.
The pre-defined kernel index can be different from fusion related prediction mode and non-fusion related prediction mode, but not limit to.
The pre-defined kernel index can be different for different CU size, but not limit to.
For another example, when the alternative transform set of LFNST/NSPT is used, LFNST/NSPT kernel index related syntax does not need to be signalled, and the kernel index is determined based on a derived syntax, intramode.
For example, for non-wide angle intramode, first kernel will be used. For wide angle intramode, second kernel will be used.
For example, for the secondary transform applying with transpose process, the first kernel will be used. Otherwise, second kernel will be used.
For another example, when the alternative transform set of LFNST/NSPT is used, LFNST/NSPT kernel index related syntax still needs to be signalled. However, the syntax design for LFNST/NSPT kernel index in this condition is different when the non-alternative transform set of LFNST/NSPT is used.
For example, the maximum value of possible kernel index for alternative transform set is smaller than the maximum value of possible kernel index for non-alternative transform set.
For another example, the context variables used to code kernel index related syntax are different from alternative transform set and non-alternative transform set.
For another example, if alternative transform set is used, EP bins will be used to code the kernel index. If the non-alternative transform set is used, context coded bins will be used to code the kernel index.
III. Using Different Samples to Derive DIMD Histogram for LFNST/NSPT Transform Set Selection
The LFNST/NSPT transform set of a TU can be determined by DIMD related method. In this case, a histogram of intra mode of the TU is derived. The intra mode with the highest amplitude value in the histogram will be used to derive the corresponding LFNST/NSPT transform set.
In one embodiment, the histogram used to derive the transform set of LFNST and NSPT is derived by collecting the statistic information of neighbouring reconstruction samples. For example, the statistic information includes x-axis and y-axis gradients.
In one embodiment, the histogram used to derive the transform set of LFNST and NSPT is derived by collecting the statistic information current predictors. For example, the statistic information includes x-axis and y-axis gradients.
In one embodiment, the derivation of histogram used for inter LFNST/NSPT and intra LFNST/NSPT can be different, but not limit to. For example, the histogram used for inter LFNST/NSPT is derived by using the statistic information of current predictors. The histogram used for intra LFNST/NSPT is derived by using neighbouring reconstruction samples.
In one embodiment, the derivation of histogram used for LFNST/NSPT can be different for different prediction modes. For example, the modes with fusion schemes will use current predictors to derive histogram (e.g. intra TMP, SGPM, TIMD) . Other modes will use neighbouring reconstruction samples to derive histogram.
In one embodiment, the derivation of histogram used for LFNST/NSPT can be determined by statistical analysis. For example, if the histogram derived by neighbouring reconstruction samples is very different from the histogram derived by current predictors, the histogram derived by current predictors will be used to derive LFNST/NSPT transform set. For example, if the histogram derived by neighbouring reconstruction samples is very different from the histogram derived by current predictors, the histogram derived by neighbouring reconstruction samples will be used to derive LFNST/NSPT transform set.
In one embodiment, the derivation of histogram used for LFNST/NSPT can be different for different TU size. For example, when width and height of TUs are both larger than 32, the current predictor is used to derive histogram. For example, when TUs width and height are both smaller than 32, the current predictor is used to derive histogram.
In another embodiment, one on/off control flag is signalled at CU level, slice level, picture level, and/or sequence level to indicate the proposed method in the above is enabled or not.
For example, only when the control flag is signalled in the bitstream and the value is true, either neighbouring reconstruction samples or the current predictor will be used to derive DIMD histogram for LFNST/NSPT transform set selection for fusion modes and non-fusion mode. Otherwise, for intra blocks, the neighbouring reconstruction samples will always be used to derive DIMD histogram for LFNST/NSPT transform set selection. For inter blocks, the current predictor will always be used to derive DIMD histogram for LFNST/NSPT transform set selection.
In another embodiment, the proposed method in the above is enabled or disabled, according to one or the combination of the selected reference pictures indices, temporal distance between the reference picture and the current picture, quantization parameter, the coded information of current CU, prediction mode, motion vectors, motion vector resolution, residual of current CU, and reference samples.
In one embodiment, the selection of using neighbouring reconstruction samples or the current predictor to derive DIMD histogram for LFNST/NSPT transform set derivation is signalled in the bitstream. It can be signalled at CU level, slice level, picture level, and/or sequence level.
In one embodiment, the current predictors mentioned above are the final predictors after blending if the prediction mode of current CU is a fusion mode. In this case, more than one predictor will be blended to generate a blended final predictor.
In one embodiment, sub-sample technology can be applied to derive the DIMD histogram for LFNST/NSPT transform set derivation. For example, current predictors (i.e., final predictors after blending) are used to derive the DIMD histogram for LFNST/NSPT transform set derivation. Step number is set to be 2.Every two samples in each row and each column are collected to do the statistical analysis. In that, the step number can be designed based on current CU size.
For another example, the step number for row and column can be different, but not limit to.
For another example, the sub-sample technique is only applied when the current predictor area is larger than a pre-defined threshold (e.g. 256) .
For another example, only the samples in the first N rows and the first K columns of current predictors (e.g. final predictors after blending) are used to derive the DIMD histogram for LFNST/NSPT transform set derivation.
IV. Other Methods to Derive LFNST/NSPT Transform Set
In one embodiment, for bi-prediction coded blocks, two DIMD histograms can be derived by using predictor L0 and predictor L1. Two DIMD histograms are referenced to derive LFNST/NSPT transform set. For example, the histogram derived based on predictor L0 is used to derived the second transform set. The histogram derived based on predictor L1 is used to derive the third transform set.
In one embodiment, the intramode of previously coded intra CU can be used to derive the LFNST/NSPT transform set of current CU.
The distance between the previously coded intra CU and the current CU can be constrained, but not limit to. For example, only if the distance between the top-left position of previously coded intra CU and the current CU is smaller than a threshold, the intramode of previously coded intra CU can be referenced by the current block.
The prediction mode of previously coded intra CU can be constrained, but not limit to. For example, only if the previously coded intra CU is not predicted by a fusion mode (i.e., DIMD, SGPM, intraTMP fusion) , the previously coded intra CU’s intramode can be referenced by current block.
In the previously mentioned method, if the previously coded intra CU cannot be referenced, a pre-defined intramode is used (i.e., planar mode) .
V. LFNST/NSPT Kernel Set Determination
In one embodiment, the LFNST/NSPT kernel sets used by inter and intra modes are different. For intra modes, the intra angular direction, intra prediction mode, transform block size, QP value of current block, absolute sum of quantized coefficients, number of non-zero quantized coefficients, boundary matching cost (e.g. the gradients at the boundary of current inter prediction samples and neighbouring reconstruction samples) , sum of gradient amplitudes of current predictor, and/or gradient directions of current predictor and/or are used to determine the LFNST/NSPT kernel set. For inter modes, an alternative IPM (Intra Prediction Mode) derived by DIMD, the inter prediction modes boundary matching cost (e.g. the gradients at the boundary of current inter prediction samples and neighbouring reconstruction samples) , QP value of current block, QP value of reference blocks, transform block size, GPM partition directions, BCW weights, absolute sum of quantized coefficients, number of non-zero quantized coefficients, MV amplitudes, sum of gradient amplitudes of current predictor, and/or gradient directions of current predictor can be used to determine the LFSNT/NSPT kernel set.
In one embodiment, the LFNST/NSPT kernel set can be determined by considering the boundary matching (BM) cost calculated between current prediction samples and neighbouring reconstruction samples on the current frame, or reference block samples and the reconstruction samples adjacent to reference block on the reference frame. The boundary matching cost is calculated as follows:

where Rx, -n indicates the neighbouring samples n column left to the prediction samples Px, 0, and Ry, -n 
indicates the neighboring samples n row above prediction samples Py, 0 As shown in Fig. 6. The boundary matching cost can be calculated in sample-based, block-based, or CU-based. In one example, the boundary matching cost is calculated in 4x1 or 1x4 block-based by the prediction and neighbouring reconstruction samples on the current frame as shown in Fig. 7A. That is, for each block, a BM cost is derived. The K (K>=1) blocks with the lowest BM costs can be used to determine the LFNST/NSPT kernel set of current coding block as shown in Table 3.
Table 3. Mapping between blocks with lowest BM cost and LFNST/NSPT kernel set
In another example, the boundary matching cost is calculated in 4x1 or 1x4 block-based by the reference block samples and the reconstruction samples adjacent to reference block on reference frame as shown in Fig. 7B. The K (K>=1) blocks with the lowest BM costs can be used to determine the LFNST/NSPT kernel set of current coding block as shown in Table 4.
Table 4. Mapping between blocks with lowest BM cost and LFNST/NSPT kernel set
In another embodiment, the LFNST/NSPT kernel set can be determined by considering the QP value of current block or the QP values of reference blocks. In one example, N (N >= 1) QP thresholds are pre-defined or determined adaptively according to base QP, picture resolution, transform block size, or other coding information. For the blocks with the QP value smaller than or equal to the first threshold, the first LFNST/NSPT kernel set is used. For the blocks with the QP value larger than the first threshold and smaller than or equal to the second threshold, the second LFNST/NSPT kernel set is used. For the blocks with the QP value larger than i-th threshold and smaller than or equal to the (i+1) -th threshold, the (i+1) -th LFNST/NSPT kernel set is used. In another example, the minimum or maximum QP value of the QP values of reference blocks or the QP value of the reference block with higher prediction blending weight or the QP value selected by the predefined method (e.g., selecting the QP value of L0 predictor or L1 predictor or j-th predictor of MHP) is set as the selected QP value. For the blocks with the selected QP value smaller than or equal to the first threshold, the first LFNST/NSPT kernel set is used. For the blocks with the selected QP value larger than first threshold and smaller than or equal to the second threshold, the second LFNST/NSPT kernel set is used. For the blocks with the selected QP value larger than i-th threshold and smaller than or equal to the (i+1) -th threshold, the (i+1) -th LFNST/NSPT kernel set is used.
In one embodiment, the LFNST/NSPT kernel set is determined by the prediction mode of current CU. In one example, the GPM partition modes can be used to determine the kernel sets for GPM-coded CUs. In another example, the CUs coded by subblock modes (e.g. affine, SBTMVP, DMVR modes) use specific kernel sets different from other inter CUs. In another example, the CUs coded by multi-hypothesis coding modes (e.g. bi-predictive inter mode, CIIP, GPM, MHP, DIMD, TIMD, etc. ) use specific kernel sets. In another example, each inter/intra mode can use a specific kernel set.
In one embodiment, the LFNST/NSPT kernel set is determined based on the absolute sum of quantized coefficients of current TB, number of non-zero quantized coefficients of current TB, MV amplitudes of current predictor and/or sum of gradient amplitudes of current predictor. In one example, N (N >= 1) thresholds are pre-defined or determined adaptively according to the base QP, picture resolution, transform block size, or other coding information. For the blocks with the absolute sum of quantized coefficients, number of non-zero quantized coefficients, MV amplitudes and/or sum of gradient amplitudes smaller than or equal to the first threshold, the first LFNST/NSPT kernel set is used. For the blocks with the absolute sum of quantized coefficients, number of non-zero quantized coefficients, MV amplitudes and/or sum of gradient amplitudes larger than the first corresponding threshold and smaller than or equal to the second corresponding threshold, the second LFNST/NSPT kernel set is used. For the blocks with the absolute sum of quantized coefficients, number of non-zero quantized coefficients, MV amplitudes and/or sum of gradient amplitudes larger than i-th threshold and smaller than or equal to the (i+1) -th corresponding threshold, the (i+1) -th LFNST/NSPT kernel set is used.
In one embodiment, the gradients of current predictor are calculated in sample-based or block-based scheme. A histogram is derived by collecting the calculated gradients and the gradient direction with highest peak can be used to determine the LFNST/NSPT kernel set.
In one embodiment, the BCW weight can be used to determine the LFNST/NSPT kernel set. The bi-predictive CUs with different BCW weights can use different kernel set. Each BCW weight correspond to a specific kernel set.
VI. Inheritance of Inter LFNST/NSPT KERNEL SET
In one invention, the inter LFNST/NSPT kernel set can be inherited from the reference blocks or neighbouring blocks. That is, the inter LFNST/NSPT kernel set of current block is set as one of the kernel sets of reference blocks or neighbouring blocks.
In one embodiment, the inter/intra LFSNT/NSPT kernel set usage of each CU is stored in the MV information (e.g. per 4x4, 8x8, or NxN subblock MV) or CU information. For an inter transform block (TB) , the LFNST/NSPT kernel set is inherited from the reference CUs by checking the stored kernel sets in reference regions. If more than one kernel set exists, an index is signalled to determine the kernel set used by the current CU or a predefined selection method (e.g. deriving a histogram to find the kernel set with highest usage over all sub-blocks) is applied to select the kernel set used by the current CU. If more than one reference region exists (e.g. for bi-predictive, GPM, or MHP CUs) , the region with higher prediction blending weight is used or M (M>=1) of the regions are used to determine the kernel set.
In another embodiment, an LFNST/NSPT kernel set list is constructed from neighbouring blocks. An index is signalled to determine the kernel set used by current CU or a predefined selection method (e.g. generating the templates at the CU boundary by different kernel sets and using BM costs to find the kernel sets with higher accuracy) is applied to select the kernel set used by current CU.
In another embodiment, multiple LFNST/NSPT kernel sets can be selected based on the current CU. Multiple kernel set candidates are determined by the strategies mentioned in section V and the remaining kernel set candidates are inherited from reference or neighbouring blocks.
In another embodiment, only P (P>=1) kernels can be used in each kernel set. In another embodiment, only P (P>=1) kernels can be used in each kernel set depending on whether the coding information is higher or lower than a threshold mentioned in this provisional. In another example, when multiple LFNST/NSPT kernel sets can be selected, only P (P>=1) kernels are used in each kernel set.
Any of the foregoing proposed methods of multiple sets of LFNST or different sets of LFNST or NSPT for different prediction modes can be implemented in encoders and/or decoders. For example, any of the proposed methods can be implemented in transform module of an encoder and/or a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to transform module of the encoder and/or the decoder.
With reference to the exemplary encoder or decoder in Fig. 1A and Fig. 1B, any of the proposed methods can be implemented in an Intra/Inter coding module (e.g. Intra Pred. 150/MC 152 in Fig. 1B) in a decoder or an Intra/Inter coding module in an encoder (e.g. Intra Pred. 110/Inter Pred. 112 in Fig. 1A) . Any of the proposed methods can also be implemented as circuits coupled to the intra coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required processing. While the Intra/Inter Pred. units (e.g. unit 110/Inter Pred. 112 in Fig. 1A and unit 150/MC 152 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .
Fig. 8 illustrates a flowchart of an exemplary video decoding system that derives multiple sets of LFNST and selecting one target set of LFNST to code a block according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data associated with a current block is received in step 810, wherein the input data comprises transformed residual data associated with the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 820. One or more target sets of LFNST are selected from said two or more sets of LFNST in step 830. Inverse transform comprising said one or more target sets of LFNST is applied to the transformed residual data to derive reconstructed residual data in step 840. The reconstructed residual data is provided in step 850.
Fig. 9 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 8. According to this method, residual data associated with a current block is received in step 910, wherein the residual data is generated by applying prediction to the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 920. One or more target sets of LFNST are selected from said two or more sets of LFNST in step 930. Transform comprising said one or more target sets of LFNST is applied to the residual data to generate transformed data in step 940. The transformed data is provided in step 950.
Fig. 10 illustrates a flowchart of an exemplary video decoding system that signals syntax to select a target set of LFNST with constraint by a sum of absolute transform coefficients of the current block according to an embodiment of the present invention. According to this method, input data associated with a current block is received in step 1010, wherein the input data comprises transformed residual data associated with the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 1020. One or more syntax elements related to selecting a target set of LFNST are parsed from said two or more sets of LFNST in step 1030, wherein said parsing said one or more syntax elements is constrained by a sum of absolute transform coefficients of the current block. Inverse transform comprising the target set of LFNST is applied to the transformed residual data to derive reconstructed residual data in step 1040. The reconstructed residual data is provided in step 1050.
Fig. 11 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 10. According to this method, residual data associated with a current block is received in step 1110, wherein the residual data is generated by applying prediction to the current block. Two or more sets of LFNST (Low-Frequency Non-Separable Transform) are determined in step 1120. One or more syntax elements related to selecting a target set of LFNST from said two or more sets of LFNST are signalled in step 1130, wherein said signalling said one or more syntax elements is constrained by sum of absolute transform coefficients of the current block. Transform comprising the target set of LFNST is applied to the residual data to generate transformed data in step 1140. The transformed data is provided in step 1150.
Fig. 12 illustrates a flowchart of an exemplary video decoding system that derives different multiple sets of LFNST or NSPT for different prediction modes according to an embodiment of the present invention. According to this method, input data associated with a current block is received in step 1210, wherein the input data comprises transformed residual data associated with the current block, wherein the current block is coded in a target prediction mode. One or more multiple sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique are derived in step 1220. A target set of LFNST or NSPT is selected from said one or more sets of LFNST or NSPT for the target prediction mode in step 1230. Inverse transform comprising the target set of LFNST to the transformed residual data is applied to derive reconstructed residual data in step 1240. The reconstructed residual data is provided in step 1050.
Fig. 13 illustrates a flowchart of an exemplary video encoding system corresponding to the decoding system in Fig. 12. According to this method, data associated with a current block is received in step 1310, wherein the residual data is generated by applying prediction to the current block, wherein the current block is coded in a target prediction mode. One or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) are derived based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique in step 1320. A target set of LFNST or NSPT is selected from said one or more sets of LFNST or NSPT for the target prediction mode in step 1330. Transform comprising the target set of LFNST or NSPT is applied to the residual data to generate transformed data in step 1340. The transformed data is provided in step 1350.
The flowcharts shown are intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims (17)

  1. A method of video decoding, the method comprising:
    receiving input data associated with a current block, wherein the input data comprises transformed residual data associated with the current block;
    determining two or more sets of LFNST (Low-Frequency Non-Separable Transform) ;
    selecting one or more target sets of LFNST from said two or more sets of LFNST;
    applying inverse transform comprising said one or more target sets of LFNST to the transformed residual data to derive reconstructed residual data; and
    providing the reconstructed residual data.
  2. The method of Claim 1, wherein at least one of said two or more sets of LFNST is derived based on a DIMD (Decoder-side Intra Mode Derivation) -related scheme.
  3. The method of Claim 1, wherein a first set of LFNST is derived based on a DIMD (Decoder-side Intra Mode Derivation) -related scheme and one or more additional sets of LFNST are derived by adding one or more predefined offsets to the first set of LFNST.
  4. The method of Claim 1, wherein an index is signalled in a bitstream to indicate the target set of LFNST selected from said two or more sets of LFNST.
  5. The method of Claim 1, wherein said two or more sets of LFNST consist of a first set of LFNST and a second set of LFNST, a flag is signalled in a bitstream to indicate whether the second set of LFNST is used for the current block.
  6. The method of Claim 5, wherein if the flag is false, the second set of LFNST is used for the current block; otherwise the first set of LFNST is used for the current block.
  7. The method of Claim 1, wherein said selecting said one or more target sets of LFNST from said two or more sets of LFNST is conditioned by one or more neighbouring blocks.
  8. The method of Claim 1, wherein said two or more sets of LFNST are determined by taking into account of boundary matching (BM) cost calculated between current prediction samples and neighbouring reconstruction samples in a current frame, or reference block samples and the reconstruction samples adjacent to a reference block in a reference frame.
  9. An apparatus for video decoding, the apparatus comprising one or more electronics or processors arranged to:
    receive input data associated with a current block, wherein the input data comprises transformed residual data associated with the current block;
    determine two or more sets of LFNST (Low-Frequency Non-Separable Transform) ;
    select one or more target sets of LFNST from said two or more sets of LFNST;
    apply inverse transform comprising said one or more target sets of LFNST to the transformed residual data to derive reconstructed residual data; and
    provide the reconstructed residual data.
  10. A method of video encoding, the method comprising:
    receiving residual data associated with a current block, wherein the residual data is generated by applying prediction to the current block;
    determining two or more sets of LFNST (Low-Frequency Non-Separable Transform) ;
    selecting one or more target sets of LFNST from said two or more sets of LFNST;
    applying transform comprising said one or more target sets of LFNST to the residual data to generate transformed data; and
    providing the transformed data.
  11. A method of video decoding, the method comprising:
    receiving input data associated with a current block, wherein the input data comprises transformed residual data associated with the current block;
    determining two or more sets of LFNST (Low-Frequency Non-Separable Transform) ;
    parsing one or more syntax elements related to selecting a target set of LFNST from said two or more sets of LFNST, wherein said signalling or parsing said one or more syntax elements is constrained by a sum of absolute transform coefficients of the current block;
    applying inverse transform comprising the target set of LFNST to the transformed residual data to derive reconstructed residual data; and
    providing the reconstructed residual data.
  12. An apparatus for video decoding, the apparatus comprising one or more electronics or processors arranged to:
    receive input data associated with a current block, wherein the input data comprises transformed residual data associated with the current block;
    determine two or more sets of LFNST (Low-Frequency Non-Separable Transform) ;
    parse one or more syntax elements related to selecting a target set of LFNST from said two or more sets of LFNST, wherein said parsing said one or more syntax elements is constrained by a sum of absolute transform coefficients of the current block;
    apply inverse transform comprising the target set of LFNST to the transformed residual data to derive reconstructed residual data; and
    provide the reconstructed residual data.
  13. A method of video encoding, the method comprising:
    receiving residual data associated with a current block, wherein the residual data is generated by applying prediction to the current block;
    determining two or more sets of LFNST (Low-Frequency Non-Separable Transform) ;
    signalling one or more syntax elements related to selecting a target set of LFNST from said two or more sets of LFNST, wherein said signalling said one or more syntax elements is constrained by sum of absolute transform coefficients of the current block;
    applying transform comprising the target set of LFNST to the residual data to generate transformed data; and
    providing the transformed data.
  14. A method of video decoding, the method comprising:
    receiving input data associated with a current block, wherein the input data comprises transformed residual data associated with the current block, wherein the current block is coded in a target prediction mode;
    deriving one or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique;
    selecting a target set of LFNST or NSPT from said one or more sets of LFNST or NSPT for the target prediction mode;
    applying inverse transform comprising the target set of LFNST to the transformed residual data to derive reconstructed residual data; and
    providing the reconstructed residual data.
  15. The method of Claim 14, wherein the sub-sample technique is only applied when the current block is larger than a pre-defined threshold.
  16. An apparatus for video decoding, the apparatus comprising one or more electronics or processors arranged to:
    receive input data associated with a current block, wherein the input data comprises transformed residual data associated with the current block, wherein the current block is coded in a target prediction mode;
    derive one or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique;
    select a target set of LFNST or NSPT from said one or more sets of LFNST or NSPT for the target prediction mode;
    apply inverse transform comprising the target set of LFNST to the transformed residual data to derive reconstructed residual data; and
    provide the reconstructed residual data.
  17. A method of video encoding, the method comprising:
    receiving residual data associated with a current block, wherein the residual data is generated by applying prediction to the current block, wherein the current block is coded in a target prediction mode;
    deriving one or more sets of LFNST (Low-Frequency Non-Separable Transform) or NSPT (Non-Separable Primary Transform) based on DIMD (Decoder-side Intra Mode Derivation) scheme with sub-sample technique;
    selecting a target set of LFNST or NSPT from said one or more sets of LFNST or NSPT for the target prediction mode;
    applying transform comprising the target set of LFNST or NSPT to the residual data to generate transformed data; and
    providing the transformed data.
PCT/CN2025/083081 2024-03-28 2025-03-18 Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system Pending WO2025201110A1 (en)

Applications Claiming Priority (6)

Application Number Priority Date Filing Date Title
US202463570842P 2024-03-28 2024-03-28
US63/570,842 2024-03-28
US202463637412P 2024-04-23 2024-04-23
US63/637,412 2024-04-23
US202463710651P 2024-10-23 2024-10-23
US63/710,651 2024-10-23

Publications (1)

Publication Number Publication Date
WO2025201110A1 true WO2025201110A1 (en) 2025-10-02

Family

ID=97216127

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/083081 Pending WO2025201110A1 (en) 2024-03-28 2025-03-18 Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system

Country Status (1)

Country Link
WO (1) WO2025201110A1 (en)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20220086449A1 (en) * 2019-06-06 2022-03-17 Lg Electronics Inc. Transform-based image coding method and device for same
CN114982240A (en) * 2020-01-08 2022-08-30 高通股份有限公司 Multiple transform set signaling for video coding
US20220329819A1 (en) * 2021-04-12 2022-10-13 Qualcomm Incorporated Low frequency non-separable transform for video coding

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20220086449A1 (en) * 2019-06-06 2022-03-17 Lg Electronics Inc. Transform-based image coding method and device for same
CN114982240A (en) * 2020-01-08 2022-08-30 高通股份有限公司 Multiple transform set signaling for video coding
US20220329819A1 (en) * 2021-04-12 2022-10-13 Qualcomm Incorporated Low frequency non-separable transform for video coding

Similar Documents

Publication Publication Date Title
WO2017084577A1 (en) Method and apparatus for intra prediction mode using intra prediction filter in video and image compression
WO2023116716A1 (en) Method and apparatus for cross component linear model for inter prediction in video coding system
WO2025016418A1 (en) Intra merge mode
WO2024174828A1 (en) Method and apparatus of transform selection depending on intra prediction mode in video coding system
WO2024131778A1 (en) Intra prediction with region-based derivation
WO2025201110A1 (en) Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system
WO2025218726A1 (en) Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system
WO2025209169A1 (en) Methods and apparatus for low-frequency non-separable transform with multiple transform sets in a video coding system
WO2025157299A1 (en) Methods and apparatus of intra merge mode for decoder side intra mode derivation
WO2025082424A1 (en) Methods and apparatus of intra fusion mode with extrapolation intra prediction
WO2025167865A1 (en) Methods and apparatus of intra merge mode for occurrence-based intra coding
WO2025082425A1 (en) Methods and apparatus of combined prediction mode with extrapolation intra prediction for video coding
WO2025153050A1 (en) Methods and apparatus of filter-based intra prediction with multiple hypotheses in video coding systems
WO2025157201A1 (en) Methods and apparatus of intra merge mode for template-based intra mode derivation
WO2025237222A1 (en) Methods and apparatus for adaptively determining transform type in image and video coding systems
WO2025157298A1 (en) Methods and apparatus of intra merge mode for reference line intra mode prediction
WO2025168021A1 (en) Methods and apparatus of intra merge mode for merged intra mode derivation
WO2024153093A1 (en) Method and apparatus of combined intra block copy prediction and syntax design for video coding
WO2026012459A1 (en) Methods and apparatus of multi-model eip in video coding
WO2025218691A1 (en) Methods and apparatus for adaptively determining selected transform type in image and video coding systems
WO2024213093A1 (en) Methods and apparatus of blending intra prediction for video coding
WO2026092625A1 (en) Methods and apparatus of combined prediction mode with intra mode derivation in video coding systems
WO2026092755A1 (en) Combined prediction mode for intra block copy with intra mode derivation
WO2025237149A1 (en) Methods and apparatus for intra prediction and transform type selection in image and video coding systems
WO2025218707A1 (en) Method and apparatus of intra merge mode for mixed modes with chroma components in video coding system

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25777485

Country of ref document: EP

Kind code of ref document: A1