WO2026007482A1 - 视频解码方法、视频编码方法及装置 - Google Patents
视频解码方法、视频编码方法及装置Info
- Publication number
- WO2026007482A1 WO2026007482A1 PCT/CN2025/086881 CN2025086881W WO2026007482A1 WO 2026007482 A1 WO2026007482 A1 WO 2026007482A1 CN 2025086881 W CN2025086881 W CN 2025086881W WO 2026007482 A1 WO2026007482 A1 WO 2026007482A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- chroma block
- syntax element
- chroma
- context model
- target
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/124—Quantisation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/13—Adaptive entropy coding, e.g. adaptive variable length coding [AVLC] or context adaptive binary arithmetic coding [CABAC]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/44—Decoders specially adapted therefor, e.g. video decoders which are asymmetric with respect to the encoder
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/46—Embedding additional information in the video signal during the compression process
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/593—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/90—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using coding techniques not provided for in groups H04N19/10-H04N19/85, e.g. fractals
- H04N19/91—Entropy coding, e.g. variable length coding [VLC] or arithmetic coding
Definitions
- Some embodiments of this application relate to the field of video encoding and decoding technology. More specifically, they relate to a video decoding method, a video encoding method, and an apparatus.
- Entropy coding is one of the main processes in video coding.
- the value of the syntax element ⁇ tu_cb_coded_flag ⁇ which indicates whether a chroma transform block is a non-zero block, is generally entropy-encoded using Context-based Adaptive Binary Arithmetic Coding (CABAC).
- CABAC Context-based Adaptive Binary Arithmetic Coding
- the context model selection is based on the initialization type of ⁇ tu_cb_coded_flag ⁇ and whether the current chroma block uses the bdpcm chroma mode (the value of the syntax element ⁇ intra_bdpcm_chroma_flag ⁇ ).
- the initialization type of ⁇ tu_cb_coded_flag ⁇ and whether the bdpcm chroma mode is used for context model selection may lead to inaccurate context models used to encode the value of ⁇ tu_cb_coded_flag ⁇ .
- Exemplary embodiments of this application provide a video decoding method, a video encoding method, and an apparatus for improving the accuracy of a context model used for value entropy encoding of tu_cb_coded_flag.
- some embodiments of this application provide a video decoding method, including:
- some embodiments of this application provide a video encoding method, including:
- the spatial information of the first chroma block is obtained from the adjacent encoded chroma blocks of the first chroma block; the target context model is determined based on the spatial information of the first chroma block; the value of the first syntax element of the first chroma block is entropy encoded based on the target context model to obtain the entropy encoded data of the first syntax element of the first chroma block; the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block.
- some embodiments of this application provide a video decoding apparatus, including:
- a first acquisition module is used to acquire the entropy-encoded data of the first syntax element of the first chroma block; the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block; a second acquisition module is used to acquire the spatial domain information of the first chroma block based on the adjacent decoded chroma blocks of the first chroma block; a determination module is used to determine the target context model based on the spatial domain information of the first chroma block; a decoding module is used to acquire the value of the first syntax element of the first chroma block based on the target context model and the entropy-encoded data of the first syntax element of the first chroma block.
- some embodiments of this application provide a video encoding apparatus, including:
- the acquisition unit is used to acquire the spatial information of the first chroma block based on the adjacent encoded chroma blocks of the first chroma block; the determination unit is used to determine the target context model based on the spatial information of the first chroma block; the entropy encoding unit is used to entropy encode the value of the first syntax element of the first chroma block based on the target context model to acquire the entropy encoded data of the first syntax element of the first chroma block; the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block.
- Figure 1 shows a schematic diagram of the structure of a video encoder in some embodiments of this application
- Figure 2 shows a schematic diagram of the structure of a video decoder in some embodiments of this application
- Figure 3 illustrates a flowchart of determining the index value of the context model in some embodiments of this application
- FIG. 4 shows a flowchart of the video decoding method in some embodiments of this application.
- FIG. 5 shows a flowchart of the video decoding method in some other embodiments of this application.
- Figure 6 shows a flowchart of the video encoding method in some embodiments of this application.
- Figure 7 shows a schematic diagram of the structure of a video decoding device in some embodiments of this application.
- Figure 8 shows a schematic diagram of the structure of a video encoding device in some embodiments of this application.
- a product or device that includes a series of components is not necessarily limited to all components explicitly listed, but may include other components not explicitly listed or inherent to such products or devices.
- the references to "some implementations,” “some embodiments,” etc. in the specification indicate that the described implementations or embodiments may include specific features, structures, or characteristics, but not every embodiment may necessarily include that specific feature, structure, or characteristic. In addition, such phrases do not necessarily refer to the same implementation.
- it is considered that implementing such feature, structure, or characteristic in connection with other implementations (whether explicitly described herein or not) is within the knowledge of those skilled in the art.
- a video can be viewed as a sequence of multiple video frames (images).
- Video playback can be viewed as the display of video frames in the order they appear in the sequence at a preset rate (e.g., 24 frames/second, 30 frames/second, 60 frames/second).
- a preset rate e.g. 24 frames/second, 30 frames/second, 60 frames/second.
- the amount of video data is positively correlated with the resolution of the video frames; the higher the resolution of the video frames, the larger the amount of video data. If the pixel data of every pixel in all video frames is directly stored in the video file, the amount of video data will be enormous, making it difficult to store and transmit the video.
- Video encoding and decoding were proposed to address this problem to some extent.
- Video decoding can include video encoding and video decoding.
- Video encoding can be understood as the process of compressing the original video frames, while video decoding can be understood as the process of reconstructing video frames based on the compressed video data.
- the encoder may include an encoding control module 101.
- the encoding control module 101 is used to perform overall control of the video encoder, including: encoding mode selection, encoding logic control, bitrate control, quantization step size control, etc.
- the encoder may include an intra-prediction module 102.
- the intra-prediction module 102 is used to perform intra-prediction based on reconstructed samples in the current image and send intra-encoding information to the entropy coding module.
- the intra-encoding information may include at least one of an intra-prediction mode, a Most Probable Mode (MPM) flag, and an MPM index.
- the intra-encoding information may also include information about reference samples.
- the encoder may include an inter-frame prediction module 103.
- the inter-frame prediction module 103 performs inter-frame prediction to predict the current image using a reference image stored in the Decoded Picture Buffer (DPB).
- the inter-frame prediction module 103 may include a motion estimation unit 103a and a motion compensation unit 103b.
- the motion estimation unit 103a performs motion estimation (ME) to obtain the motion vector of the current region by referring to a specific region of the reconstructed reference image.
- the motion compensation unit 103b performs motion compensation using the motion vector values obtained from the motion estimation unit 103a.
- the encoder may include a transform module 104.
- the transform module 104 obtains transform coefficients by transforming a residual signal, which is the difference between the input video and the predicted signal generated by the intra-frame prediction module 102 or the inter-frame prediction module 103.
- a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a wavelet transform may be used.
- DCT and DST perform the transform by dividing the input image signal into multiple blocks.
- the compilation efficiency can vary depending on the distribution and characteristics of the values in the transformed region.
- the encoder may include a quantization module 105.
- the quantization module 105 is used to quantize the transform coefficients output by the transform module 104 to obtain quantized transform coefficients.
- the encoder may include an entropy coding module 106.
- the entropy coding module 106 is used to entropy code information such as the coding mode, prediction module, motion vector, and quantization transform coefficients output by the quantization module 105 to obtain the output bitstream corresponding to the input video.
- the encoder may include an inverse quantization module 107.
- the inverse quantization module 107 is used to perform the inverse operation of the quantization module 105, inverse quantizing (scaling) the quantized transform coefficients.
- the encoder may include an inverse transform module 108.
- the inverse transform module 108 is used to reconstruct residual information using the transform coefficient values output by the inverse quantization module 107.
- the encoder may include a loop filter 109.
- the loop filter 109 is used to perform filtering operations to improve the quality of the reconstructed image and improve compilation efficiency.
- the loop filter 109 may include a deblocking filter, a Sample Adaptive Offset (SAO) filter, and an adaptive loop filter, etc.
- the image filtered by the loop filter 109 is output or stored in the decoded image buffer for use as a reference image.
- the decoder may include an entropy decoding module 201 in some embodiments.
- the entropy decoding module 201 performs entropy decoding on the input bitstream to extract transform coefficient information, intra-frame coding information, inter-frame coding information, and other information for each region.
- the entropy decoding module 201 can obtain binary codes for transform coefficient information of a specific region from the input bitstream.
- the entropy decoding module 201 obtains quantized transform coefficients by performing inverse binarization on the binary codes.
- the decoder may include an inverse transform module 203.
- the inverse transform module 203 is used to reconstruct the residual value using the inverse-quantized transform coefficients.
- the decoder may include an intra-prediction module 204.
- the intra-prediction module 204 generates prediction blocks using intra-coding information and recovered samples from the current image.
- the intra-coding information may include at least one of an intra-prediction mode, a most probable mode (MPM) flag, and an MPM index.
- the intra-prediction unit 252 predicts sample values for the current block by using recovered samples located to the left and/or above the current block as reference samples.
- the decoder may include a motion compensation module 205.
- the motion compensation module 205 generates prediction blocks using a reference image and inter-frame coding information stored in the decoded image buffer.
- the inter-frame coding information may include a set of motion information for the current block (reference image index, motion vector information, etc.) used for the reference block.
- the decoder also reconstructs the original pixel values by adding the residual values obtained by the inverse transform module 203 to the prediction blocks obtained by the intra-prediction module 204 or the motion compensation module 205.
- the decoder may include a loop filter 206.
- the loop filter 206 is used to perform filtering operations to improve the quality of the reconstructed image and improve compilation efficiency.
- the entropy coding module in the encoder performs entropy coding, a lossless coding method that follows the principle of entropy without losing any information.
- Common entropy coding methods include Shannon coding, Huffman coding, Exponential Golomb coding, and arithmetic coding.
- a suitable entropy coding model can be selected based on the distribution of the symbols to be encoded.
- the entropy coding module can be encoded using Context-based Adaptive Binary Arithmetic Coding (CABAC).
- CABAC Context-based Adaptive Binary Arithmetic Coding
- the core algorithm of CABAC is adaptive binary arithmetic coding, which combines the context model with adaptive binary arithmetic coding, performing arithmetic coding based on the probability model and the binary values of the syntax elements, and then updating the context model.
- the input to a CABAC encoder is the value of the syntax element to be encoded.
- the value of the syntax element can be binary or non-binary.
- the output of a CABAC encoder is the encoded bits.
- the encoding process of a CABAC encoder may include the following steps 1 to 3:
- Step 1 Binary conversion.
- CABAC encodes the values of syntax elements in slice data. Before performing arithmetic encoding, these syntax elements need to be converted into binary strings suitable for binary arithmetic encoding in a certain way. This conversion process is called binarization.
- step 1 is for the value of non-binary syntax elements. If the input of CABAC is the value of binary syntax elements, step 1 is skipped and execution starts directly from step 2.
- CABAC's binary conversion algorithms include: Fixed-length Binary Conversion (FL), Truncated Rice (TR), Truncated Binary (TB), and the k-th order Exp-Golomb (EGK).
- FL Fixed-length Binary Conversion
- TR Truncated Rice
- TB Truncated Binary
- EGK k-th order Exp-Golomb
- Fixed-length binary encoding is a binary algorithm that converts the values of syntax elements into fixed-length binary symbols.
- a fixed-length binary encoding scheme can be chosen. For example, given a syntax element with value x, and 0 ⁇ x ⁇ cMax, the fixed-length binary symbol string of x can be obtained directly using the decimal-to-binary conversion method, where the length of the fixed-length binary symbol string of x is... Where lFL is the length of the binary symbol string, This indicates rounding up to the nearest integer.
- Truncated Rice code is formed by concatenating a prefix string and a suffix string.
- cMax is the maximum value of the syntax element to be binaryized
- R is the Rice parameter
- V is the value of the syntax element to be binaryized.
- the binary code of the K-order exponential Golomb code is also formed by concatenating a prefix and a suffix.
- the prefix is composed of...
- the value is composed of the unary code corresponding to the value; the suffix part can be calculated by using the binary value of x+2k(1-2l(x)) with a length of k+l(x) bits.
- ⁇ 0 and ⁇ 1 are the update rates of the probabilities of the two models
- p0 (t) and p1 (t) are the prediction probabilities of the two models in the t-th iteration of the biprobability model
- x(t) is the sign of the biprobability model in the t-th iteration
- p0 (t+1) and p1 (t+1) are the prediction probabilities of the two models in the (t+1)-th iteration of the biprobability model.
- the probabilistic prediction model of the H.266/VVC standard is to obtain the prediction probability of LPS by averaging the prediction probabilities obtained from the biprobability model.
- q(t) is the b-bit integer representation of p(t).
- each context model used by a syntax element is specified by a unique context index ctxId.
- Each context model involves two types of variables: the update rate shiftIdx, which controls the rate of model probability updates, and pStateIdx0 ( q0 (t) in the probability prediction model) and pStateIdx1 ( q1 (t) in the probability prediction model), which predict the probability states.
- the probability states pStateIdx0 and pStateIdx1 are continuously updated using the probability prediction model. Furthermore, the predicted probability pState of LPS is obtained by averaging the predicted probabilities obtained from the dual-probability model.
- SliceQPy is the quantization parameter of the luminance signal.
- corresponding weights can be added to pStateIdx0 and pStateIdx1 predicted by the biprobability model, and a mechanism can be set to fine-tune shift0 and shift1 based on the encoded symbol being 0 or 1.
- ⁇ is a weight selected from the predefined set ⁇ 10,12,16,20,22 ⁇ .
- Three different weights can be predefined for each context model of I, B and P type slices.
- the weights of I type slices are only allowed to be used for intra-frame slices, while the weights of B type slices and P type slices are selected based on the initialization type initType obtained by sh_cabac_init_flag.
- I type slices are slices that use only the current image for reconstruction.
- B type slices are slices that use at most two motion vectors and reference image indices, also known as bidirectional prediction slices.
- P type slices are slices that use at most one motion vector and reference image index.
- the CABAC used in the H.266/VVC standard has two probabilistic states, which are updated using short and long windows respectively.
- the size of the long and short windows is specified by the shiftIdx item in the standard document, which controls the different update rates of the long and short windows.
- the size of the long and short windows is fine-tuned based on the initial value of shiftIdx according to the encoded sign being 0 or 1.
- the update range of both the long and short windows is -7 to 7, and the lower limit of the window size is set to 2.
- Step 3 Binary arithmetic encoding.
- Binary arithmetic coding performs arithmetic coding on each binary symbol after the value of the current syntax element is binaryized, according to its probability model parameters, to obtain the final output bitstream.
- Binary arithmetic coding is based on recursive interval partitioning, saving the coding interval and its lower bound during the recursive process.
- the H.266/VVC standard includes two coding methods: conventional coding and bypass coding.
- Conventional coding uses an adaptive probability model for encoding; bypass coding uses equal probability, and its probability state does not need to be updated.
- the input to a conventional encoder is a context model (shiftIdx, pStateIdx0, pStateIdx1) and the binary symbol to be encoded (Bin).
- the encoder's state consists of the current encoding range width (Range) and the range lower limit (Low).
- the initial value of Range is 510, and the initial value of the range lower limit (Low) is 0.
- the encoding process may include the following steps (1) to (4):
- Step 1 Calculate the interval widths R LPS and R MPS corresponding to LPS.
- pState is the prediction probability of MPS, and the XOR operation... This keeps it at the predicted probability of LPS.
- Step 2 Update the encoding range width (Range) and the lower limit (Low).
- Step 3 Renormalize the encoding range width (Range).
- the encoding range width (Range) As the encoding range width (Range) is updated, its value may become less than 256 (the range width is initialized to 510). In this case, renormalization is required.
- the renormalization method is to simultaneously shift both the lower bound (Low) and the encoding range width (Range) to the left until the value of the encoding range width (Range) is greater than or equal to 256.
- the bits shifted out by the lower bound (Low) become the encoded output bits.
- Step 4 Update the context model using the encoded binary symbols.
- the entropy decoding module in the decoder performs the inverse operation of the entropy encoding module in the encoder.
- the entropy decoding module performs the same operation as the CABAC encoding part, which may include: decoding the corresponding binary symbol by reading the lower limit of the range written to the bitstream.
- the input to the CABAC decoder is the lower bound of the interval and the index corresponding to the syntax element.
- the output of the CABAC decoder is the binary value of the corresponding syntax element.
- the encoding process of the CABAC decoder may include the following steps a and b:
- Step a Obtain the context.
- Step a (obtaining context) above may include the following steps a1 and a2:
- Step a1 Determine which syntax element is being decoded based on the standard syntax element structure. Obtain the bypass flag (bypassFlag) and possible context model index (ctxIdx) of the current symbol based on the position of the encoded binary symbol in the bitstream, the index of the corresponding binary symbol (binIdx), and the context reference information.
- bypassFlag bypass flag
- ctxIdx possible context model index
- Step a2 Based on the information obtained in step a1, obtain the update rate shiftIdx, the probability state pStateIdx0, the probability state pStateIdx1, the weights of pStateIdx0 and pStateIdx1, and use these to obtain the estimated probability of the current symbol.
- step 2 The implementation of obtaining the estimated probability of the current symbol by updating the rate shiftIdx, probability state pStateIdx0, probability state pStateIdx1, and the weights of pStateIdx0 and pStateIdx1 can refer to step 2 above. To avoid redundancy, it will not be explained in detail here.
- Step b Binary arithmetic decoding.
- step b (binary arithmetic decoding) includes two decoding methods: conventional decoding and bypass decoding.
- Conventional decoding uses an adaptive probability model for decoding; bypass decoding is performed with equal probability, and its probability state does not need to be updated.
- the two decoding methods are distinguished by the bypass flag obtained in step a.
- the encoder input for conventional decoding is the context model (update rate shiftIdx, probability state pStateIdx0, probability state pStateIdx1) and the current encoding interval width Range.
- the initial value of the encoding interval width Range is 510, and the lower limit m_value of the interval is obtained by reading bytes from the bitstream.
- the basic principle of conventional decoding is as follows: after obtaining the corresponding context used for decoding the current symbol, the size of the interval corresponding to the LPS symbol is calculated based on the estimated probability and the current encoding interval width m_Range. The lower limit of the interval is compared with the size of the current encoding interval to determine whether the symbol symbol is 1 or 0.
- the specific decoding process may include the following steps 1 to 4:
- Step 1 Calculate the interval widths R LPS and R MPS corresponding to LPS.
- pState is the prediction probability of MPS, and the XOR operation... This keeps it at the predicted probability of LPS.
- Step 2 Compare the lower limit m_Value of the interval with the corresponding sub-interval positions of MPS and LPS.
- Step 3 Renormalize the encoding range width (Range).
- the encoding range width (Range) is updated, its value may become less than 256 (the range width is initialized to 510). In this case, renormalization is required.
- the renormalization method is to simultaneously shift both the lower bound (Low) and the encoding range width (Range) to the left until the value of the encoding range width (Range) is greater than or equal to 256.
- the bits shifted out by the lower bound (Low) become the encoded output bits.
- Step 4 Reconstruct the value of the syntax element
- reconstructing the value of a syntax element may include: performing the reverse process of binarizing the binary string obtained by decoding the value of the syntax element, thereby re-adding the value of the syntax element.
- Some embodiments of this application involve entropy encoding of the value of the syntax element tu_cb_coded_flag.
- the syntax element tu_cb_coded_flag is first described below.
- the syntax element tu_cb_coded_flag is a syntax element in the H.266/VVC standard that identifies whether the transform unit (TU) of the Cb chroma component of the current coding block is a non-zero block. That is, it identifies whether the transform unit of the current Cb chroma component contains non-zero transform coefficients.
- tu_cb_coded_flag When the value of the syntax element tu_cb_coded_flag is 1, it indicates that the transform unit of the current Cb chroma component contains non-zero transform coefficients. When the value of the syntax element tu_cb_coded_flag is 0, it indicates that the transform unit of the current Cb chroma component does not contain non-zero transform coefficients. When tu_cb_coded_flag does not exist, it is assumed that the transform unit of the current Cb chroma component does not contain non-zero transform coefficients.
- the process of entropy encoding the value of tu_cb_coded_flag in related technologies may include the following steps I to IV:
- Step 1 Convert the value of tu_cb_coded_flag to binary.
- Table 1 Syntax elements and input parameters based on the basic binaryization scheme.
- Table 1 above shows the method and input parameters for binarying the value of tu_cb_coded_flag.
- the binarying of the value of tu_cb_coded_flag uses a fixed-length encoding scheme, and the maximum value of tu_cb_coded_flag is 1.
- Step II Context model selection for tu_cb_coded_flag.
- context model selection for tu_cb_coded_flag is to determine the context index value ctxIdx, which can be achieved in the following ways:
- Step 1 Determine the set of context model index values for tu_cb_coded_flag based on the initialization type of tu_cb_coded_flag.
- the H.266/VVC standard provides a table of context index values (ctxIdx) for each syntax element. These index values determine the initial value (initValue) for each syntax element.
- the H.266 standard lists a table indicating the table number for each syntax element and the set of context index values (ctxIdx) corresponding to different initialization types. Taking the H.266/VVC standard as an example, Table 2 shows the set of context index values (ctxIdx) for tu_cb_coded_flag under different initialization types.
- the initialization type (initType) is determined by the frame type and sh_cabac_init_flag.
- sh_cabac_init_flag is determined by pps_cabac_init_present_flag.
- pps_cabac_init_present_flag 1
- sh_cabac_init_flag exists in the segment header referencing the Picture Parameter Set (PPS)
- PPS Picture Parameter Set
- sh_cabac_init_flag is determined to be 1.
- a value of 0 for pps_cabac_init_present_flag indicates that sh_cabac_init_flag does not exist in the fragment header that references PPS, and the value of sh_cabac_init_flag is determined to be 0.
- initType The specific initialization type (initType) is determined by the frame type and sh_cabac_init_flag as follows:
- the initialization type initType of I type frames is 0, the initialization type initType of P type frames is 2 or 1 depending on sh_cabac_init_flag, and the initialization type initType of B type frames is 1 or 2 depending on sh_cabac_init_flag.
- Table 3 shows the initial values ⁇ initValue ⁇ and update rates ⁇ shiftIdx ⁇ for the six context models of ⁇ tu_cb_coded_flag ⁇ in the VVC reference software test platform (VVC TEST MODEL, VTM).
- the initial value ⁇ initValue ⁇ is used to calculate the probability state of the context model, and the update rate ⁇ shiftIdx ⁇ controls the update rate of the model probability.
- the inputs ( ⁇ shiftIdx ⁇ , ⁇ pStateIdx0 ⁇ , and ⁇ pStateIdx1 ⁇ ) of the regular encoder can be obtained based on these two parameters.
- the implementation method for obtaining ⁇ shiftIdx ⁇ , ⁇ pStateIdx0 ⁇ , and ⁇ pStateIdx1 ⁇ based on the initial value ⁇ initValue ⁇ and update rate ⁇ shiftIdx ⁇ is the same as in step 2 above, and will not be described in detail here to avoid redundancy.
- Step 2 Determine the index offset binIdx based on intra_bdpcm_chroma_flag.
- the index offset ⁇ binIdx ⁇ is determined based on ⁇ intra_bdpcm_chroma_flag ⁇ . This can include: when ⁇ intra_bdpcm_chroma_flag ⁇ is true, the index offset ⁇ binIdx ⁇ of ⁇ tu_cb_coded_flag ⁇ is parsed as 1; otherwise, it is parsed as 0. See Table 4 below for details.
- Table 4 assigns context model offsets to syntax elements.
- Step 3 Determine the context model index value of tu_cb_coded_flag based on the index offset binIdx of tu_cb_coded_flag.
- mapping relationship between the initialization type initType of tu_cb_coded_flag, intra_bdpcm_chroma_flag, and the context model index value ctxIdx can be shown in Table 5 below:
- mapping relationships shown in Table 5 can also be calculated using expressions.
- Step III Perform arithmetic encoding on the value of tu_cb_coded_flag.
- arithmetic coding is started based on tu_cb_coded_flag and the context model (shiftIdx, pStateIdx0, pStateIdx1) determined by the initial value initValue and the update rate shiftIdx.
- the interval is divided, and the corresponding encoded data of tu_cb_coded_flag is obtained.
- Step IV Update the probability model corresponding to the context model used by tu_cb_coded_flag.
- the probability states pStateIdx0 and pStateIdx1 of the context model used are updated according to the value of tu_cb_coded_flag.
- the entropy decoding process for the value of tu_cb_coded_flag in related technologies may include the following steps (i) to (iv):
- Step 1 Determine the set of context model index values for tu_cb_coded_flag based on the initialization type of tu_cb_coded_flag.
- the implementation method for determining the context model index value set of tu_cb_coded_flag based on the initialization type of tu_cb_coded_flag can refer to step 1 above. To avoid redundancy, it will not be explained in detail here.
- Step 2 Determine the index offset binIdx based on intra_bdpcm_chroma_flag.
- step 2 The implementation of determining the index offset binIdx based on intra_bdpcm_chroma_flag can be found in step 2 above. To avoid redundancy, it will not be explained in detail here.
- Step 3 Determine the context model index value of tu_cb_coded_flag based on the index offset binIdx of tu_cb_coded_flag, and perform arithmetic decoding.
- Arithmetic decoding may include determining the symbol corresponding to tu_cb_coded_flag based on the interval division and the value corresponding to the lower limit of the interval in the bitstream.
- the value of this symbol may be 0 or 1.
- Step 4 Assign the value of the symbol obtained by arithmetic decoding to tu_cb_coded_flag.
- the value (0 or 1) of the symbol corresponding to tu_cb_coded_flag is assigned to the value of tu_cb_coded_flag, and the value of tu_cb_coded_flag may be 0 or 1.
- Step 5 Update the probability model corresponding to the context model used by tu_cb_coded_flag.
- step IV The implementation of the probability model corresponding to the context model used to update tu_cb_coded_flag can be referred to in step IV above. To avoid redundancy, it will not be explained in detail here.
- ⁇ tu_cb_coded_flag ⁇ corresponds to two context models based on the initialization type, and then determines which of the two context models to use based on ⁇ intra_bdpcm_chroma_flag ⁇ (an identifier indicating whether the current CU uses Chroma BDPC).
- the relevant technologies only determine the context model of tu_cb_coded_flag based on the initialization type initType and the corresponding value of intra_bdpcm_chroma_flag, without considering the spatial characteristics of tu_cb_coded_flag, which may lead to an inaccurate context model of tu_cb_coded_flag.
- Some embodiments of this application modify the context model used when entropy encoding the syntax element tu_cb_coded_flag of a chroma block. Specifically, the context model index used when encoding the current tu_cb_coded_flag is determined based on the values of tu_cb_coded_flag in the chroma blocks above and to the left of the current chroma block, and the corresponding number and initial value of the context models are changed.
- the implementation of entropy encoding of tu_cb_coded_flag may include the following steps A to D:
- Step A Convert the value of tu_cb_coded_flag to binary.
- step I The implementation of binary conversion of the value of tu_cb_coded_flag can be referred to step I above. To avoid redundancy, it will not be explained in detail here.
- Step B Select the context model for entropy encoding of the tu_cb_coded_flag value.
- step B above selecting the context model for entropy encoding of the tu_cb_coded_flag value
- step B4 determines the context model ctxIdx used for entropy encoding of the tu_cb_coded_flag value. This can be achieved through steps B1 to B4 as follows:
- Step B1 Determine the set of context model index values for tu_cb_coded_flag based on the initialization type of tu_cb_coded_flag.
- the initialization type ⁇ initType ⁇ of ⁇ tu_cb_coded_flag ⁇ can include three values: 0, 1, and 2.
- the number of initialization types ⁇ initType ⁇ for each type of ⁇ tu_cb_coded_flag ⁇ changes from 2 to 6.
- the initial values of the context model parameters corresponding to index numbers 1 to 17 are preset empirical values.
- initial values of the context model parameters corresponding to indices 1 to 17 are shown in Table 7 below.
- the context model probability at the end of each frame encoding of each sequence can be recorded, and the initial values of the context model parameters corresponding to index numbers 1 to 17 can be set according to historical data.
- Step B2 Determine the value of tu_cb_coded_flag for the coded chroma block adjacent to the current chroma block.
- the encoded chroma blocks adjacent to the current chroma block may include a chroma block to the left of the current chroma block and a chroma block above the current chroma block. Therefore, based on the input encoding structure, chroma position, and chroma components, the chroma block to the left of the current chroma block and the chroma block above the current chroma block can be determined. Then, the values of ⁇ tu_cb_coded_flag ⁇ for the chroma block to the left of the current chroma block and the chroma block above the current chroma block are obtained.
- Step B3 Determine the index value of the context model of the current chroma block based on the value of tu_cb_coded_flag of the coded chroma block adjacent to the current chroma block and the value of intra_bdpcm_chroma_flag of the current chroma block.
- determining the index value of the context model of the current chroma block based on the value of tu_cb_coded_flag of the coded chroma block adjacent to the current chroma block and the value of intra_bdpcm_chroma_flag of the current chroma block may include the following steps B31 and B32:
- Step B31 Calculate the sum of the values of tu_cb_coded_flag of the coded chroma blocks adjacent to the current chroma block.
- the values of ⁇ tu_cb_coded_flag ⁇ of the chroma block to the left of the current chroma block and the chroma block above the current chroma block are accumulated to obtain the sum of the values of ⁇ tu_cb_coded_flag ⁇ of the encoded chroma blocks adjacent to the current chroma block, which is then used to calculate ⁇ Neibor_CbfcbFlag_Sum ⁇ .
- the value of ⁇ tu_cb_coded_flag ⁇ may be 0 or 1
- the value of ⁇ Neibor_CbfcbFlag_Sum ⁇ may be 0, 1, or 2.
- Step B32 Determine the index value of the context model of the current chroma block based on the sum of the values of tu_cb_coded_flag of the coded chroma blocks adjacent to the current chroma block and the value of intra_bdpcm_chroma_flag of the current chroma block.
- determining the index value of the context model of the current chroma block based on the sum of the values of tu_cb_coded_flag of the coded chroma blocks adjacent to the current chroma block and the value of intra_bdpcm_chroma_flag of the current chroma block may include: when the value of intra_bdpcm_chroma_flag is 0 (bdpcm Chroma mode is not used), the index value of the context model of the current chroma block is selected from the first three context model indices of the context index set corresponding to the initialization type.
- the index value of the current chroma block's context model is selected from the last three context model indices of the context index set corresponding to the initialization type.
- mapping relationship between the value of intra_bdpcm_chroma_flag of the current chroma block 31, the value of tu_cb_coded_flag of the chroma block 32 above the current chroma block 31, the value of tu_cb_coded_flag of the chroma block 33 to the left of the current chroma block 31, and the index value ctxIdx of the context model of the current chroma block 31 is as follows:
- the correspondence between the value of intra_bdpcm_chroma_flag of the current chroma block, the value of tu_cb_coded_flag of the coded chroma block adjacent to the current chroma block, Neibor_CbfcbFlag_Sum, and the index value ctxIdx of the context model of the current chroma block 31 can be shown in Table 8 below:
- mapping relationship between the context model index ctxId, the initialization type initType, the value of intra_bdpcm_chroma_flag, the value of tu_cb_coded_flag of the coded chroma block adjacent to the current chroma block, and Neibor_CbfcbFlag_Sum in some embodiments of this application can be shown in Table 9 below:
- Step C Perform arithmetic encoding on the value of tu_cb_coded_flag of the current chroma block according to the index value of the context model of the current chroma block, and update the model probability of the context model.
- the method may further include: obtaining the initial value of the context model of tu_cb_coded_flag when encoding a frame begins and the probability of the context model of tu_cb_coded_flag when encoding ends; if the difference between the probability of the context model when encoding begins and the probability of the context model when encoding ends is less than a first threshold, then when encoding the next frame image, the context model of tu_cb_coded_flag is not initialized, but the context model of tu_cb_coded_flag when encoding the frame image ends is used instead.
- the above embodiments classify chroma blocks using the Chroma BDPCM mode and chroma blocks not using the Chroma BDPCM mode based on the context of the tu_cb_coded_flag of the upper and left chroma blocks.
- six context models can be included.
- this context classification method can be applied only to chroma blocks not using the Chroma BDPCM mode, while only one fixed context column is set for chroma blocks using the Chroma BDPCM mode. That is, under the same initial type, four context models are set for tu_cb_coded_flag, with the first three columns representing the case without Chroma BDPCM and the last column representing the case with Chroma BDPCM. This can be specifically illustrated in Table 10 below:
- the values of the intra_bdpcm_chroma_flag of the coded chroma blocks adjacent to the current chroma block can be obtained first, and when calculating the sum of the values of tu_cb_coded_flag of the coded chroma blocks adjacent to the current chroma block (Neibor_CbfcbFlag_Sum), only the values of tu_cb_coded_flag of adjacent coded chroma blocks with the same intra_bdpcm_chroma_flag value as the current chroma block are summed.
- the video decoding method may include the following steps:
- the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block. That is, the first syntax element is tu_cb_coded_flag.
- the entropy-coded data of the first syntax element of the first chroma block can be obtained from the video bitstream according to a standard syntax element structure.
- the first chroma block is the chroma block corresponding to the blue component Cb of the current coding block. In other embodiments, the first chroma block is the chroma block corresponding to the red component Cr of the current coding block.
- the adjacent decoded chroma blocks of the first chroma block refer to: chroma blocks that are adjacent to the first chroma block and have been decoded before the first chroma block was decoded.
- the adjacent decoded chroma blocks of the first chroma block may include: a chroma block located above the first chroma block and/or a chroma block located to the left of the first chroma block; the above step S42 (obtaining the spatial information of the first chroma block based on the adjacent decoded chroma blocks of the first chroma block) may include: obtaining the spatial information of the first chroma block based on the chroma blocks located above the first chroma block and/or the chroma blocks located to the left of the first chroma block.
- step S42 (obtaining the spatial information of the first chroma block based on the adjacent decoded chroma blocks of the first chroma block) may include: obtaining the spatial information of the first chroma block based on whether the transform block of the adjacent decoded chroma block is a non-zero block.
- determining the target context model based on the spatial domain information of the first chroma block may include: selecting the target context model from a preset context model set based on the spatial domain information of the first chroma block.
- determining the target context model based on the spatial domain information of the first chroma block may include: determining the target context model based on the spatial domain information of the first chroma block, the initialization type (initType) of the first syntax element of the first chroma block, and the value of the second syntax element of the first chroma block;
- the second syntax element is a syntax element that identifies whether the chroma block uses a block-based differential pulse code modulation mode. That is, the second syntax element is intra_bdpcm_chroma_flag.
- the video decoding method when obtaining the entropy-encoded data of the first syntax element of the first chroma block (identifying whether the transform block of the first chroma block is a non-zero block) and obtaining the value of the first syntax element of the first chroma block based on the entropy-encoded data of the first syntax element of the first chroma block, firstly obtains the spatial domain information of the first chroma block based on the adjacent decoded chroma blocks, then determines the target context model based on the spatial domain information of the first chroma block, and finally obtains the value of the first syntax element of the first chroma block based on the target context model and the entropy-encoded data of the first syntax element of the first chroma block.
- the video decoding method provided in the above embodiments obtains the spatial domain information of the first chroma block based on the adjacent decoded chroma blocks and determines the context model for entropy decoding of the value of the first syntax element of the first chroma block based on the spatial domain information of the first chroma block
- the above embodiments can combine the spatial characteristics of the first syntax element to determine the context model for entropy decoding of the first syntax element, thereby improving the accuracy of the context model for entropy decoding of the first syntax element.
- the video decoding method may include the following steps:
- the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block.
- step S503 (obtaining the spatial information of the first chroma block based on the value of the first syntax element of the adjacent decoded chroma blocks) may include: summing the values of the first syntax elements of the adjacent decoded chroma blocks to obtain the spatial information of the first chroma block.
- the adjacent decoded chroma blocks may include: a chroma block above the first chroma block and a chroma block to the left of the first chroma block. If the value of the first syntax element of the chroma block above the first chroma block is 1 and the value of the first syntax element of the chroma block to the left of the first chroma block is 1, then the spatial information of the first chroma block can be determined to be 2.
- the adjacent decoded chroma blocks only include the chroma block located to the left of the first chroma block, and the value of the first syntax element of the chroma block located above the first chroma block is 1, then the spatial information of the first chroma block can be determined to be 1.
- the adjacent decoded chroma blocks may include: a chroma block above the first chroma block and a chroma block to the left of the first chroma block. If the value of the first syntax element of the chroma block above the first chroma block is 0 and the value of the first syntax element of the chroma block to the left of the first chroma block is 0, then the spatial information of the first chroma block can be determined to be 0.
- step S503 (obtaining the spatial information of the first chroma block based on the value of the first syntax element of the adjacent decoded chroma block) may include the following steps 5031 and 5032:
- Step 5031 Obtain the value of the second syntax element of the adjacent decoded chroma block.
- the second syntax element is a syntax element that identifies whether a chroma block uses a block-based differential pulse code modulation mode. That is, it determines whether the adjacent decoded chroma blocks use a block-based differential pulse code modulation mode.
- Step 5032 Summate the values of the first syntax elements of the adjacent decoded chroma blocks whose values of the second syntax elements are the same as those of the first chroma block, to obtain the spatial information of the first chroma block.
- the value of the second syntax element of the first chroma block is 0.
- the adjacent decoded chroma blocks may include: a chroma block above the first chroma block and a chroma block to the left of the first chroma block.
- the value of the first syntax element of the chroma block above the first chroma block is 1, the value of the second syntax element of the chroma block above the first chroma block is 0, the value of the first syntax element of the chroma block to the left of the first chroma block is 1, and the value of the second syntax element of the chroma block above the first chroma block is 1. Since only the chroma block above the first chroma block has the same value as the second syntax element of the first chroma block, the spatial information of the first chroma block is 1.
- obtaining the initialization type of the first syntax element of the first chroma block may include: determining the value of the syntax element sh_cabac_init_flag based on the value of the syntax element pps_cabac_init_present_flag, and determining the initialization type of the first syntax element of the first chroma block based on the value of the syntax element sh_cabac_init_flag and the frame type (I type, B type, or P type) corresponding to the first chroma block.
- the implementation of obtaining the initialization type of the first syntax element of the first chroma block may include:
- the initialization type of the first syntax element of the first chroma block is determined to be the first initialization type; if the frame type corresponding to the first chroma block is type P, and the value of sh_cabac_init_flag is 1, then the initialization type of the first syntax element of the first chroma block is determined to be the second initialization type; if the frame type corresponding to the first chroma block is type P, and the value of sh_cabac_init_flag is 0, then the initialization type of the first syntax element of the first chroma block is determined to be the third initialization type; if the frame type corresponding to the first chroma block is type B, and the value of sh_cabac_init_flag is 1, then the initialization type of the first syntax element of the first chroma block is determined to be the third initialization type; if the frame type corresponding to the first chroma block is type B, and the value of sh_cabac_init_flag is 1, then the initial
- the second syntax element is a syntax element that identifies whether the chroma block uses a block-based differential pulse code modulation mode.
- the value of the second syntax element of the first chroma block is 0, and when the first chroma block does not use a block-based differential pulse code modulation mode, the value of the second syntax element of the first chroma block is 1.
- the initialization type is a first initialization type, a second initialization type, or a third initialization type, and the first initialization type, the second initialization type, and the third initialization type each correspond to a context model index set; determining the target index set based on the initialization type may include: when the initialization type is a first initialization type, determining the context model index set corresponding to the first initialization type as the target index set; when the initialization type is a second initialization type, determining the context model index set corresponding to the second initialization type as the target index set; when the initialization type is a first initialization type, determining the context model index set corresponding to the third initialization type as the target index set.
- the context model index set corresponding to the first initialization type is ⁇ 0, 1, 2, 3, 4, 5 ⁇
- the context model index set corresponding to the second initialization type is ⁇ 6, 7, 8, 9, 10, 11 ⁇
- the context model index set corresponding to the third initialization type is ⁇ 12, 13, 14, 15, 16, 17 ⁇ . Therefore, when the initialization type is the first initialization type, the target index set is ⁇ 0, 1, 2, 3, 4, 5 ⁇ ; when the initialization type is the first initialization type, the target index set is ⁇ 6, 7, 8, 9, 10, 11 ⁇ ; and when the initialization type is the first initialization type, the target index set is ⁇ 12, 13, 14, 15, 16, 17 ⁇ .
- the context model index set corresponding to the first initialization type is ⁇ 0, 1, 2, 3 ⁇
- the context model index set corresponding to the second initialization type is ⁇ 4, 5, 6, 7 ⁇
- the context model index set corresponding to the third initialization type is ⁇ 8, 9, 10, 11 ⁇ . Therefore, when the initialization type is the first initialization type, the target index set is ⁇ 0, 1, 2, 3 ⁇ ; when the initialization type is the second initialization type, the target index set is ⁇ 4, 5, 6, 7 ⁇ ; and when the initialization type is the third initialization type, the target index set is ⁇ 8, 9, 10, 11 ⁇ .
- the target index set may include: six context model indices (as shown in Table 6 above).
- Step S507 selecting a target model index from the target index set based on the spatial domain information of the first chroma block and the value of the second syntax element of the first chroma block) may include:
- the target model index is selected from the first three context model indices of the target index set according to the spatial domain information of the first chroma block.
- the target model index is selected from the last three context model indices of the target index set according to the spatial domain information of the first chroma block.
- intra_bdpcm_chroma_flag 0
- the target model index is selected from the first three context model indices of the target index set based on the spatial domain information of the first chroma block
- intra_bdpcm_chroma_flag 1
- the target model index is selected from the last three context model indices of the target index set based on the spatial domain information of the first chroma block.
- intra_bdpcm_chroma_flag 0
- the target model index is selected from ⁇ 12, 13, 14 ⁇ according to the spatial domain information of the first chroma block
- intra_bdpcm_chroma_flag 1
- the target model index is selected from ⁇ 15, 16, 17 ⁇ according to the spatial domain information of the first chroma block.
- selecting the target model index from the first three context model indices of the target index set based on the spatial information of the first chroma block may include: when the spatial information of the first chroma block is m, then selecting the (m+1)th context model index in the target index set as the target model index; 0 ⁇ m ⁇ 2.
- 0 ⁇ m ⁇ 2 means that m ⁇ 0 and m ⁇ 2. Since the spatial information of the first chroma block is the sum of the values of at least one first syntax element, m is an integer, and therefore m can be 0, 1, or 2.
- the first context model index in the target index set is selected as the target model index
- the second context model index in the target index set is selected as the target model index
- the third context model index in the target index set is selected as the target model index
- selecting the context model index corresponding to the first syntax element of the first chroma block from the last three context model indices of the target index set based on the spatial information of the first chroma block may include: when the spatial information of the first chroma block is m, then selecting the (m+4)th context model index in the target index set as the context model index corresponding to the first syntax element of the first chroma block, where 0 ⁇ m ⁇ 2.
- the fourth context model index in the target index set is selected as the target model index; when the spatial information of the first chroma block is 1, the fifth context model index in the target index set is selected as the target model index; and when the spatial information of the first chroma block is 2, the sixth context model index in the target index set is selected as the target model index.
- the target index set may include four context model indices (as shown in Table 10 above).
- Step S507 selecting a target model index from the target index set based on the spatial domain information of the first chroma block and the value of the second syntax element of the first chroma block) may include:
- the target model index is selected from the first three context model indices of the target index set according to the spatial domain information of the first chroma block.
- the fourth context model index in the target index set is selected as the target model index.
- selecting the target model index from the first three context model indices of the target index set based on the spatial information of the first chroma block may include: when the spatial information of the first chroma block is m, then the (m+1)th context model index in the target index set is selected as the target model index; 0 ⁇ m ⁇ 2.
- Step 5081 Obtain the initial value and update rate of the target context model based on the target model index.
- the initial value (initValue) of the target context model is 35; the update rate (shiftIdx) of the target context model is 8.
- Step 5082 Construct the target context model based on the initial value and update rate of the target context model.
- SliceQPy is the quantization parameter of the luminance signal.
- determining the target context model based on the target model index may include: reading the target context model based on the target model index.
- the entropy-encoded data of the first syntax element of the first chroma block is arithmetically decoded using the CABAC algorithm to obtain the value of the first syntax element of the first chroma block.
- step b The implementation of arithmetic decoding of the entropy-encoded data of the first syntax element of the first chroma block using the CABAC algorithm can refer to step b above. To avoid redundancy, it will not be described in detail here.
- the probability states pStateIdx0 and pStateIdx1 of the target context model are updated based on the value of tu_cb_coded_flag of the first chroma block.
- the video decoding method provided in some embodiments of this application may further include: determining the context model corresponding to the second chroma block according to the target context model; wherein the first chroma block and the second chroma block are chroma blocks corresponding to two chroma components of the same coding block.
- the first chromaticity block is the chromaticity block corresponding to the Cb component
- the second chromaticity block is the chromaticity block corresponding to the Cr chromaticity component
- the target context model corresponding to the chroma block corresponding to the blue chroma component of the coded block is obtained, the value of the first syntax element of the first chroma block is obtained based on the target context model and the entropy encoding data of the first syntax element of the first chroma block, and the context model corresponding to the chroma block corresponding to the red chroma component of the coded block is determined according to the determined target context model.
- determining the context model corresponding to the second chroma block based on the target context model may include: determining the target context model as the context model corresponding to the second chroma block.
- the video encoding method may further include:
- the value of the first syntax element of the second chroma block is obtained.
- step S61 can refer to the implementation of step S42.
- step S42 the spatial information of the first chroma block is obtained based on the adjacent decoded chroma blocks of the first chroma block
- step S61 the spatial information of the first chroma block is obtained based on the adjacent encoded chroma blocks of the first chroma block.
- the adjacent decoded chroma blocks of the first chroma block during the decoding process and the adjacent encoded chroma blocks of the first chroma block during the encoding process are the same chroma blocks.
- the implementation method of determining the target context model based on the spatial domain information of the first chroma block can refer to the implementation method of determining the target context model based on the spatial domain information of the first chroma block in the above video decoding method. To avoid redundancy, it will not be described in detail here.
- entropy encoding is performed on the value of the first syntax element of the first chroma block to obtain the entropy encoded data of the first syntax element of the first chroma block.
- the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block.
- the video encoding method when performing entropy encoding on the value of the first syntax element of the first chroma block, first obtains the spatial information of the first chroma block based on its neighboring encoded chroma blocks, then determines a target context model based on the spatial information of the first chroma block, and then performs entropy encoding on the value of the first syntax element of the first chroma block based on the target context model to obtain the entropy encoded data of the first syntax element of the first chroma block.
- the video encoding method provided in the above embodiments obtains the spatial information of the first chroma block based on its neighboring encoded chroma blocks and determines the context model for entropy encoding on the value of the first syntax element of the first chroma block based on the spatial information of the first chroma block
- the above embodiments can combine the spatial characteristics of the first syntax element to determine the context model for entropy encoding on the first syntax element, thereby improving the accuracy of the context model for entropy encoding on the first syntax element.
- the video decoding apparatus 700 may include:
- the first acquisition module 71 is used to acquire the entropy encoded data of the first syntax element of the first chroma block;
- the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block;
- the second acquisition module 72 is used to acquire the spatial information of the first chroma block based on the adjacent decoded chroma blocks of the first chroma block;
- the determination module 73 is used to determine the target context model based on the spatial information of the first chroma block
- the entropy decoding module 74 is used to obtain the value of the first syntax element of the first chroma block based on the target context model and the entropy encoded data of the first syntax element of the first chroma block.
- the video decoding device provided in the above embodiments can execute the video decoding method provided in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
- the video encoding apparatus 800 may include:
- the acquisition unit 81 is used to acquire the spatial information of the first chroma block based on the adjacent encoded chroma blocks of the first chroma block;
- the determining unit 82 is used to determine the target context model based on the spatial domain information of the first chroma block
- Entropy coding unit 83 is used to entropy code the value of the first syntax element of the first chroma block based on the target context model to obtain the entropy coding data of the first syntax element of the first chroma block; the first syntax element is a syntax element that identifies whether the transform block of the chroma block is a non-zero block.
- the video encoding apparatus provided in the above embodiments can execute the video encoding method provided in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
- Memory configured to store computer programs
- the processor is configured to cause the video decoding device to implement the video decoding method or video encoding method described in any of the above embodiments when a computer program is invoked.
- Some embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the video decoding method described in any of the above embodiments.
- Some embodiments of this application provide a computer program product that, when run on a computer, enables the computer to implement the video decoding method described in any of the above embodiments.
- Some embodiments of this application provide a chip including a memory and a processor.
- the processor may be a logic circuit, an integrated circuit, or a general-purpose processor.
- the memory stores computer instructions, and the processor can implement the video decoding method described in any of the above embodiments by reading the computer instructions stored in the memory.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
本申请提供了一种视频解码方法、视频编码方法及装置,涉及视频编解码技术领域。该视频解码方法包括:获取第一色度块的第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素;根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息;根据所述第一色度块的空域信息确定目标上下文模型;基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。本申请一些实施例用于提升用于对tu_cb_coded_flag的值熵编码的上下文模型的准确率。
Description
本申请要求于2024年7月5日提交国家知识产权局、申请号为202410902887.7、申请名称为“视频解码方法、视频编码方法及装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请一些实施例涉及视频编解码技术领域。更具体地讲,涉及一种视频解码方法、视频编码方法及装置。
熵编码是视频编码中主要流程之一,目前,用于表示色度变换块是否为非零块的语法元素tu_cb_coded_flag的值一般是通过基于上下文的自适应算术编码(Context-based Adaptive Binary Arithmetic Coding,CABAC)进行熵编码的,且通过CABAC对tu_cb_coded_flag的值进行熵编码时是依据tu_cb_coded_flag的初始化类型和当前色度块是否使用了bdpcm chroma模式(语法元素intra_bdpcm_chroma_flag的值)来进行上下文模型选择的。然而,仅依据tu_cb_coded_flag的初始化类型和是否使用了bdpcm chroma模式来进行上下文模型的选择,可能会导致用于对tu_cb_coded_flag的值编码的上下文模型不准确。
本申请示例性的实施方式提供一种视频解码方法、视频编码方法及装置,用于提升用于对tu_cb_coded_flag的值熵编码的上下文模型的准确率。
第一方面,本申请一些实施例提供了一种视频解码方法,包括:
获取第一色度块的第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素;根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息;根据所述第一色度块的空域信息确定目标上下文模型;基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。
第二方面,本申请一些实施例提供了一种视频编码方法,包括:
根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息;根据所述第一色度块的空域信息确定目标上下文模型;基于所述目标上下文模型对所述第一色度块的第一语法元素的值进行熵编码,以获取所述第一色度块的所述第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。
第三方面,本申请一些实施例提供了一种视频解码装置,包括:
第一获取模块,用于获取第一色度块的第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素;第二获取模块,用于根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息;确定模块,用于根据所述第一色度块的空域信息确定目标上下文模型;解码模块,用于基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。
第四方面,本申请一些实施例提供了一种视频编码装置,包括:
获取单元,用于根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息;确定单元,用于根据所述第一色度块的空域信息确定目标上下文模型;熵编码单元,用于基于所述目标上下文模型对所述第一色度块的第一语法元素的值进行熵编码,以获取所述第一色度块的所述第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。
图1示出了本申请一些实施例中的视频编码器的结构示意图;
图2示出了本申请一些实施例中的视频解码器的结构示意图;
图3示出了本申请一些实施例中的确定上下文模型的索引值的流程示意图;
图4示出了本申请一些实施例中的视频解码方法的步骤流程图;
图5示出了本申请另一些实施例中的视频解码方法的步骤流程图;
图6示出了本申请一些实施例中的视频编码方法的步骤流程图;
图7示出了本申请一些实施例中的视频解码装置的结构示意图;
图8示出了本申请一些实施例中的视频编码装置的结构示意图。
为使本申请的目的和实施方式更加清楚,下面将结合本申请示例性实施例中的附图,对本申请示例性实施方式进行清楚、完整地描述,显然,描述的示例性实施例仅是本申请一部分实施例,而不是全部的实施例。需要说明的是,本申请中对于术语的简要说明,仅是为了方便理解接下来描述的实施方式,而不是意图限定本申请的实施方式。除非另有说明,这些术语应当按照其普通和通常的含义理解。术语“包括”和“具有”以及他们的任何变形,意图在于覆盖但不排他的包含,例如,包含了一系列组件的产品或设备不必限于清楚地列出的所有组件,而是可包括没有清楚地列出的或对于这些产品或设备固有的其它组件。说明书中提及“一些实现方式”、“一些实施例”等是表明所描述的实现方式或实施例可包括特定的特征、结构或特性,但可能不一定每个实施例都包括该特定特征、结构或特性。此外,这种短语不一定指同一实现方式。另外,当联系一实施例来描述特定的特征、结构或特性时,认为联系其他实现方式(无论本文是否明确描述)来实现这种特征、结构或特性,是在本领域技术人员的知识范围内的。
本申请实施例涉及视频编解码技术领域,以下首先对本申请一些实施例的视频编解码框架进行说明。视频可以看作是多个视频帧(图像)组成的序列。视频播放可以看作是视频帧按照在序列中的顺序以预设速率(例如:24帧/秒、30帧/秒、60帧/秒)显示。理论上,视频的数据量与视频帧的分辨率是正相关的,视频帧的分辨率越高,则视频的数据量越大。若直接在视频文件中保存全部视频帧的每一个像素点的像素数据,则视频的数据量会非常巨大,进而导致视频难以存储和传输,而视频编译码一定程度上就是为解决该问题而提出的。视频译码可包括:视频编码和视频解码。其中,视频编码可以理解为是对原始视频帧进行压缩的过程,而视频解码则可以理解为是根据压缩后的视频数据进行视频帧重建的过程。
参照图1所示,图1为本申请一些实施例提供的编码器的结构示意图。如图1所示,在一些实施例中,编码器,可包括:编码控制模块101。编码控制模块101用于对视频编码器进行整体控制,包括:进行编码模式选择、编码逻辑控制、码率控制、量化步长控制等。
如图1所示,在一些实施例中,编码器,可包括:帧内预测模块102。帧内预测模块102,用于根据当前图像中的重构样本执行帧内预测,并将帧内编译信息发送到熵编码模块。帧内编码信息可以包括帧内预测模式、最可能模式(Most Probable Mode,MPM)标志和MPM索引中的至少一种。帧内编码信息可以包括参考样本的信息。
如图1所示,在一些实施例中,编码器,可包括:帧间预测模块103。帧间预测模块103用于执行帧间预测以通过使用存储在解码图像缓存(Decoded Picture Buffer,DPB)中的参考图像来预测当前图像。帧间预测模块103可以包括运动估计单元103a和运动补偿单元103b。运动估计单元103a用于进行运动估计(Motion Estimation,ME),以参考重构的参考图像的特定区域获得当前区域的运动向量。运动补偿单元103b使用从运动估计单元103a得到的运动向量值来执行运动补偿。
如图1所示,在一些实施例中,编码器,可包括:变换模块104。变换模块104通过对残差信号进行变换来获得变换系数,该残差信号是输入的输入视频与由帧内预测模块102或帧间预测模块103生成的预测信号之间的差。在一些实施例中,可以使用离散余弦变换(Discrete Cosine Transform,DCT)、离散正弦变换(Discrete Sine Transform,DST)或小波变换。DCT和DST通过将输入图像信号分割成多个块来执行变换。在变换中,编译效率可以根据变换区域中的值的分布和特性而变化。
如图1所示,在一些实施例中,编码器,可包括:量化模块105。量化模块105用于对变换模块104输出的变换系数进行量化,以获取量化的变换系数。
如图1所示,在一些实施例中,编码器,可包括:熵编码模块106。熵编码模块106用于对编码模式、预测模块式、运动向量、量化模块105输出的量化的变换系数等信息进行熵编码,以获取输入视频对应的输出码流。
如图1所示,在一些实施例中,编码器,可包括:反量化模块107。反量化模块107用于执行量化模块105的逆操作,将量化后的变换系数进行逆量化(缩放)。
如图1所示,在一些实施例中,编码器,可包括:反变换模块108。反变换模块108用于通过反量化模块107输出的变换系数值来重构残差信息。
如图1所示,在一些实施例中,编码器,可包括:环路滤波器109。环路滤波器109用于执行滤波操作以改善重构图像的质量并改善编译效率。示例性的,环路滤波器109可以包括:去块滤波器、样本自适应偏移(Sample Adaptive Offset,SAO)和自适应环路滤波器等。经过环路滤波器109滤波后的图像被输出或存储在解码图像缓存中,以用作参考图像。
参照图2所示,图2为本申请一些实施例提供的解码器的结构示意图。如图2所示,在一些实施例中,解码器,可包括:熵解码模块201。熵解码模块201用于对输入码流进行熵解码,以提取每个区域的变换系数信息、帧内编码信息、帧间编码信息等信息。在一些实施例中,熵解码模块201可以从输入码流获得用于特定区域的变换系数信息的二进制码。此外,熵解码模块201通过对二进制码进行逆二值化来获得量化的变换系数。
如图2所示,在一些实施例中,解码器,可包括:反量化模块202。反量化模块202对熵解码模块201输出的量化的变换系数进行逆量化。
如图2所示,在一些实施例中,解码器,可包括:反变换模块203。反变换模块203用于使用反量化的变换系数来重构残差值。
如图2所示,在一些实施例中,解码器,可包括:帧内预测模块204。帧内预测模块204用于通过帧内编码信息和当前图像中的恢复的样本来生成预测块。帧内编码信息可以包括帧内预测模式、最可能模式(MPM)标志和MPM索引中的至少一种。帧内预测单元252通过使用位于当前块的左侧和/或上侧的恢复的样本作为参考样本来预测当前块的样本值。
如图2所示,在一些实施例中,解码器,可包括:运动补偿模块205。运动补偿模块205用于通过参考图片和存储在解码图像缓存中的帧间编码信息来生成预测块。帧间编译信息可以包括用于参考块的当前块的运动信息集(参考图片索引、运动向量信息等)。解码器还通过反变换模块203获取的残差值与帧内预测模块204或运动补偿模块205获取的预测块相加来重构原始像素值。
如图2所示,在一些实施例中,解码器,可包括:环路滤波器206。环路滤波器206用于执行滤波操作以改善重构图像的质量并改善编译效率。
编码器中的熵编码模块执行的熵编码操作是一种按熵原理不丢失任何信息的无损编码方式。常见的熵编码有:香农(Shannon)编码、哈夫曼(Huffman)编码,指数哥伦布编码(Exp-Golomb)和算术编码(arithmetic coding),实际使用中可以根据需要编码的符号的分布情况选择适合的熵编码模型。
在一些实施例中,熵编码模块可以通过基于上下文的自适应算术编码(Context-based Adaptive Binary Arithmetic Coding,CABAC)进行编码。CABAC的核心算法是自适应二进制算术编码,该算法将上下文模型与自适应二进制算术编码结合,根据概率模型和语法元素的二进制值进行算术编码,然后再更新上下文模型。
CABAC编码器的输入为待编码的语法元素的值,语法元素的值可以为二进制,也可以为非二进制,CABAC编码器的输出为编码比特。CABAC编码器的编码过程可包括如下步骤1至步骤3:
步骤1、二进制化。
CABAC编码的是切片(slice)数据中的语法元素的值,在进行算术编码前,需要把这些语法元素按照一定的方法转换成适合进行二进制算术编码的二进制串,这个转换的过程被称为二值化(binarization)。
需要说明的是,步骤1针对的是非二进制的语法元素的值,若CABAC的输入为二进制的语法元素的值,则跳过步骤1,直接从步骤2开始执行。
CABAC的中的二进制化算法包括:定长二进制化(Fixed-length Binary Conversion,FL)、截断莱斯码二进制化(Truncated Rice,TR)、截断二进制码二进制化(Truncated Binary,TB)、K阶指数哥伦布码二进制化(the k-th order Exp-Golomb,EGK)。CABAC会根据不同语法元素的值的不同概率分布特性来选择不同的二进制化方法。
定长二进制化是一种将语法元素的值转换为固定长度二进制符号的二进制算法。当语法元素的值的概率呈均匀分布时,可选用定长编码的二进制化方案。例如:某一给定语法元素的值为x,且0≤x≤cMax,则直接利用十进制数转换为二进制数法得到x的定长二进制符号串,x的定长二进制符号串的长度其中,lFL为二进制符号串的长度,表示向上取整。
截断莱斯码二进制码由前缀串和后缀串拼接形成。前缀值P的计算式为:P=V>>R。其中,若P值小于(cMax>>R),则前缀串由P个1和一个0组成,长度为P+1;若P值大于等于(cMax>>R),则前缀串由(cMax>>R)个1组成,长度为(cMax>>R)。后缀值S的计算式为:S=V-(P<<R),后缀串为S的二元化串,长度为R。当语法元素的值V大于等于cMax时,无后缀串。其中,cMax为待二进制化的语法元素的最大值,R为莱斯参数,V为待二进制化的语法元素的值。
K阶指数哥伦布码二进制化的二进制码同样也是由前缀和后缀拼接形成的。其中,前缀部分由的值所对应的一元码组成;后缀部分可通过使用长度为k+l(x)k+l(x)位的x+2k(1-2l(x))x+2k(1-2l(x))的二进制值来计算。
步骤2、上下文建模。
在多功能视频编码(Versatile Video Coding,H.266/VVC)标准中,采用了双概率模型预测每个上下文的最小概率符号(Least Probable Symbol,LPS)符号概率。H.266/VVC标准的概率预测模型如下所示:p(t+1)=(p0(t+1)+p1(t+1))/2 (1)p0(t+1)=p0(t)·(1-α0)+x(t)·α0 (2)p1(t+1)=p1(t)·(1-α1)+x(t)·ɑ1 (3)
其中,ɑ0、α1为两个模型的概率的更新速率,p0(t)、p1(t)分别为双概率模型的两个模型第t次的预测概率,x(t)双概率模型第t次的符号,p0(t+1)、p1(t+1)分别双概率模型的两个模型第t+1次的预测概率。
即,H.266/VVC标准的概率预测模型为对双概率模型得到的预测概率求均值得到LPS的预测概率。
为了避免乘法运算,α0、α1被限定为α=2-β(β∈N+),则有:q(t+1)=q(t)-(q(t)>>β+x(t)·((2b-1)>>β) (4)
其中,q(t)为p(t)的b比特整数化表示。
H.266/VVC标准中b0=10、b1=14,因此H.266/VVC标准的概率预测模型可以为:p(t)=q(t)·2b+2b-1 (5)
在H.266/VVC标准中,语法元素使用的每个上下文模型都由唯一的上下文索引ctxId指定。每个上下文模型涉及两个类型的变量:控制模型概率更新速率的更新速率shiftIdx,预测概率状态的pStateIdx0(概率预测模型中的q0(t))和pStateIdx1(概率预测模型中的q1(t))。
根据已编码二进制符号的值和更新速率shiftIdx,利用概率预测模型不断更新概率状态pStateIdx0和pStateIdx1。并进一步对双概率模型得到的预测概率求均值可得到LPS的预测概率pState,具体计算式如下所示:pState=pStateldx1+16·pStateldx0 (6)pStateldx0=pStateIdx0-(pStateIdx0>>shift0)+((210-1)·Bin>>shift0) (7)pStateldx1=pStateIdx1-(pStateIdx1>>shift1)+((214-1)·Bin>>shift1) (8)shift0=(shiftIdx>>2)+2 (9)shift1=(shiftIdx&3)+3+shift0 (10)
其中,shift0(概率预测模型中的α0)和shift1(概率预测模型中的α1)分别为预测概率pStateldx0和pStateldx1的更新速率;Bin为己编码的二进制符号;预测概率pStateldx0和pStateldx1分别为10bit和14bit整数化表示。
此外,在编码第一个二进制符号前,如何为该符号初始化其上下文模型是上下文建模的关键技术之一。在H.266/VVC标准中,为每个上下文索引分配了初始initValue和更新速率shiftIdx。initValue用来计算模型概率状态,具体算式如下所示:slopeIdx=initValue>>3 (11)offsetIdx=initValue (12)m=slopeIdx-4 (13)n=(offsetIdx·18)+1 (14)preCtxState=Clip3(1,127,((m·(Clip3(0,63,SliceQPy)-16))>>1)+n) (14)pStateIdx0=preCtxState<<3 (15)pStateIdx1=preCtxState7 (16)
其中,SliceQPy为亮度信号的量化参数。
在一些实施例中,还可以为双概率模型预测的pStateIdx0和pStateIdx1增加对应的权重,同时设置shift0和shift1根据编码符号为0或1进行微调的机制。
通过引入权重来导出用于二进制算数编码的结果概率的操作如下:p=((32-ω)·p0+ω·p1)>>5 (17)
其中,其中ω是从预定义的集合ω∈{10,12,16,20,22}中选择的权重,可以为I、B和P类型切片(slice)的每个上下文模型预先定义三个不同的权重,且I类型的slice的权重仅允许用于帧内切片,而B类型的slice和P类型的slice的权重则基于sh_cabac_init_flag得到的初始化类型initType选取,I类型的slice为仅将当前图像用于重构的切片,B类型的slice使用最多两个运动向量和参考图片索引的切片,也称为双向预测切片,P类型的slice为使用最多一个运动向量和参考图片索引的切片。
H.266/VVC标准中使用的CABAC有两个概率状态,分别用短窗口和长窗口进行更新,长窗口和短窗口的大小由标准文档中的shiftIdx项指定,控制长短窗口不同的更新速率,上述实施例根据编码的符号为0或1针对shiftIdx的初始值对长短窗口的大小进行微调,长短窗口的更新范围都是-7到7,同时会将窗口大小的下限设置为2。
步骤3、二进制算术编码。
二进制算术编码对当前语法元素的值二进制化后的每个二进制符号根据其概率模型参数进行算术编码,得到最后的输出码流。
二进制算术编码基于递归区间划分的方式,在递归过程中保存编码区间和区间下限。H.266/VVC标准中包括两种编码方式:常规编码和旁路编码。常规编码利用自适应的概率模型进行编码;旁路编码则以等概率的方式进行编码,其概率状态无须更新。
H.266/VVC标准中常规编码器的输入是上下文模型(shiftIdx、pStateIdx0、pStateIdx1)和待编码的二进制符号(Bin),编码器的状态是当前编码区间宽度Range和区间下限Low。Range的初始值为510,区间下限Low的初始值为0。编码流程可包括如下步骤⑴至步骤⑷:
步骤⑴、计算LPS对应的区间宽度RLPS和RMPS。
计算LPS对应的区间宽度RLPS和RMPS的算式如下所示:RangeIdx=Range>> 5 (18)pState=pStateIdx1+16· pStateldx0 (19)RLPS=(pStateIdx·pState5)>>1+4 (21)RMPS=Range-RLPS (22)
其中,pState5表示5bit精度的预测概率,当pState>>14=1时,预测概率大于0.5,pState为MPS的预测概率,异或操作使其保持为LPS的预测概率。
步骤⑵、更新编码区间宽度Range和区间下限Low。
更新编码区间宽度Range和区间下限Low的算式如下所示:MPS=pState>>14 (23)
若Bin=LPS,则Low=Low+RMPS,Range=RLPS;
若Bin=MPS,则Low保持不变,Range=RMPS。
步骤⑶、编码区间宽度Range重归一化。
随着编码区间宽度Range的更新,编码区间宽度Range的值可能会小于256(区间宽度初始化为510),这时需要进行重归一化。重归一化的方式为:同时对区间下限Low和编码区间宽度Range进行左移操作,直到编码区间宽度Range的值大于或等于256,区间下限Low左移出的比特即编码输出比特。
步骤⑷、使用已编码的二进制符号更新上下文模型。
至此,完成了CABAC的编码实现方式的说明。
解码器中的熵解码模块执行的操作为编码器中的熵编码模块执行的操作的逆操作,当熵编码模块通过CABAC进行编码时,熵解码模块执行的操作同样与CABAC编码部分对应,可包括:通过读取写入码流的区间下限解码出对应的二进制符号。
CABAC解码器的输入为区间下限和语法元素对应的索引,CABAC解码器的输出为对应的语法元素的二进制化值。CABAC解码器的编码过程可包括如下步骤a和步骤b:
步骤a、获取上下文。
上述步骤a(获取上下文)可包括如下步骤a1和a2:
步骤a1、根据标准的语法元素结构确定当前解码的是哪一个语法元素,根据已编码的二进制符号在比特流中的位置、对应的二进制符号的索引(binIdx)以及上下文参考信息获取当前符号的旁路标志(bypassFlag)和可能存在的上下文模型索引(ctxIdx)。
步骤a2、根据步骤a1中获取的信息,获取更新速率shiftIdx、概率状态pStateIdx0、概率状态pStateIdx1、pStateIdx0的权重以及pStateIdx1的权重,并以此获取对当前符号的估计概率。
更新速率shiftIdx、概率状态pStateIdx0、概率状态pStateIdx1以及pStateIdx0和pStateIdx1的权重获取当前符号的估计概率的实现方式可以参照上述步骤2,为避免赘述,此处不再详细说明。
步骤b、二进制算术解码。
对应于上述步骤3,上述步骤b(二进制算术解码)包括两种解码方式:常规解码和旁路解码。常规解码利用自适应的概率模型进行解码;旁路解码以等概率的方式进行解码,其概率状态无须更新,通过步骤a中获取的旁路标志(bypassFlag)区分这两种解码方式。
常规解码的编码器的输入是上下文模型(更新速率shiftIdx、概率状态pStateIdx0、概率状态pStateIdx1)和当前编码区间宽度Range,编码区间宽度Range的初始值为510,区间下限m_value从码流中读取字节获得。常规解码的基本原理为:在获取了当前符号解码使用的对应上下文后,根据估计的概率和当前编码区间宽度m_Range计算LPS符号对应区间的大小,对比区间下限和当前编码区间的大小,判决symbol符号是1还是0。具体解码流程可包括如下步骤①至步骤④:
步骤①、计算LPS对应的区间宽度RLPS和RMPS。
计算LPS对应的区间宽度RLPS和RMPS的算式如下所示:qRangeIdx=Range>>5 (18)pState=pStateldx1+16·pStateldx0 (19)RLPS=(qRangeIdx·pState5)>>1+4 (21)RMPS=Range-RLPS (22)
其中,pState5表示5bit精度的预测概率,当pState>>14=1时,预测概率大于0.5,pState为MPS的预测概率,异或操作使其保持为LPS的预测概率。
步骤②、比较区间下限m_Value和MPS与LPS对应子区间位置
若m_Value<RMPS,则m_Value保持不变,Range=RMPS,对应的语法元素的二进制值Bin=MPS;
若m_Value>=RMPS,则m_Value=m_Value-RLPS,Range=RLPS,对应的语法元素的二进制值Bin=LPS。
步骤③、编码区间宽度Range重归一化。
同样,随着编码区间宽度Range的更新,编码区间宽度Range的值可能会小于256(区间宽度初始化为510),这时需要进行重归一化。重归一化的方式为:同时对区间下限Low和编码区间宽度Range进行左移操作,直到编码区间宽度Range的值大于或等于256,区间下限Low左移出的比特即编码输出比特。
步骤④、重建语法元素的值
在一些实施例中,重建语法元素的值可包括:对语法元素的值解码得到的二进制字符串进行二值化的逆过程,从而重加语法元素的值。
至此,完成了CABAC的解码实现方式的说明。
本申请一些实施例涉及对语法元素tu_cb_coded_flag的值进行熵编码。以下首先对语法元素tu_cb_coded_flag进行说明。语法元素tu_cb_coded_flag是H.266/VVC标准中标识当前编码块的Cb色度分量的变换块(Transform Unit,TU)是否为非零块的语法元素,即,标识当前Cb色度分量的变换块是否包含不为零的变换系数语法元素,当语法元素tu_cb_coded_flag的值为1,则表示当前Cb色度分量的变换块包含不为零的变换系数,当语法元素tu_cb_coded_flag的值为0,则表示当前Cb色度分量的变换块不包含不为零的变换系数,当tu_cb_coded_flag不存在时,默认当前Cb色度分量的变换块不包含不为零的变换系数。
相关技术中对tu_cb_coded_flag的值进行熵编码的过程可包括如下步骤Ⅰ至步骤Ⅳ:
步骤Ⅰ、对tu_cb_coded_flag的值进行二进制化。
对tu_cb_coded_flag的值进行二进制化的实现方式可以参照下表1所示:
表1基于基本二进制化方案的语法元素及其输入参数
上表1显示了对tu_cb_coded_flag的值进行二进制化的方式和输入参数,由上表1可知,对tu_cb_coded_flag的值进行二进制化选用的是定长编码的二进制化方案,且tu_cb_coded_flag的最大值为1。
步骤Ⅱ、tu_cb_coded_flag的上下文模型选择。
对tu_cb_coded_flag进行上下文模型选择的目的是确定上下文索引值ctxIdx,实现方式可包括:
步骤⒈根据tu_cb_coded_flag的初始化类型确定tu_cb_coded_flag的上下文模型索引值集合。
H.266/VVC标准中为每个语法元素给出一个上下文索引值ctxIdx的表格,由索引值决定每个语法元素的初始值initValue。H.266标准列出一个表格,注明了每个语法元素表格编号和不同的初始化类型对应的上下文索引值ctxIdx集合。以H.266/VVC标准为例,表2给出了tu_cb_coded_flag在不同初始化类型的上下文索引值ctxIdx集合。初始化类型initType由帧类型以及sh_cabac_init_flag确定,sh_cabac_init_flag由pps_cabac_init_present_flag确定,pps_cabac_init_present_flag等于1表示在引用图像参数集(Picture Parameter Set,PPS)的片段标头中存在sh_cabac_init_flag,sh_cabac_init_flag的值被确定为1。pps_cabac_init_present_flag等于0表示在引用PPS的片段标头中不存在sh_cabac_init_flag,sh_cabac_init_flag的值被确定为0。
由帧类型以及sh_cabac_init_flag确定具体确定初始化类型initType的方式如下:
即,I类型的帧的初始化类型initType为0,P类型的帧的初始化类型initType根据sh_cabac_init_flag为2或1,B类型的帧的初始化类型initType根据sh_cabac_init_flag为1或2。
对于每种类型的帧,如表2所示,tu_cb_coded_flag使用2种上下文模型,initType=0时,上下文索引ctxIdx候选集为0和1;initType=1时,上下文索引ctxIdx候选集为2和3;initType=2时,上下文索引ctxIdx候选集为4和5。
表2语法元素的初始化类型表格号及其不同初始化类型的上下文索引集合
表3示出VVC的参考软件测试平台(VVC TEST MODEL,VTM)中tu_cb_coded_flag的六种上下文模型的初始值initValue和更新速率shiftIdx。初始值initValue用于计算上下文模型的概率状态,更新速率shiftIdx用于控制模型概率的更新速率。根据初始值initValue和更新速率shiftIdx这两个参数即可获取常规编码器的输入(shiftIdx、pStateIdx0、pStateIdx1),根据初始值initValue和更新速率shiftIdx获取shiftIdx、pStateIdx0、pStateIdx1的实现方式参照上述步骤2,为避免赘述此处不再详细说明。
表3 tu_cb_coded_flag上下文模型参数初始化值
步骤⒉根据intra_bdpcm_chroma_flag确定索引偏移binIdx。
对于tu_cb_coded_flag,根据intra_bdpcm_chroma_flag确定索引偏移binIdx,可包括:当intra_bdpcm_chroma_flag为true时,tu_cb_coded_flag的索引偏移binIdx解析为1,否则解析为0。具体见如下表4所示:
表4将上下文模型偏移量分配给语法元素
步骤⒊根据tu_cb_coded_flag的索引偏移binIdx确定tu_cb_coded_flag的上下文模型索引值。
即,先根据tu_cb_coded_flag的初始化类型选取tu_cb_coded_flag的上下文模型索引值集合,然后再根据tu_cb_coded_flag的索引偏移binIdx从tu_cb_coded_flag的上下文模型索引值集合选取tu_cb_coded_flag的上下文模型索引值。
综上,tu_cb_coded_flag的初始化类型initType、intra_bdpcm_chroma_flag以及上下文模型索引值ctxIdx的映射关系可以如下表5所示:
表5 tu_cb_coded_flag的上下文模型选择
在一些实施例中,表5所示映射关系也可以通过表达式计算得出。
步骤Ⅲ、对tu_cb_coded_flag的值进行算术编码。
得到tu_cb_coded_flag的上下文模型索引ctxIdx的值后,根据tu_cb_coded_flag和由初始值initValue和更新速率shiftIdx确定的上下文模型(shiftIdx、pStateIdx0、pStateIdx1)启动算术编码,进行区间划分,并获取tu_cb_coded_flag的对应的编码数据。
步骤Ⅳ、更新tu_cb_coded_flag所使用的上下文模型对应的概率模型。
即,根据tu_cb_coded_flag的值对使用的上下文模型的概率状态pStateIdx0和pStateIdx1进行更新。
相关技术中对tu_cb_coded_flag的值进行熵解码的过程可包括如下步骤㈠至步骤㈣:
步骤㈠、根据tu_cb_coded_flag的初始化类型确定tu_cb_coded_flag的上下文模型索引值集合。
根据tu_cb_coded_flag的初始化类型确定tu_cb_coded_flag的上下文模型索引值集合的实现方式可以参照上述步骤⒈,为避免赘述,此处不再详细说明。
步骤㈡、根据intra_bdpcm_chroma_flag确定索引偏移binIdx。
根据intra_bdpcm_chroma_flag确定索引偏移binIdx的实现方式可以参照上述步骤⒉,为避免赘述,此处不再详细说明。
步骤㈢、根据tu_cb_coded_flag的索引偏移binIdx确定tu_cb_coded_flag的上下文模型索引值,并进行算术解码。
进行算术解码可包括:根据区间划分情况和码流中区间下限对应的值,确定tu_cb_coded_flag对应的符号(symbol)。该符号的值可能为0或1。
步骤㈣、根据算术解码得到的符号的值为tu_cb_coded_flag赋值。
即,将tu_cb_coded_flag对应的符号的值(0或1)赋值为tu_cb_coded_flag的值,tu_cb_coded_flag的值可能为0或1。
步骤㈤、更新tu_cb_coded_flag所使用的上下文模型对应的概率模型。
更新tu_cb_coded_flag所使用的上下文模型对应的概率模型的实现方式可以参照上述步骤Ⅳ,为避免赘述,此处不再详细说明。
如上所述,相关技术中tu_cb_coded_flag会根据初始化类型对应两种上下文模型,然后再根据intra_bdpcm_chroma_flag(标识当前CU是否使用chroma bdpcm的标识)确定使用初始化类型对应的两种上下文模型中的哪一个。tu_cb_coded_flag的上下文索引可以由下式计算得到:ctxIdx=intra_bdpcm_chroma_flag?1:0
然而,相关技术中仅仅是依据tu_cb_coded_flag的初始化类型initType和对应的intra_bdpcm_chroma_flag的值来确定tu_cb_coded_flag的上下文模型,并没有考虑tu_cb_coded_flag的空间特性,进而可能导致tu_cb_coded_flag的上下文模型不准确。
针对tu_cb_coded_flag进行上下文模型选取时并没有考虑tu_cb_coded_flag的空间特性,进而可能导致tu_cb_coded_flag的上下文模型不准确的问题,本申请一些实施例进一步提供了如下技术方案:
本申请一些实施例更改色度块的语法元素tu_cb_coded_flag熵编码时使用的上下文模型。具体为:根据当前色度块的上侧和左侧相邻色度块中tu_cb_coded_flag的值决定编码当前tu_cb_coded_flag时使用的上下文模型索引,并更改相应的上下文模型数量及初始值。
在一些实施例中,更改色度块的语法元素tu_cb_coded_flag熵编码时使用的上下文模型后,对tu_cb_coded_flag进行熵编码的实现方式可包括如下步骤A至步骤D:
步骤A、对tu_cb_coded_flag的值进行二进制化。
对tu_cb_coded_flag的值进行二进制化的实现方式可以参照上述步骤Ⅰ,为避免赘述,此处不再详细说明。
步骤B、选取对tu_cb_coded_flag值进行熵编码的上下文模型。
上述步骤B(选取对tu_cb_coded_flag值进行熵编码的上下文模型)的目的是确定用于对tu_cb_coded_flag值进行熵编码的上下文模型ctxIdx,实现方式可包括如下步骤B1至步骤B4:
步骤B1、根据tu_cb_coded_flag的初始化类型确定tu_cb_coded_flag的上下文模型索引值集合。
本申请一些实施例中,tu_cb_coded_flag的初始化类型initType可包括三种值,分别为:0、1、2,每一种tu_cb_coded_flag的初始化类型initType由2个变为6个。tu_cb_coded_flag的初始化类型initType与上下文模型索引值集合的映射关系可包括:当初始化类型initType=0时,tu_cb_coded_flag的上下文模型索引值集合为:{0,1,2,3,4,5};当初始化类型initType=1时,tu_cb_coded_flag的上下文模型索引值集合为:{6,7,8,9,10,11};当初始化类型initType=2时,tu_cb_coded_flag的上下文模型索引值集合为:{12,13,14,15,16,17}。即,本申请一些实施例中,tu_cb_coded_flag的初始化类型initType与上下文模型索引值集合的映射关系如下表6所示:
表6 tu_cb_coded_flag的不同初始化类型的上下文索引集合
在一些实施例中,索引号1至17对应的上下文模型参数的初始化值为预设经验值。
在一些实施例中,索引号1至17对应的上下文模型参数的初始化值可包括:初始值initValue设置为CNU=35,更新速率shiftIdx设置为DWS=8,区间宽度的初始值设置为DWE=18,长短窗口偏移量的初始值设为DWO=119。索引号1至17对应的上下文模型参数初始化值如下表7所示:
表7 tu_cb_coded_flag的上下文模型参数的初始化值
在另一些实施例中,可以记录每个序列每帧编码结束后的上下文模型概率情况,并根据历史数据设置索引号1至17对应的上下文模型参数的初始化值。
步骤B2、确定与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值。
在一些实施例中,与当前色度块相邻的已编码色度块可包括:位于当前色度块左侧的色度块和位于当前色度块上方的色度块,因此可以根据传入的编码结构和色度位置、色度分量,确定位于当前色度块左侧的色度块和位于当前色度块上方的色度块。进而获取位于当前色度块左侧的色度块的tu_cb_coded_flag的值以及位于当前色度块上方的色度块的tu_cb_coded_flag的值。
步骤B3、根据与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值以及当前色度块的intra_bdpcm_chroma_flag的值,确定当前色度块的上下文模型的索引值。
在一些实施例中,根据与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值以及当前色度块的intra_bdpcm_chroma_flag的值,确定当前色度块的上下文模型的索引值,可包括如下步骤B31和B32:
步骤B31、计算与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和。
即,对位于当前色度块左侧的色度块的tu_cb_coded_flag的值以及位于当前色度块上方的色度块的tu_cb_coded_flag的值进行累加,获取计算与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和Neibor_CbfcbFlag_Sum。由于位于当前色度块左侧的色度块和/或位于当前色度块左侧的色度块可能会不存在,且tu_cb_coded_flag的值为0或1,因此Neibor_CbfcbFlag_Sum的值为0或1或2。
步骤B32、根据与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和以及当前色度块的intra_bdpcm_chroma_flag的值,确定当前色度块的上下文模型的索引值。
在一些实施例中,根据与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和以及当前色度块的intra_bdpcm_chroma_flag的值,确定当前色度块的上下文模型的索引值,可包括:当intra_bdpcm_chroma_flag的值为0(未使用bdpcm Chroma模式)时,当前色度块的上下文模型的索引值从初始化类型对应的上下文索引集合的前三个上下文模型索引中选取,即,若初始化类型initType=0,则上下文模型的索引值ctxIdx的选取范围为{0,1,2},若初始化类型initType=1,则上下文模型的索引值ctxIdx的选取范围为{6,7,8},若初始化类型initType=2,则上下文模型的索引值ctxIdx的选取范围为{12,13,14};当intra_bdpcm_chroma_flag的值为1(使用了bdpcm Chroma模式)时,当前色度块的上下文模型的索引值从初始化类型对应的上下文索引集合的后三个上下文模型索引中选取,即,若初始化类型initType=0,则上下文模型的索引值ctxIdx的选取范围为{3,4,5},若初始化类型initType=1,则上下文模型的索引值ctxIdx的选取范围为{9,10,11},若初始化类型initType=2,则上下文模型的索引值ctxIdx的选取范围为{15,16,17}。
在一些实施例中,从初始化类型对应的上下文索引集合的前三个上下文模型索引中选取当前色度块的上下文模型的索引值,可包括:若初始化类型initType=0,则上下文模型的索引值ctxIdx等于与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和Neibor_CbfcbFlag_Sum,即,ctxIdx=Neibor_CbfcbFlag_Sum;若初始化类型initType=1,则上下文模型的索引值ctxIdx等于与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值与6的累加和,即,ctxIdx=Neibor_CbfcbFlag_Sum+6;若初始化类型initType=2,则上下文模型的索引值ctxIdx等于与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值与Neibor_CbfcbFlag_Sum与12的累加和,即,ctxIdx=Neibor_CbfcbFlag_Sum+12;从初始化类型对应的上下文索引集合的后三个上下文模型索引中选取当前色度块的上下文模型的索引值,可包括:若初始化类型initType=0,则上下文模型的索引值ctxIdx等于与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值与3的累加和,即,ctxIdx=Neibor_CbfcbFlag_Sum+3;若初始化类型initType=1,则上下文模型的索引值ctxIdx等于与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和与9的累加和,即,ctxIdx=Neibor_CbfcbFlag_Sum+9;若初始化类型initType=2,则上下文模型的索引值ctxIdx等于与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值与15的累加和,即,ctxIdx=Neibor_CbfcbFlag_Sum+15。
参照图3所示,当前色度块31的intra_bdpcm_chroma_flag的值、位于当前色度块31上方的色度块32的tu_cb_coded_flag的值、位于当前色度块31左侧的色度块33的tu_cb_coded_flag的值,以及当前色度块31的上下文模型的索引值ctxIdx的映射关系为:
ctxIdx=!intra_bdpcm_chroma_flag?Neibor_CbfcbFlag_Sum:Neibor_CbfcbFlag_Sum+3
即,本申请一些实施例中,当前色度块的intra_bdpcm_chroma_flag的值、与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和Neibor_CbfcbFlag_Sum以及当前色度块31的上下文模型的索引值ctxIdx的对应关系可以如下表8所示:
表8上下文模型索引值取值标准
综上,本申请一些实施例中上下文模型索引ctxId、初始化类型initType、intra_bdpcm_chroma_flag的值、与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和Neibor_CbfcbFlag_Sum之间的映射关系可以如下表9所示:
表9 tu_cb_coded_flag的上下文模型选择
步骤C、根据当前色度块的上下文模型的索引值对当前色度块的tu_cb_coded_flag的值进行算术编码,并更新上下文模型的模型概率。
根据当前色度块的上下文模型的索引值对当前色度块的tu_cb_coded_flag的值进行算术编码,并更新上下文模型的模型概率的实现方式可以参照上述步骤⑶和步骤⑷,为避免赘述,此处不再详细说明。
在一些实施例中,所述方法还可包括:获取一帧图像开始编码时tu_cb_coded_flag的上下文模型的初始值和结束编码时tu_cb_coded_flag的上下文模型的概率,若开始编码时上下文模型的概率和结束编码时上下文模型的概率的差值小于第一阈值,则在开始编码下帧图像时,不对tu_cb_coded_flag的上下文模型进行初始化,而是沿用该帧图像结束编码时tu_cb_coded_flag的上下文模型。
上述实施例对使用chroma bdpcm模式的色度块和不使用chroma bdpcm模式的色度块都依据上方的色度块和左侧的色度块的tu_cb_coded_flag进行了上下文分类,同一种初始类型下可包括6种上下文模型。而在另一些实施例中,也可以仅对不使用chroma bdpcm模式的色度块采取这种上下文分类方式,而对使用chroma bdpcm模式的色度块,只设置一列固定的上下文。即,同一种初始类型下,对tu_cb_coded_flag设置4列上下文模型,前3列为不使用chroma bdpcm的情况,后1列为使用chroma bdpcm的情况。具体可以如下表10所示:
表10 tu_cb_coded_flag的上下文模型选择
上述实施例在计算与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和Neibor_CbfcbFlag_Sum时,没有与当前色度块相邻的已编码色度块的intra_bdpcm_chroma_flag的值,在另一些实施例中,可以先获取与当前色度块相邻的已编码色度块的intra_bdpcm_chroma_flag的值,且在计算与当前色度块相邻的已编码色度块的tu_cb_coded_flag的值的和Neibor_CbfcbFlag_Sum时,只对与当前色度块的intra_bdpcm_chroma_flag的值相同的相邻已编码色度块的tu_cb_coded_flag的值求和。
本申请一些实施例提供了一种视频解码方法,参照图4所示,该视频解码方法可包括如下步骤:
[根据细则91更正 06.11.2025]
S41、获取第一色度块的第一语法元素的熵编码数据。
S41、获取第一色度块的第一语法元素的熵编码数据。
其中,所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。即,所述第一语法元素为tu_cb_coded_flag。
在一些实施例中,可以根据标准的语法元素结构从视频码流中获取第一色度块的第一语法元素的熵编码数据。在一些实施例中,所述第一色度块为当前编码块的蓝色分量Cb对应的色度块。在另一些实施例中,所述第一色度块为当前编码块的红色分量Cr对应的色度块。
S42、根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息。
本申请一些实施例中所述第一色度块的相邻已解码色度块是指:与所述第一色度块相邻且在解码第一色度块之前已完成了解码的色度块。
在一些实施例中,所述第一色度块的相邻已解码色度块,可包括:位于所述第一色度块上方的色度块和/或位于所述第一色度块左侧的色度块;上述步骤S42(根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息),可包括:根据位于所述第一色度块上方的色度块和/或位于所述第一色度块左侧的色度块获取所述第一色度块的空域信息。
在一些实施例中,上述步骤S42(根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息),可包括:根据所述相邻已解码色度块的变换块是否为非零块获取所述第一色度块的空域信息。
S43、根据所述第一色度块的空域信息确定目标上下文模型。
在一些实施例中,根据所述第一色度块的空域信息确定目标上下文模型,可包括:根据所述第一色度块的空域信息从预设上下文模型集合中选取所述目标上下文模型。
在一些实施例中,根据所述第一色度块的空域信息确定目标上下文模型,可包括:根据所述第一色度块的空域信息、所述第一色度块的所述第一语法元素的初始化类型(initType)以及所述第一色度块的所述第二语法元素的值,确定所述目标上下文模型;
其中,所述第二语法元素为标识色度块是否使用了基于块的差分脉冲编码调制模式的语法元素。即,所述第二语法元素为intra_bdpcm_chroma_flag。
S44、基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。
上述实施例提供的视频解码方法在获取第一色度块的标识色度块的变换块是否为非零块的第一语法元素的熵编码数据,并根据第一色度块的第一语法元素的熵编码数据获取第一色度块的第一语法元素的值时,首先根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息,然后根据所述第一色度块的空域信息确定目标上下文模型,再基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。由于上述实施例提供的视频解码方法会根据相邻已解码色度块获取所述第一色度块的空域信息,并根据第一色度块的空域信息确定用于对第一色度块的第一语法元素的值进行熵解码的上下文模型,因此上述实施例可以结合第一语法元素的空间特性确定用于对第一语法元素进行熵解码的上下文模型,进而提升用于对第一语法元素进行熵解码的上下文模型的准确率。
作为对上述实施例的扩展和细化,本申请一些实施例提供了另一种视频解码方法,参照图5所示,该视频解码方法可包括如下步骤:
S501、获取第一色度块的第一语法元素的熵编码数据。其中,所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。
S502、获取所述相邻已解码色度块的所述第一语法元素的值。即,获取所述相邻已解码色度块的tu_cb_coded_flag的值。
S503、根据所述相邻已解码色度块的所述第一语法元素的值,获取所述第一色度块的空域信息。
在一些实施例中,上述步骤S503(根据所述相邻已解码色度块的所述第一语法元素的值,获取所述第一色度块的空域信息),可包括:对所述相邻已解码色度块的所述第一语法元素的值求和,以获取所述第一色度块的空域信息。
示例性的,所述相邻已解码色度块可包括:位于第一色度块上方的色度块和位于第一色度块左侧的色度块,且位于第一色度块上方的色度块的第一语法元素的值为1,位于第一色度块左侧的色度块的第一语法元素的值为1,则可以确定所述第一色度块的空域信息为2。
示例性的,所述相邻已解码色度块仅包括位于第一色度块左侧的色度块,且位于第一色度块上方的色度块的第一语法元素的值为1,则可以确定所述第一色度块的空域信息为1。
示例性的,所述相邻已解码色度块可包括:位于第一色度块上方的色度块和位于第一色度块左侧的色度块,且位于第一色度块上方的色度块的第一语法元素的值为0,位于第一色度块左侧的色度块的第一语法元素的值为0,则可以确定所述第一色度块的空域信息为0。
在一些实施例中,上述步骤S503(根据所述相邻已解码色度块的所述第一语法元素的值,获取所述第一色度块的空域信息),可包括如下步骤5031和5032:
步骤5031、获取所述相邻已解码色度块的所述第二语法元素的值。
其中,所述第二语法元素为标识色度块是否使用了基于块的差分脉冲编码调制模式的语法元素。即,确定所述相邻已解码色度块是否使用了基于块的差分脉冲编码调制模式。
步骤5032、对所述相邻已解码色度块中所述第二语法元素的值与所述第一色度块相同的色度块的所述第一语法元素的值求和,以获取所述第一色度块的空域信息。
示例性的,所述第一色度块的第二语法元素的值0,所述相邻已解码色度块可包括:位于第一色度块上方的色度块和位于第一色度块左侧的色度块,且位于第一色度块上方的色度块的第一语法元素的值为1,位于第一色度块上方的色度块的第二语法元素的值0,位于第一色度块左侧的色度块的第一语法元素的值为1,位于第一色度块上方的色度块的第二语法元素的值1,由于只有位于第一色度块上方的色度与所述第一色度块的第二语法元素的值相同,因此所述第一色度块的空域信息为1。
S504、获取所述第一色度块的所述第一语法元素的初始化类型(initType)。
在一些实施例中,获取所述第一色度块的第一语法元素的初始化类型的实现方式可包括:根据语法元素pps_cabac_init_present_flag的值确定语法元素sh_cabac_init_flag的值,根据语法元素sh_cabac_init_flag的值以及所述第一色度块对应的帧类型(I类型或B类型或P类型)确定所述第一色度块的第一语法元素的初始化类型。
在一些实施例中,获取所述第一色度块的第一语法元素的初始化类型的实现方式可包括:
若所述第一色度块对应帧类型为I类型,则确定所述第一色度块的第一语法元素的初始化类型为第一初始化类型;若所述第一色度块对应帧类型为P类型,且sh_cabac_init_flag的值为1,则确定所述第一色度块的第一语法元素的初始化类型为第二初始化类型;若所述第一色度块对应帧类型为P类型,且sh_cabac_init_flag的值为0,则确定所述第一色度块的第一语法元素的初始化类型为第三初始化类型,若所述第一色度块对应帧类型为B类型,且sh_cabac_init_flag的值为1,则确定所述第一色度块的第一语法元素的初始化类型为第三初始化类型;若所述第一色度块对应帧类型为B类型,且sh_cabac_init_flag的值为0,则确定所述第一色度块的第一语法元素的初始化类型为第二初始化类型。
S505、获取所述第一色度块的第二语法元素的值。
其中,所述第二语法元素为标识色度块是否使用了基于块的差分脉冲编码调制模式的语法元素。
在一些实施例中,当所述第一色度块未使用基于块的差分脉冲编码调制模式,则所述第一色度块的第二语法元素的值为0,当所述第一色度块未使用基于块的差分脉冲编码调制模式,则所述第一色度块的第二语法元素的值为1。
S506、根据所述初始化类型确定目标索引集合。
在一些实施例中,所述初始化类型为第一初始化类型或者第二初始化类型或者第三初始化类型,所述第一初始化类型、所述第二初始化类型以及所述第三初始化类型分别对应一个上下文模型索引集合;所述根据所述初始化类型确定目标索引集合,可包括:当所述初始化类型为第一初始化类型时,将所述第一初始化类型对应的上下文模型索引集合确定为所述目标索引集合;当所述初始化类型为第二初始化类型时,将所述第二初始化类型对应的上下文模型索引集合确定为所述目标索引集合;当所述初始化类型为第一初始化类型时,将所述第三初始化类型对应的上下文模型索引集合确定为所述目标索引集合。
如上表6所示,第一初始化类型(initType=0)对应的上下文模型索引集合为{0,1,2,3,4,5},第二初始化类型(initType=1)对应的上下文模型索引集合为{6,7,8,9,10,11},第三初始化类型(initType=3)对应的上下文模型索引集合为{12,13,14,15,16,17},因此当所述初始化类型为第一初始化类型,则所述目标索引集合为{0,1,2,3,4,5},当所述初始化类型为第一初始化类型,则所述目标索引集合为{6,7,8,9,10,11},当所述初始化类型为第一初始化类型,则所述目标索引集合为{12,13,14,15,16,17}。
如上表10所示,第一初始化类型(initType=0)对应的上下文模型索引集合为{0,1,2,3},第二初始化类型(initType=1)对应的上下文模型索引集合为{4,5,6,7},第三初始化类型(initType=2)对应的上下文模型索引集合为{8,9,10,11},因此当所述初始化类型为第一初始化类型,则所述目标索引集合为{0,1,2,3},当所述初始化类型为第二初始化类型,则所述目标索引集合为{4,5,6,7},当所述初始化类型为第三初始化类型,则所述目标索引集合为{8,9,10,11}。
S507、根据所述第一色度块的空域信息和所述第一色度块的所述第二语法元素的值,从所述目标索引集合中选取目标模型索引。
在一些实施例中,所述目标索引集合,可包括:六个上下文模型索引(上表6所示)。上述步骤S507(根据所述第一色度块的空域信息和所述第一色度块的所述第二语法元素的值,从所述目标索引集合中选取目标模型索引),可包括:
当所述第一色度块的所述第二语法元素的值为标识所述第一色度块未使用基于块的差分脉冲编码调制模式的值时,根据所述第一色度块的空域信息从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引;
当所述第一色度块的所述第二语法元素的值为标识所述第一色度块使用了基于块的差分脉冲编码调制模式的值时,根据所述第一色度块的空域信息从所述目标索引集合的后三个上下文模型索引中选取所述目标模型索引。
即,若intra_bdpcm_chroma_flag=0,则根据所述第一色度块的空域信息从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引,若intra_bdpcm_chroma_flag=1,则根据所述第一色度块的空域信息从所述目标索引集合的后三个上下文模型索引中选取所述目标模型索引。
示例性的,当initType=2时,若intra_bdpcm_chroma_flag=0,则根据所述第一色度块的空域信息从{12,13,14}中选取所述目标模型索引,若intra_bdpcm_chroma_flag=1,则根据所述第一色度块的空域信息从{15,16,17}中选取所述目标模型索引。
在一些实施例中,根据所述第一色度块的空域信息,从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引,可包括:当所述第一色度块的空域信息为m,则选取所述目标索引集合中的第m+1个上下文模型索引作为所述目标模型索引;0≤m≤2。
其中,0≤m≤2是指:m≥0,且m≤2。因为第一色度块的空域信息为至少一个第一语法元素的值的和,因此m为整数,因此m为0或1或2。
即,当第一色度块的空域信息为0,则选取所述目标索引集合中的第1个上下文模型索引作为所述目标模型索引,当第一色度块的空域信息为1,则选取所述目标索引集合中的第2个上下文模型索引作为所述目标模型索引,当第一色度块的空域信息为2,则选取所述目标索引集合中的第3个上下文模型索引作为所述目标模型索引。
在一些实施例中,根据所述第一色度块的空域信息,从所述目标索引集合的后三个上下文模型索引中选取所述第一色度块的所述第一语法元素对应的上下文模型索引,可包括:当所述第一色度块的空域信息为m,则选取所述目标索引集合中的第m+4个上下文模型索引作为所述第一色度块的所述第一语法元素对应的上下文模型索引,0≤m≤2。
即,当第一色度块的空域信息为0,则选取所述目标索引集合中的第4个上下文模型索引作为所述目标模型索引,当第一色度块的空域信息为1,则选取所述目标索引集合中的第5个上下文模型索引作为所述目标模型索引,当第一色度块的空域信息为2,则选取所述目标索引集合中的第6个上下文模型索引作为所述目标模型索引。
在一些实施例中,所述目标索引集合,可包括:四个上下文模型索引(上表10所示)。上述步骤S507(根据所述第一色度块的空域信息和所述第一色度块的所述第二语法元素的值,从所述目标索引集合中选取目标模型索引),可包括:
当所述第一色度块的所述第二语法元素的值为标识所述第一色度块未使用基于块的差分脉冲编码调制模式的值时,根据所述第一色度块的空域信息从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引;
当所述第一色度块的所述第二语法元素的值为标识所述第一色度块使用了基于块的差分脉冲编码调制模式的值时,选取所述目标索引集合中的第四个上下文模型索引作为所述目标模型索引。
同样,根据所述第一色度块的空域信息从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引,可包括:当所述第一色度块的空域信息为m,则选取所述目标索引集合中的第m+1个上下文模型索引作为所述目标模型索引;0≤m≤2。
即,在若intra_bdpcm_chroma_flag=0的情况下,当第一色度块的空域信息为0,则选取所述目标索引集合中的第1个上下文模型索引作为所述目标模型索引,当第一色度块的空域信息为1,则选取所述目标索引集合中的第2个上下文模型索引作为所述目标模型索引,当第一色度块的空域信息为2,则选取所述目标索引集合中的第3个上下文模型索引作为所述目标模型索引;在若intra_bdpcm_chroma_flag=1的情况下,当第一色度块的空域信息为0或1或2,则选取所述目标索引集合中的第4个上下文模型索引作为所述目标模型索引。
S508、根据所述目标模型索引确定所述目标上下文模型。
在一些实施例中,上述S508(根据所述目标模型索引确定所述目标上下文模型),可包括如下步骤5801和步骤5082:
步骤5081、根据所述目标模型索引获取所述目标上下文模型的初始值和更新速率。
在一些实施例中,所述目标上下文模型的初始值(initValue)为35;所述目标上下文模型的更新速率(shiftIdx)为8。
步骤5082、根据所述目标上下文模型的初始值和更新速率构建所述目标上下文模型。
在一些实施例中,根据所述目标上下文模型的初始值和更新速率构建所述目标上下文模型,可包括根据所述目标上下文模型的初始值(initValue)计算所述目标上下文模型的模型概率pStateIdx0和pStateIdx1:slopeIdx=initValue>>3 (11)offsetIdx=initValue (12)m=slopeIdx-4 (13)n=(offsetIdx·18)+1 (14)preCtxState=Clip3(1,127,((m·(Clip3(0,63,SliceQPy)-16))>>1)+n) (14)pStateIdx0=preCtxState3 (15)pStateIdx1=preCtxState7 (16)
其中,SliceQPy为亮度信号的量化参数。
在一些实施例中,根据所述目标模型索引确定所述目标上下文模型,可包括:根据所述目标模型索引读取所述目标上下文模型。
S509、将所述目标上下文模型作为CABAC算法的上下文模型,通过所述CABAC算法对所述第一色度块的第一语法元素的熵编码数据进行算术解码,以获取所述第一色度块的所述第一语法元素的值。
通过所述CABAC算法对所述第一色度块的第一语法元素的熵编码数据进行算术解码的实现方式可以参照上述步骤b,为避免赘述,此处不再详细说明。
S510、根据所述第一色度块的所述第一语法元素的值对所述目标上下文模型进行更新。
即,根据第一色度块的tu_cb_coded_flag的值对使用的目标上下文模型的概率状态pStateIdx0和pStateIdx1进行更新。
在一些实施例中,本申请一些实施例提供的视频解码方法还可包括:根据所述目标上下文模型确定第二色度块对应的上下文模型;其中,所述第一色度块和所述第二色度块分别为同一编码块的两个色度分量对应的色度块。
在一些实施例中,所述第一色度块为Cb分量对应的色度块,所述第二色度块为Cr色度分量对应的色度块。
即,根据本申请一些实施例提供的视频解码方法获取编码块的蓝色色度分量对应的色度块对应的目标上下文模型,基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值,并根据所述确定目标上下文模型确定该编码块红色色度分量对应的色度块对应的上下文模型。
在一些实施例中,所述根据所述目标上下文模型确定第二色度块对应的上下文模型,可包括:将所述目标上下文模型确定为所述第二色度块对应的上下文模型。
在根据所述目标上下文模型确定第二色度块对应的上下文模型之后,本申请一些实施例提供的视频编码方法还可包括:
基于第二色度块对应的上下文模型和所述第二色度块的第一语法元素的熵编码数据,获取所述第二色度块的所述第一语法元素的值。
本申请一些实施例提供了一种视频编码方法,参照图6所示,该视频解码方法可包括如下步骤:
S61、根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息。
上述步骤S61的实现方式可以参照上述步骤S42的实现方式,不同之处在于,步骤S42中是根据第一色度块的相邻已解码色度块获取所述第一色度块的空域信息,而步骤S61中是根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息,此外,基于编解码的顺序,解码过程中第一色度块的相邻已解码色度块和编码过程中第一色度块的相邻已编码色度块是相同的色度块。
S62、根据所述第一色度块的空域信息确定目标上下文模型。
根据所述第一色度块的空域信息确定目标上下文模型的实现方式可以参照上述视频解码方法中根据所述第一色度块的空域信息确定目标上下文模型的实现方式,为避免赘述,此处不再详细说明。
S63、基于所述目标上下文模型对所述第一色度块的第一语法元素的值进行熵编码,以获取所述第一色度块的所述第一语法元素的熵编码数据。
其中,所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。
上述实施例提供的视频编码方法在对第一色度块的第一语法元素的值进行熵编码时,首先根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息,然后根据所述第一色度块的空域信息确定目标上下文模型,再基于所述目标上下文模型对所述第一色度块的第一语法元素的值进行熵编码,以获取所述第一色度块的所述第一语法元素的熵编码数据。由于上述实施例提供的视频编码方法会根据相邻已编码色度块获取所述第一色度块的空域信息,并根据第一色度块的空域信息确定用于对第一色度块的第一语法元素的值进行熵编码的上下文模型,因此上述实施例可以结合第一语法元素的空间特性确定用于对第一语法元素进行熵编码的上下文模型,进而提升用于对第一语法元素进行熵编码的上下文模型的准确率。
本申请一些实施例还提供了一种视频解码装置,参照图7所示,该视频解码装置700可包括:
第一获取模块71,用于获取第一色度块的第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素;
第二获取模块72,用于根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息;
确定模块73,用于根据所述第一色度块的空域信息确定目标上下文模型;
熵解码模块74,用于基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。
上述实施例提供的视频解码装置可以执行上述任一实施例提供的视频解码方法,其实现原理与技术效果类似,此处不再赘述。
本申请一些实施例提供了一种视频编码装置,参照图8所示,该视频编码装置800可包括:
获取单元81,用于根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息;
确定单元82,用于根据所述第一色度块的空域信息确定目标上下文模型;
熵编码单元83,用于基于所述目标上下文模型对所述第一色度块的第一语法元素的值进行熵编码,以获取所述第一色度块的所述第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。
上述实施例提供的视频编码装置可以执行上述任一实施例提供的视频编码方法,其实现原理与技术效果类似,此处不再赘述。
本申请一些实施例提供了一种电子设备,该电子设备,可包括:
存储器,被配置为存储计算机程序;
处理器,被配置为用于在调用计算机程序时,使得所述视频解码装置实现上述任一实施例所述的视频解码方法或视频编码方法。
本申请一些实施例提供了一种计算机可读存储介质,所述计算机可读存储介质上存储有计算机程序,当所述计算机程序被计算设备执行时,使得所述计算设备实现上述任一实施例所述的视频解码方法或上述任一实施例所述的视频解码方法。
本申请一些实施例提供了一种计算机程序产品,当所述计算机程序产品在计算机上运行时,使得所述计算机实现上述任一实施例所述的视频解码方法或上述任一实施例所述的视频解码方法。
本申请一些实施例提供了一种芯片,该芯片包括存储器和处理器,该处理器可以是逻辑电路、集成电路或者通用处理器,该存储器上存储有计算机指令,该处理器可以通过读取存储器中存储的计算机指令来实现上述任一实施例所述的视频解码方法或上述任一实施例所述的视频解码方法。
最后应说明的是:以上各实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述各实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分或者全部技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的范围。
为了方便解释,已经结合具体的实施方式进行了上述说明。但是,上述示例性的讨论不是意图穷尽或者将实施方式限定到上述公开的具体形式。根据上述的教导,可以得到多种修改和变形。上述实施方式的选择和描述是为了更好的解释原理以及实际的应用,从而使得本领域技术人员更好的使用所述实施方式以及适于具体使用考虑的各种不同的变形的实施方式。
Claims (19)
- 一种视频解码方法,其特征在于,包括:获取第一色度块的第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素;根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息;根据所述第一色度块的空域信息确定目标上下文模型;基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。
- 根据权利要求1所述的方法,其特征在于,所述根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息,包括:根据位于所述第一色度块上方的色度块和/或位于所述第一色度块左侧的色度块获取所述第一色度块的空域信息。
- 根据权利要求1所述的方法,其特征在于,所述根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息,包括:获取所述相邻已解码色度块的所述第一语法元素的值;根据所述相邻已解码色度块的所述第一语法元素的值,获取所述第一色度块的空域信息。
- 根据权利要求3所述的方法,其特征在于,所述根据所述相邻已解码色度块的所述第一语法元素的值,获取所述第一色度块的空域信息,包括:对所述相邻已解码色度块的所述第一语法元素的值求和,以获取所述第一色度块的空域信息。
- 根据权利要求3所述的方法,其特征在于,所述根据所述相邻已解码色度块的所述第一语法元素的值,获取所述第一色度块的空域信息,包括:获取所述相邻已解码色度块的第二语法元素的值;所述第二语法元素为标识色度块是否使用了基于块的差分脉冲编码调制模式的语法元素;对所述相邻已解码色度块中所述第二语法元素的值与所述第一色度块相同的色度块的所述第一语法元素的值求和,以获取所述第一色度块的空域信息。
- 根据权利要求3所述的方法,其特征在于,所述根据所述第一色度块的空域信息确定目标上下文模型,包括:根据所述第一色度块的空域信息、所述第一色度块的所述第一语法元素的初始化类型以及所述第一色度块的第二语法元素的值,确定所述目标上下文模型;其中,所述第二语法元素为标识色度块是否使用了基于块的差分脉冲编码调制模式的语法元素。
- 根据权利要求6所述的方法,其特征在于,所述根据所述第一色度块的空域信息、所述第一色度块的所述第一语法元素的初始化类型以及所述第一色度块的所述第二语法元素的值,确定所述目标上下文模型,包括:根据所述初始化类型确定目标索引集合;根据所述第一色度块的空域信息和所述第一色度块的所述第二语法元素的值,从所述目标索引集合中选取目标模型索引;根据所述目标模型索引确定所述目标上下文模型。
- 根据权利要求7所述的方法,其特征在于,所述目标索引集合,包括:六个上下文模型索引;所述根据所述第一色度块的空域信息和所述第一色度块的所述第二语法元素的值,从所述目标索引集合中选取目标模型索引,包括:当所述第一色度块的所述第二语法元素的值为标识所述第一色度块未使用基于块的差分脉冲编码调制模式的值时,根据所述第一色度块的空域信息从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引;当所述第一色度块的所述第二语法元素的值为标识所述第一色度块使用了基于块的差分脉冲编码调制模式的值时,根据所述第一色度块的空域信息从所述目标索引集合的后三个上下文模型索引中选取所述目标模型索引。
- 根据权利要求7所述的方法,其特征在于,所述目标索引集合,包括:四个上下文模型索引;所述根据所述第一色度块的空域信息和所述第一色度块的所述第二语法元素的值,从所述目标索引集合选取所述第一色度块的所述第一语法元素对应的上下文模型索引,包括:当所述第一色度块的所述第二语法元素的值为标识所述第一色度块未使用基于块的差分脉冲编码调制模式的值时,根据所述第一色度块的空域信息从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引;当所述第一色度块的所述第二语法元素的值为标识所述第一色度块使用了基于块的差分脉冲编码调制模式的值时,选取所述目标索引集合中的第四个上下文模型索引作为所述目标模型索引。
- 根据权利要求8或9所述的方法,其特征在于,所述根据所述第一色度块的空域信息,从所述目标索引集合的前三个上下文模型索引中选取所述目标模型索引,包括:当所述第一色度块的空域信息为m,则选取所述目标索引集合中的第m+1个上下文模型索引作为所述目标模型索引;0≤m≤2。
- 根据权利要求9所述的方法,其特征在于,所述根据所述第一色度块的空域信息,从所述目标索引集合的后三个上下文模型索引中选取所述第一色度块的所述第一语法元素对应的上下文模型索引,包括:当所述第一色度块的空域信息为m,则选取所述目标索引集合中的第m+4个上下文模型索引作为所述第一色度块的所述第一语法元素对应的上下文模型索引,0≤m≤2。
- 根据权利要求7所述的方法,其特征在于,所述初始化类型为第一初始化类型或者第二初始化类型或者第三初始化类型,所述第一初始化类型、所述第二初始化类型以及所述第三初始化类型分别对应一个上下文模型索引集合;所述根据所述初始化类型确定目标索引集合,包括:当所述初始化类型为第一初始化类型时,将所述第一初始化类型对应的上下文模型索引集合确定为所述目标索引集合;当所述初始化类型为第二初始化类型时,将所述第二初始化类型对应的上下文模型索引集合确定为所述目标索引集合;当所述初始化类型为第一初始化类型时,将所述第三初始化类型对应的上下文模型索引集合确定为所述目标索引集合。
- 根据权利要求7所述的方法,其特征在于,所述根据所述目标模型索引确定所述目标上下文模型,包括:根据所述目标模型索引获取所述目标上下文模型的初始值和更新速率;根据所述目标上下文模型的初始值和更新速率构建所述目标上下文模型。
- 根据权利要求13所述的方法,其特征在于,所述目标上下文模型的初始值为35,所述目标上下文模型的更新速率为8。
- 根据权利要求1所述的方法,其特征在于,所述根据所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值,包括:将所述目标上下文模型作为基于上下文的自适应算术编码CABAC算法的上下文模型,通过所述CABAC算法对所述第一色度块的第一语法元素的熵编码数据进行算术解码,以获取所述第一色度块的所述第一语法元素的值。
- 根据权利要求15所述的方法,其特征在于,在获取所述第一色度块的所述第一语法元素的值之后,所述方法还包括:根据所述第一色度块的所述第一语法元素的值对所述目标上下文模型进行更新。
- 根据权利要求1所述的方法,其特征在于,所述方法还包括:根据所述目标上下文模型确定第二色度块对应的上下文模型;其中,所述第一色度块和所述第二色度块分别为同一编码块的两个色度分量对应的色度块。
- 一种视频编码方法,其特征在于,包括:根据第一色度块的相邻已编码色度块获取所述第一色度块的空域信息;根据所述第一色度块的空域信息确定目标上下文模型;基于所述目标上下文模型对所述第一色度块的第一语法元素的值进行熵编码,以获取所述第一色度块的所述第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素。
- 一种视频解码装置,其特征在于,包括:第一获取模块,用于获取第一色度块的第一语法元素的熵编码数据;所述第一语法元素为标识色度块的变换块是否为非零块的语法元素;第二获取模块,用于根据所述第一色度块的相邻已解码色度块获取所述第一色度块的空域信息;确定模块,用于根据所述第一色度块的空域信息确定目标上下文模型;熵解码模块,用于基于所述目标上下文模型和所述第一色度块的第一语法元素的熵编码数据,获取所述第一色度块的所述第一语法元素的值。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410902887.7A CN121334401A (zh) | 2024-07-05 | 2024-07-05 | 视频解码方法、视频编码方法及装置 |
| CN202410902887.7 | 2024-07-05 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026007482A1 true WO2026007482A1 (zh) | 2026-01-08 |
Family
ID=98317597
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/086881 Pending WO2026007482A1 (zh) | 2024-07-05 | 2025-04-02 | 视频解码方法、视频编码方法及装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121334401A (zh) |
| WO (1) | WO2026007482A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050169374A1 (en) * | 2004-01-30 | 2005-08-04 | Detlev Marpe | Video frame encoding and decoding |
| US20110249754A1 (en) * | 2010-04-12 | 2011-10-13 | Qualcomm Incorporated | Variable length coding of coded block pattern (cbp) in video compression |
| US20130195182A1 (en) * | 2012-02-01 | 2013-08-01 | General Instrument Corporation | Simplification of significance map coding |
| CN113632471A (zh) * | 2019-08-23 | 2021-11-09 | 腾讯美国有限责任公司 | 视频编解码的方法和装置 |
| CN117203961A (zh) * | 2022-04-06 | 2023-12-08 | 腾讯美国有限责任公司 | 利用来自空间邻居的信息的cabac上下文建模 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11470329B2 (en) * | 2018-12-26 | 2022-10-11 | Tencent America LLC | Method and apparatus for video coding |
| KR20250011723A (ko) * | 2019-07-10 | 2025-01-21 | 엘지전자 주식회사 | 영상 코딩 시스템에서 영상 코딩 방법 및 장치 |
| CN113132731B (zh) * | 2019-12-31 | 2025-07-22 | 腾讯科技(深圳)有限公司 | 视频解码方法、装置、设备及存储介质 |
-
2024
- 2024-07-05 CN CN202410902887.7A patent/CN121334401A/zh active Pending
-
2025
- 2025-04-02 WO PCT/CN2025/086881 patent/WO2026007482A1/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050169374A1 (en) * | 2004-01-30 | 2005-08-04 | Detlev Marpe | Video frame encoding and decoding |
| US20110249754A1 (en) * | 2010-04-12 | 2011-10-13 | Qualcomm Incorporated | Variable length coding of coded block pattern (cbp) in video compression |
| US20130195182A1 (en) * | 2012-02-01 | 2013-08-01 | General Instrument Corporation | Simplification of significance map coding |
| CN113632471A (zh) * | 2019-08-23 | 2021-11-09 | 腾讯美国有限责任公司 | 视频编解码的方法和装置 |
| CN117203961A (zh) * | 2022-04-06 | 2023-12-08 | 腾讯美国有限责任公司 | 利用来自空间邻居的信息的cabac上下文建模 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121334401A (zh) | 2026-01-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112640457B (zh) | 使用阈值和莱斯参数的用于系数编译码的常规编译码二进制数缩减 | |
| CN101502123B (zh) | 编码装置 | |
| CN113170138B (zh) | 使用阈值和莱斯参数进行系数解码的常规编译码二进制位缩减 | |
| KR102462386B1 (ko) | 데이터 인코딩 및 디코딩 | |
| US8401321B2 (en) | Method and apparatus for context adaptive binary arithmetic coding and decoding | |
| RU2637879C2 (ru) | Кодирование и декодирование значащих коэффициентов в зависимости от параметра указанных значащих коэффициентов | |
| US6900748B2 (en) | Method and apparatus for binarization and arithmetic coding of a data value | |
| TWI750624B (zh) | 編解碼變換係數的方法及裝置 | |
| CN104205831B (zh) | 用于对与变换系数相关联的比特流进行编码和解码方法 | |
| TWI827662B (zh) | 用於係數寫碼之規則寫碼位元子之減少 | |
| JP6526099B2 (ja) | Hevcにおけるcabacのための変換スキップされたブロックのための修正コーディング | |
| JP6532467B2 (ja) | ビデオ符号化および復号におけるシンタックス要素符号化方法および装置 | |
| JP7509784B2 (ja) | 係数レベルのためのエスケープコーディング | |
| US9544599B2 (en) | Context adaptive data encoding | |
| US20210243446A1 (en) | Context initialization in entropy coding | |
| JP2009021775A (ja) | 符号化装置及び符号化方法 | |
| CN121334401A (zh) | 视频解码方法、视频编码方法及装置 | |
| JP2007074337A (ja) | 符号化装置及び符号化方法 | |
| HK40045696B (zh) | 用於對視頻數據進行編解碼的方法和設備 | |
| HK40051990A (zh) | 使用阈值和莱斯参数进行系数解码的常规编译码二进制位缩减 | |
| GB2496193A (en) | Context adaptive data encoding and decoding | |
| CN116982314A (zh) | 系数编解码方法、编解码设备、终端及存储介质 | |
| HK40051990B (zh) | 使用阈值和莱斯参数进行系数解码的常规编译码二进制位缩减 | |
| JP2014068247A (ja) | 符号量計算装置及びプログラム、並びに、動画像符号化装置及びプログラム | |
| HK40058472B (zh) | 系数级别转义编解码 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25831841 Country of ref document: EP Kind code of ref document: A1 |