EP4696018A1 - An apparatus, a method and a computer program for video - Google Patents
An apparatus, a method and a computer program for videoInfo
- Publication number
- EP4696018A1 EP4696018A1 EP24788298.8A EP24788298A EP4696018A1 EP 4696018 A1 EP4696018 A1 EP 4696018A1 EP 24788298 A EP24788298 A EP 24788298A EP 4696018 A1 EP4696018 A1 EP 4696018A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- pixels
- blocks
- reconstruction
- visible
- decoding
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/167—Position within a video image, e.g. region of interest [ROI]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/597—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
- H04N19/82—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation involving filtering within a prediction loop
Definitions
- the present invention relates to an apparatus, a method and a computer program for video coding.
- the video frames representing content may be provided with an accompanying reconstruction signal for controlling the reconstruction or visibility of parts of the frame, or even single pixels, upon decoding.
- an accompanying reconstruction signal is an occupancy picture (a.k.a. occupancy map or occupancy component) used in MPEG Visual Volumetric Video-based coding (V3C).
- the occupancy component can inform a V3C decoding and/or rendering system of which samples in the 2D components are associated with data in the final 3D representation.
- an alpha channel a.k.a. alpha mask
- Modern video codecs utilise in-loop filtering to reduce coding artefacts caused e.g. by quantised transform coefficients.
- In-loop filtering such as deblocking filtering (DBF), adaptive loop filtering (ALF), sample adaptive offset (SAG), work well on traditional 2D video content.
- DPF deblocking filtering
- ALF adaptive loop filtering
- SAG sample adaptive offset
- an apparatus comprising means for obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and means for disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following: an occupancy map video; an occupancy signal; an alpha map video; an alpha channel.
- the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
- CU coding unit
- PU prediction unit
- TU transform unit
- CTU coding tree unit
- the apparatus comprises means for excluding samples corresponding pixels not becoming reconstructed or visible upon decoding from a filter parameter derivation phase.
- the apparatus comprises means for including only samples corresponding to pixels in a boundary area of pixels being reconstructed or becoming visible and non-visible upon decoding in a filter parameter derivation phase. [0012] According to an embodiment, the apparatus comprises means for adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non-visible upon decoding. [0013] According to an embodiment, the apparatus comprises means for calculating a dedicated class of filters for a block comprising pixels becoming both visible and non- visible upon decoding.
- the apparatus comprises means for using only a subset of in-loop filters for a block comprising pixels becoming both visible.
- the apparatus comprises means for applying a combination of in-loop filters block-wise and/or pixel-wise.
- the apparatus comprises means for fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
- the apparatus comprises means for signaling one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
- the apparatus comprises means for carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
- RBSP video parameter set raw byte sequence payload
- RBSP sequence parameter set raw byte sequence payload
- An apparatus comprises at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtain information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disable, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- a method comprises obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- An apparatus comprises means for receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and means for decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling inloop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
- An apparatus comprises at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receive, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and decode the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
- a method comprises receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
- Computer readable storage media comprise code for use by an apparatus, which when executed by a processor, causes the apparatus to perform the above methods.
- Figs, la and lb show an encoder and decoder for encoding and decoding 2D pictures
- Fig. 2 shows an example of an in-loop filter chain in VCC
- Figs. 3a and 3b show a compression and a decompression process for V3C volumetric video
- Fig. 4 shows an example of different stages in reconstructing a V3C coded volumetric video
- FIG. 5 shows a flow chart for a sending/ encoding method according to an embodiment
- Fig. 6 shows an example of different states for encoding a coding unit according to occupancy information
- Fig. 7 shows a flow chart for a receiving/ decoding method according to an embodiment.
- a video codec comprises an encoder that transforms the input video into a compressed representation suited for storage/transmission, and a decoder that can uncompress the compressed video representation back into a viewable form.
- An encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e. at lower bitrate).
- FIGs, la and lb show an encoder and decoder for encoding and decoding 2D pictures.
- a video codec consists of an encoder that transforms an input video into a compressed representation suited for storage/transmission and a decoder that can uncompress the compressed video representation back into a viewable form.
- the encoder discards and/or loses some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
- An example of an encoding process is illustrated in Figure la.
- Figure la illustrates an image to be encoded (I n ); a predicted representation of an image block (P' n ); a prediction error signal (D n ); a reconstructed prediction error signal (D' n ); a preliminary reconstructed image (I' n ); a final reconstructed image (R' n ); a transform (T) and inverse transform (T 1 ); a quantization (Q) and inverse quantization (Q 1 ); entropy encoding (E); a reference frame memory (RFM); inter prediction (Pinter); intra prediction (Pintra); mode selection (MS) and filtering (F).
- An example of a decoding process is illustrated in Figure lb.
- Figure lb illustrates a predicted representation of an image block (P' n ); a reconstructed prediction error signal (D' n ); a preliminary reconstructed image (I' n ); a final reconstructed image (R' n ); an inverse transform (T 1 ); an inverse quantization (Q 1 ); an entropy decoding (E 1 ); a reference frame memory (RFM); a prediction (either inter or intra) (P); and filtering (F).
- Many hybrid video encoders for example ITU-T H.263, H.264/AVC, HEVC, and WC, encode the video information in two phases.
- pixel values in a certain picture area are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner).
- the prediction error i.e. the difference between the predicted block of pixels and the original block of pixels. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients.
- DCT Discrete Cosine Transform
- Video codecs may also provide a transform skip mode, which the encoders may choose to use.
- the prediction error is coded in a sample domain, for example by deriving a sample-wise difference value relative to certain adjacent samples and coding the sample-wise difference value with an entropy coder.
- a coding block may be defined as an NxN block of samples for some value of N such that the division of a coding tree block into coding blocks is a partitioning.
- a coding tree block may be defined as an NxN block of samples for some value of N such that the division of a component into coding tree blocks is a partitioning.
- a coding tree unit may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture that has three sample arrays, or a coding tree block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples.
- a coding unit may be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture that has three sample arrays, or a coding block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples.
- a CU with the maximum allowed size may be named as LCU (largest coding unit) or coding tree unit (CTU) and the video picture is divided into non-overlapping LCUs.
- a picture can be partitioned in tiles, which are rectangular and contain an integer number of LCUs.
- the partitioning to tiles forms a regular grid, where heights and widths of tiles differ from each other by one LCU at the maximum.
- a slice is defined to be an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit.
- a slice segment is defined to be an integer number of coding tree units ordered consecutively in the tile scan and contained in a single NAL unit. The division of each picture into slice segments is a partitioning.
- an independent slice segment is defined to be a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values for a preceding slice segment
- a dependent slice segment is defined to be a slice segment for which the values of some syntax elements of the slice segment header are inferred from the values for the preceding independent slice segment in decoding order.
- a slice header is defined to be the slice segment header of the independent slice segment that is a current slice segment or is the independent slice segment that precedes a current dependent slice segment
- a slice segment header is defined to be a part of a coded slice segment containing the data elements pertaining to the first or all coding tree units represented in the slice segment.
- the CUs are scanned in the raster scan order of LCUs within tiles or within a picture, if tiles are not in use. Within an LCU, the CUs have a specific scan order.
- Entropy coding/decoding may be performed in many ways. For example, context-based coding/decoding may be applied, where in both the encoder and the decoder modify the context state of a coding parameter based on previously coded/ decoded coding parameters.
- Context-based coding may for example be context adaptive binary arithmetic coding (CABAC) or context-adaptive variable length coding (CAVLC) or any similar entropy coding.
- Entropy coding/decoding may alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding/decoding or Exp-Golomb coding/decoding. Decoding of coding parameters from an entropy-coded bitstream or codewords may be referred to as parsing.
- in-loop filtering The purpose of in-loop filtering is to reduce artifacts and distortions that can occur during the compression process.
- Compression techniques such as block-based motion compensation and discrete cosine transform (DCT) can introduce artifacts such as blocking, ringing, and blurring in the decoded video.
- DCT discrete cosine transform
- In-loop filtering is designed to reduce these artifacts and improve the perceived visual quality of the video.
- In-loop filters play a critical role in the maintenance of compressed video quality, since they can not only improve the quality of the current frame but can also provide a higher quality reference for subsequent frames.
- the in-loop filters in WC are depicted in Figure 2.
- Four processing steps namely a luma mapping with chroma scaling (LMCS) process, followed by a deblocking filter (DBF), an SAG filter, and an adaptive loop filter (ALF) are applied to the reconstructed samples before writing them into the decoded picture buffer.
- LMCS luma mapping with chroma scaling
- DBF deblocking filter
- SAG filter an adaptive loop filter
- ALF adaptive loop filter
- in-loop filtering consists of a fixed-order chain of three filters including deblocking filter (DBF), sample adaptive offset (SAG) filter, and adaptive loop filter (ALF).
- a block-based ALF is used in WC, which comprises luma ALF, chroma ALF and cross-component ALF (CC-ALF).
- the ALF filter coefficients are either pre-defined and fixed in both encoder and decoder or adaptively signaled on a picture basis using adaptation parameter set (APS).
- APS adaptation parameter set
- Solution 1 (Disabled ALF): As a naive method where no coordinated coding is required, ALF (i.e., both fixed ALF and ALF APS) is disabled to facilitate merging of different subpictures representations into a single picture at the cost of losing a substantial coding efficiency.
- ALF i.e., both fixed ALF and ALF APS
- Solution 2 (Disabled ALF APS): the pre-defined fixed ALF is enabled, while the ALF APS usage is disabled for each subpicture, at the cost of disregarding the benefit of ALF parameter adaptation which comes with a coding efficiency.
- the phrase along the bitstream may be defined to refer to out-of-band transmission, signaling, or storage in a manner that the out- of-band data is associated with the bitstream.
- the phrase decoding along the bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream.
- an indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.
- the video frames may be provided with an accompanying reconstruction signal for controlling the reconstruction or visibility of parts of the frame, or even single pixels, upon decoding.
- an accompanying reconstruction signal is an occupancy picture (a.k.a. occupancy map or occupancy component) used in MPEG Visual Volumetric Video-based coding (V3C).
- the occupancy component can inform a V3C decoding and/or rendering system of which samples in the 2D components are associated with data in the final 3D representation.
- Another example is an alpha channel (a.k.a. alpha mask) providing parts of the frame, or even single pixels, with an alpha value defining its level of transparency.
- V3C video coding is depicted more in detail further below to illustrate how the occupancy is used at the decoder to reconstruct an original 3D model.
- a first texture picture may be encoded into a bitstream, and the first texture picture may comprise a first projection of texture data of a first source volume of a scene model onto a first projection surface.
- the scene model may comprise a number of further source volumes.
- data on the position of the originating geometry primitive may also be determined, and based on this determination, a geometry picture may be formed. This may happen for example so that depth data is determined for each or some of the texture pixels of the texture picture. Depth data is formed such that the distance from the originating geometry primitive such as a point to the projection surface is determined for the pixels.
- Such depth data may be represented as a depth picture, and similarly to the texture picture, such geometry picture (such as a depth picture) may be encoded and decoded with a video codec.
- This first geometry picture may be seen to represent a mapping of the first projection surface to the first source volume, and the decoder may use this information to determine the location of geometry primitives in the model to be reconstructed.
- encoding a geometry (or depth) picture into or along the bitstream with the texture picture is only optional and arbitrary for example in the cases where the distance of all texture pixels to the projection surface is the same or there is no change in said distance between a plurality of texture pictures.
- a geometry (or depth) picture may be encoded into or along the bitstream with the texture picture, for example, only when there is a change in the distance of texture pixels to the projection surface.
- An attribute picture may be defined as a picture that comprises additional information related to an associated texture picture.
- An attribute picture may for example comprise surface normal, opacity, or reflectance information for a texture picture.
- a geometry picture may be regarded as one type of an attribute picture, although a geometry picture may be treated as its own picture type, separate from an attribute picture.
- Texture picture(s) and the respective geometry picture(s), if any, and the respective attribute picture(s) may have the same or different chroma format.
- Terms texture (component) image and texture (component) picture may be used interchangeably.
- Terms geometry (component) image and geometry (component) picture may be used interchangeably.
- a specific type of a geometry image is a depth image.
- Embodiments described in relation to a geometry (component) image equally apply to a depth (component) image, and embodiments described in relation to a depth (component) image equally apply to a geometry (component) image.
- Terms attribute image and attribute picture may be used interchangeably.
- a geometry picture and/or an attribute picture may be treated as an auxiliary picture in video/image encoding and/or decoding.
- FIGS 3a and 3b illustrate an overview of exemplified compression/ decompression processes.
- the processes may be applied, for example, in MPEG visual volumetric video-based coding (V3C), defined currently in ISO/IEC DIS 23090-5: “Visual Volumetric Video-based Coding and Video-based Point Cloud Compression”, 2nd Edition.
- V3C specification enables the encoding and decoding processes of a variety of volumetric media by using video and image coding technologies. This is achieved through first a conversion of such media from their corresponding 3D representation to multiple 2D representations, also referred to as V3C components, before coding such information.
- Such representations may include occupancy, geometry, and attribute components.
- the occupancy component can inform a V3C decoding and/or rendering system of which samples in the 2D components are associated with data in the final 3D representation.
- the geometry component contains information about the precise location of 3D data in space, while attribute components can provide additional properties, e.g. texture or material information, of such 3D data.
- An example of volumetric media conversion at an encoder is shown in Figure 2a and an example of a 3D reconstruction at a decoder is shown in Figure 3b.
- a V3C decoder receives the three video bitstreams, alongside a V3C atlas bitstream and then reconstructs the volumetric video frame by frame as follows:
- the decoder reconstructs the patch positions in 3D space based on the atlas bitstream
- the decoder reconstructs the patch shape according to the occupancy map signal
- the decoder reconstructs the 3D positions of each point per patch based on the geometry video
- the decoder applies attributes to each point, e.g. texture color.
- attributes e.g. texture color.
- volumetric frame can be represented as a point cloud.
- a point cloud is a set of unstructured points in 3D space, where each point is characterized by its position in a 3D coordinate system (e.g., Euclidean), and some corresponding attributes (e.g., color information provided as RGBA value, or normal vectors).
- a volumetric frame can be represented as images, with or without depth, captured from multiple viewpoints in 3D space.
- the volumetric video can be represented by one or more view frames (where a view is a projection of a volumetric scene on to a plane (the camera plane) using a real or virtual camera with known/ computed extrinsic and intrinsic).
- Each view may be represented by a number of components (e.g., geometry, color, transparency, and occupancy picture), which may be part of the geometry picture or represented separately.
- a volumetric frame can be represented as a mesh.
- Mesh is a collection of points, called vertices, and connectivity information between vertices, called edges. Vertices along with edges form faces. The combination of vertices, edges and faces can uniquely approximate shapes of objects.
- a volumetric frame can provide viewers the ability to navigate a scene with six degrees of freedom, i.e., both translational and rotational movement of their viewing pose (which includes yaw, pitch, and roll).
- the data to be coded for a volumetric frame can also be significant, as a volumetric frame can contain many numbers of objects, and the positioning and movement of these objects in the scene can result in many dis-occluded regions.
- the interaction of the light and materials in objects and surfaces in a volumetric frame can generate complex light fields that can produce texture variations for even a slight change of pose.
- a sequence of volumetric frames is a volumetric video. Due to large amount of information, storage and transmission of a volumetric video requires compression.
- a way to compress a volumetric frame can be to project the 3D geometry and related attributes into a collection of 2D images along with additional associated metadata.
- the projected 2D images can then be coded using 2D video and image coding technologies, for example ISO/IEC 14496-10 (H.264/AVC) and ISO/IEC 23008-2 (H.265/HEVC).
- the metadata can be coded with technologies specified in specification such as ISO/IEC 23090-5.
- the coded images and the associated metadata can be stored or transmitted to a client that can decode and render the 3D volumetric frame.
- In-loop filtering such as deblocking filtering (DBF), adaptive loop filtering (ALF), sample adaptive offset (SAG), work well on traditional 2D video content.
- DPF deblocking filtering
- ALF adaptive loop filtering
- SAG sample adaptive offset
- the method comprises obtaining (500) a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtaining (502) information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disabling (504), based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- the method enables to remove at least a part of distortion and increase the coding efficiency through reduced bitrate by activating and deactivating some or all video codec in-loop filters, based on a secondary input indicating information about reconstruction of said pixels in said blocks upon decoding, e.g. an occupancy signal. Based on this information, the encoder may disable in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- subset of pixels refers to a block where at least one but not all pixels of the block are indicated to be reconstructed or visible upon decoding.
- FIG. 6 shows three possible states (A, B, C) for a block, such as an 8 x 8 pixel coding unit (CU).
- the upper block represents the video, i.e. the coding unit
- the lower block i.e. a reconstruction information CU
- White values in the reconstruction information CU indicate that the corresponding pixel is visible/reconstructed at the decoder and dark values indicate that the corresponding pixel is not visible/reconstructed at the decoder.
- Some (a subset) of the pixels in the CU will be reconstructed/visible at the decoder.
- the encoder adaptively guides the in-loop filtering process for each state, such as follows:
- the RDO calculations may be optimized to favor less reduced rate over increased distortion.
- the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following:
- the occupancy map video addresses the V3C V-PCC use case by indicating which pixel positions will be reconstructed / visible at the decoder.
- the occupancy signal which may be multiplexed in the same or yet another video, addresses the V3C MIV use case by indicating which pixel positions will be reconstructed / visible at the decoder, wherein, for example, values below a certain offset are considered unoccupied.
- the alpha map video and alpha channel correspondingly addresses the Alpha use case.
- the video signal comprises one or more of the following: individual video streams; separate layers for multi-layer coding; separate sublayers; separate pictures in a temporally interleaved manner in the same coded layer video sequence; separate constituent frames in spatially frame-packed video.
- the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
- CU coding unit
- PU prediction unit
- TU transform unit
- CTU coding tree unit
- the size of the block to be processed is irrespective for the implementation.
- the method comprises excluding samples corresponding pixels not becoming visible upon decoding from a filter parameter derivation phase.
- the corresponding samples in the encoder side will be excluded from the filter parameter derivation phase.
- the adaptive loop filter (ALF) training in the encoder side may use only samples that will be visible at the decoder side.
- other filters such as luma mapping with chroma scaling (LMCS) and sample adaptive offset filter (SAO) may follow the same principles.
- the method comprises including only samples corresponding to pixels in a boundary area of pixels becoming visible and non- visible upon decoding in a filter parameter derivation phase.
- the method comprises adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels becoming visible and non-visible upon decoding.
- the filter strength may be adjusted in such a way that it has weaker impact to the samples in the boundary areas of the two sample types.
- a weaker deblocking filter may be considered in such areas.
- the weaker filtering may still reduce the coding artifacts while preserving the object edges in the boundary regions of the two sample types.
- the method comprises calculating a dedicated class of filters for a block comprising pixels becoming both visible and non-visible upon decoding.
- a dedicated class of filters may be calculated for the blocks that include both visible and non-visible samples for the decoder (e.g., state C of Figure 6).
- the ALF filter may contain one or more classes of filters trained for such cases.
- the available occupancy information may be used for indicating the filter class based on the content.
- the method comprises disabling all in-loop filters for a block comprising pixels becoming both visible and non-visible upon decoding.
- the encoder may disable all in-loop filter for blocks, such as CUs, containing partially reconstructed/visible content.
- the method comprises using only a subset of inloop filters for a block comprising pixels becoming both visible and non-visible upon decoding.
- the encoder may only use a subset of in-loop filters for blocks, such as CUs, containing partially reconstructed/visible content.
- the encoder may only use LMCS and SAO, while disabling DBF and ALF.
- the method comprises applying at least one inloop filter for only pixels of a block becoming visible upon decoding.
- the encoder does not disable in-loop filters, however, the filter is only applied on pixel marked as visible/reconstructed at the decoder, while other pixels retain their original value.
- the method comprises applying a combination of in-loop filters block- wise and/or pixel- wise.
- the encoder may apply or disable one or more in-loop filters on a block level, and at the same time, apply or disable one or more in-loop filters on a pixel level.
- DBF may be disabled per-CU, whereas ALF is applied per-pixel.
- the method comprises fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
- the one or more of the in-loop filters may be fully or partially deactivated.
- the method comprises skipping a signaling of filter information for said fully or partially deactivated one or more in-loop filters.
- signaling of the filter information for the one or more in-loop filters of the underlying codec may be also skipped and a pre-defined values may be assigned for the inloop filters.
- the filter information may comprise one or more of the following: filter activation, filter type, filter class and index, etc.
- the method comprises signaling one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
- the method comprises carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
- RBSP video parameter set raw byte sequence payload
- RBSP sequence parameter set raw byte sequence payload
- the method comprises signaling one or more parameters about applicability of the in-loop filters as a supplemental enhancement information (SEI) message.
- SEI Supplemental Enhancement Information
- said signaling comprises a reference to a reconstruction guided in-loop filtering as part of a multi-layer encoding structure.
- a new syntax element vps_reconstruction_layer_id specifies the nuh layer id value of the layer carrying the reconstruction information relevant for reconstruction guided in-loop filter.
- the signaling comprises indicating a layer using the reconstruction guided in-loop filtering per a dependent layer.
- a new syntax element vps_reconstruction_guided_inloop_layer_flag[ i ] having a value equal to 1 specifies that the i-th layer uses reconstruction guided in-loop filtering
- vps reconstruction guided inloop layer _flag[ i ] 0 specifies that the i-th layer does not use reconstruction guided in-loop filtering
- a new syntax element vps_reconstruction_layer_idx[ i ] specifies that the layer index for the reconstuction information for the i-th layer is ReferenceLayerIdx[ i ] [ vps_reconstruction_layer_idx[ i ] ] .
- the signaling comprises indicating an activation of one or more in-loop filters individually.
- vps_reconstruction_lmcs_flag 0 specifies that LMCS in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction lmcs flag 1 specifies that LMCS in-loop filter is applied if partial reconstruction information is available.
- vps reconstruction dbf flag 0 specifies that DBF in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction dbf flag 1 specifies that DBF in-loop filter is applied if partial reconstruction information is available.
- vps_reconstruction_sao_flag 0 specifies that SAG in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction sao flag 1 specifies that SAG in-loop filter is applied if partial reconstruction information is available.
- vps reconstruction alf flag 0 specifies that ALF in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction alf flag 1 specifies that ALF in-loop filter is applied if partial reconstruction information is available.
- the signaling comprises both a flag indicating activation of reconstruction guided in-loop filtering and indication for an activation of one or more in-loop filters individually.
- the flag values are indicated by two bits, thereby enabling more versatile signaling.
- New syntax elements are introduced as follows:
- vps_reconstruction_lmcs_flag 0 specifies that LMCS in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction lmcs flag 1 specifies that LMCS in-loop filter is applied if partial reconstruction information is available per CU.
- vps_reconstruction_lmcs_flag 2 specifies that LMCS in-loop filter is applied only to sample with reconstruction information available.
- vps reconstruction dbf flag 0 specifies that DBF in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction dbf flag 1 specifies that DBF in-loop filter is applied if partial reconstruction information is available per CU.
- vps reconstruction lmcs flag 2 specifies that DBF in-loop filter is applied only to sample with reconstruction information available.
- vps_reconstruction_sao_flag 0 specifies that SAG in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction sao flag 1 specifies that SAG in-loop filter is applied if partial reconstruction information is available per CU.
- vps_reconstruction_lmcs_flag 2 specifies that SAG in-loop filter is applied only to sample with reconstruction information available.
- vps reconstruction alf flag 0 specifies that ALF in-loop filter is not applied if partial reconstruction information is available
- vps reconstruction alf flag 1 specifies that ALF in-loop filter is applied if partial reconstruction information is available per CU.
- vps reconstruction lmcs flag 2 specifies that ALF in-loop filter is applied only to sample with reconstruction information available.
- one or more threshold syntax elements that control the use of the reconstruction information for in-loop filtering are included by an encoder in or along a bitstream, such as in a VPS, and/or decoded by a decoder from or along a bitstream, such as from a VPS.
- a threshold syntax element may for example define the sample value range in the reconstruction information that indicates a non- visible pixel for in-loop filtering.
- numReconstructionlnformationLayers is a variable indicating the number of layers that contain reconsruction information to be used for guided in-loop filtering:
- the syntax element vps_reconstruction_luma_threshold[ i ] specifies that samples with luma sample value in the range of 0 to vps_reconstruction_luma_threshold[ i ], inclusive, indicate non-visible samples for in-loop filtering.
- FIG. 7 Another aspect relates to the operation of a decoder (or a Tenderer/ receiver/player/ client).
- the method which is disclosed in Figure 7 as illustrating the operation of a decoder, comprises receiving (700) a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receiving (702), in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; decoding (704) the an encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
- the decoding apparatus receives a bitstream, containing reconstruction guided in-loop filter signalling as disclosed above, and comprising encoded video data, and in or along the bitstream the reconstruction information.
- the reconstruction information may be provided in the form of a bitstream comprising an encoded occupancy (V3C use case) or alpha map video, either as individual videos (V3C V-PCC use case) or as layer in a multi-layer encoding, or a reconstruction signal, multiplexed in the same or yet another video (V3C MIV use case).
- the decoder first decodes the reconstruction information, then the decoder decodes the video data, activating/disabling in-loop filters according to said information indicative about reconstruction or visibility of said pixels in said blocks, i.e. reconstruction information state (full, none, partial) and the received signaling.
- any of the embodiments disclosed above, relating to either encoding or decoding may be applied similarly to a post-processing filter.
- an input tensor may be formed.
- it may be indicated by an encoder, e.g. in a neural-network post-filter characteristics (NNPFC) SEI message, that the post-filter expects auxiliary input in the input tensor for the reconstruction information and/or variables controlling the reconstruction information guided filtering, such as one or more threshold values defining the sample value range in the reconstruction information that indicates a non-visible pixel for filtering.
- NNPFC neural-network post-filter characteristics
- it may be decoded by a decoder, e.g.
- the decoder forms an input tensor with reconstruction information and/or variables controlling the reconstruction information guided filtering.
- inventions relating to the sending/encoding aspects may be implemented in an apparatus comprising: means for obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and means for disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following: an occupancy map video; an occupancy signal; an alpha map video; an alpha channel.
- the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
- the apparatus comprises means for adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non-visible upon decoding.
- the apparatus comprises means for calculating a dedicated class of filters for a block comprising pixels becoming both visible and non- visible upon decoding.
- the apparatus comprises means for applying a combination of in-loop filters block-wise and/or pixel-wise.
- the apparatus comprises means for fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
- the apparatus comprises means for carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
- RBSP video parameter set raw byte sequence payload
- RBSP sequence parameter set raw byte sequence payload
- the embodiments relating to the sending/ encoding aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtain information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disable, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
- the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following: an occupancy map video; an occupancy signal; an alpha map video; an alpha channel.
- the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
- CU coding unit
- PU prediction unit
- TU transform unit
- CTU coding tree unit
- the apparatus comprises code configured to cause the apparatus to exclude samples corresponding pixels not becoming reconstructed or visible upon decoding from a filter parameter derivation phase.
- the apparatus comprises code configured to cause the apparatus to include only samples corresponding to pixels in a boundary area of pixels being reconstructed or becoming visible and non-visible upon decoding in a filter parameter derivation phase.
- the apparatus comprises code configured to cause the apparatus to adjust a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non-visible upon decoding.
- the apparatus comprises code configured to cause the apparatus to calculate a dedicated class of filters for a block comprising pixels becoming both visible and non-visible upon decoding.
- the apparatus comprises code configured to cause the apparatus to use only a subset of in-loop filters for a block comprising pixels becoming both visible.
- the apparatus comprises code configured to cause the apparatus to apply a combination of in-loop filters block-wise and/or pixel-wise.
- the apparatus comprises code configured to cause the apparatus to signal one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
- the apparatus comprises code configured to cause the apparatus to carry out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
- RBSP video parameter set raw byte sequence payload
- RBSP sequence parameter set raw byte sequence payload
- the receiving/decoding/rendering aspects may be implemented by an apparatus comprising means for receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and means for decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
- the embodiments relating to the receiving/decoding/rendering aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receive, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and decode the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
- said encoding may comprise one or more of the following: encoding source image data into a bitstream, encapsulating the encoded bitstream in a container file and/or in packet(s) or stream(s) of a communication protocol, and announcing or describing the bitstream in a content description, such as the Media Presentation Description (MPD) of ISO/IEC 23009-1 (known as MPEG-DASH) or the IETF Session Description Protocol (SDP).
- MPD Media Presentation Description
- SDP IETF Session Description Protocol
- Programs such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre stored design modules.
- the resultant design in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
A method comprising: obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels (500); obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding (502); and disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding (504).
Description
AN APPARATUS, A METHOD AND A COMPUTER PROGRAM FOR VIDEO
TECHNICAE FIEED
[0001 ] The present invention relates to an apparatus, a method and a computer program for video coding.
BACKGROUND
[0002] The video frames representing content may be provided with an accompanying reconstruction signal for controlling the reconstruction or visibility of parts of the frame, or even single pixels, upon decoding. One example of such an accompanying reconstruction signal is an occupancy picture (a.k.a. occupancy map or occupancy component) used in MPEG Visual Volumetric Video-based coding (V3C). The occupancy component can inform a V3C decoding and/or rendering system of which samples in the 2D components are associated with data in the final 3D representation. Another example is an alpha channel (a.k.a. alpha mask) providing parts of the frame, or even single pixels, with an alpha value defining its level of transparency.
[0003] Modern video codecs utilise in-loop filtering to reduce coding artefacts caused e.g. by quantised transform coefficients. In-loop filtering, such as deblocking filtering (DBF), adaptive loop filtering (ALF), sample adaptive offset (SAG), work well on traditional 2D video content.
[0004] However, for video content provided with an accompanying reconstruction signal such filtering may lead to unwanted distortions and decreased coding efficiency when applied on areas not intended for reconstruction, for example, areas in geometry and attribute videos with a corresponding occupancy map video area set to 0 for V3C video, or areas with an alpha mask set to fully transparent.
SUMMARY
[0005] Now, an improved method and technical equipment implementing the method has been invented, by which the above problems are alleviated. Various aspects include a method, an apparatus and a computer readable medium comprising a computer program, or a signal stored therein, which are characterized by what is stated in the independent claims.
Various details of the embodiments are disclosed in the dependent claims and in the corresponding images and description.
[0006] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.
[0007] According to a first aspect, there is provided an apparatus comprising means for obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and means for disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0008] According to an embodiment, the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following: an occupancy map video; an occupancy signal; an alpha map video; an alpha channel.
[0009] According to an embodiment, the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
[0010] According to an embodiment, the apparatus comprises means for excluding samples corresponding pixels not becoming reconstructed or visible upon decoding from a filter parameter derivation phase.
[001 1] According to an embodiment, the apparatus comprises means for including only samples corresponding to pixels in a boundary area of pixels being reconstructed or becoming visible and non-visible upon decoding in a filter parameter derivation phase. [0012] According to an embodiment, the apparatus comprises means for adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non-visible upon decoding.
[0013] According to an embodiment, the apparatus comprises means for calculating a dedicated class of filters for a block comprising pixels becoming both visible and non- visible upon decoding.
[0014] According to an embodiment, the apparatus comprises means for using only a subset of in-loop filters for a block comprising pixels becoming both visible.
[0015] According to an embodiment, the apparatus comprises means for applying a combination of in-loop filters block-wise and/or pixel-wise.
[0016] According to an embodiment, the apparatus comprises means for fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
[0017] According to an embodiment, the apparatus comprises means for signaling one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
[0018] According to an embodiment, the apparatus comprises means for carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
[0019] An apparatus according to a second aspect comprises at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtain information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disable, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0020] A method according to a third aspect comprises obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disabling, based on said information indicative about reconstruction or
visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0021] An apparatus according to a fourth aspect comprises means for receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and means for decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling inloop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
[0022] An apparatus according to a fifth aspect comprises at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receive, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and decode the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
[0023] A method according to a sixth aspect comprises receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
[0024] Computer readable storage media according to further aspects comprise code for use by an apparatus, which when executed by a processor, causes the apparatus to perform the above methods.
BRIEF DESCRIPTION OF THE DRAWINGS
[0025] For a more complete understanding of the example embodiments, reference is now made to the following descriptions taken in connection with the accompanying drawings in which:
[0026] Figs, la and lb show an encoder and decoder for encoding and decoding 2D pictures;
[0027] Fig. 2 shows an example of an in-loop filter chain in VCC;
[0028] Figs. 3a and 3b show a compression and a decompression process for V3C volumetric video;
[0029] Fig. 4 shows an example of different stages in reconstructing a V3C coded volumetric video;
[0030] Fig. 5 shows a flow chart for a sending/ encoding method according to an embodiment;
[0031] Fig. 6 shows an example of different states for encoding a coding unit according to occupancy information; and
[0032] Fig. 7 shows a flow chart for a receiving/ decoding method according to an embodiment.
DETAILED DESCRIPTON OF SOME EXAMPLE EMBODIMENTS
[0033] A video codec comprises an encoder that transforms the input video into a compressed representation suited for storage/transmission, and a decoder that can uncompress the compressed video representation back into a viewable form. An encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e. at lower bitrate).
[0034] Figs, la and lb show an encoder and decoder for encoding and decoding 2D pictures. A video codec consists of an encoder that transforms an input video into a compressed representation suited for storage/transmission and a decoder that can uncompress the compressed video representation back into a viewable form. Typically, the encoder discards and/or loses some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate). An example of an encoding process is illustrated in Figure la. Figure la illustrates an image to be encoded
(In); a predicted representation of an image block (P'n); a prediction error signal (Dn); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (I'n); a final reconstructed image (R'n); a transform (T) and inverse transform (T 1); a quantization (Q) and inverse quantization (Q 1); entropy encoding (E); a reference frame memory (RFM); inter prediction (Pinter); intra prediction (Pintra); mode selection (MS) and filtering (F). [0035] An example of a decoding process is illustrated in Figure lb. Figure lb illustrates a predicted representation of an image block (P'n); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (I'n); a final reconstructed image (R'n); an inverse transform (T 1); an inverse quantization (Q 1); an entropy decoding (E 1); a reference frame memory (RFM); a prediction (either inter or intra) (P); and filtering (F). [0036] Many hybrid video encoders, for example ITU-T H.263, H.264/AVC, HEVC, and WC, encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate). Video codecs may also provide a transform skip mode, which the encoders may choose to use. In the transform skip mode, the prediction error is coded in a sample domain, for example by deriving a sample-wise difference value relative to certain adjacent samples and coding the sample-wise difference value with an entropy coder.
[0037] Many video encoders partition a picture into blocks along a block grid. For example, in the High Efficiency Video Coding (HEVC) standard, the following partitioning and definitions are used. A coding block may be defined as an NxN block of samples for some value of N such that the division of a coding tree block into coding
blocks is a partitioning. A coding tree block (CTB) may be defined as an NxN block of samples for some value of N such that the division of a component into coding tree blocks is a partitioning. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture that has three sample arrays, or a coding tree block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture that has three sample arrays, or a coding block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples. A CU with the maximum allowed size may be named as LCU (largest coding unit) or coding tree unit (CTU) and the video picture is divided into non-overlapping LCUs.
[0038] In HEVC, a picture can be partitioned in tiles, which are rectangular and contain an integer number of LCUs. In HEVC, the partitioning to tiles forms a regular grid, where heights and widths of tiles differ from each other by one LCU at the maximum. In HEVC, a slice is defined to be an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit. In HEVC, a slice segment is defined to be an integer number of coding tree units ordered consecutively in the tile scan and contained in a single NAL unit. The division of each picture into slice segments is a partitioning. In HEVC, an independent slice segment is defined to be a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values for a preceding slice segment, and a dependent slice segment is defined to be a slice segment for which the values of some syntax elements of the slice segment header are inferred from the values for the preceding independent slice segment in decoding order. In HEVC, a slice header is defined to be the slice segment header of the independent slice segment that is a current slice segment or is the independent slice segment that precedes a current dependent slice segment, and a slice segment header is defined to be a part of a coded slice segment containing the data elements pertaining to the first or all coding tree units represented in the slice segment. The CUs are scanned in the raster scan order of
LCUs within tiles or within a picture, if tiles are not in use. Within an LCU, the CUs have a specific scan order.
[0039] Entropy coding/decoding may be performed in many ways. For example, context-based coding/decoding may be applied, where in both the encoder and the decoder modify the context state of a coding parameter based on previously coded/ decoded coding parameters. Context-based coding may for example be context adaptive binary arithmetic coding (CABAC) or context-adaptive variable length coding (CAVLC) or any similar entropy coding. Entropy coding/decoding may alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding/decoding or Exp-Golomb coding/decoding. Decoding of coding parameters from an entropy-coded bitstream or codewords may be referred to as parsing.
[0040] The purpose of in-loop filtering is to reduce artifacts and distortions that can occur during the compression process. Compression techniques such as block-based motion compensation and discrete cosine transform (DCT) can introduce artifacts such as blocking, ringing, and blurring in the decoded video. In-loop filtering is designed to reduce these artifacts and improve the perceived visual quality of the video.
[00 1] In-loop filters play a critical role in the maintenance of compressed video quality, since they can not only improve the quality of the current frame but can also provide a higher quality reference for subsequent frames. The in-loop filters in WC are depicted in Figure 2. Four processing steps, namely a luma mapping with chroma scaling (LMCS) process, followed by a deblocking filter (DBF), an SAG filter, and an adaptive loop filter (ALF) are applied to the reconstructed samples before writing them into the decoded picture buffer. The DBF and SAG are similar to that of the HEVC standard, whereas LMCS and ALF are newly introduced in WC.
[0042] In WC, in-loop filtering consists of a fixed-order chain of three filters including deblocking filter (DBF), sample adaptive offset (SAG) filter, and adaptive loop filter (ALF). A block-based ALF is used in WC, which comprises luma ALF, chroma ALF and cross-component ALF (CC-ALF). The ALF filter coefficients are either pre-defined and fixed in both encoder and decoder or adaptively signaled on a picture basis using adaptation parameter set (APS). In order to enable merging of subpictures bitstreams
(coded using independent encoder instances) into a single WC-compliant picture without any ALF-APS identifiers conflict issue, the following solutions have been proposed:
Solution 1 (Disabled ALF): As a naive method where no coordinated coding is required, ALF (i.e., both fixed ALF and ALF APS) is disabled to facilitate merging of different subpictures representations into a single picture at the cost of losing a substantial coding efficiency.
Solution 2 (Disabled ALF APS): the pre-defined fixed ALF is enabled, while the ALF APS usage is disabled for each subpicture, at the cost of disregarding the benefit of ALF parameter adaptation which comes with a coding efficiency.
[0043] The phrase along the bitstream (e.g. indicating along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that the out- of-band data is associated with the bitstream. The phrase decoding along the bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream. For example, an indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.
[0044] The video frames may be provided with an accompanying reconstruction signal for controlling the reconstruction or visibility of parts of the frame, or even single pixels, upon decoding. One example of such an accompanying reconstruction signal is an occupancy picture (a.k.a. occupancy map or occupancy component) used in MPEG Visual Volumetric Video-based coding (V3C). The occupancy component can inform a V3C decoding and/or rendering system of which samples in the 2D components are associated with data in the final 3D representation. Another example is an alpha channel (a.k.a. alpha mask) providing parts of the frame, or even single pixels, with an alpha value defining its level of transparency. As a more detailed example, V3C video coding is depicted more in detail further below to illustrate how the occupancy is used at the decoder to reconstruct an original 3D model.
[0045] A first texture picture may be encoded into a bitstream, and the first texture picture may comprise a first projection of texture data of a first source volume of a scene model onto a first projection surface. The scene model may comprise a number of further source volumes.
[0046] In the projection, data on the position of the originating geometry primitive may also be determined, and based on this determination, a geometry picture may be formed. This may happen for example so that depth data is determined for each or some of the texture pixels of the texture picture. Depth data is formed such that the distance from the originating geometry primitive such as a point to the projection surface is determined for the pixels. Such depth data may be represented as a depth picture, and similarly to the texture picture, such geometry picture (such as a depth picture) may be encoded and decoded with a video codec. This first geometry picture may be seen to represent a mapping of the first projection surface to the first source volume, and the decoder may use this information to determine the location of geometry primitives in the model to be reconstructed. In order to determine the position of the first source volume and/or the first projection surface and/or the first projection in the scene model, there may be first geometry information encoded into or along the bitstream. It is noted that encoding a geometry (or depth) picture into or along the bitstream with the texture picture is only optional and arbitrary for example in the cases where the distance of all texture pixels to the projection surface is the same or there is no change in said distance between a plurality of texture pictures. Thus, a geometry (or depth) picture may be encoded into or along the bitstream with the texture picture, for example, only when there is a change in the distance of texture pixels to the projection surface.
[0047] An attribute picture may be defined as a picture that comprises additional information related to an associated texture picture. An attribute picture may for example comprise surface normal, opacity, or reflectance information for a texture picture. A geometry picture may be regarded as one type of an attribute picture, although a geometry picture may be treated as its own picture type, separate from an attribute picture.
[0048] Texture picture(s) and the respective geometry picture(s), if any, and the respective attribute picture(s) may have the same or different chroma format.
[0049] Terms texture (component) image and texture (component) picture may be used interchangeably. Terms geometry (component) image and geometry (component) picture may be used interchangeably. A specific type of a geometry image is a depth image. Embodiments described in relation to a geometry (component) image equally apply to a depth (component) image, and embodiments described in relation to a depth (component)
image equally apply to a geometry (component) image. Terms attribute image and attribute picture may be used interchangeably. A geometry picture and/or an attribute picture may be treated as an auxiliary picture in video/image encoding and/or decoding.
[0050] Figures 3a and 3b illustrate an overview of exemplified compression/ decompression processes. The processes may be applied, for example, in MPEG visual volumetric video-based coding (V3C), defined currently in ISO/IEC DIS 23090-5: “Visual Volumetric Video-based Coding and Video-based Point Cloud Compression”, 2nd Edition. [0051 ] V3C specification enables the encoding and decoding processes of a variety of volumetric media by using video and image coding technologies. This is achieved through first a conversion of such media from their corresponding 3D representation to multiple 2D representations, also referred to as V3C components, before coding such information. Such representations may include occupancy, geometry, and attribute components. The occupancy component can inform a V3C decoding and/or rendering system of which samples in the 2D components are associated with data in the final 3D representation. The geometry component contains information about the precise location of 3D data in space, while attribute components can provide additional properties, e.g. texture or material information, of such 3D data. An example of volumetric media conversion at an encoder is shown in Figure 2a and an example of a 3D reconstruction at a decoder is shown in Figure 3b.
[0052] A V3C decoder receives the three video bitstreams, alongside a V3C atlas bitstream and then reconstructs the volumetric video frame by frame as follows:
1. The decoder reconstructs the patch positions in 3D space based on the atlas bitstream;
2. The decoder reconstructs the patch shape according to the occupancy map signal;
3. The decoder reconstructs the 3D positions of each point per patch based on the geometry video;
4. The decoder applies attributes to each point, e.g. texture color. [0053] The different stages of this process are visualised in Figure 4.
[0054] There are alternatives to capture and represent a volumetric frame. The format used to capture and represent the volumetric frame depends on the process to be performed on it, and the target application using the volumetric frame. As a first example a volumetric
frame can be represented as a point cloud. A point cloud is a set of unstructured points in 3D space, where each point is characterized by its position in a 3D coordinate system (e.g., Euclidean), and some corresponding attributes (e.g., color information provided as RGBA value, or normal vectors). As a second example, a volumetric frame can be represented as images, with or without depth, captured from multiple viewpoints in 3D space. In other words, the volumetric video can be represented by one or more view frames (where a view is a projection of a volumetric scene on to a plane (the camera plane) using a real or virtual camera with known/ computed extrinsic and intrinsic). Each view may be represented by a number of components (e.g., geometry, color, transparency, and occupancy picture), which may be part of the geometry picture or represented separately. As a third example, a volumetric frame can be represented as a mesh. Mesh is a collection of points, called vertices, and connectivity information between vertices, called edges. Vertices along with edges form faces. The combination of vertices, edges and faces can uniquely approximate shapes of objects.
[0055] Depending on the capture, a volumetric frame can provide viewers the ability to navigate a scene with six degrees of freedom, i.e., both translational and rotational movement of their viewing pose (which includes yaw, pitch, and roll). The data to be coded for a volumetric frame can also be significant, as a volumetric frame can contain many numbers of objects, and the positioning and movement of these objects in the scene can result in many dis-occluded regions. Furthermore, the interaction of the light and materials in objects and surfaces in a volumetric frame can generate complex light fields that can produce texture variations for even a slight change of pose.
[0056] A sequence of volumetric frames is a volumetric video. Due to large amount of information, storage and transmission of a volumetric video requires compression. A way to compress a volumetric frame can be to project the 3D geometry and related attributes into a collection of 2D images along with additional associated metadata. The projected 2D images can then be coded using 2D video and image coding technologies, for example ISO/IEC 14496-10 (H.264/AVC) and ISO/IEC 23008-2 (H.265/HEVC). The metadata can be coded with technologies specified in specification such as ISO/IEC 23090-5. The coded images and the associated metadata can be stored or transmitted to a client that can decode and render the 3D volumetric frame.
[0057] As described above, modern video codecs utilise in-loop filtering to reduce coding artefacts caused e.g. by quantised transform coefficients. In-loop filtering, such as deblocking filtering (DBF), adaptive loop filtering (ALF), sample adaptive offset (SAG), work well on traditional 2D video content.
[0058] However, for video content accompanied by a reconstruction signal, e.g. occupancy in V3C video or an alpha channel, such filtering may lead to unwanted distortions and decreased coding efficiency when applied on areas not intended for reconstruction at the decoder, i.e. areas in geometry and attribute videos with a corresponding occupancy map video area set to 0 for V3C or areas with a corresponding alpha channel set to fully transparent.
[0059] In the following, an enhanced method for will be described in more detail, in accordance with various embodiments.
[0060] The method, which is disclosed in Figure 5, comprises obtaining (500) a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtaining (502) information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disabling (504), based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0061] Thus, the method enables to remove at least a part of distortion and increase the coding efficiency through reduced bitrate by activating and deactivating some or all video codec in-loop filters, based on a secondary input indicating information about reconstruction of said pixels in said blocks upon decoding, e.g. an occupancy signal. Based on this information, the encoder may disable in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0062] It is noted that herein the term “subset of pixels” refers to a block where at least one but not all pixels of the block are indicated to be reconstructed or visible upon decoding.
[0063] The principle underlying the method may be illustrated by an example shown in Figure 6, which shows three possible states (A, B, C) for a block, such as an 8 x 8 pixel coding unit (CU). For each state, the upper block represents the video, i.e. the coding unit, whereas the lower block, i.e. a reconstruction information CU, represents the information
indicative about reconstruction of said pixels in said blocks. White values in the reconstruction information CU indicate that the corresponding pixel is visible/reconstructed at the decoder and dark values indicate that the corresponding pixel is not visible/reconstructed at the decoder.
[0064] The three states may be summarized as follows:
A. All pixels (pel, sample-values) in the CU will be reconstructed/visible at the decoder;
B. None of the pixels in the CU will be reconstructed/visible at the decoder;
C. Some (a subset) of the pixels in the CU will be reconstructed/visible at the decoder.
[0065] Based on these states, the encoder adaptively guides the in-loop filtering process for each state, such as follows:
A. Visual distortion will be visible at the decoder, but CU content represents “typical” video content and in-loop filter are expected to perform optimal. No changes are made to the in-loop filtering process, and the RDO calculations may be kept original.
B. Visual distortion will not be visible at the decoder, thus no changes to the in-loop filtering process are required. The RDO calculations may be optimized to favor less reduced rate over increased distortion.
C. A subset of the pixels in the CU will be reconstructed/visible at the decoder, whereupon the visual distortion will be visible. The boundaries introduced due to occupancy information within the CU change the video content characteristic, as a result, in-loop filter(s) fail to perform optimally, and they are therefore disabled.
[0066] According to an embodiment, the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following:
- an occupancy map video;
- an occupancy signal;
- an alpha map video;
- an alpha channel.
[0067] Herein, the occupancy map video addresses the V3C V-PCC use case by indicating which pixel positions will be reconstructed / visible at the decoder. The
occupancy signal, which may be multiplexed in the same or yet another video, addresses the V3C MIV use case by indicating which pixel positions will be reconstructed / visible at the decoder, wherein, for example, values below a certain offset are considered unoccupied. The alpha map video and alpha channel correspondingly addresses the Alpha use case.
[0068] According to an embodiment, the video signal comprises one or more of the following: individual video streams; separate layers for multi-layer coding; separate sublayers; separate pictures in a temporally interleaved manner in the same coded layer video sequence; separate constituent frames in spatially frame-packed video.
[0069] According to an embodiment, the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
[0070] Thus, the size of the block to be processed is irrespective for the implementation. [0071 ] According to an embodiment, the method comprises excluding samples corresponding pixels not becoming visible upon decoding from a filter parameter derivation phase.
[0072] Accordingly, when the visual distortion will not be visible at the decoder side, i.e., case B in above, the corresponding samples in the encoder side will be excluded from the filter parameter derivation phase. For example, the adaptive loop filter (ALF) training in the encoder side may use only samples that will be visible at the decoder side. In another examples, other filters such as luma mapping with chroma scaling (LMCS) and sample adaptive offset filter (SAO) may follow the same principles.
[0073] According to an embodiment, the method comprises including only samples corresponding to pixels in a boundary area of pixels becoming visible and non- visible upon decoding in a filter parameter derivation phase.
[0074] Thus, among the samples whose visual distortion will not be visible at decoder side, only those samples that are in the boundary areas of the visible and non-visible
samples for the decoder may be used in the encoder side for training the in-loop filters, such as ALF and SAO. An example of such boundary area is demonstrated in state C of Figure 6. Including such samples may guide the in-loop filters by classifying those boundary areas more accurately and hence the resulting filtering process may perform more optimally.
[0075] According to an embodiment, the method comprises adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels becoming visible and non-visible upon decoding.
[0076] When the filtering is applied to blocks containing both visible and non-visible samples at the decoder (e.g., state C of Figure 6), the filter strength may be adjusted in such a way that it has weaker impact to the samples in the boundary areas of the two sample types. For example, a weaker deblocking filter may be considered in such areas. As a result, the weaker filtering may still reduce the coding artifacts while preserving the object edges in the boundary regions of the two sample types.
[0077] According to an embodiment, the method comprises calculating a dedicated class of filters for a block comprising pixels becoming both visible and non-visible upon decoding.
[0078] Consequently, a dedicated class of filters may be calculated for the blocks that include both visible and non-visible samples for the decoder (e.g., state C of Figure 6). For example, the ALF filter may contain one or more classes of filters trained for such cases. In the decoder side, the available occupancy information may be used for indicating the filter class based on the content.
[0079] According to an embodiment, the method comprises disabling all in-loop filters for a block comprising pixels becoming both visible and non-visible upon decoding.
[0080] Thus, the encoder may disable all in-loop filter for blocks, such as CUs, containing partially reconstructed/visible content.
[0081] According to an embodiment, the method comprises using only a subset of inloop filters for a block comprising pixels becoming both visible and non-visible upon decoding.
[0082] Hence, the encoder may only use a subset of in-loop filters for blocks, such as CUs, containing partially reconstructed/visible content. For example, the encoder may only use LMCS and SAO, while disabling DBF and ALF.
[0083] According to an embodiment, the method comprises applying at least one inloop filter for only pixels of a block becoming visible upon decoding.
[0084] As an alternative, the encoder does not disable in-loop filters, however, the filter is only applied on pixel marked as visible/reconstructed at the decoder, while other pixels retain their original value.
[0085] According to an embodiment, the method comprises applying a combination of in-loop filters block- wise and/or pixel- wise.
[0086] Thus, the encoder may apply or disable one or more in-loop filters on a block level, and at the same time, apply or disable one or more in-loop filters on a pixel level. For example, DBF may be disabled per-CU, whereas ALF is applied per-pixel.
[0087] According to an embodiment, the method comprises fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
[0088] Accordingly, for the blocks that contain only the non-visible sample types (e.g., state B of Figure 6) at the decoder, the one or more of the in-loop filters may be fully or partially deactivated.
[0089] According to an embodiment, the method comprises skipping a signaling of filter information for said fully or partially deactivated one or more in-loop filters.
[0090] Thus, signaling of the filter information for the one or more in-loop filters of the underlying codec may be also skipped and a pre-defined values may be assigned for the inloop filters. The filter information, as such, may comprise one or more of the following: filter activation, filter type, filter class and index, etc.
[0091 ] According to an embodiment, the method comprises signaling one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
[0092] Hence, further aspects of the invention relate to signaling of reconstruction guided in-loop filtering, which may be performed, for example, in or along the bitstream of
the video signal. In the following, various example embodiments for carrying the signaling are given.
[0093] According to an embodiment, the method comprises carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
[0094] Examples of the signaling are given below:
[0095] In the examples above, a new syntax element vps/sps_reconstruction_guided_inloop_present_flag is introduced, wherein the flag value equal to 0 specifies that reconstruction guided in-loop filtering is not activated, and the flag value equal to 1 specifies that reconstruction guided in-loop filtering is activated.
[0096] According to an alternative embodiment, the method comprises signaling one or more parameters about applicability of the in-loop filters as a supplemental enhancement information (SEI) message.
[0097] According to an embodiment, said signaling comprises a reference to a reconstruction guided in-loop filtering as part of a multi-layer encoding structure.
[0098] An example of such signaling is given below:
[0099] Herein, a new syntax element vps_reconstruction_layer_id specifies the nuh layer id value of the layer carrying the reconstruction information relevant for reconstruction guided in-loop filter. [0100] According to an embodiment, the signaling comprises indicating a layer using the reconstruction guided in-loop filtering per a dependent layer.
[0101] An example of such signaling is given below:
[0102] Herein, a new syntax element vps_reconstruction_guided_inloop_layer_flag[ i ] having a value equal to 1 specifies that the i-th layer uses reconstruction guided in-loop filtering, vps reconstruction guided inloop layer _flag[ i ] equal to 0 specifies that the i-th layer does not use reconstruction guided in-loop filtering. A new syntax element vps_reconstruction_layer_idx[ i ] specifies that the layer index for the reconstuction information for the i-th layer is ReferenceLayerIdx[ i ] [ vps_reconstruction_layer_idx[ i ] ] . [0103] According to an embodiment, the signaling comprises indicating an activation of one or more in-loop filters individually.
[0104] An example of such signaling is given below:
[0105] Herein, new syntax elements are introduced:
[0106] vps_reconstruction_lmcs_flag equal to 0 specifies that LMCS in-loop filter is not applied if partial reconstruction information is available, vps reconstruction lmcs flag equal to 1 specifies that LMCS in-loop filter is applied if partial reconstruction information is available.
[0107] vps reconstruction dbf flag equal to 0 specifies that DBF in-loop filter is not applied if partial reconstruction information is available, vps reconstruction dbf flag equal to 1 specifies that DBF in-loop filter is applied if partial reconstruction information is available.
[0108] vps_reconstruction_sao_flag equal to 0 specifies that SAG in-loop filter is not applied if partial reconstruction information is available, vps reconstruction sao flag equal to 1 specifies that SAG in-loop filter is applied if partial reconstruction information is available.
[0109] vps reconstruction alf flag equal to 0 specifies that ALF in-loop filter is not applied if partial reconstruction information is available, vps reconstruction alf flag equal to 1 specifies that ALF in-loop filter is applied if partial reconstruction information is available.
[0110] According to an embodiment, the signaling comprises both a flag indicating activation of reconstruction guided in-loop filtering and indication for an activation of one or more in-loop filters individually.
[01 11] An example of such signaling is given below:
[01 12] Herein, the flag values are indicated by two bits, thereby enabling more versatile signaling. New syntax elements are introduced as follows:
[0113] vps_reconstruction_lmcs_flag equal to 0 specifies that LMCS in-loop filter is not applied if partial reconstruction information is available, vps reconstruction lmcs flag equal to 1 specifies that LMCS in-loop filter is applied if partial reconstruction information is available per CU. vps_reconstruction_lmcs_flag equal to 2 specifies that LMCS in-loop filter is applied only to sample with reconstruction information available.
[0114] vps reconstruction dbf flag equal to 0 specifies that DBF in-loop filter is not applied if partial reconstruction information is available, vps reconstruction dbf flag equal to 1 specifies that DBF in-loop filter is applied if partial reconstruction information is available per CU. vps reconstruction lmcs flag equal to 2 specifies that DBF in-loop filter is applied only to sample with reconstruction information available.
[0115] vps_reconstruction_sao_flag equal to 0 specifies that SAG in-loop filter is not applied if partial reconstruction information is available, vps reconstruction sao flag equal to 1 specifies that SAG in-loop filter is applied if partial reconstruction information is available per CU. vps_reconstruction_lmcs_flag equal to 2 specifies that SAG in-loop filter is applied only to sample with reconstruction information available.
[0116] vps reconstruction alf flag equal to 0 specifies that ALF in-loop filter is not applied if partial reconstruction information is available, vps reconstruction alf flag equal to 1 specifies that ALF in-loop filter is applied if partial reconstruction information is available per CU. vps reconstruction lmcs flag equal to 2 specifies that ALF in-loop filter is applied only to sample with reconstruction information available.
[0117] In another embodiment, one or more threshold syntax elements that control the use of the reconstruction information for in-loop filtering are included by an encoder in or along a bitstream, such as in a VPS, and/or decoded by a decoder from or along a bitstream, such as from a VPS. A threshold syntax element may for example define the
sample value range in the reconstruction information that indicates a non- visible pixel for in-loop filtering.
[0118] The following syntax table provides an example, where numReconstructionlnformationLayers is a variable indicating the number of layers that contain reconsruction information to be used for guided in-loop filtering:
[0119] Herein, the syntax element vps_reconstruction_luma_threshold[ i ] specifies that samples with luma sample value in the range of 0 to vps_reconstruction_luma_threshold[ i ], inclusive, indicate non-visible samples for in-loop filtering.
[0120] Another aspect relates to the operation of a decoder (or a Tenderer/ receiver/player/ client). The method, which is disclosed in Figure 7 as illustrating the operation of a decoder, comprises receiving (700) a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receiving (702), in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; decoding (704) the an encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
[0121] Hence, the decoding apparatus receives a bitstream, containing reconstruction guided in-loop filter signalling as disclosed above, and comprising encoded video data, and in or along the bitstream the reconstruction information. The reconstruction information may be provided in the form of a bitstream comprising an encoded occupancy (V3C use case) or alpha map video, either as individual videos (V3C V-PCC use case) or as layer in a multi-layer encoding, or a reconstruction signal, multiplexed in the same or yet another
video (V3C MIV use case). If necessary, the decoder first decodes the reconstruction information, then the decoder decodes the video data, activating/disabling in-loop filters according to said information indicative about reconstruction or visibility of said pixels in said blocks, i.e. reconstruction information state (full, none, partial) and the received signaling.
[0122] According to an embodiment, any of the embodiments disclosed above, relating to either encoding or decoding, may be applied similarly to a post-processing filter. For a neural-network post-processing filter, an input tensor may be formed. In an embodiment, it may be indicated by an encoder, e.g. in a neural-network post-filter characteristics (NNPFC) SEI message, that the post-filter expects auxiliary input in the input tensor for the reconstruction information and/or variables controlling the reconstruction information guided filtering, such as one or more threshold values defining the sample value range in the reconstruction information that indicates a non-visible pixel for filtering. In an embodiment, it may be decoded by a decoder, e.g. from a NNPFC SEI message, that the post-filter expects auxiliary input in the input tensor for the reconstruction information and/or variables controlling the reconstruction information guided filtering. Accordingly, the decoder forms an input tensor with reconstruction information and/or variables controlling the reconstruction information guided filtering.
[0123] The embodiments relating to the sending/encoding aspects may be implemented in an apparatus comprising: means for obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and means for disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0124] According to an embodiment, the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following: an occupancy map video; an occupancy signal; an alpha map video; an alpha channel.
[0125] According to an embodiment, the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
[0126] According to an embodiment, the apparatus comprises means for excluding samples corresponding pixels not becoming reconstructed or visible upon decoding from a filter parameter derivation phase.
[0127] According to an embodiment, the apparatus comprises means for including only samples corresponding to pixels in a boundary area of pixels being reconstructed or becoming visible and non-visible upon decoding in a filter parameter derivation phase.
[0128] According to an embodiment, the apparatus comprises means for adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non-visible upon decoding.
[0129] According to an embodiment, the apparatus comprises means for calculating a dedicated class of filters for a block comprising pixels becoming both visible and non- visible upon decoding.
[0130] According to an embodiment, the apparatus comprises means for using only a subset of in-loop filters for a block comprising pixels becoming both visible.
[0131] According to an embodiment, the apparatus comprises means for applying a combination of in-loop filters block-wise and/or pixel-wise.
[0132] According to an embodiment, the apparatus comprises means for fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
[0133] According to an embodiment, the apparatus comprises means for signaling one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
[0134] According to an embodiment, the apparatus comprises means for carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
[0135] The embodiments relating to the sending/ encoding aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory,
said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtain information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disable, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
[0136] According to an embodiment, the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following: an occupancy map video; an occupancy signal; an alpha map video; an alpha channel.
[0137] According to an embodiment, the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a subpicture.
[0138] According to an embodiment, the apparatus comprises code configured to cause the apparatus to exclude samples corresponding pixels not becoming reconstructed or visible upon decoding from a filter parameter derivation phase.
[0139] According to an embodiment, the apparatus comprises code configured to cause the apparatus to include only samples corresponding to pixels in a boundary area of pixels being reconstructed or becoming visible and non-visible upon decoding in a filter parameter derivation phase.
[0140] According to an embodiment, the apparatus comprises code configured to cause the apparatus to adjust a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non-visible upon decoding.
[0141] According to an embodiment, the apparatus comprises code configured to cause the apparatus to calculate a dedicated class of filters for a block comprising pixels becoming both visible and non-visible upon decoding.
[0142] According to an embodiment, the apparatus comprises code configured to cause the apparatus to use only a subset of in-loop filters for a block comprising pixels becoming both visible.
[0143] According to an embodiment, the apparatus comprises code configured to cause the apparatus to apply a combination of in-loop filters block-wise and/or pixel-wise.
[01 4] According to an embodiment, the apparatus comprises code configured to cause the apparatus to fully or partially deactivate one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
[0145] According to an embodiment, the apparatus comprises code configured to cause the apparatus to signal one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
[0146] According to an embodiment, the apparatus comprises code configured to cause the apparatus to carry out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided in-loop filtering.
[0147] The receiving/decoding/rendering aspects may be implemented by an apparatus comprising means for receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and means for decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
[0148] The embodiments relating to the receiving/decoding/rendering aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receive, in or along said bitstream, information indicative about reconstruction or visibility
of said pixels in said blocks; and decode the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
[0149] Such apparatuses may comprise e.g. the functional units disclosed in any of the Figures la, lb, 2, 3a and 3b for implementing the embodiments.
[0150] In the above, some embodiments have been described with reference to encoding. It needs to be understood that said encoding may comprise one or more of the following: encoding source image data into a bitstream, encapsulating the encoded bitstream in a container file and/or in packet(s) or stream(s) of a communication protocol, and announcing or describing the bitstream in a content description, such as the Media Presentation Description (MPD) of ISO/IEC 23009-1 (known as MPEG-DASH) or the IETF Session Description Protocol (SDP). Similarly, some embodiments have been described with reference to decoding. It needs to be understood that said decoding may comprise one or more of the following: decoding image data from a bitstream, decapsulating the bitstream from a container file and/or from packet(s) or stream(s) of a communication protocol, and parsing a content description of the bitstream, [0151] In the above, where the example embodiments have been described with reference to an encoder or an encoding method, it needs to be understood that the resulting bitstream and the decoder or the decoding method may have corresponding elements in them. Likewise, where the example embodiments have been described with reference to a decoder, it needs to be understood that the encoder may have structure and/or computer program for generating the bitstream to be decoded by the decoder.
[0152] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits or any combination thereof. While various aspects of the invention may be illustrated and described as block diagrams or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0153] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0154] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0155] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended examples. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention.
Claims
1. An apparatus comprising: means for obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and means for disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
2. The apparatus according to claim 1, wherein the information indicative about reconstruction or visibility of said pixels in said blocks upon decoding is one or more of the following:
- an occupancy map video;
- an occupancy signal;
- an alpha map video;
- an alpha channel.
3. The apparatus according to claim 1 or 2, wherein the block is one of a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice or a sub-picture.
4. The apparatus according to any preceding claim, comprising: means for excluding samples corresponding pixels not becoming reconstructed or visible upon decoding from a filter parameter derivation phase.
5. The apparatus according to any of claims 1 - 3, comprising: means for including only samples corresponding to pixels in a boundary area of pixels being reconstructed or becoming visible and non-visible upon decoding in a filter parameter derivation phase.
6. The apparatus according to claim 5, comprising means for adjusting a filter strength for said samples corresponding to pixels in said boundary area of the pixels being reconstructed or becoming visible and non- visible upon decoding.
7. The apparatus according to any preceding claim, comprising means for calculating a dedicated class of filters for a block comprising pixels becoming both visible and non- visible upon decoding.
8. The apparatus according to any preceding claim, comprising: means for using only a subset of in-loop filters for a block comprising pixels becoming both visible.
9. The apparatus according to any preceding claim, comprising: means for applying a combination of in-loop filters block-wise and/or pixel-wise.
10. The apparatus according to any preceding claim, comprising means for fully or partially deactivating one or more in-loop filters for a block comprising only pixels not becoming visible upon decoding.
11. The apparatus according to any preceding claim, comprising means for signaling one or more parameters about applicability of the in-loop filters in or along a bitstream comprising said video signal.
12. The apparatus according to claim 11, comprising means for carrying out said signalling in a video parameter set raw byte sequence payload (RBSP) syntax or a sequence parameter set raw byte sequence payload (RBSP) syntax, wherein said signaling comprises a flag indicating activation of reconstruction guided inloop filtering.
13. A method comprising: obtaining a video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; obtaining information indicative about reconstruction or visibility of said pixels in said blocks upon decoding; and disabling, based on said information indicative about reconstruction or visibility of said pixels in said blocks, in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible upon decoding.
14. An apparatus comprising: means for receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; means for receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and means for decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling in-loop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
15. A method comprising: receiving a bitstream comprising an encoded video signal comprising blocks of video frames, said blocks comprising a plurality of pixels; receiving, in or along said bitstream, information indicative about reconstruction or visibility of said pixels in said blocks; and decoding the encoded video signal according to at least said information indicative about reconstruction or visibility of said pixels in said blocks by disabling inloop filtering for such blocks having only a subset of the pixels to be reconstructed or visible.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FI20235408 | 2023-04-11 | ||
| PCT/FI2024/050102 WO2024213824A1 (en) | 2023-04-11 | 2024-03-11 | An apparatus, a method and a computer program for video |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4696018A1 true EP4696018A1 (en) | 2026-02-18 |
Family
ID=93058855
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24788298.8A Pending EP4696018A1 (en) | 2023-04-11 | 2024-03-11 | An apparatus, a method and a computer program for video |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4696018A1 (en) |
| CN (1) | CN120898428A (en) |
| WO (1) | WO2024213824A1 (en) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11017566B1 (en) * | 2018-07-02 | 2021-05-25 | Apple Inc. | Point cloud compression with adaptive filtering |
| CN114830662B (en) * | 2019-12-27 | 2023-04-14 | 阿里巴巴(中国)有限公司 | Method and system for performing progressive decode refresh processing on images |
| EP4268457A1 (en) * | 2020-12-28 | 2023-11-01 | Koninklijke KPN N.V. | Partial output of a decoded picture buffer in video coding |
-
2024
- 2024-03-11 CN CN202480024440.2A patent/CN120898428A/en active Pending
- 2024-03-11 WO PCT/FI2024/050102 patent/WO2024213824A1/en not_active Ceased
- 2024-03-11 EP EP24788298.8A patent/EP4696018A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024213824A1 (en) | 2024-10-17 |
| CN120898428A (en) | 2025-11-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250008126A1 (en) | Cross-component quantization in video coding | |
| JP2022179505A (en) | Video decoding method and video decoder | |
| KR20210095958A (en) | Improved flexible tiling in video coding | |
| TWI856336B (en) | Methods for applying film grain, extracting and merging a bitstream, method performed by an encoder, and related computer program, carrier containing the same, and computing apparatus | |
| JP7684472B2 (en) | Image or video coding based on NAL unit related information - Patents.com | |
| KR20150128915A (en) | Device and method for scalable coding of video information | |
| AU2020352513A1 (en) | Method for signaling output layer set with sub picture | |
| JP7684470B2 (en) | VIDEO CODING APPARATUS AND METHOD FOR CONTROLLING LOOP FILTERING - Patent application | |
| CN114762339B (en) | Image or video encoding based on high-level syntax elements related to transform skipping and palette encoding | |
| EP4364415A1 (en) | Applying an overlay process to a picture | |
| JP2024100892A (en) | Filtering-based image coding apparatus and method | |
| US12047604B2 (en) | Apparatus, a method and a computer program for volumetric video | |
| US20240357181A1 (en) | Transmission of volumetric images in multiplane imaging format | |
| US12069314B2 (en) | Apparatus, a method and a computer program for volumetric video | |
| KR20220097997A (en) | Video coding apparatus and method for controlling loop filtering | |
| KR102918038B1 (en) | A method for coding images based on information related to tiles and information related to slices in a video or image coding system. | |
| KR20220097996A (en) | Signaling-based video coding apparatus and method of information for filtering | |
| KR20220100700A (en) | Subpicture-based video coding apparatus and method | |
| WO2024213824A1 (en) | An apparatus, a method and a computer program for video | |
| WO2024079383A1 (en) | An apparatus, a method and a computer program for volumetric video | |
| KR20220082081A (en) | Method and apparatus for signaling picture segmentation information | |
| KR20220100702A (en) | Picture segmentation-based video coding apparatus and method | |
| KR20220083818A (en) | Method and apparatus for signaling slice-related information | |
| KR102925116B1 (en) | Signaling-based image or video coding of recovery point-related information for GDR | |
| WO2025108607A1 (en) | An apparatus, a method and a computer program for video coding and decoding |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251111 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |