WO2014106641A1 - Video coding and decoding methods and corresponding devices including flag for enhancement layer - Google Patents

Video coding and decoding methods and corresponding devices including flag for enhancement layer Download PDF

Info

Publication number
WO2014106641A1
WO2014106641A1 PCT/EP2014/050040 EP2014050040W WO2014106641A1 WO 2014106641 A1 WO2014106641 A1 WO 2014106641A1 EP 2014050040 W EP2014050040 W EP 2014050040W WO 2014106641 A1 WO2014106641 A1 WO 2014106641A1
Authority
WO
WIPO (PCT)
Prior art keywords
block
base layer
layer block
enhancement layer
enhancement
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2014/050040
Other languages
French (fr)
Inventor
Patrick Lopez
Philippe Bordes
Pierre Andrivon
Philippe Salmon
Franck Hiron
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Thomson Licensing SAS
Original Assignee
Thomson Licensing SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Thomson Licensing SAS filed Critical Thomson Licensing SAS
Publication of WO2014106641A1 publication Critical patent/WO2014106641A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/30Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using hierarchical techniques, e.g. scalability
    • H04N19/33Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using hierarchical techniques, e.g. scalability in the spatial domain
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/46Embedding additional information in the video signal during the compression process
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • H04N19/105Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/59Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial sub-sampling or interpolation, e.g. alteration of picture size or resolution

Definitions

  • the invention relates to a decoding method for decoding a scalable bitstream comprising one base layer and at least one enhancement layer.
  • the invention further relates to corresponding coding method, coding and decoding devices.
  • the scalable bitstream usually comprises a base layer and at least one enhancement layer.
  • a low resolution video content is obtained.
  • a video content of higher resolution is obtained.
  • some coding information e.g. motion vectors
  • an enhancement layer block uses the motion vector of a spatially corresponding base layer block as a motion vector predictor.
  • the motion vector of the enhancement layer block is either equal to the motion vector predictor or is equal to the motion vector predictor to which a residual vector is added.
  • the residual vector is encoded in the scalable bitstream (respectively decoded from the scalable bitstream).
  • a predictor P is determined.
  • the reference block is identified in a reference picture Iref by the motion vector MV.
  • the predictor P is then subtracted from the enhancement layer block.
  • the residue thus obtained is encoded in the scalable bitstream.
  • the residue is decoded from the scalable bitstream and added to the predictor in order to reconstruct the enhancement layer block.
  • the purpose of the invention is to overcome at least one of the disadvantages of the prior art.
  • a method for decoding a scalable bitstream comprising a base layer and at least one enhancement layer.
  • the method comprises for an enhancement layer block, decoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of the enhancement layer block and reconstructing the enhancement layer block depending on the flag, wherein the flag is decoded only when the enhancement block has a single spatially corresponding base layer block.
  • reconstructing the enhancement layer block depending on the flag comprises:
  • determining the predictor comprises adding a part of a scaled residue of the spatially corresponding base layer block co-located to the enhancement layer block to the motion compensated reference block when the flag indicates that a residue of a spatially corresponding base layer block is used for the reconstruction.
  • the motion vector is a motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
  • the flag is decoded only when the enhancement block has a single spatially corresponding base layer block and when the motion vector is the motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
  • a method for encoding picture blocks in a scalable bitstream comprising a base layer and at least one enhancement layer comprises for an enhancement layer block, encoding the enhancement layer block and encoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding of the enhancement layer block, wherein the flag is encoded only when the enhancement block has a single spatially corresponding base layer block.
  • encoding the enhancement layer block comprises:
  • determining the predictor comprises adding a part of a scaled residue of the spatially corresponding base layer block co-located to the enhancement layer block to the motion compensated reference block when the flag indicates that a residue of a spatially corresponding base layer block is used for the encoding.
  • the motion vector is a motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
  • the flag is encoded only when the enhancement block has a single spatially corresponding base layer block and when the motion vector is the motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
  • a scalable video decoder for decoding a scalable bitstream comprising a base layer and at least one enhancement layer comprises means for decoding, for an enhancement layer block, a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of the enhancement layer block and means for reconstructing the enhancement layer block depending on the flag, wherein the flag is decoded only when the enhancement block has a single spatially corresponding base layer block.
  • a scalable video encoder for encoding picture blocks in a scalable bitstream comprising a base layer and at least one enhancement layer is further disclosed.
  • the encoder comprises means for encoding an enhancement layer block and means for encoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding of the enhancement layer block, wherein the flag is encoded only when the enhancement block has a single spatially corresponding base layer block.
  • a scalable video bitstream comprising a base layer and at least one enhancement layer is disclosed.
  • the bitstream encodes for an enhancement layer block a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding or reconstruction of the enhancement layer block, wherein the flag is encoded only when the enhancement block has a single spatially corresponding base layer block.
  • a computer program product comprising program code instructions to execute of the steps of the decoding method according to the invention when this program is executed on a computer is also disclosed.
  • - Figure 1 represents a quadtree structure as defined in HEVC coding standard
  • - Figure 2 represents enhancement layer blocks, e.g. A and B, and corresponding base layer blocks A', B' and C;
  • FIG. 3 represents the flowchart of a decoding method of a scalable bitstream F with a view to the reconstruction of an enhancement layer block Be according to the invention
  • FIG. 4 depicts in detail a step of the decoding method according to the invention.
  • FIG. 5 represents the flowchart of a coding method according to the invention.
  • FIG. 6 depicts in detail a step of the coding method according to the invention.
  • FIG. 7 depicts a scalable video decoder according to the invention.
  • FIG. 8 depicts a scalable video encoder according to the invention.
  • the invention relates to a method for reconstructing a block of pixels of a sequence of pictures and a method for coding such a block.
  • a picture sequence is a series of several pictures. Each picture comprises pixels or picture points with each of which at least one item of picture data is associated.
  • An item of picture data is for example an item of luminance data or an item of chrominance data.
  • the coding and decoding methods are described with reference to a picture block. It is clear that these methods can be applied on several picture blocks of a picture and on several pictures of a sequence with a view to the coding respectively the decoding of one or more pictures.
  • a picture block is a set of pixels of any form. It can be a square, a rectangle. But the invention is not limited to such forms. In the following section the word block is used for picture block.
  • the "predictor” term designates data used to predict other data.
  • a predictor is used to predict a picture block.
  • a predictor or prediction block is obtained from a block or several blocks of the same picture as the picture to which belongs the block that it predicts (spatial prediction or intra-picture prediction) or from one (mono-directional prediction) or several reference blocks (bi-directional prediction or bi-predicted) of a different picture (temporal prediction or inter- picture prediction) of the picture to which the block that it predicts belongs.
  • Predictors are also used to predict motion vectors.
  • the predictor is a motion vector predictor.
  • a motion vector predictor is for example a motion vector associated with a block located in the causal neighborhood of a current block.
  • the motion vector predictor for an enhancement layer block can be the motion vector associated with a spatially corresponding base layer block after scaling, e.g. the motion vector associated with a base layer block covering the central position of the enhancement layer block.
  • a "causal neighbourhood" of a current block designates a neighbourhood of this block that comprises pixels coded/reconstructed before the current block.
  • the term "residue” signifies data obtained after extraction of other data.
  • the extraction is generally a subtraction pixel by pixel of prediction data from source data.
  • the extraction is more general and notably comprises a weighted subtraction in order for example to account for an illumination variation model.
  • a residual motion vector AMV is obtained by subtracting a motion vector predictor MVP from a motion vector MVc.
  • reconstruction designates data (e.g. pixels, blocks) obtained after merging residues with a predictor.
  • the merging is generally a sum of prediction data with residues. However, the merging is more general and notably comprises a weighted sum in order for example to account for an illumination variation model.
  • a reconstructed block is a block of reconstructed pixels.
  • the term coding is to be taken in the widest sense.
  • the coding can possibly comprise the transformation and/or the quantization of picture data. It can also designate only the entropy coding.
  • aspects of the present principles can be embodied as a system, method or computer readable medium. Accordingly, aspects of the present principles can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, and so forth), or an embodiment combining software and hardware_aspects that can all generally be referred to herein as a "circuit,” “module”, or “system.” Furthermore, aspects of the present principles can take the form of a computer readable storage medium. Any combination of one or more computer readable storage medium(s) may be utilized.
  • each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s).
  • the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, or blocks may be executed in an alternative order, depending upon the functionality involved.
  • HEVC video coding standard I SO/I EC ISO/I EC 23008-2 MPEG-H Part 2 defines a quadtree structure.
  • a picture is divided into coding units CU.
  • Each coding unit is divided into prediction units (PU) and each PU is further divided into transform units (TU).
  • Each TU is encoded by applying an appropriate transform.
  • the CU is divided into 4 PUs.
  • a PU can then be encoded by applying a single transform. In this case a single TU corresponds to the PU.
  • a PU can be encoded by dividing it horizontally or vertically into two TUs and applying a transform on each of them.
  • a PU can also be encoded by dividing it in 4 TUs and applying a transform on each of them.
  • an enhancement layer block has spatially corresponding base layer blocks.
  • the spatially corresponding base layer blocks are the base layer blocks that cover the enhancement layer blocks when the scaled base layer picture is superposed on the enhancement layer picture.
  • the scaled base layer picture also called upsampled base layer picture is obtained by scaling/upsampling the base layer picture according to the spatial resolution ratio between the two layers.
  • Figure 2 represents enhancement layer blocks, e.g. A and B, on the top right and base layer blocks A', B', C on the top left.
  • the scaled base layer blocks are represented with thick lines on the bottom left part of the figure.
  • the scaled base layer blocks are superposed on the enhancement layer blocks.
  • To the enhancement layer block A corresponds a single base layer block A' while two base layer blocks B' and C correspond to the enhancement layer block B.
  • the enhancement layer block A is totally covered by the single scaled base layer block A'.
  • the top part of enhancement layer block B is covered by the scaled base layer block B' and the bottom part is covered by the scaled base layer block C.
  • Figure 3 represents the flowchart of a decoding method of a scalable bitstream F with a view to the reconstruction of an enhancement layer block Be according to the invention.
  • a step E10 it is checked whether the enhancement layer block Be has a single spatially corresponding base layer block or a plurality of them. If the enhancement layer block Be has a single spatially corresponding base layer block, then the method continues to step E12, otherwise it goes to step E16.
  • a flag is decoded for the enhancement layer block Be.
  • the flag indicates whether a spatially corresponding base layer block residue is used for the reconstruction of the enhancement layer block Be, i.e. in the prediction process.
  • This flag is only decoded for the enhancement layer block Be when the block Be has a single spatially corresponding base layer block as block A on figure 2. Consequently no such flag is decoded for an enhancement layer block which has more than one spatially corresponding base layer block as block B on figure 2.
  • less bits are used since the flag is not decoded systematically for each block of the enhancement layer.
  • the enhancement layer block is reconstructed depending on the value of the flag.
  • the flag takes the value ⁇ ' then the residue of the spatially corresponding base layer block is not used for the reconstruction of the enhancement layer block A and when the flag takes the value then the residue of the spatially corresponding base layer block is used for the reconstruction of the enhancement layer block A.
  • the flag takes the value then the residue of the spatially corresponding base layer block is not used for the reconstruction of the enhancement layer block A and when the flag takes the value ⁇ ' then the residue of the spatially corresponding base layer block is used for the reconstruction of the enhancement layer block A.
  • the enhancement layer block is reconstructed without using the residue of spatially corresponding base layer blocks.
  • the step E16 comprises decoding a residue for the block Be from the scalable bitstream and adding the residue to a predictor P.
  • the predictor is usually a motion compensated reference block.
  • the motion compensated reference block is determined in a reference picture from a motion vector decoded for the block Be.
  • a plurality of base layer blocks corresponds spatially to the enhancement layer block as for block B on figure 2.
  • the use of residues of at least two different base layer blocks for the reconstruction of the enhancement layer block is probably not efficient. Indeed, when the residues of the two different base layer blocks are obtained with different motion vectors using them for the prediction of an enhancement layer block whose motion vector is also probably different is not appropriate.
  • Figure 4 depicts in detail the step E14 of the decoding method according to the invention.
  • a motion vector MVc is determined for the enhancement layer block Be.
  • the motion vector MVc is directly decoded from the enhancement layer.
  • the motion vector is directly derived from the single spatially corresponding base layer block. More precisely, the motion vector MVc is equal to the motion vector MV B L_U PSC of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • the MVc for the enhancement layer block A is for example the motion vector of the base layer block A' scaled according to the spatial resolution ratio between the two layers.
  • MV C MVBL_U P SC, where MV B L_U P SC is the scaled base layer motion vector.
  • the scaled motion vector is only used as a predictor and thus a residual motion vector AMV decoded from the scalable bitstream is added to the scaled base layer motion vector.
  • a motion compensated reference block MC(MVc, Iref) is determined from the motion vector MV C determined for the enhancement layer block In a step E140.
  • the motion compensated reference block is determined in a reference picture Iref.
  • step E144 the value of the flag is checked. If the flag takes the value ⁇ ' then the method continues to step E146, otherwise it goes to step E148.
  • a step E146 the predictor P of the enhancement layer block Be is set equal to the motion compensated reference block MC(MVc, Iref).
  • the predictor P of the enhancement layer block Be is set equal to the sum of the motion compensated reference block MC(MVc, Iref) and the co-located part ResBL of the scaled/upsampled residue of the spatially corresponding base layer block.
  • the co-located part is the part of the scaled/upsampled residue of the spatially corresponding base layer block A' that covers the enhancement layer block A.
  • the enhancement layer block Be is reconstructed by adding the predictor P to the residue, if any, decoded for the enhancement layer block Be from the scalable bitstream.
  • the decoding of the residue usually comprises entropy decoding, scaling (in the sense of inverse quantizing) and transform. According to a variant no such residue is decoded. In this case, the reconstructed block Be is equal to its predictor P.
  • the enhancement layer block is reconstructed according to a skip mode.
  • the flag is only decoded from the scalable bitstream for the blocks Be having a single spatially corresponding base layer block and whose motion vector predictor MVP is the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • This variant encompasses the specific embodiment wherein the motion vector of the enhancement layer block is equal to the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • FIG. 5 represents the flowchart of a coding method according to the invention.
  • a step E50 it is checked whether the enhancement layer block Be has a single spatially corresponding base layer block or a plurality of them. If the enhancement layer block Be has a single spatially corresponding base layer block, then the method continues to step E52, otherwise it goes to step E56.
  • the enhancement layer block Be is encoded.
  • the way the enhancement layer block Be is encoded depends on a coding mode selected according to a selection criterion. Indeed, on the encoder side several coding modes are checked. The best mode from a rate distortion point of view is usually selected.
  • the enhancement layer block Be is encoded with or without using the residue of the spatially corresponding base layer block for the prediction of block Be depending on the selected coding mode.
  • the invention is not limited by the way the coding mode is selected. One way to select a coding mode is to check all the coding modes and to select the one that reaches the best rate-distortion compromise.
  • a flag is encoded for the enhancement layer block Be.
  • the flag indicates whether a spatially corresponding base layer block residue is used for the encoding (resp. reconstruction on the decoder side) of the enhancement layer block Be.
  • This flag is only encoded for the enhancement layer block Be when the block Be has a single spatially corresponding base layer block as block A depicted on figure 2. Consequently no such flag is encoded for an enhancement layer block which has more than one spatially corresponding base layer block as block B on figure 2.
  • less bits are used since the flag is not encoded systematically for each block of the enhancement layer.
  • the enhancement layer block is encoded without using any base layer block residue.
  • the step E56 comprises encoding a residue for the block Be in the scalable bitstream F.
  • the residue is obtained by subtracting a predictor P from the block Be.
  • the predictor is usually a motion compensated reference block.
  • the motion compensated reference block is determined in a reference picture from a motion vector encoded for the block Be.
  • Figure 6 depicts detail the step E52 of the coding method according to the invention.
  • a motion vector MVc is determined for the enhancement layer block.
  • the motion vector is determined by motion estimation.
  • the motion vector MVc equals the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • the MVc for the enhancement layer block A is for example the motion vector of the base layer block A' scaled according to the spatial resolution ratio between the two layers.
  • the ratio is 2 and the scaled motion vectors components equal the components of the base layer motion vector multiplied by 2.
  • MV C MVBL_U P SC, where MVBL_U P SC is the scaled base layer motion vector.
  • a motion compensated reference block MC(MVc, Iref) is determined from the motion vector MV C determined for the enhancement layer block in the step E520.
  • the motion compensated reference block is determined in a reference picture Iref.
  • a step E524 it is checked whether the base layer block residue is used for the prediction of block Be. For a block Be having a single spatially corresponding base layer block, this choice is encoder dependent. The invention is not at all limited by the way the choice is made. If the encoding method chooses to use the residue of the spatially corresponding base layer block for the prediction of block Be then the method continues to step E526, otherwise it goes to step E528.
  • the predictor P of the enhancement layer block Be is set equal to the sum of the motion compensated reference block MC(MVc, Iref) and the co-located part ResBL of the scaled/upsampled residue of the spatially corresponding base layer block.
  • the co-located part is the part of the scaled/upsampled residue of the spatially corresponding base layer block A' that covers the enhancement layer block A.
  • the flag for the block Be is set equal to the value ⁇ ' if this value indicates that the base layer block residue is used for the encoding of the enhancement layer block Be.
  • the predictor P of the enhancement layer block Be is set equal to the motion compensated reference block MC(MVc, Iref). The flag for the block Be is thus step equal to the value ⁇ ' if this value indicates that the base layer block residue is not used for the encoding of the enhancement layer block Be.
  • the enhancement layer block Be is encoded from the predictor P.
  • a residue is possibly encoded for the block Be by subtracting from the block Be, the predictor P to get a residue.
  • the residue is usually transformed and quantized before being entropy coded. According to a variant no such residue is decoded.
  • a residual motion vector is also possibly encoded.
  • the residual motion vector is equal to (MVc-MVP), where MVP is a predictor.
  • the predictor is for example MV B L_upsc. According to a variant no such residual motion vector is encoded.
  • the enhancement layer block is encoded in a skip mode.
  • the flag is only encoded in the scalable bitstream for the blocks Be having a single spatially corresponding base layer block and whose motion vector predictor MVP is the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • This variant encompasses the specific embodiment wherein the motion vector of the enhancement layer block is equal to the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • a scalable video bitstream comprises a base layer and at least one enhancement layer.
  • the video bitstream comprises for each enhancement layer block Be having a single spatially corresponding base layer block a flag specifying whether the residue of said single spatially corresponding base layer block is used for the encoding/reconstruction of the block Be, more precisely for its prediction.
  • the bitstream comprises such a flag only for the enhancement layer block Be having a single spatially corresponding base layer block and whose motion vector predictor MVP is the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • This variant encompasses the specific embodiment wherein the motion vector of the enhancement layer block is equal to the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
  • the enhancement layer block is encoded in a skip mode.
  • idx_mv_baselayer is defined as the index in the merging candidate list corresponding to the motion vector candidate derived from the base layer colocated PU, i.e. the spatially corresponding base layer block. If such a candidate is not present in the merging candidate list, for instance if the base layer PU was coded in Intra mode, idx_mv_baselayer is set to -1 .
  • a new flag use_base_residue_flag is conditionnally read from the bitstream if mergejdx corresponds to idx_mv_baselayer and if cu_single_basis[ xO ][ yO ] is true.
  • cu_single_basis[ xO ][ yO ] indicates that the current enhancement layer PU is included in one single base-layer PU, i.e. has a single spatially corresponding base layer block. Depending on its value, the scaled/upsampled residue will be added to the predictor.
  • Figure 7 depicts a scalable video decoder DEC according to the invention.
  • the decoder DEC receives on a first input IN a scalable bitstream.
  • the input is connected to a demultiplexer DEMUX configured to demultiplex the bitstream into a base layer BL and at least one enhancement layer EL.
  • the decoder DEC comprises two inputs and receives on a first input the base layer BL and on the second input the enhancement layer EL.
  • the decoder DEC does not comprise DEMUX.
  • the DEMUX if present is connected to a base layer decoder 'BL decoder' adapted to decode base layer pictures.
  • the DEMUX if present is further connected to an enhancement layer decoder ⁇ decoder' adapted to decode enhancement layer pictures.
  • the enhancement layer decoder is adapted to implement the steps E10 to E16 of the decoding according to one of the various embodiments of the decoding method.
  • the decoded base layer pictures are then possibly transmitted on a first output of the decoder OUT1 and the decoded enhancement layer pictures are then transmitted on a second output of the decoder OUT2.
  • only decoded enhancement layer pictures are transmitted on the second output OUT2 for being displayed on a 2D display such as an UHDTV display.
  • the decoder DEC only comprises the second output of the decoder OUT2.
  • Figure 8 depicts a scalable video encoder according to the invention.
  • the encoder ENC receives on a first input IN1 base layer pictures and on a second input enhancement layer pictures.
  • the first input IN1 is connected to a base layer encoder adapted to encode and reconstruct the base layer pictures.
  • the second input IN2 is connected to an enhancement layer encoder adapted to implement the steps E50 to E56 and thus to encode the enhancement layer pictures according to one of the various embodiments of the encoding method.
  • the scalable video coders and decoders according to the invention are for example implemented in various forms of hardware, software, firmware, special purpose processors, or a combination thereof.
  • the present principles may be implemented as a combination of hardware and software.
  • the software is preferably implemented as an application program tangibly embodied on a program storage device.
  • the application program may be uploaded to, and executed by, a machine comprising any suitable architecture.
  • the machine is implemented on a computer platform having hardware such as one or more central processing units (CPU), a random access memory (RAM), and input/output (I/O) interface(s).
  • the computer platform also includes an operating system and microinstruction code.
  • the various processes and functions described herein may either be part of the microinstruction code or part of the application program (or a combination thereof) that is executed via the operating system.
  • various other peripheral devices may be connected to the computer platform such as an enhancement data storage device and a printing device.
  • the coding and decoding devices according to the invention are implemented according to a purely hardware realisation, for example in the form of a dedicated component (for example in an ASIC (Application Specific Integrated Circuit) or FPGA (Field-Programmable Gate Array) or VLSI (Very Large Scale Integration) or of several electronic components integrated into a device or even in a form of a mix of hardware elements and software elements.
  • a dedicated component for example in an ASIC (Application Specific Integrated Circuit) or FPGA (Field-Programmable Gate Array) or VLSI (Very Large Scale Integration) or of several electronic components integrated into a device or even in a form of a mix of hardware elements and software elements.
  • ASIC Application Specific Integrated Circuit
  • FPGA Field-Programmable Gate Array
  • VLSI Very Large Scale Integration

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

A method for decoding a scalable bitstream comprising a base layer and at least one enhancement layer is disclosed. The method comprises for an enhancement layer block, decoding (E10, E12) a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of the enhancement layer block and reconstructing (E14, E16) the enhancement layer block depending on the flag, wherein the flag is decoded (E10, E12) only when the enhancement block has a single spatially corresponding base layer block.

Description

VIDEO CODING AND DECODING METHODS AND CORRESPONDING DEVICES INCLUDING FLAG FOR ENHANCEMENT LAYER
1 . FIELD OF THE INVENTION
The invention relates to a decoding method for decoding a scalable bitstream comprising one base layer and at least one enhancement layer. The invention further relates to corresponding coding method, coding and decoding devices. 2. BACKGROUND OF THE INVENTION
Spatial scalability makes it possible to encode in a single bitstream, i.e. the scalable bitstream, a video content at multiple spatial resolutions. The scalable bitstream usually comprises a base layer and at least one enhancement layer. When decoding only the base layer a low resolution video content is obtained. When decoding the enhancement layer a video content of higher resolution is obtained. In order to increase coding efficiency, it is known to deduce some coding information (e.g. motion vectors) for the enhancement layer from coding information of a lower layer, e.g. the based layer. As an example, it is known for an enhancement layer block to use the motion vector of a spatially corresponding base layer block as a motion vector predictor. The motion vector of the enhancement layer block is either equal to the motion vector predictor or is equal to the motion vector predictor to which a residual vector is added. The residual vector is encoded in the scalable bitstream (respectively decoded from the scalable bitstream).
On the other hand, when encoding an enhancement layer block Be, a predictor P is determined. The predictor is usually a motion compensated reference block, i.e. P=MC(lref, MV). The reference block is identified in a reference picture Iref by the motion vector MV. The predictor P is then subtracted from the enhancement layer block. The residue thus obtained is encoded in the scalable bitstream. On the decoder side, the residue is decoded from the scalable bitstream and added to the predictor in order to reconstruct the enhancement layer block. In order to further increase coding efficiency, it is known to add the residue Res(UpscBL) of upscaled spatially corresponding base layer blocks to the motion compensated reference block in order to get the predictor. In this case, P=MC(lref, MV)+Res(UpscBL). Adding the residue of upscaled spatially corresponding base layer blocks to the motion compensated reference block in order to get the predictor is not systematic and thus requires signalling. However, signalling whether the residue of upscaled spatially corresponding base layer blocks is used for the prediction of the enhancement layer block is bit consuming.
3. BRIEF SUMMARY OF THE INVENTION
The purpose of the invention is to overcome at least one of the disadvantages of the prior art. For this purpose, a method for decoding a scalable bitstream comprising a base layer and at least one enhancement layer is disclosed. The method comprises for an enhancement layer block, decoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of the enhancement layer block and reconstructing the enhancement layer block depending on the flag, wherein the flag is decoded only when the enhancement block has a single spatially corresponding base layer block.
According to a specific embodiment, reconstructing the enhancement layer block depending on the flag comprises:
- determining a motion vector for the enhancement layer block;
- determining a motion compensated reference block from the motion vector;
- determining a predictor from the motion compensated reference block and depending on the flag from the residue of the spatially corresponding base layer block; and
- reconstructing the enhancement block from the predictor.
Advantageously, determining the predictor comprises adding a part of a scaled residue of the spatially corresponding base layer block co-located to the enhancement layer block to the motion compensated reference block when the flag indicates that a residue of a spatially corresponding base layer block is used for the reconstruction.
According to a specific characteristic, the motion vector is a motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer. Advantageously, the flag is decoded only when the enhancement block has a single spatially corresponding base layer block and when the motion vector is the motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
A method for encoding picture blocks in a scalable bitstream comprising a base layer and at least one enhancement layer is further disclosed. The method comprises for an enhancement layer block, encoding the enhancement layer block and encoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding of the enhancement layer block, wherein the flag is encoded only when the enhancement block has a single spatially corresponding base layer block. According to a specific embodiment, encoding the enhancement layer block comprises:
- determining a motion vector for the enhancement layer block;
- determining a motion compensated reference block from the motion vector;
- determining a predictor from at least the motion compensated reference block; and
- encoding the enhancement block from the predictor and encoding the flag whose value depends on whether a residue of a spatially corresponding base layer block is used in determining the predictor.
Advantageously, determining the predictor comprises adding a part of a scaled residue of the spatially corresponding base layer block co-located to the enhancement layer block to the motion compensated reference block when the flag indicates that a residue of a spatially corresponding base layer block is used for the encoding.
According to a specific characteristic, the motion vector is a motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
Advantageously, the flag is encoded only when the enhancement block has a single spatially corresponding base layer block and when the motion vector is the motion vector of the spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
A scalable video decoder for decoding a scalable bitstream comprising a base layer and at least one enhancement layer is also disclosed. The decoder comprises means for decoding, for an enhancement layer block, a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of the enhancement layer block and means for reconstructing the enhancement layer block depending on the flag, wherein the flag is decoded only when the enhancement block has a single spatially corresponding base layer block.
A scalable video encoder for encoding picture blocks in a scalable bitstream comprising a base layer and at least one enhancement layer is further disclosed. The encoder comprises means for encoding an enhancement layer block and means for encoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding of the enhancement layer block, wherein the flag is encoded only when the enhancement block has a single spatially corresponding base layer block. A scalable video bitstream comprising a base layer and at least one enhancement layer is disclosed. The bitstream encodes for an enhancement layer block a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding or reconstruction of the enhancement layer block, wherein the flag is encoded only when the enhancement block has a single spatially corresponding base layer block. A computer program product comprising program code instructions to execute of the steps of the decoding method according to the invention when this program is executed on a computer is also disclosed.
4. BRIEF DESCRIPTION OF THE DRAWINGS
Other features and advantages of the invention will appear with the following description of some of its embodiments, this description being made in connection with the drawings in which:
- Figure 1 represents a quadtree structure as defined in HEVC coding standard; - Figure 2 represents enhancement layer blocks, e.g. A and B, and corresponding base layer blocks A', B' and C;
- Figure 3 represents the flowchart of a decoding method of a scalable bitstream F with a view to the reconstruction of an enhancement layer block Be according to the invention;
- Figure 4 depicts in detail a step of the decoding method according to the invention;
- Figure 5 represents the flowchart of a coding method according to the invention;
- Figure 6 depicts in detail a step of the coding method according to the invention;
- Figure 7 depicts a scalable video decoder according to the invention; and
- Figure 8 depicts a scalable video encoder according to the invention.
5. DETAILED DESCRIPTION OF THE INVENTION
The invention relates to a method for reconstructing a block of pixels of a sequence of pictures and a method for coding such a block. A picture sequence is a series of several pictures. Each picture comprises pixels or picture points with each of which at least one item of picture data is associated. An item of picture data is for example an item of luminance data or an item of chrominance data. Hereafter, the coding and decoding methods are described with reference to a picture block. It is clear that these methods can be applied on several picture blocks of a picture and on several pictures of a sequence with a view to the coding respectively the decoding of one or more pictures. A picture block is a set of pixels of any form. It can be a square, a rectangle. But the invention is not limited to such forms. In the following section the word block is used for picture block.
The "predictor" term designates data used to predict other data. A predictor is used to predict a picture block. A predictor or prediction block is obtained from a block or several blocks of the same picture as the picture to which belongs the block that it predicts (spatial prediction or intra-picture prediction) or from one (mono-directional prediction) or several reference blocks (bi-directional prediction or bi-predicted) of a different picture (temporal prediction or inter- picture prediction) of the picture to which the block that it predicts belongs. Predictors are also used to predict motion vectors. In this case, the predictor is a motion vector predictor. A motion vector predictor is for example a motion vector associated with a block located in the causal neighborhood of a current block. In case of spatial scalability, the motion vector predictor for an enhancement layer block can be the motion vector associated with a spatially corresponding base layer block after scaling, e.g. the motion vector associated with a base layer block covering the central position of the enhancement layer block. A "causal neighbourhood" of a current block designates a neighbourhood of this block that comprises pixels coded/reconstructed before the current block.
The term "residue" signifies data obtained after extraction of other data. The extraction is generally a subtraction pixel by pixel of prediction data from source data. However, the extraction is more general and notably comprises a weighted subtraction in order for example to account for an illumination variation model. In the case of motion vector, a residual motion vector AMV is obtained by subtracting a motion vector predictor MVP from a motion vector MVc.
The term "reconstruction" designates data (e.g. pixels, blocks) obtained after merging residues with a predictor. The merging is generally a sum of prediction data with residues. However, the merging is more general and notably comprises a weighted sum in order for example to account for an illumination variation model. A reconstructed block is a block of reconstructed pixels.
In reference to the decoding of pictures, the terms "reconstruction" and "decoding" are very often used as synonyms. Hence, a "reconstructed block" is also designated under the terminology "decoded block".
The term coding is to be taken in the widest sense. The coding can possibly comprise the transformation and/or the quantization of picture data. It can also designate only the entropy coding.
In the Figures, the represented boxes are purely functional entities, which do not necessarily correspond to physical separated entities. As will be appreciated by one skilled in the art, aspects of the present principles can be embodied as a system, method or computer readable medium. Accordingly, aspects of the present principles can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, and so forth), or an embodiment combining software and hardware_aspects that can all generally be referred to herein as a "circuit," "module", or "system." Furthermore, aspects of the present principles can take the form of a computer readable storage medium. Any combination of one or more computer readable storage medium(s) may be utilized.
The flowchart and/or block diagrams in the Figures illustrate the configuration, operation and functionality of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, or blocks may be executed in an alternative order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and/or flowchart illustration, and combinations of the blocks in the block diagrams and/or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions. While not explicitly described, the present embodiments may be employed in any combination or subcombination.
As depicted on Figure 1 , HEVC video coding standard I SO/I EC ISO/I EC 23008-2 MPEG-H Part 2 defines a quadtree structure. A picture is divided into coding units CU. Each coding unit is divided into prediction units (PU) and each PU is further divided into transform units (TU). Each TU is encoded by applying an appropriate transform. On Figure 1 , the CU is divided into 4 PUs. A PU can then be encoded by applying a single transform. In this case a single TU corresponds to the PU. A PU can be encoded by dividing it horizontally or vertically into two TUs and applying a transform on each of them. A PU can also be encoded by dividing it in 4 TUs and applying a transform on each of them. In spatial scalability, an enhancement layer block has spatially corresponding base layer blocks. The spatially corresponding base layer blocks are the base layer blocks that cover the enhancement layer blocks when the scaled base layer picture is superposed on the enhancement layer picture. The scaled base layer picture also called upsampled base layer picture is obtained by scaling/upsampling the base layer picture according to the spatial resolution ratio between the two layers. Figure 2 represents enhancement layer blocks, e.g. A and B, on the top right and base layer blocks A', B', C on the top left. The scaled base layer blocks are represented with thick lines on the bottom left part of the figure. On the bottom right part of the figure the scaled base layer blocks are superposed on the enhancement layer blocks. To the enhancement layer block A corresponds a single base layer block A' while two base layer blocks B' and C correspond to the enhancement layer block B. On the bottom right part of the figure, the enhancement layer block A is totally covered by the single scaled base layer block A'. On the contrary, the top part of enhancement layer block B is covered by the scaled base layer block B' and the bottom part is covered by the scaled base layer block C. To find the spatially corresponding base layer blocks for a given enhancement layer block it is not necessary to effectively do the scaling/upsampling and the superposition.
Figure 3 represents the flowchart of a decoding method of a scalable bitstream F with a view to the reconstruction of an enhancement layer block Be according to the invention. In a step E10, it is checked whether the enhancement layer block Be has a single spatially corresponding base layer block or a plurality of them. If the enhancement layer block Be has a single spatially corresponding base layer block, then the method continues to step E12, otherwise it goes to step E16.
In a step E12, a flag is decoded for the enhancement layer block Be. The flag indicates whether a spatially corresponding base layer block residue is used for the reconstruction of the enhancement layer block Be, i.e. in the prediction process. This flag is only decoded for the enhancement layer block Be when the block Be has a single spatially corresponding base layer block as block A on figure 2. Consequently no such flag is decoded for an enhancement layer block which has more than one spatially corresponding base layer block as block B on figure 2. Advantageously, less bits are used since the flag is not decoded systematically for each block of the enhancement layer.
In a step E14, the enhancement layer block is reconstructed depending on the value of the flag. As an example, when the flag takes the value Ό' then the residue of the spatially corresponding base layer block is not used for the reconstruction of the enhancement layer block A and when the flag takes the value then the residue of the spatially corresponding base layer block is used for the reconstruction of the enhancement layer block A. According to a variant, when the flag takes the value then the residue of the spatially corresponding base layer block is not used for the reconstruction of the enhancement layer block A and when the flag takes the value Ό' then the residue of the spatially corresponding base layer block is used for the reconstruction of the enhancement layer block A. In the following sections, we suppose that the value Ό' indicates that the residue of the spatially corresponding base layer block is not used for the reconstruction of the enhancement layer block while the value indicates that the residue of the spatially corresponding base layer block is used for the reconstruction of the enhancement layer block.
In a step E16, the enhancement layer block is reconstructed without using the residue of spatially corresponding base layer blocks. Classically, the step E16 comprises decoding a residue for the block Be from the scalable bitstream and adding the residue to a predictor P. The predictor is usually a motion compensated reference block. The motion compensated reference block is determined in a reference picture from a motion vector decoded for the block Be. Indeed, in this case a plurality of base layer blocks corresponds spatially to the enhancement layer block as for block B on figure 2. In this case, the use of residues of at least two different base layer blocks for the reconstruction of the enhancement layer block is probably not efficient. Indeed, when the residues of the two different base layer blocks are obtained with different motion vectors using them for the prediction of an enhancement layer block whose motion vector is also probably different is not appropriate.
Figure 4 depicts in detail the step E14 of the decoding method according to the invention. In a step E140, a motion vector MVc is determined for the enhancement layer block Be. According to a specific embodiment, the motion vector MVc is directly decoded from the enhancement layer. According to a first variant, the motion vector is directly derived from the single spatially corresponding base layer block. More precisely, the motion vector MVc is equal to the motion vector MVBL_UPSC of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers. With respect to figure 2, the MVc for the enhancement layer block A is for example the motion vector of the base layer block A' scaled according to the spatial resolution ratio between the two layers. As an example, if the enhancement layer pictures dimensions are twice the dimensions of the base layer pictures then the ratio is 2 and the scaled base layer motion vector components are equal to the components of the base layer motion vector multiplied by 2. In this case, MVC= MVBL_UPSC, where MVBL_UPSC is the scaled base layer motion vector.
According to another variant, the scaled motion vector is only used as a predictor and thus a residual motion vector AMV decoded from the scalable bitstream is added to the scaled base layer motion vector. In this case, MVC=MVP+AMV=MVBL_UPSC+ AMV.
In a step E142, a motion compensated reference block MC(MVc, Iref) is determined from the motion vector MVC determined for the enhancement layer block In a step E140. The motion compensated reference block is determined in a reference picture Iref.
In a step E144, the value of the flag is checked. If the flag takes the value Ό' then the method continues to step E146, otherwise it goes to step E148.
In a step E146, the predictor P of the enhancement layer block Be is set equal to the motion compensated reference block MC(MVc, Iref).
In a step E148, the predictor P of the enhancement layer block Be is set equal to the sum of the motion compensated reference block MC(MVc, Iref) and the co-located part ResBL of the scaled/upsampled residue of the spatially corresponding base layer block. With respect to figure 2, the co-located part is the part of the scaled/upsampled residue of the spatially corresponding base layer block A' that covers the enhancement layer block A.
In a step E150, the enhancement layer block Be is reconstructed by adding the predictor P to the residue, if any, decoded for the enhancement layer block Be from the scalable bitstream. The decoding of the residue usually comprises entropy decoding, scaling (in the sense of inverse quantizing) and transform. According to a variant no such residue is decoded. In this case, the reconstructed block Be is equal to its predictor P. When neither a residue nor a residual motion vector is decoded for the enhancement layer block, the enhancement layer block is reconstructed according to a skip mode.
According to another embodiment, the flag is only decoded from the scalable bitstream for the blocks Be having a single spatially corresponding base layer block and whose motion vector predictor MVP is the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers. This variant encompasses the specific embodiment wherein the motion vector of the enhancement layer block is equal to the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
Figure 5 represents the flowchart of a coding method according to the invention. In a step E50, it is checked whether the enhancement layer block Be has a single spatially corresponding base layer block or a plurality of them. If the enhancement layer block Be has a single spatially corresponding base layer block, then the method continues to step E52, otherwise it goes to step E56.
In a step E52, the enhancement layer block Be is encoded. The way the enhancement layer block Be is encoded depends on a coding mode selected according to a selection criterion. Indeed, on the encoder side several coding modes are checked. The best mode from a rate distortion point of view is usually selected. In the step E52, the enhancement layer block Be is encoded with or without using the residue of the spatially corresponding base layer block for the prediction of block Be depending on the selected coding mode. The invention is not limited by the way the coding mode is selected. One way to select a coding mode is to check all the coding modes and to select the one that reaches the best rate-distortion compromise.
In a step E54, a flag is encoded for the enhancement layer block Be. The flag indicates whether a spatially corresponding base layer block residue is used for the encoding (resp. reconstruction on the decoder side) of the enhancement layer block Be. This flag is only encoded for the enhancement layer block Be when the block Be has a single spatially corresponding base layer block as block A depicted on figure 2. Consequently no such flag is encoded for an enhancement layer block which has more than one spatially corresponding base layer block as block B on figure 2. Advantageously, less bits are used since the flag is not encoded systematically for each block of the enhancement layer.
In a step E56, the enhancement layer block is encoded without using any base layer block residue. Classically, the step E56 comprises encoding a residue for the block Be in the scalable bitstream F. The residue is obtained by subtracting a predictor P from the block Be. The predictor is usually a motion compensated reference block. The motion compensated reference block is determined in a reference picture from a motion vector encoded for the block Be.
Figure 6 depicts detail the step E52 of the coding method according to the invention. In a step E520, a motion vector MVc is determined for the enhancement layer block. According to a specific embodiment, the motion vector is determined by motion estimation. According to a variant, the motion vector MVc equals the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers. With respect to figure 2, the MVc for the enhancement layer block A is for example the motion vector of the base layer block A' scaled according to the spatial resolution ratio between the two layers. As an example, if the enhancement layer pictures dimensions are twice the dimensions of the base layer pictures then the ratio is 2 and the scaled motion vectors components equal the components of the base layer motion vector multiplied by 2. In this case, MVC= MVBL_UPSC, where MVBL_UPSC is the scaled base layer motion vector. In a step E522, a motion compensated reference block MC(MVc, Iref) is determined from the motion vector MVC determined for the enhancement layer block in the step E520. The motion compensated reference block is determined in a reference picture Iref.
In a step E524, it is checked whether the base layer block residue is used for the prediction of block Be. For a block Be having a single spatially corresponding base layer block, this choice is encoder dependent. The invention is not at all limited by the way the choice is made. If the encoding method chooses to use the residue of the spatially corresponding base layer block for the prediction of block Be then the method continues to step E526, otherwise it goes to step E528.
In a step E526, the predictor P of the enhancement layer block Be is set equal to the sum of the motion compensated reference block MC(MVc, Iref) and the co-located part ResBL of the scaled/upsampled residue of the spatially corresponding base layer block. With respect to figure 2, the co-located part is the part of the scaled/upsampled residue of the spatially corresponding base layer block A' that covers the enhancement layer block A. The flag for the block Be is set equal to the value Ί ' if this value indicates that the base layer block residue is used for the encoding of the enhancement layer block Be. In a step E528, the predictor P of the enhancement layer block Be is set equal to the motion compensated reference block MC(MVc, Iref). The flag for the block Be is thus step equal to the value Ό' if this value indicates that the base layer block residue is not used for the encoding of the enhancement layer block Be.
In a step E530, the enhancement layer block Be is encoded from the predictor P. A residue is possibly encoded for the block Be by subtracting from the block Be, the predictor P to get a residue. The residue is usually transformed and quantized before being entropy coded. According to a variant no such residue is decoded. A residual motion vector is also possibly encoded. The residual motion vector is equal to (MVc-MVP), where MVP is a predictor. The predictor is for example MVBL_upsc. According to a variant no such residual motion vector is encoded. When neither a residue nor a residual motion vector is encoded for the enhancement layer block, the enhancement layer block is encoded in a skip mode.
According to another embodiment, the flag is only encoded in the scalable bitstream for the blocks Be having a single spatially corresponding base layer block and whose motion vector predictor MVP is the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers. This variant encompasses the specific embodiment wherein the motion vector of the enhancement layer block is equal to the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
A scalable video bitstream is also disclosed that comprises a base layer and at least one enhancement layer. The video bitstream comprises for each enhancement layer block Be having a single spatially corresponding base layer block a flag specifying whether the residue of said single spatially corresponding base layer block is used for the encoding/reconstruction of the block Be, more precisely for its prediction. According to a variant, the bitstream comprises such a flag only for the enhancement layer block Be having a single spatially corresponding base layer block and whose motion vector predictor MVP is the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers. This variant encompasses the specific embodiment wherein the motion vector of the enhancement layer block is equal to the motion vector of the single spatially corresponding base layer block scaled according to the spatial resolution ratio between the two layers.
When neither a residue nor a residual motion vector is encoded in the scalable bitstream for the enhancement layer block, the enhancement layer block is encoded in a skip mode.
An example of a syntax modified according to the invention is provided for the HEVC video coding standard. All the modifications are provided in italics. The syntax as defined in document JCT-VC K1003 "High Efficiency Video Coding (HEVC) text specification draft" is modified as follows at the prediction unit level:
idx_mv_baselayer is defined as the index in the merging candidate list corresponding to the motion vector candidate derived from the base layer colocated PU, i.e. the spatially corresponding base layer block. If such a candidate is not present in the merging candidate list, for instance if the base layer PU was coded in Intra mode, idx_mv_baselayer is set to -1 .
A new flag use_base_residue_flag is conditionnally read from the bitstream if mergejdx corresponds to idx_mv_baselayer and if cu_single_basis[ xO ][ yO ] is true. cu_single_basis[ xO ][ yO ] indicates that the current enhancement layer PU is included in one single base-layer PU, i.e. has a single spatially corresponding base layer block. Depending on its value, the scaled/upsampled residue will be added to the predictor.
Figure imgf000018_0001
The decoding process (see section 8.5.4 of JCT-VC K1003) is thus modified as follows:
Depending on rqt_root_cbf, the following applies:
- If rqt_root_cbf is equal to 0, If use_base_residue_flag is equal to 0, all samples of the (nCSL)x(nCSL) array resSamplesL and all samples of the two (nCSc)x(nCSc) arrays resSamplescb and resSamplescr are set equal to 0.
- Otherwise (use_base_residue_flag is equal to 1) , all samples of the (nCSL)x(nCSi_) array resSamplesL and all samples of the two
(nCSc)x(nCSc) arrays resSamplescb and resSamplescr are derived from the colocated base layer residue after upsampling.
Figure 7 depicts a scalable video decoder DEC according to the invention. The decoder DEC receives on a first input IN a scalable bitstream. The input is connected to a demultiplexer DEMUX configured to demultiplex the bitstream into a base layer BL and at least one enhancement layer EL. According to a variant not represented on Figure 7, the decoder DEC comprises two inputs and receives on a first input the base layer BL and on the second input the enhancement layer EL. In this case, the decoder DEC does not comprise DEMUX. The DEMUX if present is connected to a base layer decoder 'BL decoder' adapted to decode base layer pictures. The DEMUX if present is further connected to an enhancement layer decoder ΈΙ decoder' adapted to decode enhancement layer pictures. The enhancement layer decoder is adapted to implement the steps E10 to E16 of the decoding according to one of the various embodiments of the decoding method. The decoded base layer pictures are then possibly transmitted on a first output of the decoder OUT1 and the decoded enhancement layer pictures are then transmitted on a second output of the decoder OUT2. As another example, only decoded enhancement layer pictures are transmitted on the second output OUT2 for being displayed on a 2D display such as an UHDTV display. According to a variant not represented on Figure 7, the decoder DEC only comprises the second output of the decoder OUT2. Figure 8 depicts a scalable video encoder according to the invention.
The encoder ENC receives on a first input IN1 base layer pictures and on a second input enhancement layer pictures. The first input IN1 is connected to a base layer encoder adapted to encode and reconstruct the base layer pictures. The second input IN2 is connected to an enhancement layer encoder adapted to implement the steps E50 to E56 and thus to encode the enhancement layer pictures according to one of the various embodiments of the encoding method. The scalable video coders and decoders according to the inventionare for example implemented in various forms of hardware, software, firmware, special purpose processors, or a combination thereof. Preferably, the present principles may be implemented as a combination of hardware and software. Moreover, the software is preferably implemented as an application program tangibly embodied on a program storage device. The application program may be uploaded to, and executed by, a machine comprising any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units (CPU), a random access memory (RAM), and input/output (I/O) interface(s). The computer platform also includes an operating system and microinstruction code. The various processes and functions described herein may either be part of the microinstruction code or part of the application program (or a combination thereof) that is executed via the operating system. In addition, various other peripheral devices may be connected to the computer platform such as an enhancement data storage device and a printing device.
According to variants, the coding and decoding devices according to the invention are implemented according to a purely hardware realisation, for example in the form of a dedicated component (for example in an ASIC (Application Specific Integrated Circuit) or FPGA (Field-Programmable Gate Array) or VLSI (Very Large Scale Integration) or of several electronic components integrated into a device or even in a form of a mix of hardware elements and software elements.

Claims

Claims
1 . A method for decoding a scalable bitstream comprising a base layer and at least one enhancement layer, the method comprising for an enhancement layer block, decoding (E10, E12) a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of said enhancement layer block and reconstructing (E14, E16) said enhancement layer block depending on said flag, wherein said flag is decoded (E10, E12) only when said enhancement block has a single spatially corresponding base layer block.
2. The method according to claim 1 , wherein reconstructing (E14, E16) said enhancement layer block depending on said flag comprises:
- determining (E140) a motion vector for said enhancement layer block;
- determining (E142) a motion compensated reference block from said motion vector;
- determining (E144, E146, E148) a predictor from said motion compensated reference block and depending on said flag from said residue of said spatially corresponding base layer block; and
- reconstructing (E150) said enhancement block from said predictor.
3. The method according to claim 2, wherein determining (E144, E146, E148) the predictor comprises adding a part of a scaled residue of said spatially corresponding base layer block co-located to said enhancement layer block to said motion compensated reference block when said flag indicates that a residue of a spatially corresponding base layer block is used for the reconstruction.
4. The method according to claim 2 or 3, wherein said motion vector is a motion vector of said spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
5. The method according to claim 4, wherein said flag is decoded only when said enhancement block has a single spatially corresponding base layer block and when said motion vector is the motion vector of said spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
6. A method for encoding picture blocks in a scalable bitstream comprising a base layer and at least one enhancement layer, the method comprising for an enhancement layer block, encoding said enhancement layer block (E52, E56) and encoding (E50, E54) a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding of said enhancement layer block, wherein said flag is encoded (E50, E54) only when said enhancement block has a single spatially corresponding base layer block.
7. The method according to claim 6, wherein encoding (E52, E56) said enhancement layer block comprises:
- determining (E140) a motion vector for said enhancement layer block;
- determining (E142) a motion compensated reference block from said motion vector;
- determining (E144, E146, E148) a predictor from at least said motion compensated reference block; and
- encoding (E150) said enhancement block from said predictor and encoding (E54) said flag whose value depends on whether a residue of a spatially corresponding base layer block is used in determining the predictor.
8. The method according to claim 7, wherein determining (E144, E146, E148) the predictor comprises adding a part of a scaled residue of said spatially corresponding base layer block co-located to said enhancement layer block to said motion compensated reference block when said flag indicates that a residue of a spatially corresponding base layer block is used for the encoding.
9. The method according to claim 7 or 8, wherein said motion vector is a motion vector of said spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
10. The method according to claim 7, wherein said flag is encoded only when said enhancement block has a single spatially corresponding base layer block and when said motion vector is the motion vector of said spatially corresponding base layer block whose components are scaled according a spatial resolution ratio between the base layer and the enhancement layer.
1 1 . A scalable video decoder for decoding a scalable bitstream comprising a base layer and at least one enhancement layer, the decoder comprising means for decoding, for an enhancement layer block, a flag indicating whether a residue of a spatially corresponding base layer block is used for the reconstruction of said enhancement layer block and means for reconstructing said enhancement layer block depending on said flag, wherein said flag is decoded only when said enhancement block has a single spatially corresponding base layer block.
12. A scalable video decoder according to claim 1 1 , wherein said decoder is adapted to execute the steps of the method according to any of claims 1 to 5.
13. A scalable video encoder for encoding picture blocks in a scalable bitstream comprising a base layer and at least one enhancement layer, the encoder comprising means for encoding an enhancement layer block and means for encoding a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding of said enhancement layer block, wherein said flag is encoded only when said enhancement block has a single spatially corresponding base layer block.
14. A scalable video encoder according to claim 13, wherein said encoder is adapted to execute the steps of the method according to any of claims 6 to 10.
15. A scalable video bitstream comprising a base layer and at least one enhancement layer encoding for an enhancement layer block a flag indicating whether a residue of a spatially corresponding base layer block is used for the encoding or reconstruction of said enhancement layer block, wherein said flag is encoded only when said enhancement block has a single spatially corresponding base layer block.
16. A computer program product comprising program code instructions to execute the steps of the decoding method according to any of claims 1 to 5 when this program is executed on a computer.
17. A computer program product comprising program code instructions to execute the steps of the encoding method according to any of claims 6 to 10 when this program is executed on a computer.
PCT/EP2014/050040 2013-01-07 2014-01-03 Video coding and decoding methods and corresponding devices including flag for enhancement layer Ceased WO2014106641A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP13305007.0 2013-01-07
EP13305007 2013-01-07

Publications (1)

Publication Number Publication Date
WO2014106641A1 true WO2014106641A1 (en) 2014-07-10

Family

ID=47559358

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2014/050040 Ceased WO2014106641A1 (en) 2013-01-07 2014-01-03 Video coding and decoding methods and corresponding devices including flag for enhancement layer

Country Status (1)

Country Link
WO (1) WO2014106641A1 (en)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1694074A1 (en) * 2005-02-18 2006-08-23 Thomson Licensing Process for scalable coding of images
US20080165850A1 (en) * 2007-01-08 2008-07-10 Qualcomm Incorporated Extended inter-layer coding for spatial scability
US20100158116A1 (en) * 2006-11-17 2010-06-24 Byeong Moon Jeon Method and apparatus for decoding/encoding a video signal

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1694074A1 (en) * 2005-02-18 2006-08-23 Thomson Licensing Process for scalable coding of images
US20100158116A1 (en) * 2006-11-17 2010-06-24 Byeong Moon Jeon Method and apparatus for decoding/encoding a video signal
US20080165850A1 (en) * 2007-01-08 2008-07-10 Qualcomm Incorporated Extended inter-layer coding for spatial scability

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
SEGALL C A ET AL: "Spatial Scalability Within the H.264/AVC Scalable Video Coding Extension", IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, IEEE SERVICE CENTER, PISCATAWAY, NJ, US, vol. 17, no. 9, 1 September 2007 (2007-09-01), pages 1121 - 1135, XP011193020, ISSN: 1051-8215, DOI: 10.1109/TCSVT.2007.906824 *

Similar Documents

Publication Publication Date Title
KR102493516B1 (en) Intra prediction method based on CCLM and its device
KR102543468B1 (en) Intra prediction method based on CCLM and its device
CN108632628B (en) Method for deriving reference prediction mode values
US9826244B2 (en) Device and method for scalable coding of video information based on high efficiency video coding
KR102301450B1 (en) Device and method for scalable coding of video information
US9560358B2 (en) Device and method for scalable coding of video information
US9584808B2 (en) Device and method for scalable coding of video information
US20140092956A1 (en) Adaptive transform options for scalable extension
CN112806008A (en) Video coding and decoding method and device
CN110719485A (en) Video decoding method, device and storage medium
KR102314587B1 (en) Device and method for scalable coding of video information
JP6393323B2 (en) Method and device for decoding a scalable stream representing an image sequence and corresponding encoding method and device
HUE032327T2 (en) Conditional indication of timing information for video sequencing for video timing in video encoding
EP3854081A1 (en) A method and an apparatus for encoding and decoding of digital image/video material
WO2014197595A1 (en) Dynamic range control of intermediate data in resampling process
JP2025521822A (en) METHOD AND APPARATUS FOR IMAGE ENCODING/DECODING BASED ON ILLUMINATION COMPENSATION, AND RECORDING MEDIUM FOR STORING BITSTREAM
KR20150026924A (en) Methods and Apparatus for depth quadtree prediction
KR20160064843A (en) Method and apparatus for intra-coding in depth map
EP2741487A1 (en) Video coding and decoding methods and corresponding devices

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 14700141

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 14700141

Country of ref document: EP

Kind code of ref document: A1