WO2012147966A1 - 画像復号装置、画像符号化装置、および符号化データのデータ構造 - Google Patents

画像復号装置、画像符号化装置、および符号化データのデータ構造 Download PDF

Info

Publication number
WO2012147966A1
WO2012147966A1 PCT/JP2012/061478 JP2012061478W WO2012147966A1 WO 2012147966 A1 WO2012147966 A1 WO 2012147966A1 JP 2012061478 W JP2012061478 W JP 2012061478W WO 2012147966 A1 WO2012147966 A1 WO 2012147966A1
Authority
WO
WIPO (PCT)
Prior art keywords
decoding
unit
coefficient
encoding
transform
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2012/061478
Other languages
English (en)
French (fr)
Inventor
将伸 八杉
知宏 猪飼
山本 智幸
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sharp Corp
Original Assignee
Sharp Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sharp Corp filed Critical Sharp Corp
Priority to JP2013512497A priority Critical patent/JP6051156B2/ja
Priority to RU2013152171A priority patent/RU2609096C2/ru
Priority to CN201280020067.0A priority patent/CN103493494A/zh
Publication of WO2012147966A1 publication Critical patent/WO2012147966A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T9/00Image coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/129Scanning of coding units, e.g. zig-zag scan of transform coefficients or flexible macroblock ordering [FMO]
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/157Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
    • H04N19/159Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/18Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being a set of transform coefficients
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/44Decoders specially adapted therefor, e.g. video decoders which are asymmetric with respect to the encoder
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/46Embedding additional information in the video signal during the compression process
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/70Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/85Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression
    • H04N19/88Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression involving rearrangement of data among different coding units, e.g. shuffling, interleaving, scrambling or permutation of pixel data or permutation of transform coefficient data among different blocks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/13Adaptive entropy coding, e.g. adaptive variable length coding [AVLC] or context adaptive binary arithmetic coding [CABAC]

Definitions

  • the present invention relates to an image decoding apparatus that decodes transform coefficients, an image encoding apparatus that encodes transform coefficients, and a data structure of encoded data in which transform coefficients are encoded.
  • a moving image encoding device that generates encoded data by encoding the moving image, and a moving image that generates a decoded image by decoding the encoded data
  • An image decoding device is used.
  • Non-Patent Documents 1 and 2 A method adopted in KTA software, which is a codec for joint development in AVC and VCEG (Video Coding Expert Group), a method adopted in TMuC (Test Model Under Consulation) software, and HEVC (High-Efficiency (Video Coding) is proposed (Non-Patent Documents 1 and 2).
  • an image is divided into blocks of a predetermined size, and a transform coefficient is derived by frequency-converting a pixel value for each block, and an encoding process is performed on the derived transform coefficient.
  • a transform coefficient is derived by frequency-converting a pixel value for each block, and an encoding process is performed on the derived transform coefficient.
  • HM (HEVC Test Model) 2.0 has proposed a technique for reducing the code amount by encoding only 64 coefficients at the maximum in a block having a size of 16 ⁇ 16 pixels or more. For example, a technique has been proposed in which a block of 8 ⁇ 8 or more encoded is encoded only up to a maximum 8 ⁇ 8 region on the low frequency component side (Non-patent Document 3).
  • Non-Patent Document 4 a technique for reducing the code amount by deriving a value indicating a combination of ⁇ run, level ⁇ for a block having a size of 8 ⁇ 8 pixels or more by calculation.
  • run is the number of zero coefficients (0 run) continuous in a predetermined scan order
  • level means the absolute value of the coefficient.
  • HM3.0 which is the successor of HM2.0, adopts the proposal of Non-Patent Document 4 in order to reduce the code amount, and also has a size of 16 ⁇ 16 or more for improving the image quality Even in the block (transform unit) size, all transform coefficients are encoded.
  • ⁇ WD2 Working Draft 2 of the High-Efficiency Video Video Coding (JCTVC-D503), Join Collaborative Team Video Coding (JCT-VC) of ITU-T SG16, WP3, ISO / IEC, JTC1 / SC29 / WG11, 4th KR, 1/2011 (released in January 2011)
  • ⁇ WD3 Working Draft 3 of High-Efficiency Video Coding (JCTVC-E603) '', Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 5thevaMeeting: CH, 3/2011 (released in March 2011) ⁇ Samsung's Response to the Call for Proposals on Video Compression Technology (JCTVC-A124) '', Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SCdenWGing
  • the maximum number of coefficients is 256.
  • the maximum number of coefficients is 1024.
  • examples of the table that tends to increase in size include a scan table that specifies a scan order, a VLC table for run-level encoding, and the like.
  • the scan table needs a size proportional to the number of coefficients
  • the run-level encoding VLC table needs a size corresponding to the maximum run length
  • the present invention has been made in view of the above-described problems, and an object of the present invention is to reduce the amount of decoding information for obtaining transform coefficients from encoded data and the amount of calculation based on the decoding information.
  • An object is to realize an image decoding device, an image encoding device, and a data structure of encoded data.
  • the image decoding apparatus uses the encoded data obtained by encoding the transform coefficient obtained by frequency-converting the pixel value of the target image for each transform unit.
  • transform unit dividing means for dividing the transform unit into a plurality of sub-units, and decoding information for obtaining the transform coefficients from the encoded data, each sub-unit
  • transform coefficient decoding means for decoding transform coefficients included in the sub-unit with reference to the decoding information assigned to the sub-unit.
  • an image encoding apparatus is an image encoding apparatus that encodes a transform coefficient obtained by frequency-converting a pixel value of a target image for each transform unit.
  • a transform unit dividing unit that divides a transform unit into a plurality of subunits, encoding information for encoding the transform coefficient, and referring to the encoding information assigned to each subunit, Transform coefficient coding means for coding the transform coefficient included in the transform unit.
  • the data structure of the encoded data according to the present invention is generated by encoding a transform coefficient obtained by frequency-converting the pixel value of the target image for each transform unit in order to solve the above problem.
  • an image decoding apparatus that includes a relative position with respect to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded, thereby decoding the encoded data.
  • the position of the transform coefficient to be decoded is specified from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the conversion unit to be decoded is divided into a plurality of sub-units.
  • the conversion unit is a unit for converting pixel values into the frequency domain. Examples of the conversion unit include a size of 64 ⁇ 64 pixels, 32 ⁇ 32 pixels, and 16 ⁇ 16 pixels.
  • the sub unit may be, for example, an 8 ⁇ 8 size area.
  • a plurality of subunits obtained by division are processed one by one, and transform coefficients included in the subunit are decoded.
  • the decoding process can be performed in any order.
  • the decoding information assigned to each of the plurality of sub-units is referred to when transform coefficients are decoded.
  • Decoding information is information for reproducing a predetermined parameter value of a transform coefficient from a code (bit string) of encoded data.
  • the decoding information is a table indicating association for reproducing a predetermined parameter value of the transform coefficient from the code of the encoded data.
  • the decoding information is a calculation formula for deriving a predetermined parameter value of the transform coefficient from the code of the encoded data.
  • the transform coefficient is decoded using the decoding information defined for a sub-unit smaller than the size of the original transform unit.
  • the size of the scan table that defines the scan order of the transform coefficients can also be reduced.
  • the amount of memory and processing capacity required for the decoding process can be kept low.
  • the sub unit may coincide with any of the encoding units in the techniques of Non-Patent Documents 1 and 2.
  • a VLC table defined in advance in the coding unit that is, decoding information can be used.
  • the image decoding apparatus uses the encoded data obtained by encoding the transform coefficient obtained by frequency-converting the pixel value of the target image for each transform unit.
  • relative position decoding means for decoding a relative position from the transform coefficient decoded immediately before the transform coefficient to be decoded, and the transform unit of the transform coefficient decoded immediately before
  • position specifying means for specifying the position of the transform coefficient to be decoded from the relative position.
  • an image encoding apparatus is an image encoding apparatus that encodes a transform coefficient obtained by frequency-converting a pixel value of a target image for each transform unit.
  • Relative position encoding means for encoding a relative position with respect to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be converted is provided.
  • the data structure of the encoded data according to the present invention is generated by encoding the transform coefficient obtained by converting the pixel value of the target image into the frequency transform for each transform unit in order to solve the above problem.
  • the image decoding includes a relative position to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded, thereby decoding the encoded data
  • the apparatus is characterized in that the position of the transform coefficient to be decoded is specified from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the position of the transform coefficient to be decoded is specified from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the position of the conversion coefficient can be specified in a chain manner based on the relative position.
  • the conversion unit is a predetermined unit for conversion.
  • the length of the run is counted according to a predetermined scan order, so that the relative position of the two-dimensional coordinate in the conversion unit between the reference non-zero coefficient and the next non-zero coefficient is Even if they are close to each other, as a result, the run may become longer, which may increase the code amount.
  • the code amount can be reduced in such a case.
  • the amount of memory and processing capacity required for the decoding process can be kept low.
  • An image decoding apparatus includes transform unit dividing means for dividing a transform unit into a plurality of subunits, and decoding information for obtaining the transform coefficient from encoded data, which is assigned to each subunit. And transform coefficient decoding means for decoding transform coefficients included in the sub-unit with reference to the decoded information.
  • the image coding apparatus includes transform unit dividing means for dividing a transform unit into a plurality of subunits, and encoding information for encoding the transform coefficient, and each of the subunits is encoded. It is a structure provided with the conversion coefficient encoding means which encodes the conversion coefficient contained in the said conversion unit with reference to the encoding information allocated.
  • the data structure of the encoded data according to the present invention includes a relative position with respect to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded, thereby decoding the encoded data.
  • the image decoding apparatus having a data structure that specifies the position of the transform coefficient to be decoded from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the image decoding apparatus includes a relative position decoding unit that decodes a relative position from a transform coefficient decoded immediately before a transform coefficient to be decoded, and a position in the transform unit of the transform coefficient decoded immediately before And position specifying means for specifying the position of the transform coefficient to be decoded from the relative position.
  • the image encoding apparatus includes a relative position encoding unit that encodes a relative position with respect to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded. .
  • the data structure of the encoded data according to the present invention includes a relative position with respect to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded, thereby decoding the encoded data.
  • the image decoding device that performs the data structure specifies the position of the transform coefficient to be decoded from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the image decoding device it is possible to reduce the information amount of the decoding information for obtaining the transform coefficient from the encoded data and the calculation amount based on the decoding information. Moreover, according to the data structure of the said image coding apparatus or coded data, there exists an effect similar to the said image decoding apparatus.
  • FIG. 3 is a diagram illustrating a data configuration of encoded data generated by a video encoding device according to an embodiment of the present invention and decoded by the video decoding device, wherein (a) to (d) are pictures, respectively. It is a figure which shows a layer, a slice layer, a tree block layer, and a CU layer. It is the flowchart which illustrated about the flow of the process which divides
  • FIG. 1 It is a figure which shows the example of a division
  • FIG. 20 is a diagram illustrating an example in which the high frequency component side region in the target block illustrated in FIG. 17 and FIG. 19 is further subdivided into three regions. An example of a VLC table associated with the above three areas is shown.
  • (A) of the same figure shows the VLC table referred when the value of x or y does not become a positive value more than predetermined. Further, (b) in the figure shows a VLC table that is referred to when the value of x or y does not become a predetermined negative value or less. It is a flowchart shown about an example of the flow of the process encoded / decoded while switching two processes. It is a figure which illustrates about the process which encodes / decodes only 64 coefficients by the side of the low frequency component in an object block. (A) of the figure shows the case of inter prediction, and (b) shows the case of intra prediction. It is a figure shown about an example of the data structure of coefficient coding data.
  • FIG. 27 shows a flag tree representation example (quadtree representation) representing the division status and coefficient distribution status of the target block shown in FIG. It is a flowchart shown about an example of the flow of a decoding process of a recursive area
  • FIG. 2 is a functional block diagram showing a schematic configuration of the moving picture decoding apparatus 1.
  • VCEG Video Coding Expert Group
  • TMuC Transmission Model Underside
  • the moving picture decoding apparatus 1 receives encoded data # 1 obtained by encoding a moving picture by the moving picture encoding apparatus 2.
  • the video decoding device 1 decodes the input encoded data # 1 and outputs the video # 2 to the outside.
  • the configuration of the encoded data # 1 will be described below.
  • the encoded data # 1 exemplarily includes a sequence and a plurality of pictures constituting the sequence.
  • FIG. 3 shows the hierarchical structure below the picture layer in the encoded data # 1.
  • 3A to 3D are included in the picture layer that defines the picture PICT, the slice layer that defines the slice S, the tree block layer that defines the tree block TBLK, and the tree block TBLK, respectively.
  • Picture layer In the picture layer, a set of data referred to by the video decoding device 1 for decoding a picture PICT to be processed (hereinafter also referred to as a target picture) is defined. As shown in FIG. 3A, the picture PICT includes a picture header PH and slices S 1 to S NS (NS is the total number of slices included in the picture PICT).
  • the picture header PH includes a coding parameter group referred to by the video decoding device 1 in order to determine a decoding method of the target picture.
  • the encoding mode information (entropy_coding_mode_flag) indicating the variable length encoding mode used in encoding by the moving image encoding device 2 is an example of an encoding parameter included in the picture header PH.
  • the picture PICT is encoded by LCEC (Low Complexity Entropy Coding) or CAVLC (Context-based Adaptive Variable Length Coding).
  • LCEC Low Complexity Entropy Coding
  • CAVLC Context-based Adaptive Variable Length Coding
  • CABAC Context-based Adaptive Binary Arithmetic Coding
  • picture header PH is also referred to as a picture parameter set (PPS).
  • PPS picture parameter set
  • slice layer In the slice layer, a set of data referred to by the video decoding device 1 for decoding the slice S to be processed (also referred to as a target slice) is defined. As shown in FIG. 3B, the slice S includes a slice header SH and a sequence of tree blocks TBLK 1 to TBLK NC (NC is the total number of tree blocks included in the slice S).
  • the slice header SH includes a coding parameter group that the moving image decoding apparatus 1 refers to in order to determine a decoding method of the target slice.
  • Slice type designation information (slice_type) for designating a slice type is an example of an encoding parameter included in the slice header SH.
  • I slice that uses only intra prediction at the time of encoding (2) P slice that uses unidirectional prediction or intra prediction at the time of encoding, (3) B-slice using unidirectional prediction, bidirectional prediction, or intra prediction at the time of encoding may be used.
  • the slice header SH may include a filter parameter referred to by a loop filter (not shown) included in the video decoding device 1.
  • Tree block layer In the tree block layer, a set of data referred to by the video decoding device 1 for decoding a processing target tree block TBLK (hereinafter also referred to as a target tree block) is defined.
  • the tree block TBLK includes a tree block header TBLKH and coding unit information CU 1 to CU NL (NL is the total number of coding unit information included in the tree block TBLK).
  • NL is the total number of coding unit information included in the tree block TBLK.
  • the tree block TBLK is divided into units for specifying a block size for each process of intra prediction or inter prediction and conversion.
  • the above unit of the tree block TBLK is divided by recursive quadtree partitioning.
  • the tree structure obtained by this recursive quadtree partitioning is hereinafter referred to as a coding tree.
  • a unit corresponding to a leaf that is a node at the end of the coding tree is referred to as a coding node.
  • the encoding node is a basic unit of the encoding process, hereinafter, the encoding node is also referred to as an encoding unit (CU).
  • CU encoding unit
  • coding unit information (hereinafter referred to as CU information)
  • CU 1 to CU NL is information corresponding to each coding node (coding unit) obtained by recursively dividing the tree block TBLK into quadtrees. is there.
  • the root of the coding tree is associated with the tree block TBLK.
  • the tree block TBLK is associated with the highest node of the tree structure of the quadtree partition that recursively includes a plurality of encoding nodes.
  • each coding node is half the size of the coding node to which the coding node directly belongs (that is, the unit of the node one layer higher than the coding node).
  • the size that each coding node can take depends on the size designation information of the coding node and the maximum hierarchy depth (maximum hierarchical depth) included in the sequence parameter set SPS of the coded data # 1. For example, when the size of the tree block TBLK is 64 ⁇ 64 pixels and the maximum hierarchical depth is 3, the encoding nodes in the hierarchy below the tree block TBLK have four sizes, that is, 64 ⁇ 64. It can take any of a pixel, 32 ⁇ 32 pixel, 16 ⁇ 16 pixel, and 8 ⁇ 8 pixel.
  • the tree block header TBLKH includes an encoding parameter referred to by the video decoding device 1 in order to determine a decoding method of the target tree block. Specifically, as shown in (c) of FIG. 3, tree block division information SP_TBLK that designates a division pattern of the target tree block into each CU, and a quantization parameter difference that designates the size of the quantization step ⁇ qp (qp_delta) is included.
  • the tree block division information SP_TBLK is information representing a coding tree for dividing the tree block. Specifically, the shape and size of each CU included in the target tree block, and the position in the target tree block Is information to specify.
  • the tree block division information SP_TBLK may not explicitly include the shape or size of the CU.
  • the tree block division information SP_TBLK may be a set of flags (split_coding_unit_flag) indicating whether or not the entire target tree block or a partial area of the tree block is divided into four.
  • the shape and size of each CU can be specified by using the shape and size of the tree block together.
  • the quantization parameter difference ⁇ qp is a difference qp ⁇ qp ′ between the quantization parameter qp in the target tree block and the quantization parameter qp ′ in the tree block encoded immediately before the target tree block.
  • CU layer In the CU layer, a set of data referred to by the video decoding device 1 for decoding a CU to be processed (hereinafter also referred to as a target CU) is defined.
  • the encoding node is a node at the root of a prediction tree (PT) and a transformation tree (TT).
  • PT prediction tree
  • TT transformation tree
  • the encoding node is divided into one or a plurality of prediction blocks, and the position and size of each prediction block are defined.
  • the prediction block is one or a plurality of non-overlapping areas constituting the encoding node.
  • the prediction tree includes one or a plurality of prediction blocks obtained by the above division.
  • Prediction processing is performed for each prediction block.
  • a prediction block that is a unit of prediction is also referred to as a prediction unit (PU).
  • intra prediction There are roughly two types of division in the prediction tree: intra prediction and inter prediction.
  • inter prediction there are 2N ⁇ 2N (the same size as the encoding node), 2N ⁇ N, N ⁇ 2N, N ⁇ N, and the like.
  • the encoding node is divided into one or a plurality of transform blocks, and the position and size of each transform block are defined.
  • the transform block is one or a plurality of non-overlapping areas constituting the encoding node.
  • the conversion tree includes one or a plurality of conversion blocks obtained by the above division.
  • transform processing is performed for each conversion block.
  • the transform block which is a unit of transform is also referred to as a transform unit (TU).
  • the CU information CU specifically includes a skip flag SKIP, PT information PTI, and TT information TTI.
  • the skip flag SKIP is a flag indicating whether or not the skip mode is applied to the target PU.
  • the value of the skip flag SKIP is 1, that is, when the skip mode is applied to the target CU, PT information PTI and TT information TTI in the CU information CU are omitted. Note that the skip flag SKIP is omitted for the I slice.
  • the PT information PTI is information regarding the PT included in the CU.
  • the PT information PTI is a set of information related to each of one or more PUs included in the PT, and is referred to when the moving image decoding apparatus 1 generates a predicted image.
  • the PT information PTI includes prediction type information PType and prediction information PInfo.
  • Prediction type information PType is information that specifies whether intra prediction or inter prediction is used as a prediction image generation method for the target PU.
  • the prediction information PInfo is composed of intra prediction information or inter prediction information depending on which prediction method is specified by the prediction type information PType.
  • a PU to which intra prediction is applied is also referred to as an intra PU
  • a PU to which inter prediction is applied is also referred to as an inter PU.
  • the prediction information PInfo includes information specifying the shape, size, and position of the target PU. As described above, the generation of the predicted image is performed in units of PU. Details of the prediction information PInfo will be described later.
  • TT information TTI is information related to TT included in the CU.
  • the TT information TTI is a set of information regarding each of one or a plurality of TUs included in the TT, and is referred to when the moving image decoding apparatus 1 decodes residual data.
  • a TU may be referred to as a block.
  • the TT information TTI includes TT division information SP_TU for designating a division pattern of the target CU into each conversion block, and TU information TUI 1 to TUI NT (NT is assigned to the target CU. The total number of blocks included.
  • TT division information SP_TU is information for determining the shape and size of each TU included in the target CU and the position in the target CU.
  • the TT division information SP_TU can be realized from information (split_transform_unit_flag) indicating whether or not the target node is divided and information (trafoDepth) indicating the depth of the division.
  • each TU obtained by the division can take a size from 32 ⁇ 32 pixels to 2 ⁇ 2 pixels.
  • the TU information TUI 1 to TUI NT are individual information regarding one or more TUs included in the TT.
  • the TU information TUI includes a quantized prediction residual.
  • Each quantized prediction residual is encoded data generated by the video encoding device 2 performing the following processes 1 to 3 on a target block that is a processing target block.
  • Process 1 DCT transform (Discrete Cosine Transform) of the prediction residual obtained by subtracting the prediction image from the encoding target image;
  • Process 2 Quantize the transform coefficient obtained in Process 1;
  • Process 3 Variable length coding is performed on the transform coefficient quantized in Process 2;
  • prediction information PInfo As described above, there are two types of prediction information PInfo: inter prediction information and intra prediction information.
  • the inter prediction information includes an encoding parameter that is referred to when the video decoding device 1 generates an inter predicted image by inter prediction. More specifically, the inter prediction information includes inter PU division information that specifies a division pattern of the target CU into each inter PU, and inter prediction parameters for each inter PU.
  • the inter prediction parameters include a reference image index, an estimated motion vector index, and a motion vector residual.
  • the intra prediction information includes an encoding parameter that is referred to when the video decoding device 1 generates an intra predicted image by intra prediction. More specifically, the intra prediction information includes intra PU division information that specifies a division pattern of the target CU into each intra PU, and intra prediction parameters for each intra PU.
  • the intra prediction parameter is a parameter for designating an intra prediction method (prediction mode) for each intra PU.
  • the video decoding device 1 generates a predicted image for each PU, generates a decoded image # 2 by adding the generated predicted image and a prediction residual decoded from the encoded data # 1, and generates The decoded image # 2 is output to the outside.
  • An encoding parameter is a parameter referred in order to generate a prediction image.
  • the encoding parameters include PU size and shape, block size and shape, and original image and Residual data with the predicted image is included.
  • side information a set of all information excluding the residual data among the information included in the encoding parameter.
  • the present invention is not limited to this, and the present invention can also be applied to a case where PU and TU are units smaller than CU.
  • a picture (frame), a slice, a tree block, a block, and a PU to be decoded are referred to as a target picture, a target slice, a target tree block, a target block, and a target PU, respectively.
  • the size of the tree block is, for example, 64 ⁇ 64 pixels
  • the size of the PU is, for example, 64 ⁇ 64 pixels, 32 ⁇ 32 pixels, 16 ⁇ 16 pixels, 8 ⁇ 8 pixels, 4 ⁇ 4 pixels, or the like.
  • these sizes are merely examples, and the sizes of the tree block and PU may be other than the sizes shown above.
  • FIG. 2 is a functional block diagram showing a schematic configuration of the moving picture decoding apparatus 1.
  • the moving image decoding apparatus 1 includes a variable length code demultiplexing unit 11, a TU information decoding unit 12, an inverse quantization / inverse transform unit 13, a predicted image generation unit 14, an adder 15, and a frame memory 16. It has.
  • variable-length code demultiplexing unit 11 demultiplexes the encoded data # 1 for one frame input to the video decoding device 1 to obtain various kinds of information included in the hierarchical structure shown in FIG. To separate.
  • the variable length code demultiplexer 11 refers to information included in various headers and sequentially separates the encoded data # 1 into slices and tree blocks.
  • the various headers include (1) information about the method of dividing the target picture into slices, and (2) information about the size, shape, and position of the tree block belonging to the target slice. .
  • variable length code demultiplexing unit 11 refers to the tree block division information SP_TBLK included in the tree block header TBLKH, and divides the target tree block into CUs. In addition, the variable-length code demultiplexer 11 acquires TT information TTI and PT information PTI for the target CU.
  • variable length code demultiplexing unit 11 supplies the TU information TUI included in the TT information TTI obtained for the target CU to the TU information decoding unit 12 in a predetermined order. In addition, the variable-length code demultiplexing unit 11 supplies the PT information PTI obtained for the target CU to the predicted image generation unit 14.
  • the TU information decoding unit 12 decodes the TU information TUI supplied from the variable length code demultiplexing unit 11 to generate decoded TU information TUI ′.
  • the TU information decoding unit 12 decodes the quantized prediction residual from the TU information TUI for the target block.
  • the quantized prediction residual for the target block can be expressed in a form in which quantized transform coefficients are arranged in a two-dimensional matrix.
  • a two-dimensional matrix representation of quantized transform coefficients is referred to as a coefficient matrix.
  • the TU information decoding unit 12 supplies the decoded TU information TUI ′ to the inverse quantization / inverse transform unit 13. Details of the operation of the TU information decoding unit 12 will be described later.
  • the inverse quantization / inverse transform unit 13 performs inverse quantization / inverse transform of the quantized prediction residual for each block for the target CU based on the decoded decoded TU information TUI ′ supplied from the TU information decoding unit 12. I do.
  • the inverse quantization / inverse transform unit 13 performs inverse quantization and inverse DCT transform (Inverse Discrete Cosine Transform) on the quantized prediction residual included in the decoded TU information TUI ′, so that for each target PU, for each pixel.
  • the prediction residual D is restored.
  • the inverse quantization / inverse transform unit 13 supplies the restored prediction residual D to the adder 15.
  • the predicted image generation unit 14 For each PU included in the target CU, the predicted image generation unit 14 refers to a local decoded image P ′ that is a decoded image around the PU, and generates a predicted image Pred by intra prediction or inter prediction. The predicted image generation unit 14 supplies the predicted image Pred generated for the target CU to the adder 15.
  • the adder 15 adds the predicted image Pred supplied from the predicted image generation unit 14 and the prediction residual D supplied from the inverse quantization / inverse transform unit 13, thereby obtaining the decoded image P for the target CU. Generate.
  • the decoded image P that has been decoded is sequentially recorded in the frame memory 16.
  • decoded images corresponding to all tree blocks decoded before the target tree block are stored. It is recorded.
  • Decoded image # 2 corresponding to # 1 is output to the outside.
  • the moving picture coding apparatus 2 scans a coefficient matrix representing a set of coefficients in the target block in a predetermined order.
  • Scan is a process of converting the coordinates of the coefficients in the coefficient matrix into a one-dimensional scan order index.
  • the converted coefficients are stored and held in a one-dimensional array.
  • For scanning a conventionally known zigzag scanning order or the like can be used.
  • the DC coefficient at the upper left position to the coefficient of the highest frequency component at the lower right position are scanned based on the zigzag scan order.
  • the coefficient value of the high frequency component tends to be zero or close to zero, whereas the coefficient value of the low frequency component has a characteristic that it is likely that the value is not zero or is large. For this reason, the zigzag scan has a scan order in which the low-frequency component coefficients are scanned quickly.
  • a coefficient coding process is performed after scanning.
  • the coefficient encoding process is performed in the reverse order of the scan.
  • this processing order is also called reverse zigzag scanning. That is, the encoding process is performed by a reverse zigzag scan on the coefficient matrix. Further, when describing the encoding processing order, for convenience, the description will be made based on a coefficient matrix.
  • the scan order may be adaptively changed to a method other than the zigzag scan, for example, the same order as the raster scan.
  • coefficient encoded data encoded by the moving image encoding apparatus 2 will be described.
  • the last non-zero coefficient is encoded.
  • encoding in the run mode is performed under a predetermined condition.
  • the run mode ends by satisfying the run mode end condition the remaining coefficients are encoded in the level mode. In this way, all the coefficients are encoded.
  • the last non-zero coefficient is encoded by the last non-zero coefficient position last_pos, the last non-zero coefficient level level, and the last non-zero coefficient sign sign.
  • the coefficient level means the absolute value of the coefficient.
  • last_pos takes a value in the form of a scan order index. When the last non-zero coefficient is the 10th in the scanning order (the DC coefficient is the first) and the value is “ ⁇ 1”, it is as follows.
  • the run mode is a mode for encoding the number of consecutive zero coefficients (0 run). Under what conditions the encoding in the run mode is performed will be described later.
  • Non-Patent Documents 1 and 2 the case where a coefficient whose level is 2 or more appears is used as one of the end conditions of the run mode.
  • Level mode is a mode in which even zero coefficients are encoded one by one. In the level mode, for each coefficient, the coefficient level and the sign of the coefficient sign are encoded.
  • the case where the value of the coefficient to be encoded is “ ⁇ 6” and the case where the coefficient to be encoded is a zero coefficient (value is “0”) are as follows.
  • the syntax element may be encoded with a name or definition different from the above.
  • the level level of the coefficient of the run mode is the syntax of isLevelOne (level is 1), and in addition to this, the level_magnitude_minus2 (when the coefficient value is 2 or more, encoding is a level obtained by subtracting 2). It may be encoded as an element (see “level_magnitude_minus2” in Non-Patent Documents 1 and 2).
  • the level may be encoded as a syntax element of last_pos_level for the last non-zero coefficient.
  • level of the non-run mode may be encoded as a syntax element of level_magnitude.
  • FIG. 1 is a functional block diagram illustrating a configuration example of the TU information decoding unit 12.
  • the TU information decoding unit 12 decodes data related to coefficients among the encoded data included in the TU information TUI.
  • a configuration for the TU information decoding unit 12 to decode the encoded quantized prediction residual, that is, coefficient encoded data will be described below.
  • the TT information decoding unit 12 is not limited to this, and can decode data other than the coefficient encoded data included in the encoded data, for example, side information.
  • the TU information decoding unit 12 includes a VLC table TBL11, a region dividing unit (transform unit dividing unit) 121, and a region decoding unit (transform coefficient decoding unit) 122.
  • the TU information decoding unit 12 decodes the TU information TUI for a 16 ⁇ 16 size target block and outputs the decoded TU information TUI ′.
  • the present invention is not limited to this, and the size of the target block decoded by the TU information decoding unit 12 may be 64 ⁇ 64, 32 ⁇ 32, or the like.
  • the VLC table TBL11 is a table in which a bit number (code), a code number that can be mutually converted, and a parameter to be decoded are associated with each other.
  • the VLC table TBL11 is referred to in the decoding process of the area decoding unit 122.
  • VLC table TBL11 for example, a table defined for an 8 ⁇ 8 size encoding unit is used. That is, for the VLC table TBL11, an 8 ⁇ 8 size smaller than the input conversion unit (or encoding unit) 16 ⁇ 16 size is used.
  • VLC tables TBL11 are defined according to the context in the decoding process so that an adaptive decoding process can be performed.
  • the context for example, the position of the coefficient being processed, the attribute of the target block (pixel type such as luminance / color difference, and prediction method) and the like can be used.
  • the area dividing unit 121 divides the target block into a plurality of areas.
  • each area obtained by the division by the area dividing unit 121 is referred to as a decoding area.
  • the region dividing unit 121 divides a 16 ⁇ 16 size target block into four 8 ⁇ 8 size decoding regions.
  • the present invention is not limited to this, and various methods can be adopted as the dividing method of the area dividing unit 121. The modification will be described in detail later.
  • the area decoding unit 122 performs a decoding process for each of the decoding areas obtained by dividing the target block by the area dividing unit 121 while referring to the VLC table TBL11 defined according to the size of the decoding area. Do.
  • the region decoding unit 122 performs a decoding process while referring to the VLC table TBL11 defined for an 8 ⁇ 8 size encoding unit according to a context.
  • the region decoding unit 122 includes a final non-zero coefficient decoding unit 101, a run mode decoding unit 102, and a level mode decoding unit 103.
  • the last non-zero coefficient decoding unit 101 decodes the last non-zero coefficient for a decoding area to be decoded (hereinafter referred to as a target decoding area). More specifically, the last non-zero coefficient decoding unit 101 decodes last_pos, level, and sign included in the coefficient encoded data corresponding to the target decoding area.
  • the run mode decoding unit 102 decodes the coefficient encoded in the run mode from the coefficient encoded data for the target decoding area. That is, the run mode decoding unit 102 decodes run, level, and sign encoded in the run mode from the coefficient encoded data while referring to the VLC table TBL11 corresponding to the context.
  • the decoding process performed by the run mode decoding unit 102 is referred to as a run mode decoding process.
  • the run mode decoding unit 102 repeatedly performs the run mode decoding process until the run mode end condition is satisfied.
  • the run mode end condition include a case where the value of the decoded coefficient exceeds a threshold value and a case where a predetermined number of coefficients are decoded.
  • the run mode decoding unit 102 causes the level mode decoding unit 103 to start decoding in the level mode.
  • the level mode decoding unit 103 decodes the coefficient encoded in the level mode from the coefficient encoded data for the target decoding area. That is, the level mode decoding unit 103 decodes level and sign encoded in the level mode from the coefficient encoded data while referring to the VLC table TBL11 corresponding to the context.
  • the decoding process performed by the level mode decoding unit 103 is referred to as a level mode decoding process. Further, the level mode decoding unit 103 repeatedly performs the level mode decoding process until the DC component coefficient is decoded.
  • the area decoding unit 122 outputs the decoded TU information TUI ′ including the coefficient obtained by the decoding process.
  • FIG. 4 is a flowchart exemplifying the flow of the process S10 in which the target block is divided into regions and the coefficients are encoded / decoded.
  • the difference between the encoding process of the moving image encoding apparatus 2 and the decoding process of the moving image decoding apparatus 1 is whether the encoding process or the decoding process is performed.
  • the decoding process is almost the same. Therefore, in FIG. 4, the encoding process of the moving image encoding apparatus 2 and the decoding process of the moving image decoding apparatus 1 are collectively shown.
  • the area dividing unit 121 divides the target block (S11).
  • the loop LP1 for each divided decoding area is entered (S12).
  • the region decoding unit 122 decodes the coefficient included in the coefficient encoded data for the target decoding region (S13).
  • the last non-zero coefficient decoding unit 101 decodes the last non-zero coefficient in the target decoding area.
  • the run mode decoding unit 102 performs the run mode decoding process until the run mode end condition is satisfied.
  • the level mode decoding unit 103 After the run mode ends, the level mode decoding unit 103 performs a level mode decoding process. When the decoding process for the target decoding area is completed in this way, the process returns to the beginning of the loop LP1 (from S14 to S12), and the decoding process is further performed for the next target decoding area.
  • the loop LP1 ends. Thereafter, the process S10 for dividing the region and decoding the coefficients ends.
  • FIG. 5 shows a 16 ⁇ 16 size target block (conversion unit) BLK.
  • FIG. 6 shows an example in which the decoding process is performed by dividing the target block BLK into four 8 ⁇ 8 size areas.
  • the region dividing unit 121 divides the 16 ⁇ 16 size target block BLK into four 8 ⁇ 8 size decoding regions (sub units) R11 to R14 as shown in FIG.
  • the area decoding unit 122 processes each decoding area shown in FIG. 6 in the order of decoding areas R11, R12, R13 and R14.
  • the arrows shown in the decoding regions R11 to R14 indicate the scan order. That is, the scan order of the decoding regions R11 to R14 is a zigzag scan.
  • the last non-zero coefficient decoding unit 101 first decodes the last non-zero coefficient in the scan order. Then, by the run mode decoding process and the level mode decoding process, the last non-zero coefficient to the DC component coefficient are decoded by reverse zigzag scanning.
  • the decoding process of the decoding area R11 is completed, the decoding process is similarly performed on the decoding areas R12, R13, and R14.
  • the decoding process is performed by the conventional decoding method for the 8 ⁇ 8 size encoding unit. That is, in the decoding process of the decoding regions R11 to R14, a conventional decoding method that decodes the last non-zero coefficient and performs the run mode decoding process and the level mode decoding process can be used.
  • Modification 1-1 [Determination of Existence of Non-zero Coefficient]
  • the area decoding unit 122 may perform a decoding process according to the presence or absence of a non-zero coefficient.
  • the region decoding unit 122 may be configured as follows.
  • the region decoding unit 122 decodes the non-zero coefficient flag and determines whether or not there is a non-zero coefficient. If the non-zero coefficient flag indicates that there is no non-zero coefficient in the target decoding area, the area decoding unit 122 skips the decoding process in the target area.
  • FIG. 7 shows an example in which the decoding regions R11, R13, and R13 have non-zero coefficients and the decoding region R12 has no non-zero coefficients. That is, “0” shown in the decoding region R12 indicates that there is no non-zero coefficient.
  • non-zero coefficient flag for example, “0” is encoded when there is no non-zero coefficient, and “1” is encoded when there is a non-zero coefficient.
  • the regions R11, R13, and R14 have non-zero coefficients, and the region R12 has no non-zero coefficients. Therefore, “1011” is encoded as the non-zero coefficient flag. Note that a pattern such as “1011” of the non-zero coefficient flag may be subjected to variable length encoding according to the appearance frequency of the pattern without being fixed length encoded with 4 bits as it is.
  • the region decoding unit 122 performs a decoding process on the decoding region R11.
  • the region decoding unit 122 skips the decoding process for the decoding region R12.
  • the region decoding unit 122 For the remaining decoding regions R13 and R14, as in the decoding region R11, the non-zero coefficient flag “1” indicating that there is a non-zero coefficient is encoded, so the region decoding unit 122 performs the decoding regions R13 and R14. Decoding processing is performed for each of the above.
  • Modification 1-2 [Change the processing method according to the position of the decoding area] [1] Change of scanning method In the above description, zigzag scanning is adopted as the scanning method of each decoding area, but the scanning method may be changed according to the position of the decoding area.
  • horizontal scanning may be employed in the decoding region R12 located at the upper right in the target block BLK.
  • vertical scanning may be employed in the decoding region R13 located at the lower left in the target block BLK.
  • zigzag scanning may be employed in the decoding regions R11 and R14.
  • the run mode decoding process and the level mode decoding process are performed in the decoding process of each decoding area.
  • the run mode decoding process or the level mode decoding process is performed depending on the position of the decoding area. It is good also as a structure which performs either.
  • the decoding process may be performed only by the run mode decoding process.
  • VLC table to be referenced and the code number calculation method may be changed according to the position of the decoding area.
  • a VLC table for converting a bit string indicating a code number into a set of ⁇ run, level ⁇ parameters may be changed according to the position of each decoding area.
  • the zero coefficient tends to increase, and in the decoding region on the low frequency component side, the non-zero coefficient increases and the run tends to be shortened.
  • FIG. 12 shows an example of the VLC table TBL11 for converting a set of ⁇ run, level ⁇ parameters into code numbers.
  • FIG. 13 shows an example in which the VLC tables T1 and T2 are associated with each decoding area. As shown in FIG. 13, the decoding area R11 is associated with the VLC table T2, and the decoding areas R12 to R14 are associated with the VLC table T1.
  • the run mode decoding unit 102 refers to the VLC table T2 in the run mode decoding process of the decoding region R11 according to the association shown in FIG.
  • the run tends to be short, and therefore, by assigning a smaller code number to a combination of ⁇ run, level ⁇ having a short run, the coding efficiency can be improved.
  • the run mode decoding unit 102 refers to the VLC table T1 in the run mode decoding process of the decoding regions R12 to 14 in accordance with the association shown in FIG.
  • the present invention is not limited to this, and a conversion process equivalent to the VLC tables T1 and T2 is realized by calculation. It doesn't matter.
  • the run mode decoding unit 102 may be configured to refer to mutually different VLC tables in the decoding regions R11 to R14.
  • the area decoding unit 122 may count the appearance frequency of parameter values and rewrite the code number of the VLC table according to the appearance frequency.
  • the area decoding unit 122 rewrites the code number of the parameter value with a high appearance frequency to a smaller one (with a shorter code), and increases the code number of the parameter value with a lower appearance frequency (with a longer code). You may rewrite it as a thing.
  • the code number CN-1 is assigned to the value y of a certain parameter, and the code number CN is assigned to the value x of another parameter.
  • region decoding unit 122 refers to x in the VLC table T3 before optimization and the corresponding code number CN.
  • the area decoding unit 122 decrements the code number CN by 1 and increments the code number corresponding to x to CN-1.
  • the area decoding unit 122 increments the code number of y originally corresponding to CN-1 by 1 and associates it with CN.
  • optimization such a code number carry-up process is referred to as optimization.
  • a VLC table T4 shown in FIG. 14 is a table obtained after the optimization.
  • the region decoding unit 122 sets the code number corresponding to x to CN-2.
  • the region decoding unit 122 dynamically updates the VLC table during the decoding process so that a short code number is assigned according to the appearance frequency of the parameter value to be decoded.
  • the speed of the adaptation speed can be expressed as a size that decrements the code number, for example. That is, it can be said that the adaptation speed is faster when the code number decrement is “2” than when the code number is decremented by “1”.
  • the amount of increment at the time of optimization is set to “0” in the decoding region on the high frequency component side, and “1” on the low frequency component side. Good.
  • a configuration may be adopted in which optimization is not performed in the decoding region on the high frequency component side.
  • the reason for adopting this configuration is as follows. In other words, the run tends to be long on the high frequency component side, and there is a low possibility that a specific ⁇ run, level ⁇ will occur frequently. For this reason, according to the said structure, the increase in the computational complexity by performing an optimization whenever the value of a parameter appears can be prevented.
  • the run mode end condition may be changed according to the position of the decoding area.
  • the run mode end condition may be tightened in the decoding region on the high frequency component side, and the run mode end condition may be relaxed in the decoding region on the low frequency component side. That is, it is only necessary to make the run mode easier to end in the decoding area located at the upper left in the target block. Further, the level mode decoding process may not be performed in the decoding region on the high frequency component side.
  • the absolute value of the coefficient tends to be small. For this reason, there is a high possibility that the run will be long in the decoding region on the high frequency component side. Therefore, it is preferable to execute the run mode decoding process as much as possible.
  • the run mode decoding unit 102 may determine whether or not to start the run mode decoding process. For example, the run mode decoding unit 102 starts the run mode decoding process when “the absolute value of the coefficient can be determined to be small overall” based on data that can be referred to when the last non-zero coefficient is decoded. Then, it may be determined.
  • the run mode decoding unit 102 executes the run mode decoding process because there is a high possibility that the run will be long.
  • data that can be referred to when the last non-zero coefficient is decoded include, for example, prediction information of the target block, information about the last non-zero coefficient, and other encoded and already decoded on the encoder side. Flag.
  • the run mode decoding unit 102 may skip the run mode decoding process.
  • threshold values for various determinations in the coefficient decoding process may be changed according to the position of the decoding area.
  • Modified example 1-3 [Determining whether to divide according to the prediction mode of the target block]
  • the region dividing unit 121 may be configured to determine the presence or absence of division according to the prediction mode of the target block.
  • the region dividing unit 121 may perform division when the prediction mode of the target block is intra prediction, and may not perform division when the prediction mode is inter prediction.
  • the region decoding unit 122 decodes only the first to 64th coefficients in the scan order in the target block. You may make it do. That is, in this case, the region decoding unit 122 may decode only 64 coefficients on the upper left side of the target block.
  • Modification 1-4 [Number and Size of Regions]
  • the region dividing unit 121 has been described as dividing the target block into four decoding regions.
  • the present invention is not limited to this, and the area dividing unit 121 may divide the target block into more than four areas.
  • the target block is not limited to 16 ⁇ 16 size.
  • the target block may be 32 ⁇ 32 size or 64 ⁇ 64 size.
  • the area dividing unit 121 may divide the target block into four or more decoding areas as follows.
  • the area dividing unit 121 may divide the target block into 64 8 ⁇ 8 size decoding areas.
  • the area dividing unit 121 may divide the target block into 16 16 ⁇ 16 size decoding areas.
  • the area dividing unit 121 may divide the target block into four 32 ⁇ 32 size decoding areas.
  • the decoding area obtained by dividing the target block by the area dividing unit 121 is not limited to a square area.
  • the decoding area may be a rectangle.
  • the target block is divided into two, an upper left 8 ⁇ 8 size decoding area (corresponding to decoding area R11 in FIG. 5) and other decoding areas (corresponding to decoding areas R12 to R14 in FIG. 5) May be.
  • each decoding area may not be the same.
  • the target block may be divided into 64 coefficients in the zigzag scan order to obtain a decoding area.
  • the target block may be divided into decoding regions R21 to R24.
  • the number “64” shown in the decoding areas R21 to R24 indicates that 64 coefficients are included in the area.
  • the decoding region R21 is a region including DC coefficients, and the shape thereof is represented by a right triangle in the figure.
  • the shapes of the decoding regions R22 and R23 are each represented by a trapezoid.
  • the decoding region R24 is a region on the most high frequency component side among the regions, and the shape thereof is represented by a right triangle.
  • the shapes of the regions R21 and R24 are shown as right triangles, and the shapes of the regions R22 and R23 are shown as trapezoids. Please note that it does not become trapezoid.
  • each decoding area and the number of coefficients included in each decoding area may not be the same.
  • the area dividing unit 121 can arbitrarily divide the target block into a plurality of decoding areas having a size smaller than the size of the target block.
  • Modification 1-5 [Processing Order of Decoding Area]
  • the area decoding unit 122 performs the decoding process in the order of the decoding areas R11, R12, R13, and R14 (so-called raster scan order).
  • the present invention is not limited to this, and the area decoding unit 122 may perform the decoding process in an order other than this.
  • the region decoding unit 122 may perform the decoding process in the order of the decoding regions R14, R13, R12, and R11, or as another example, in the order of the decoding regions R11, R13, R12, and R14. Further, when the number of decoding areas is more than four, for example, when the number is 16, the processing may be performed in the zigzag scan order.
  • the moving image encoding device 2 is a device that generates and outputs encoded data # 1 by encoding the input image # 10.
  • FIG. 10 is a functional block diagram showing the configuration of the moving image encoding device 2.
  • the moving image encoding device 2 includes an encoding setting unit 21, an inverse quantization / inverse conversion unit 22, a predicted image generation unit 23, an adder 24, a frame memory 25, a subtractor 26, a conversion / A quantization unit 27 and a variable length coding unit 28 are provided.
  • the encoding setting unit 21 generates image data related to encoding and various setting information based on the input image # 10.
  • the encoding setting unit 21 generates the next image data and setting information.
  • the encoding setting unit 21 generates the CU image # 100 for the target CU by sequentially dividing the input image # 10 into slice units and tree block units.
  • the encoding setting unit 21 generates header information H ′ based on the result of the division process.
  • the header information H ′ includes (1) information about the size and shape of the tree block belonging to the target slice and the position in the target slice, and (2) the size, shape and shape of the CU belonging to each tree block.
  • the encoding setting unit 21 refers to the CU image # 100 and the CU information CU 'to generate PT setting information PTI'.
  • the PT setting information PTI ' includes information on all combinations of (1) possible division patterns of the target CU for each PU and (2) prediction modes that can be assigned to each PU.
  • the encoding setting unit 21 supplies the CU image # 100 to the subtractor 26. In addition, the encoding setting unit 21 supplies the header information H ′ to the variable length encoding unit 28. Also, the encoding setting unit 21 supplies the PT setting information PTI ′ to the predicted image generation unit 23.
  • the inverse quantization / inverse transform unit 22 performs inverse quantization and inverse DCT transform (Inverse Discrete Cosine Transform) on the quantization prediction residual for each block supplied from the transform / quantization unit 27, Restore the prediction residual for each block.
  • inverse quantization and inverse DCT transform Inverse Discrete Cosine Transform
  • the inverse quantization / inverse transform unit 22 integrates the prediction residual for each block according to the division pattern specified by the TT division information (described later), and generates the prediction residual D for the target CU.
  • the inverse quantization / inverse transform unit 22 supplies the prediction residual D for the generated target CU to the adder 24.
  • the predicted image generation unit 23 refers to the locally decoded image P ′ and the PT setting information PTI ′ recorded in the frame memory 25 to generate a predicted image Pred for the target CU.
  • the predicted image generation unit 23 sets the prediction parameter obtained by the predicted image generation process in the PT setting information PTI ′, and transfers the set PT setting information PTI ′ to the variable length encoding unit 28. Note that the predicted image generation process performed by the predicted image generation unit 23 is the same as that performed by the predicted image generation unit 14 included in the video decoding device 1, and thus the description thereof is omitted here.
  • the adder 24 adds the predicted image Pred supplied from the predicted image generation unit 23 and the prediction residual D supplied from the inverse quantization / inverse transform unit 22 to thereby obtain the decoded image P for the target CU. Generate.
  • Decoded decoded image P is sequentially recorded in the frame memory 25.
  • decoded images corresponding to all tree blocks decoded before the target tree block for example, all tree blocks preceding in the raster scan order
  • the time of decoding the target tree block It is recorded.
  • the subtractor 26 generates a prediction residual D for the target CU by subtracting the prediction image Pred from the CU image # 100.
  • the subtractor 26 supplies the generated prediction residual D to the transform / quantization unit 27.
  • the transform / quantization unit 27 performs a DCT transform (Discrete Cosine Transform) and quantization on the prediction residual D to generate a quantized prediction residual.
  • DCT transform Discrete Cosine Transform
  • the transform / quantization unit 27 refers to the CU image # 100 and the CU information CU 'and determines a division pattern of the target CU into one or a plurality of blocks. Further, according to the determined division pattern, the prediction residual D is divided into prediction residuals for each block.
  • the transform / quantization unit 27 generates a prediction residual in the frequency domain by performing DCT transform (DiscretecreCosine Transform) on the prediction residual for each block, and then quantizes the prediction residual in the frequency domain. Thus, a quantized prediction residual for each block is generated.
  • DCT transform DiscretecreCosine Transform
  • the transform / quantization unit 27 generates the quantization prediction residual for each block, TT division information that specifies the division pattern of the target CU, information about all possible division patterns for each block of the target CU, and TT setting information TTI ′ including is generated.
  • the transform / quantization unit 27 supplies the generated TT setting information TTI ′ to the inverse quantization / inverse transform unit 22.
  • the transform / quantization unit 27 generates TU setting information TUI ′ including the quantization prediction residual of the target block, and supplies the TU setting information TUI ′ to the variable length encoding unit 28.
  • variable length encoding unit 28 generates and outputs encoded data # 1 based on the TU setting information TUI ', the PT setting information PTI', and the header information H '. Details of the variable length coding unit 28 will be described below.
  • FIG. 11 is a block diagram illustrating a configuration example of the variable length coding unit 28.
  • variable length coding unit 28 includes a TU information coding unit 280, a header information coding unit 40, a PTI information coding unit 41, and a coded data multiplexing unit 42.
  • the TU information encoding unit 280 encodes the TU setting information TUI ′ and supplies it to the encoded data multiplexing unit 42.
  • the header information encoding unit 40 encodes the header information H ′ and supplies the encoded header information H ′ to the encoded data multiplexing unit 42.
  • the PTI information encoding unit 41 encodes the PTI information PTI ′ and supplies the encoded PTI information PTI ′ to the encoded data multiplexing unit 42.
  • the encoded data multiplexing unit 42 multiplexes the TU setting information TUI ′, the header information H ′, and the PTI information PTI ′ to generate encoded data # 1, and outputs it.
  • the configuration of the TU information encoding unit 280 will be described in more detail as follows.
  • a configuration for encoding coefficient prediction data by encoding a quantized prediction residual, that is, a coefficient matrix, included in the TU setting information TUI ' will be described.
  • the TU information encoding unit 280 encodes the TU information TUI for the 16 ⁇ 16 size target block and outputs the encoded TU information TUI ′.
  • the present invention is not limited to this, and the size of the target block encoded by the TU information encoding unit 280 may be 64 ⁇ 64, 32 ⁇ 32, or the like.
  • the TU information encoding unit 280 includes a VLC table TBL21, a region dividing unit (transform unit dividing unit) 281 and a region encoding unit (transform coefficient coding unit) 282.
  • the VLC table TBL21 is a table in which the correspondence between each parameter and a code that is a bit string of encoded data is defined.
  • VLC table TBL21 for example, a table defined for an 8 ⁇ 8 size encoding unit is used. That is, for the VLC table TBL21, an 8 ⁇ 8 size smaller than the input conversion unit (or encoding unit) 16 ⁇ 16 size is used.
  • VLC tables TBL21 are defined according to the context in the encoding process so that an adaptive encoding process can be performed.
  • the area dividing unit 281 divides the target block into a plurality of areas.
  • each area obtained by the division of the area dividing unit 281 is referred to as an encoding area.
  • the area dividing unit 281 divides a 16 ⁇ 16 size target block into four coding areas. That is, the region dividing unit 281 divides a 16 ⁇ 16 size target block into four 8 ⁇ 8 size coding regions.
  • the area dividing unit 281 can divide the target block, for example, as shown in FIG. In the following description, the decoding regions R11 to R14 shown in FIG. 5 are read as coding regions R11 to R14.
  • the area dividing unit 281 divides the target block BLK into coding areas R11 to R14.
  • the present invention is not limited to this, and various methods can be applied to the dividing method of the region dividing unit 281. The modification will be described in detail later.
  • the area coding unit 282 performs coding for each of the coding areas obtained by dividing the target block by the area dividing unit 281 while referring to the VLC table TBL21 defined according to the size of the coding area. Process. That is, in the following, an example in which the region encoding unit 282 performs the encoding process while referring to the VLC table TBL21 defined for the 8 ⁇ 8 size according to the context will be described.
  • the region encoding unit 282 includes a final non-zero coefficient encoding unit 201, a run mode encoding unit 202, and a level mode encoding unit 203.
  • the last non-zero coefficient encoding unit 201 encodes the last non-zero coefficient for a coding area to be encoded (hereinafter referred to as a target coding area). More specifically, the last non-zero coefficient coding unit 201 codes last_pos, level, and sign for the last non-zero coefficient included in the coefficient matrix corresponding to the target coding region.
  • the run mode encoding unit 202 performs a run mode encoding process on the coefficient matrix for the target encoding region. That is, the run mode encoding unit 202 sequentially encodes run, level, and sign for the non-zero coefficients included in the coefficient matrix by the run mode while referring to the VLC table TBL21 corresponding to the context.
  • the encoding process performed by the run mode encoding unit 202 is referred to as a run mode encoding process.
  • the run mode encoding unit 202 repeats the run mode encoding process until the run mode end condition is satisfied.
  • the run mode end condition include a case where the value of the encoded coefficient exceeds a threshold value and a case where a predetermined number of coefficients are encoded.
  • the run mode encoding unit 202 causes the level mode encoding unit 203 to start encoding in the level mode.
  • the level mode encoding unit 203 encodes the coefficients included in the coefficient matrix in the target encoding area in the level mode. That is, the level mode encoding unit 203 sequentially encodes the level and sign of the coefficients included in the coefficient matrix in the level mode while referring to the VLC table TBL21 corresponding to the context.
  • the encoding process performed by the level mode encoding unit 203 is referred to as a level mode encoding process. Further, the level mode encoding unit 203 repeatedly performs the level mode encoding process until the first coefficient is encoded in the scan order of the target encoding region.
  • the area encoding unit 282 outputs encoded TU setting information TUI ′ including the coefficient obtained by the encoding process.
  • region encoding unit 282 can perform the encoding process as shown in FIG.
  • the decoding regions R11 to R14 shown in FIG. 6 are read as coding regions R11 to R14.
  • the region encoding unit 282 performs encoding processing in the order of the encoding regions R11, R12, R13, and R14 shown in FIG.
  • the last non-zero coefficient decoding unit 101 first decodes the last non-zero coefficient in the scan order. Then, the run mode decoding process and the level mode decoding process decode from the last non-zero coefficient to the first coefficient in the scan order of the target coding region in the reverse order of the scan order.
  • Modified example 1-1 ′ [Determination of presence / absence of non-zero coefficient]
  • the region encoding unit 282 included in the moving image encoding device 2 determines whether or not there is a non-zero coefficient for each encoding region, encodes a non-zero coefficient flag indicating the determination result, and includes it in the coefficient encoded data It is good. Further, in such a configuration, the region encoding unit 282 can omit the encoding of the coefficient for the encoding region where there is no non-zero coefficient.
  • coding region and the non-zero coefficient flag are substantially the same as those already described in Modification 1-1 of the video decoding device 1, and thus description thereof is omitted here.
  • “decoding region”, “decoding process”, and “region decoding unit 122” in the description of Modification 1-1 are respectively referred to as “encoding region”, “encoding process”, and “region encoding unit 282”. ".”
  • Modification 1-2 ′ [Processing method is changed according to the position of the coding area] [1] Change of scanning method
  • the zigzag scanning is adopted as the scanning method of each coding area.
  • the scanning method may be changed according to the position of the coding area.
  • the specific scanning method is the same as that described in, for example, the modified example 1-2 [1] of the video decoding device 1, and thus the description thereof is omitted here.
  • the “decoding region”, “decoding process”, and “region decoding unit 122” in the description of the modified example 1-2 [1] are respectively referred to as “encoding region”, “encoding process”, and “region code”. It shall be read as “the conversion unit 282”.
  • the run mode encoding process and the level mode encoding process are performed in the encoding process of each encoding area.
  • the run mode encoding process is performed according to the position of the encoding area. It is good also as composition which performs.
  • a specific processing example is the same as that described in Modification 1-2 [2] of the video decoding device 1, for example, and will not be described here.
  • the “decoding region”, “decoding process”, and “region decoding unit 122” in the description of the modified example 1-2 [2] are respectively referred to as “encoding region”, “encoding process”, and “region code”. It shall be read as “the conversion unit 282”.
  • VLC table to be referred to and the code number calculation method may be changed according to the position of the encoding area.
  • a VLC table for converting a bit string indicating a code number into a set of ⁇ run, level ⁇ parameters may be changed according to the position of each coding area.
  • a specific processing example is the same as that described in Modification 1-2 [3] of the moving picture decoding apparatus 1, for example, and thus the description thereof is omitted here.
  • the “decoding region”, “decoding process”, and “region decoding unit 122” in the description of the modified example 1-2 [3] are respectively referred to as “encoding region”, “encoding process”, and “region code”. It shall be read as “the conversion unit 282”.
  • the area encoding unit 282 may count the appearance frequency of parameter values and rewrite the code number of the VLC table according to the appearance frequency.
  • the area encoding unit 282 rewrites the code number of the parameter value having a high appearance frequency to a smaller one (with a shorter code), and increases the code number of the parameter value with a lower appearance frequency (having a longer code). ) You may rewrite it to something.
  • a specific processing example is the same as that described in, for example, the modification 1-2 [4-1] of the video decoding device 1, and thus the description thereof is omitted here.
  • the “decoding region”, “decoding process”, and “region decoding unit 122” in the description of the modified example 1-2 [4-1] are respectively referred to as “encoding region”, “encoding process”, and “ It shall be read as “region encoding unit 282”.
  • the run mode end condition may be changed according to the position of the coding area. Further, the run mode start condition may be determined according to the position of the coding region. Details thereof are the same as those described in the description of Modification 1-2 [5] of the video decoding device 1, and thus detailed description thereof is omitted here.
  • Modified example 1-3 ′ [Determining whether to divide according to the prediction mode of the target block]
  • the region dividing unit 281 may be configured to determine the presence or absence of division according to the prediction mode of the target block.
  • the specific processing is the same as that described in the modification 1-3 of the video decoding device 1, and thus the description thereof is omitted here.
  • Modified example 1-4 ′ [Number and size of regions]
  • the area dividing unit 281 may divide the target block into more than four areas.
  • the specific processing is the same as that described in, for example, the modified example 1-4 of the video decoding device 1, and thus the description thereof is omitted here.
  • Modification 1-5 ′ [Processing order of coding area]
  • the scan order of the encoding processing by the region encoding unit 282 is not limited to the above-described example.
  • the region encoding unit 282 may be configured to perform the same processing as the processing described in Modification 1-5 of the video decoding device 1.
  • [Modification] of the video decoding device 1 can be applied to the video encoding device 2.
  • [Modification] of the moving picture encoding apparatus 2 shown here can be adapted to the moving picture decoding apparatus 1 by changing the encoding process to the decoding process.
  • the video decoding device 1 uses the TU information TUI of the encoded data obtained by encoding the transform coefficient obtained by frequency transforming the pixel value of the target image for each transform unit.
  • a region division unit 121 that divides the target block that is the transform unit into a plurality of decoding regions, and decoding information for obtaining the transform coefficient from the TU information TUI
  • An area decoding unit 122 that decodes transform coefficients included in the decoding area with reference to the VLC table TBL11 assigned to each decoding area.
  • the moving image encoding device 2 encodes the transform coefficient obtained by frequency-converting the pixel value of the target image for each conversion unit.
  • An area dividing unit 281 that divides the image into a plurality of encoding areas and a VLC table TBL21 for encoding the conversion coefficient, and the conversion is performed with reference to the VLC table TBL21 assigned to each encoding area.
  • a region encoding unit 282 that encodes transform coefficients included in a unit.
  • the size of the original target block (16 ⁇ 16) is defined.
  • the size of the VLC table can be reduced as compared with the case where the decoding process is performed based on the existing VLC table.
  • the table representing the scan order is subjected to the decoding process based on the table defined for the 8 ⁇ 8 size decoding area, so that the size of the table can be reduced.
  • Embodiment 2 Another embodiment of the present invention will be described below with reference to FIGS. For convenience of explanation, members having the same functions as those in the drawings described in the first embodiment are denoted by the same reference numerals and description thereof is omitted.
  • the size of the target block is assumed to be 16 ⁇ 16 as an example.
  • the TU information decoding unit 12A shown in FIG. 15 will be described as follows. That is, as shown in FIG. 15, the TU information decoding unit 12A includes a VLC table TBL30, a run-level mode decoding unit 310, a relative position mode decoding unit 320, and a processing mode control unit 330.
  • the VLC table TBL30 is a table in which a bit number (code), a code number that can be mutually converted, and a parameter to be decoded are associated with each other.
  • the VLC table TBL30 includes a run-level mode table TBL31 referred to by a later-described run-level mode decoding unit 310 and a relative position mode table TBL32 referred to by a relative position mode decoding unit 320.
  • the run-level mode table TBL31 can be the same as the VLC table TBL11 of the TU information decoding unit 12 shown in FIG. Therefore, the description is omitted here.
  • the definition of the relative position mode table TBL32 will be described later.
  • the run-level mode decoding unit 310 performs a run mode decoding process and a level mode decoding process under the control of the processing mode control unit 330.
  • the decoding process performed by the run-level mode decoding unit 310 is referred to as a run-level mode decoding process. Since the run mode decoding process and the level mode decoding process have already been described in the first embodiment, the description thereof is omitted here.
  • the run-level mode decoding unit (decoding unit) 310 includes a final non-zero coefficient decoding unit 311, a run mode decoding unit 312, and a level mode decoding unit 313.
  • the last non-zero coefficient decoding unit 311, the run mode decoding unit 312, and the level mode decoding unit 313 are the last non-zero coefficient decoding unit 101, the run mode decoding unit 102, and the level shown in FIG. It has the same function as the mode decoding unit 103. Therefore, since the function has already been described, the description thereof is omitted here.
  • the run mode decoding unit 102 and the level mode decoding unit 103 are configured to refer to the run-level mode table TBL31 in the run-level mode decoding process.
  • the relative position mode decoding unit 320 performs the decoding process of the coefficient encoded data in which the relative position of the non-zero coefficient is encoded. Specifically, the relative position mode decoding unit 320 includes a final non-zero coefficient decoding unit 321, a relative position decoding unit (relative position decoding unit) 322, and a coefficient position determination unit (position specifying unit) 323.
  • the last non-zero coefficient decoding unit 321 decodes the last non-zero coefficient in the target block.
  • the last non-zero coefficient decoding unit 321 decodes last_pos, level, and sign included in the coefficient encoded data for the target block.
  • the last non-zero coefficient decoding unit 321 converts last_pos into a coordinate display (lastx, lasty) with the DC coefficient as the origin (0, 0).
  • a position indicated by coordinate display with the DC coefficient as the origin (0, 0) is referred to as an absolute position.
  • the relative position decoding unit 322 calculates the relative position (dx, dy) of the non-zero coefficient to be decoded and the value of the non-zero coefficient from the coefficient encoded data obtained by encoding the relative position of the non-zero coefficient for the target block. (Level and sign).
  • the relative position of the non-zero coefficient refers to the relative position of the non-zero coefficient to be decoded, as viewed from the absolute position of the non-zero coefficient decoded immediately before.
  • the expression of the position of the non-zero coefficient based on such a relative position is referred to as relative position designation.
  • the coefficient position determination unit 323 becomes a decoding target from the relative position (dx, dy) of the non-zero coefficient decoded by the relative position decoding unit 322 and the absolute position (x, y) of the non-zero coefficient decoded immediately before. Determine the absolute position of the non-zero coefficient.
  • the relative position from the last non-zero coefficient C 0 to the (n + 1) th non-zero coefficient C n + 1 can be expressed by, for example, the following relational expressions (1-1) to (1-3).
  • the coefficient position determination unit 323 determines the absolute position of the non-zero coefficient to be decoded using the relational expressions (1-1) to (1-3).
  • the processing mode control unit 330 determines whether or not the non-zero coefficient to be decoded is located within a predetermined region, and whether the run-level mode decoding unit 310 performs the decoding process according to the determination result
  • the relative position mode decoding unit 320 controls whether the decoding process is performed.
  • the processing mode control unit 330 determines whether or not the position of the non-zero coefficient to be decoded is within the 8 ⁇ 8 size region on the low frequency component side.
  • the position (x n , y n ) of the non-zero coefficient C n to be decoded is “x n ⁇ 8 && y n ⁇ 8 It is determined whether or not. “&&” is an operator indicating a logical product.
  • the processing mode control unit 330 sends the next non-zero coefficient C n + 1 to the run-level mode decoding unit 310. Decryption processing is executed.
  • the processing mode control unit 330 causes the relative position mode decoding unit 320 to perform decoding processing.
  • FIG. 16 is a flowchart illustrating the flow of the process S20 for encoding / decoding a non-zero coefficient by specifying a relative position.
  • FIG. 16 the encoding process of the moving image encoding apparatus 2 and the decoding process of the moving image decoding apparatus 1 are collectively shown.
  • the processing mode control unit 330 performs a decoding process on the relative position mode decoding unit 320. Let it run.
  • the last non-zero coefficient decoding unit 321 decodes the last non-zero coefficient of the target block (S21).
  • the processing mode control unit 330 determines whether or not the position of the last non-zero coefficient decoded in S21 is within the region for decoding the coefficient for which the relative position is specified (S22).
  • the processing mode control unit 330 causes the relative position mode decoding unit 320 to execute the decoding process.
  • the relative position decoding unit 322 decodes the position of the coefficient whose relative position is designated, and the coefficient position determination unit 323 decodes the coefficient to be decoded by determining the absolute position of the coefficient ( S23).
  • the processing mode control unit 330 causes the run-level mode decoding unit 310 to execute the decoding process.
  • the run mode decoding unit 312 executes the run mode decoding process, and then the level mode decoding unit 313 executes the level mode decoding process (S24). In this way, all the coefficients of the target block are decoded, and the process of decoding the coefficients by specifying the relative position ends.
  • FIG. 17 is a diagram illustrating an execution example of the decoding process of the TU information decoding unit 12A.
  • the target block BLK is composed of two regions R1 and R2.
  • the target block BLK is a 16 ⁇ 16 size block
  • the region R1 is an 8 ⁇ 8 size region on the low frequency component side, as described above.
  • the arrow shown in the region R1 indicates the scan order in the region R1, and the scan order is a zigzag scan.
  • the region R2 is a region for decoding coefficients by specifying a relative position.
  • the last non-zero coefficient decoding unit 321 decodes the last non-zero coefficient C N in the target block BLK (S21).
  • the processing mode control unit 330 causes the relative position mode decoding unit 320 to perform decoding processing.
  • the relative position decoding unit 322 decodes the relative position of the non-zero coefficient C N ⁇ 1 , and the coefficient position determination unit 323 determines from the decoded relative position and the absolute position of the last non-zero coefficient C N. Determine the absolute position of the non-zero coefficient C N ⁇ 1 . Further, the non-zero coefficient C N ⁇ 1 is decoded by decoding the level and the sign (S23).
  • the relative position mode decoding unit 320 executes the decoding process in the same manner as described above (S22, S23).
  • the non-zero coefficients to C 0, the relative position mode decoding unit 320 completes the decoding process, the non-zero coefficients C 0, so located in the region R1, (NO in S22), the processing mode controller 330, the run -Let the level mode decoding unit 310 execute a decoding process.
  • Run - Level mode decoding unit 310 the run mode decoding unit 312, the non-zero coefficients C 0 as a reference point, to perform the run mode decoding process in the region R1.
  • the level mode decoding unit 313 executes the level mode decoding process.
  • the decoding process of the run-level mode decoding unit 310 may be the same as the process of decoding an 8 ⁇ 8 size encoding unit by (in reverse order) zigzag scanning.
  • the non-zero coefficient C 0 that is first decoded in the region R1 is preferably the last non-zero coefficient in the 8 ⁇ 8 size region R1.
  • the relative position mode table TBL32 used for decoding the relative position (dx, dy) can be configured as follows.
  • a shorter code may be associated with the relative position mode table TBL32.
  • the relative distance between the non-zero coefficient to be processed and the previous non-zero coefficient can be derived, for example, from the relative position (dx, dy) of the non-zero coefficient to be processed.
  • a shorter code in the relative position mode table TBL32 may be associated with a pair (dx, dy) indicating a relative distance having a higher appearance frequency.
  • the maximum absolute value of dx and dy is the side length of the encoding target block minus one. That is, when the encoding target block is 16 ⁇ 16, the maximum absolute value of dx and dy is 15.
  • the processing mode control unit 330 determines the size of the target block, and controls whether the run-level mode decoding unit 310 performs the decoding process or the relative position mode decoding unit 320 performs the decoding process according to the determination. May be.
  • the processing mode control unit 330 causes the run-level mode decoding unit 310 to perform decoding processing. Further, if the size of the target block is 16 ⁇ 16 or more, the processing mode control unit 330 causes the relative position mode decoding unit 320 to perform decoding processing.
  • processing mode control unit 330 may cause the run-level mode decoding unit 310 to perform decoding processing if the number of non-zero coefficients to be decoded in the target block is a predetermined number, for example, 64 or less. .
  • all non-zero coefficients may be expressed by specifying relative positions.
  • the relative position mode decoding unit 330 decodes all non-zero coefficients.
  • processing mode control unit 330 may change the predetermined area according to the slice type, the prediction mode, and the size of the target block.
  • the processing mode control unit 330 sets an 8 ⁇ 8 size area on the low frequency component side as a predetermined area. For example, when inter prediction is encoded in the target block, the processing mode control unit 330 sets a 4 ⁇ 4 size area on the low frequency component side as a predetermined area.
  • a plurality of relative position mode tables TBL32 may be prepared in accordance with the absolute position (x n , y n ) of the non-zero coefficient.
  • the relative position mode table TBL32 is preferably optimized based on the range of values that dx and dy can take at the absolute position of each non-zero coefficient.
  • the relative position mode decoding unit 420 may change the relative position mode table TBL32 to be referred to according to the absolute position (x n , y n ) of the non-zero coefficient.
  • the relative position mode table TBL42 to be referred to adaptively can be switched, so that the amount of code to be decoded can be reduced.
  • the relative position of the non-zero coefficient is expressed in the form of (dx, dy), but is not limited thereto.
  • the relative position may be expressed by a direction and a distance.
  • the relative position mode table TBL32 may be configured as shown in FIG. In the relative position mode table TBL32 shown in FIG. 20, the smaller the dx or dy value, the smaller the code number is associated with.
  • a plurality of relative position mode tables TBL32 may be prepared in accordance with the absolute position (x n , y n ) of the coefficient.
  • the relative position mode table TBL32 is preferably optimized based on a range of values that dx and dy can take at the absolute position of each coefficient.
  • FIG. 21 shows an example in which the region R2 in the target block BLK shown in FIG. 17 and FIG. 19 (described later) is further divided into three regions, a region R2a, a region R2b, and a region R2c.
  • FIG. 22 shows an example of a VLC table associated with the region R2a, the region R2b, and the region R2c.
  • (A) of the figure shows a relative position mode table Td1 to which the relative position determining unit 323 of the relative position mode decoding unit 320 refers when the value of x or y does not become a positive value greater than or equal to a predetermined value. .
  • (b) in the figure shows a relative position mode table Td2 that is referred to by the relative position determining unit 323 of the relative position mode decoding unit 320 when the value of x or y does not become a predetermined negative value or less. Show.
  • the target block BLK shown in FIG. 21 includes a region R1 and regions R2a to R2c.
  • the position of the DC coefficient of the target block BLK is represented by (0, 0) coordinate display.
  • the position of the last coefficient in the zigzag scan order that is, the position of the highest frequency component in the target block BLK is represented by (15, 15).
  • the region R1 is a square inner region having (0, 0) as the upper left vertex.
  • the region R2a is a square inner region with (8, 0) as the upper left vertex.
  • the region R2b is a square inner region with (0, 8) as the upper left vertex, and the region R2c is a square inner region with (8, 8) as the upper left vertex.
  • the code number “0” is assigned to the coefficient value “0”.
  • code numbers “1”, “2”,..., “13”, “14” are assigned to the coefficient values “ ⁇ 1”, “1”, “ ⁇ 7”, “7”, respectively. .
  • the smaller code number is assigned to the smaller absolute value of the coefficient.
  • a negative code number is assigned to a negative value that is smaller than a positive value.
  • the code values “1”, “2”,..., “13”, “14” are assigned to the coefficient values “ ⁇ 1”, “1”, “ ⁇ 7”, “7”, respectively. Yes.
  • the relative position mode table Td2 shown in FIG. 22B has the same definition as the relative position mode table Td1 in the range where the absolute value of the coefficient is “0” to “7”.
  • the relative position determination unit 323 switches the VLC table referred to as follows according to the possible range of the position (dx, dy).
  • the relative position determination unit 323 refers to the relative position mode table Td1. Further, since the range that dy can take is ⁇ 7 ⁇ dy ⁇ 15, the relative position determination unit 323 refers to the relative position mode table Td2.
  • the range that dx can take is ⁇ 7 ⁇ dx ⁇ 15, so the relative position determination unit 323 refers to the relative position mode table Td2. Since the range that dy can take is ⁇ 15 ⁇ dy ⁇ 7, the relative position determination unit 323 refers to the relative position mode table Td1.
  • the relative position determination unit 323 refers to the relative position mode table Td1. Since the range that dy can take is ⁇ 15 ⁇ dy ⁇ 7, the relative position determination unit 323 refers to the relative position mode table Td1.
  • the coding efficiency can be improved by optimizing the VLC table based on the possible range of the values of dx and dy.
  • the configuration of the video encoding device 2 will be described with reference to FIG.
  • the TU information encoding unit 280 is replaced with the TU information encoding unit shown in FIG. Change to 280A.
  • the TU information encoding unit 280A illustrated in FIG. 18 will be described as follows. That is, as shown in FIG. 18, the TU information encoding unit 280A includes a VLC table TBL40, a run-level mode encoding unit 410, a relative position mode encoding unit 420, and a processing mode control unit 430.
  • the VLC table TBL40 is a table in which the correspondence between each parameter and a code that is a bit string of encoded data is defined.
  • the VLC table TBL40 includes a run-level mode table TBL41 referred to by a later-described run-level mode encoding unit 410 and a relative position mode table TBL42 referred to by a relative position mode encoding unit 420.
  • the run-level mode table TBL41 can be the same as the VLC table TBL21 of the TU information encoding unit 280 shown in FIG. Therefore, the description is omitted here.
  • the definition of the relative position mode table TBL42 will be described later.
  • the run-level mode encoding unit 410 performs a run mode encoding process and a level mode encoding process under the control of the processing mode control unit 330 (hereinafter referred to as a run-level mode encoding process). Since the run mode encoding process and the level mode encoding process have already been described in the first embodiment, the description thereof is omitted here.
  • the run-level mode encoding unit 410 includes a final non-zero coefficient encoding unit 411, a run mode encoding unit 412, and a level mode encoding unit 413.
  • the last non-zero coefficient encoding unit 411, the run mode encoding unit 412, and the level mode encoding unit 413 are respectively the last non-zero coefficient encoding unit 201, the run mode encoding shown in FIG. Unit 202 and level mode encoding unit 203 have the same functions. Therefore, since the function has already been described, the description thereof is omitted here.
  • the run mode encoding unit 202 and the level mode encoding unit 203 are configured to refer to the run-level mode table TBL41.
  • the relative position mode encoding unit 420 encodes the relative position in the target block of the non-zero coefficient to be encoded, and generates coefficient encoded data.
  • the relative position mode encoding unit 420 includes a final non-zero coefficient encoding unit 421, a relative position calculation unit (relative position encoding unit) 422, and a relative position encoding unit (relative position encoding unit). 423.
  • the last non-zero coefficient encoding unit 421 encodes the last non-zero coefficient in the target block.
  • the last non-zero coefficient encoding unit 421 encodes the last non-zero coefficient last_pos, level, and sign in the target block according to the reverse zigzag scan.
  • the last non-zero coefficient encoding unit 421 converts last_pos into a coordinate display (lastx, lasty) having the DC coefficient as the origin (0, 0).
  • a position indicated by coordinate display with the DC coefficient as the origin (0, 0) is referred to as an absolute position.
  • the relative position calculation unit 422 calculates the relative position of the coefficient to be encoded from the absolute position of the coefficient and the absolute position of the non-zero coefficient encoded immediately before.
  • the relative position calculation unit 422 can calculate the relative position (dx, dy) of the non-zero coefficient to be encoded based on the relational expressions (1-1) to (1-3) described above.
  • the relative position encoding unit 423 refers to the relative position mode table TBL42, and in a predetermined order, the relative position (dx, dy) of the non-zero coefficient calculated by the relative position calculation unit 422 and the level of the non-zero coefficient. Coding data and sign are generated to generate coefficient encoded data.
  • the processing mode control unit 430 determines whether or not the non-zero coefficient to be encoded is located within a predetermined area, and the run-level mode encoding unit 410 performs encoding processing according to the determination result. Or the relative position mode encoding unit 420 controls whether to perform the encoding process.
  • the control method of the processing mode control unit 430 is the same as that described for the processing mode control unit 330 of the TU information decoding unit 12A, and the description thereof is omitted here.
  • FIG. 19 is a diagram illustrating an example of encoding of non-zero coefficients by specifying a relative position.
  • the target block BLK shown in FIG. 19 is illustratively 16 ⁇ 16 size, and the region R1 is an 8 ⁇ 8 size region on the low frequency component side in the target block.
  • the region R2 is a region other than the region R1 in the target block.
  • processing mode control unit 430 causes the run-level mode encoding unit 410 to execute the encoding process for the region R1, and causes the relative position mode encoding unit 420 to execute the encoding process for the region R2.
  • N non-zero coefficients C (1) to C (N) in the region R2 are detected in advance, and the last non-zero coefficient C 0 in the region R1 is detected.
  • the relative position encoding unit 423 sequentially encodes non-zero coefficients through the following steps (or may be expressed by encoding the relative positions of non-zero coefficients in a daisy chain).
  • Step [1] the relative position encoding unit 423 uses the last non-zero coefficient C 0 in the region R1 as a base point, and calculates non-zero coefficients C (1) to C (N) in the region R2 according to a predetermined selection criterion. From there, the next non-zero coefficient C 1 to be chained to the non-zero coefficient C 0 is selected.
  • the relative position encoding unit 423 selects a non-zero coefficient with a small code amount as the next non-zero coefficient C 1 .
  • the relative position encoding unit 423 repeats the selection process based on the above selection criterion based on the selected non-zero coefficient until the non-selected non-zero coefficient in the region R2 disappears.
  • ) and the Euclidean distance between C n and C n + 1 (dx 2 + dy 2) or the like may be used.
  • the position may be specified by, for example, a relative position (dx N , dy N ) from the lower right of the encoding target block, or may be specified by LastPos according to the scan order.
  • the non-zero coefficient C N does not necessarily have to be the last non-zero coefficient when the encoding target block is scanned in some scan order.
  • Step [4] The relative position encoding unit 423 performs the relative position (dx, dy) and absolute value (level) of the non-zero coefficients C N ⁇ 1 to C 0 in the reverse order to the selection step in the step [2]. ), A positive / negative sign (sign) is encoded.
  • the processing mode control unit 430 causes the run-level mode encoding unit 410 to perform encoding processing. Run - level mode encoding unit 410, run - the level mode encoding process to encode the coefficients of non-zero coefficients C 0 other region R1.
  • the relative position mode table TBL42 may be configured as shown in FIG. Since the relative position mode table shown in FIG. 20 has already been described, the description thereof is omitted here.
  • a plurality of relative position mode tables TBL42 may be prepared in accordance with the absolute positions (x n , y n ) of the coefficients.
  • the relative position mode table TBL42 is preferably optimized based on the range of values that dx and dy can take at the absolute position of each coefficient.
  • [Modification] of the video decoding device 1 according to the present embodiment can also be applied to the video encoding device 2.
  • the video decoding device 1 uses the transform coefficient from the TU information TUI of the encoded data obtained by encoding the transform coefficient obtained by frequency transforming the pixel value of the target image for each transform unit.
  • the relative position decoding unit 320 that decodes the relative position from the transform coefficient decoded immediately before the transform coefficient to be decoded, and the conversion of the conversion coefficient decoded immediately before
  • a coefficient position determination unit 323 that specifies the position of the transform coefficient to be decoded from the position in the unit and the relative position.
  • the moving image encoding apparatus 2 encodes the transform coefficient obtained by frequency-converting the pixel value of the target image for each conversion unit.
  • This is a configuration including a relative position encoding unit 423 that encodes a relative position with respect to the position of the transform coefficient encoded immediately before the position.
  • the coefficient tends to be sparse in the region on the high frequency component side. For this reason, when a run is encoded according to the scan order, the length of run tends to be very long. For this reason, there is a tendency that a large table must be used for encoding and decoding, and the amount of codes increases. Moreover, these tendencies appear more prominently as the size of the target block is larger.
  • the size of the VLC table representing the combination of ⁇ run, level ⁇ is basically proportional to the maximum value of the run length in the scan order, that is, the area of the target block.
  • Embodiment 3 Still another embodiment of the present invention will be described with reference to FIGS. 23 to 25 as follows. For convenience of explanation, members having the same functions as those in the drawings described in the first embodiment are denoted by the same reference numerals and description thereof is omitted.
  • processing B “region division A method of performing the encoding process / decoding process while switching between the process of encoding / decoding coefficients (see S10 in FIG. 4) will be described.
  • the encoded data is encoded with an encoding method identifier for switching between process A and process B. That is, in the following example, one of process A and process B is designated as the encoding method identifier.
  • the TU information decoding unit 12 performs the process B, in the following example, the TU information decoding unit 12 executes the process A in addition to the process B.
  • the size of the target block is assumed to be 16 ⁇ 16 as an example.
  • FIG. 23 is a flowchart illustrating an example of the flow of processing for encoding / decoding while switching between processing A and processing B.
  • FIG. 23 the encoding process of the moving image encoding apparatus 2 and the decoding process of the moving image decoding apparatus 1 are shown together. In the following, the operation on the video decoding device 1 side will be described, but the operation on the video encoding device 2 side is also substantially the same.
  • the TU information decoding unit 12 of the video decoding device 1 determines the encoding method identifier.
  • the TU information decoding unit 12 determines an encoding method identifier (S101).
  • the TU information decoding unit 12 executes process A (S102). That is, the TU information decoding unit 12 decodes only the upper left 64 coefficients located on the low frequency component side of the target block.
  • the coefficients to be decoded in the process A will be described in detail with reference to FIG.
  • the coefficient to be decoded in process A may be changed depending on whether the target block is a block that performs intra prediction or a block that performs inter prediction.
  • the TU information decoding unit 12 sets a coefficient located in the upper left 8 ⁇ 8 size region RInter on the low frequency component side as a decoding target. .
  • the TU information decoding unit 12 sets the coefficient located in the region RIntra illustrated in (b) of FIG. 24 as a decoding target. That is, the TU information decoding unit 12 sets decoding targets from the first DC coefficient to the 64th coefficient in the zigzag scan order.
  • the number “64” shown in the region RIntra indicates the number of coefficients included in the region.
  • the shape of the region RIntra is shown by a right triangle. However, in reality, strictly speaking, it should be noted that the region RIntra does not become a right triangle. I want to be.
  • the TU information decoding unit 12 executes process B (S10, see FIG. 4).
  • FIG. 25 is a diagram illustrating a data structure of coefficient encoded data.
  • the encoding method identifier FLG is a flag in which the process A or the process B is designated as described above.
  • the coefficient encoded data can adopt the data structure shown in DATA1 (hereinafter referred to as coefficient encoded data DATA1).
  • the encoded data DATA1 includes 64 upper left coefficient data (run, level, sign) located on the low frequency component side of the target block.
  • the coefficient encoded data can adopt the data structure shown in DATA2 (hereinafter referred to as coefficient encoded data DATA2).
  • the non-zero coefficient flag indicates the presence or absence of non-zero coefficients in the decoding area as described above. Since the non-zero coefficient flag is encoded by the number of decoding regions, it is indicated by “ ⁇ n”.
  • non-zero coefficient flag For example, if the non-zero coefficient flag is “1 (true)”, it indicates that there is a non-zero coefficient in the decoding area, and if the non-zero coefficient flag is “0 (false)”, the decoding area is not non-zero. Indicates that there is no zero coefficient.
  • the coefficient data [region x] is omitted.
  • the unit for encoding the encoding scheme identifier is arbitrary.
  • encoding can be performed in units of LCUs.
  • the encoding method identifier can be omitted because the processing method to be executed implicitly is determined.
  • block attributes and states, and predetermined parameters can be used.
  • the processing A can be executed if it is 16 ⁇ 16, and the processing B can be executed if it is 32 ⁇ 32.
  • the determination may be made in the prediction mode of the target block.
  • the process A can be executed in the intra prediction mode
  • the process B can be executed in the inter prediction mode.
  • the determination may be made based on the conversion unit size of the adjacent block of the target block. For example, if the conversion unit of the left adjacent block of the target block is equal to or larger than a predetermined size (for example, 32 ⁇ 32 size), the process A is executed in the target block, and the process B is executed otherwise. It may be configured.
  • a predetermined size for example, 32 ⁇ 32 size
  • determination conditions such as “more than the size of the target block”, “less than a predetermined size”, and “equal to a predetermined size” can also be applied.
  • the process B is executed when the size of the conversion unit is small.
  • predetermined parameters other than the encoding method identifier may be used for determining the processing method.
  • the processing method may be switched based on the profile identifier added to the header or the like. Also, the processing method may be switched based on the level of the decoder that is added to the header or the like and the level that defines the complexity of the bit stream.
  • process B is the process of dividing and coding / decoding coefficients shown in FIG. 4, but is not limited thereto.
  • the process B may be the process S20 shown in FIG. 16 for encoding / decoding a coefficient by specifying a relative position.
  • the moving image encoding apparatus 2 may try the encoding of the processing S10 and the processing S20, or may estimate the amount of code and specify a processing with high encoding efficiency as the encoding method identifier.
  • Embodiment 4 Still another embodiment of the present invention will be described with reference to FIGS. 26 to 36 as follows. For convenience of explanation, members having the same functions as those in the drawings described in the first embodiment are denoted by the same reference numerals and description thereof is omitted.
  • the video encoding device 2 may hierarchically divide the target block BLK in encoding the target block BLK.
  • the TU information encoding unit 280 of the video encoding device 2 repeats the following processes ENC and DIV.
  • Process ENC Encode without dividing the area to be processed
  • Process DIV Divide the area to be processed, and set each area obtained by the division as the next area to be processed.
  • the target block is the first area to be processed.
  • the TU information encoding unit 280 selects one of the processes P and Q that improves the encoding efficiency as a whole.
  • the TU information encoding unit 280 compares the code amount between the case where only the process ENC is executed for the processing target and the case where the process ENC is executed after trying the process DIV.
  • the processing procedure can be determined so as to reduce the number.
  • FIG. 26 exemplifies the case where the division of two layers is performed.
  • the target block BLK includes first layer coding regions R10, R20, R30, and R40 when one layer is divided.
  • first layer coding regions R30 and R40 are divided into two layers. That is, the first layer coding region R30 includes second layer coding regions R31 to R34, and the first layer coding region R40 includes second layer coding regions R41 to R44.
  • the TU information encoding unit 280 divides the target block BLK shown in FIG. 26 by the following process.
  • the TU information encoding unit 280 uses the target block BLK as a processing target area, and executes processing ENC and processing DIV. That is, the area encoding unit 282 executes the process ENC, and the area encoding unit 282 stores the code amount by the process ENC for the target block BLK. In addition, the region dividing unit 281 performs the process DIV on the target block BLK to obtain the first hierarchical coding regions R10 to R40.
  • Step [2] The region encoding unit 282 executes the processing ENC for each of the first layer encoding regions R10 to R40 and stores the code amount of the processing ENC for the entire first layer encoding regions R10 to R40.
  • Step [3] The TU information encoding unit 280 compares the code amount of the process [1] with the code amount of the step [2]. Here, it is assumed that the code amount of the process [2] is smaller.
  • Step [4] The region dividing unit 281 further executes the process DIV on the first layer coding regions R10 to R40 to obtain the second layer coding region.
  • Step [5] The TU information encoding unit 280 compares the code amount obtained by encoding up to the first layer with the code amount obtained by encoding up to the second layer, and adopts the processing procedure with the smaller code amount. .
  • Step [6] Based on the comparison result of step [5], the first layer coding regions R10 and R20 have a larger amount of code when they are coded up to the second layer, and the first layer coding regions R30 and R40. For, it is assumed that the amount of code is smaller when encoding up to the second layer.
  • the first layer coding regions R10 and R20 and the second layer coding regions R31 to R34 and R41 to R44 are finally determined.
  • the region encoding unit 282 encodes a division flag indicating the state of division by the region dividing unit 281.
  • the region encoding unit 282 encodes the division flag “1” for the region for which the division by the region dividing unit 281 is determined. In addition, the region encoding unit 282 encodes the division flag “0” for a region that is determined not to be divided by the region dividing unit 281.
  • the shaded area indicates that there is even one non-zero coefficient, and the area that is not shaded is an area that does not have any non-zero coefficient. Is shown.
  • the region encoding unit 282 may encode a non-zero coefficient flag indicating the presence or absence of a non-zero coefficient in the processing target region. For example, when there is a non-zero coefficient, the non-zero coefficient flag is “1”, and when there is no non-zero coefficient, the non-zero coefficient flag is “0”.
  • FIG. 27 shows a representation example (quadtree representation) of the flag tree FT representing the division status and coefficient distribution status of the target block BLK shown in FIG.
  • the flag tree FT has a hierarchical structure of ROOT level, LEVEL 1 and LEVEL 2.
  • the ROOT level and LEVEL1 of the flag tree FT correspond to the division flag.
  • LEVEL2 of the flag tree FT corresponds to a non-zero coefficient flag.
  • the division flag FRoot of the target block BLK is encoded.
  • the division flags F10, F20, F30, and F40 of the first layer encoding regions R10, R20, R30, and R40 are encoded.
  • LEVEL2 which is a leaf (terminal node) of the flag tree FT
  • a non-zero coefficient flag is encoded.
  • the circled leaves indicate the non-zero coefficient flags of the second layer coding regions R41, R42, R43, and R44 in order from the left, and “1”, “0”, “ “0” and “1” are encoded.
  • the length of the run is averagely shortened, so that the coding efficiency tends to be improved without dividing.
  • FIG. 28 is a flowchart illustrating an example of the flow of the region decoding process S200 in the video decoding device 1.
  • each procedure included in the region decoding process S200 is executed with the target block as the processing target region.
  • the region dividing unit 121 refers to the division flag to determine whether or not the region to be processed is divided (S201).
  • the coefficient decoding process S210 is executed.
  • the region decoding unit 122 determines whether there is a non-zero coefficient in the region by referring to the non-zero coefficient flag (S211).
  • the area decoding unit 122 collectively decodes all the coefficients included in the processing target area by the run-level mode decoding process (S212).
  • the area decoding unit 122 skips the decoding process of the processing target area.
  • the coefficient decoding process S210 ends, the process area decoding process ends, and the area decoding process S200 for the next process target area is executed.
  • the region decoding unit 122 decodes the division flag encoded for each region (S221), and enters the loop LP200 for each divided region.
  • the area decoding process S200 is recursively executed (S223), and the process returns to the top of the loop LP200 (from S224 to S223). Thereafter, the processes in the loop are sequentially executed for each divided area. Is done.
  • the loop LP200 is exited, and the divided area decoding process S220 ends. Thereafter, the decoding process S200 for the area being executed ends.
  • the region decoding process S200 is recursively called, control is returned to the caller.
  • the division flag may be always determined to be false in a block having a size that is not divided any more. Further, the non-zero coefficient flag may always be determined to be true when the processing target area is the target block itself.
  • Example of storing flags in a distributed manner A data structure in the case of storing flags in a distributed manner will be described with reference to FIG. As shown in FIG. 29, a division flag FRoot is stored at the head of the coefficient encoded data. If the division flag FRoot is “0”, the coefficient encoded data adopts the data structure shown in DATA 11 as an example (hereinafter referred to as coefficient encoded data DATA 11).
  • the encoded data DATA11 includes 16 ⁇ 16 coefficient data (run, level, sign) of the target block.
  • the coefficient encoded data adopts the data structure shown in DATA12 as an example (hereinafter referred to as coefficient encoded data DATA12).
  • the coefficient encoded data DATA12 includes area information [area 1] F1 to area information [area n] Fn.
  • the “region n” indicates a “region” in the first hierarchy. In the example using the target block BLK shown in FIG. 26, the regions R10 to R40 correspond to the “region”.
  • the area information [area 1] F1 includes a division flag [area 1] F10.
  • the non-zero coefficient flag and the coefficient data include coefficient information F12 when the division flag [region 1] F10 is “0”, and the division flag [region 1] F10 is “1”. Includes coefficient information F11.
  • the coefficient information F11 includes a non-zero coefficient flag [1-1], coefficient data [1-1] to non-zero coefficient flag [1-n], and coefficient data [1-n].
  • FIG. 30 is a diagram illustrating an example of coefficient encoded data indicating the target block BLK described with reference to FIG.
  • quadtree partitioning is designated at the ROOT level, so the partitioning flag Root is “1”.
  • the coefficient encoded data DATA12 includes region information F1 to F4 corresponding to the decoding regions R10 to R40 included in the target block BLK.
  • Data is stored in the order of area information F1, F2, F3 and F4.
  • area information F1 to F4 will be described in order.
  • the area information F1 includes coefficient data [1].
  • the data included in the area information F3 will be described as follows.
  • the decoding regions R31 and R33 include non-zero coefficients, and the decoding regions R32 and R34 do not include any non-zero coefficients.
  • the non-zero coefficient flag [3-1] 1 and coefficient data [3-1] are included for the decoding area R31. The same applies to the decoding area R33.
  • the data included in the area information F4 will be described as follows.
  • the decoding regions R41 and R44 include non-zero coefficients, and the decoding regions R42 and R43 do not include any non-zero coefficients.
  • the non-zero coefficient flag [4-1] 1 and coefficient data [4-1] are included for the decoding area R41. The same applies to the decoding area R44.
  • the encoded data DATA21 includes 16 ⁇ 16 coefficient data (run, level, sign) of the target block.
  • the coefficient encoded data DATA22 includes coefficient data [area 1] to coefficient data [area n].
  • the “region n” indicates a “region” that is not further divided.
  • the region R10, the region R31, and the like correspond to the “region”.
  • FIG. 32 is a diagram illustrating an example of coefficient encoded data indicating the target block BLK described with reference to FIG.
  • a flag tree FT is stored in the coefficient encoded data DATA22.
  • a division flag FRoot is first stored. Subsequently, flags relating to the decoding regions R10 to R40 are stored in order.
  • coefficient data of each decoding area is stored after the flag tree FT.
  • Coefficient data [1], coefficient data [3-1], coefficient data [3-3], coefficient data [4-1] corresponding to decoding regions R10, R31, R33, R41, and R44 having one or more coefficient data ] And coefficient data [4-4] are stored in order.
  • the target block BLK is 16 ⁇ 16 size by way of example. It is assumed that the first layer coding region R30 has an 8 ⁇ 8 size, and the second layer coding regions R31 to R34 have a 4 ⁇ 4 size.
  • the TU information encoding unit 280 performs region division and determination of the presence / absence of a non-zero coefficient when the encoding region to be processed is a predetermined size (for example, 8 ⁇ 8 size) or less. Encoding is performed in 8 ⁇ 8 units.
  • FIG. 33 is a diagram showing in detail the first layer coding region R30 of FIG. As shown in FIG. 33, region dividing section 281 divides first layer coding region R30 into second layer coding regions R31 to R34.
  • the region encoding unit 282 encodes the division flag “1” and further encodes the non-zero coefficient flag “1001” for the first layer encoding region R30.
  • the region encoding unit 282 performs the encoding process using the first layer encoding region R30 as the target encoding region.
  • the arrows shown throughout the first layer coding region R30 indicate the scan order.
  • the run mode encoding unit 202 does not count run in the encoding target regions R32 and R34 having the non-zero coefficient flag “0 (false)”.
  • the run mode encoding unit 202 performs the run mode encoding process as follows.
  • the run mode encoding unit 202 reads the coefficients in the reverse zigzag scan order indicated by the arrows in FIG.
  • the reverse zigzag scan order and the next non-zero coefficient are A3 (hereinafter referred to as non-zero coefficient A3).
  • non-zero coefficient A3 There are nine zero coefficients between the non-zero coefficient A1 and the non-zero coefficient A3.
  • the run mode encoding unit 202 may determine whether or not the non-zero coefficient flag is “0” from the two-dimensional coordinate value of each coefficient.
  • the TU information decoding unit 12 may decode the coefficients by dividing the predetermined decoding target area as determined in advance.
  • FIG. 35 and FIG. 36 show the division method in the target block BLK of 16 ⁇ 16 size.
  • the target block BLK is divided into first layer decoding regions R101 to R104.
  • the first hierarchy decoding area that can be divided up to the second hierarchy decoding area is indicated by a dotted line (for example, first hierarchy decoding areas R102 to R104 in FIG. 35).
  • the region dividing unit 121 of the TU information decoding unit 12 may be set to always divide a predetermined decoding target region (can be divided), or may be set not to always divide (division). Impossible).
  • the upper left 8 ⁇ 8 first layer decoding region R101 may not be divided.
  • the region dividing unit 121 can divide the first hierarchy decoding regions R102 to R104.
  • the encoding efficiency on the moving image encoding device 2 side tends to be improved when not dividing than when dividing.
  • the region dividing unit 121 can divide the first layer decoding region R104.
  • the region on the high frequency component side is likely to have zero coefficients, so that the division efficiency tends to improve on the moving image coding device 2 side.
  • the area dividing unit 121 may always divide a decoding area larger than a predetermined size.
  • the area dividing unit 281 may not always divide a decoding area having a size smaller than a predetermined size. For example, it can be configured as follows.
  • the area dividing unit 121 may always divide a decoding area having a size larger than 8 ⁇ 8 size.
  • the area dividing unit 121 may always set the highest-level (that is, target block) division flag to “1”.
  • the region dividing unit 121 may not further divide a 4 ⁇ 4 size region included in a target block having a size of 4 ⁇ 4 size or more.
  • the area dividing unit 121 may not perform division of three or more layers.
  • the present modification can also be applied to the TU information encoding unit 280 of the moving image encoding device 2.
  • the selection of the encoding method shown in the third embodiment and the area division according to the regulations shown in the fourth embodiment may be applied only to a target block having a specific size and shape or a specific slice type.
  • B slices tend to have a relatively large number of zero coefficients. Therefore, in the B slice, only the 64 upper left coefficients of the target block may be encoded without performing division.
  • the coefficients may be encoded by the relative position designation shown in the second embodiment.
  • the present invention can also be expressed as follows.
  • the transform unit dividing means divides the transform unit into a plurality of sub-units, and the transform coefficient decoding means is decoding information for obtaining the transform coefficients from the encoded data. Then, the transform coefficient included in the sub unit is decoded with reference to the decoding information assigned to the sub unit.
  • the length of the zero coefficient continuous in a predetermined scan order is encoded.
  • a shorter code is assigned as the length of the continuous zero coefficient is shorter.
  • a shorter code is assigned as the length of consecutive zero coefficients is shorter, so that an efficient decoding process can be performed in consideration of the tendency according to the position of the region. Thereby, the code amount to decode can be reduced.
  • a shorter code is assigned to the parameter set having the absolute value of 1 for the parameter set including the absolute value of the transform coefficient. Yes.
  • the absolute value of the conversion coefficient tends to decrease overall. For this reason, if the conversion coefficient is not zero, the absolute value tends to be 1.
  • a shorter code can be assigned to a set of parameters whose appearance frequency is high at a position on the high frequency component side.
  • the parameter set is, for example, the ⁇ run, level ⁇ set in the run mode, and the absolute value here corresponds to level.
  • the order is specified according to the length of the code
  • the decoding information updating unit moves up the order according to the position of the subunit in the conversion unit.
  • the order determined according to the length of the code assigned to the parameter is advanced according to the position of the subunit in the conversion unit.
  • the order determined according to the length of the code refers to, for example, a code number. That is, in the above configuration, the code number assigned to the parameter is updated to a shorter one by incrementing the code number.
  • the code length can be updated by a relatively simple procedure of increasing the code number.
  • the conversion unit dividing means performs the recursive division according to at least one of the position and size of the area to be divided.
  • the image decoding apparatus uses the encoded data obtained by encoding the transform coefficient obtained by frequency-converting the pixel value of the target image for each transform unit.
  • transform unit dividing means for dividing the transform unit into a plurality of sub-units, and decoding information for obtaining the transform coefficients from the encoded data, each sub-unit
  • transform coefficient decoding means for decoding transform coefficients included in the sub-unit with reference to decoding information assigned to the sub-unit.
  • an image encoding apparatus is an image encoding apparatus that encodes a transform coefficient obtained by frequency-converting a pixel value of a target image for each transform unit.
  • a transform unit dividing unit that divides a transform unit into a plurality of subunits, and encoding information for encoding the transform coefficient, with reference to the encoding information assigned to each subunit, Transform coefficient coding means for coding the transform coefficient included in the transform unit.
  • the data structure of the encoded data according to the present invention is generated by encoding a transform coefficient obtained by frequency-converting the pixel value of the target image for each transform unit in order to solve the above problem.
  • an image decoding device that includes a relative position with respect to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded, thereby decoding the encoded data.
  • This is a data structure that specifies the position of the transform coefficient to be decoded from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the conversion unit to be decoded is divided into a plurality of sub-units.
  • the conversion unit is a unit for converting pixel values into the frequency domain. Examples of the conversion unit include a size of 64 ⁇ 64 pixels, 32 ⁇ 32 pixels, and 16 ⁇ 16 pixels.
  • the sub unit may be, for example, an 8 ⁇ 8 size area.
  • a plurality of subunits obtained by division are processed one by one, and transform coefficients included in the subunit are decoded.
  • the decoding process can be performed in any order.
  • the decoding information assigned to each of the plurality of sub-units is referred to when transform coefficients are decoded.
  • Decoding information is information for reproducing a predetermined parameter value of a transform coefficient from a code (bit string) of encoded data.
  • the decoding information is a table indicating association for reproducing a predetermined parameter value of the transform coefficient from the code of the encoded data.
  • the decoding information is a calculation formula for deriving a predetermined parameter value of the transform coefficient from the code of the encoded data.
  • the transform coefficient is decoded using the decoding information defined for a sub-unit smaller than the size of the original transform unit.
  • the size of the scan table that defines the scan order of the transform coefficients can also be reduced.
  • the amount of memory and processing capacity required for the decoding process can be kept low.
  • the sub unit may coincide with any of the encoding units in the techniques of Non-Patent Documents 1 and 2.
  • a VLC table defined in advance in the coding unit that is, decoding information can be used.
  • the transform coefficient decoding means refers to non-zero information indicating the presence or absence of a non-zero transform coefficient in the sub unit, and the non-zero information is a non-zero information in the sub unit.
  • the transform coefficient decoding means refers to non-zero information indicating the presence or absence of a non-zero transform coefficient in the sub unit, and the non-zero information is a non-zero information in the sub unit.
  • the decoding information is adaptively defined according to the position of the subunit in the conversion unit.
  • the appearance tendency of the value of the conversion coefficient is different between the low frequency component side including the DC component and the high frequency component side. For example, a non-zero coefficient is likely to appear on the low frequency component side near the DC component. Further, there is a high possibility that a zero coefficient appears on the high frequency component side.
  • the adaptive definition depending on the position means, for example, that the decoding information is adaptively defined depending on whether the position of the subunit is on the low frequency component side or the low frequency component side. .
  • adaptive means that a code is assigned according to the above-mentioned appearance tendency. For example, on the high frequency component side, a shorter code is assigned to the zero coefficient or to the one having a small absolute value of the conversion coefficient.
  • a shorter code is assigned to a non-zero coefficient or a coefficient having a large absolute value of a transform coefficient.
  • the image decoding apparatus further includes decoding information updating means for updating a code assigned to the parameter in the decoding information to a shorter one according to the appearance frequency of the parameter indicating the transform coefficient. Is preferred.
  • a shorter code can be assigned to a parameter having a high appearance frequency. That is, the appearance frequency of the parameter can be dynamically reflected in the shortness of the code.
  • the transform coefficient decoding means performs a first mode decoding procedure for decoding a length of continuous non-zero coefficients, an absolute value of the transform coefficient, and a sign of the transform coefficient under a predetermined condition. After performing below, it is preferable to perform a decoding process that executes a second mode decoding procedure for decoding the absolute value of the transform coefficient and the sign of the transform coefficient.
  • the first mode decoding procedure for decoding the length (run) of continuous non-zero coefficients, the absolute value of the transform coefficient (level), and the sign (sign) of the transform coefficient is performed.
  • a second mode decoding procedure for decoding the absolute value (level) of the transform coefficient and the sign (sign) of the transform coefficient is executed.
  • the first mode decoding procedure is a so-called run mode
  • the second mode decoding procedure is a so-called level mode.
  • the predetermined condition includes, for example, the number of decoded transform coefficients, the absolute value of the transform coefficients, and the like.
  • the predetermined condition may be a condition corresponding to the position of the sub-unit in the conversion unit or the appearance tendency of the coefficient.
  • the run mode and the level mode are technologies adopted in Non-Patent Documents 1 and 2, for example. Such a conventional configuration can be used in the decoding process in each region. Thereby, high encoding efficiency is realizable.
  • the transform coefficient decoding unit changes the predetermined condition according to the position of the subunit in the transform unit.
  • the predetermined condition is changed according to the position of the subunit in the conversion unit. This is to make it easier to end the first mode decoding procedure.
  • the first mode decoding procedure is preferentially used when the length of consecutive non-zero coefficients becomes long.
  • changing the predetermined condition includes executing only the first mode decoding procedure and not executing the second mode decoding procedure.
  • limited region decoding means for decoding transform coefficients only in a predetermined region on the low frequency component side in the transform unit, decoding processing by the transform unit dividing means and transform coefficient decoding means,
  • switching means for switching between decoding processing by the limited area decoding means is provided.
  • any one of the decoding processing method by the transform unit dividing unit and the transform coefficient decoding unit and the decoding processing method for decoding the transform coefficient only in a predetermined region on the low frequency component side in the transform unit It is possible to appropriately switch to an efficient decoding processing method.
  • the conversion unit dividing unit recursively divides the plurality of divided areas.
  • the amount of code to be decoded is smaller when the decoding process is performed on a smaller size area. According to the above configuration, when the decoding process is performed for a smaller size area, the decoding process can be efficiently performed when the amount of code to be decoded is small.
  • the image decoding apparatus uses the encoded data obtained by encoding the transform coefficient obtained by frequency-converting the pixel value of the target image for each transform unit.
  • relative position decoding means for decoding a relative position from the transform coefficient decoded immediately before the transform coefficient to be decoded, and the transform unit of the transform coefficient decoded immediately before
  • position specifying means for specifying the position of the transform coefficient to be decoded from the relative position and the relative position.
  • an image encoding apparatus is an image encoding apparatus that encodes a transform coefficient obtained by frequency-converting a pixel value of a target image for each transform unit. It is a structure provided with the relative position encoding means which encodes the relative position with respect to the position of the said conversion coefficient encoded immediately before the position of the said conversion coefficient used as conversion object.
  • the data structure of the encoded data according to the present invention is generated by encoding the transform coefficient obtained by converting the pixel value of the target image into the frequency transform for each transform unit in order to solve the above problem.
  • the image decoding includes a relative position to the position of the transform coefficient encoded immediately before the position of the transform coefficient to be encoded, thereby decoding the encoded data
  • the apparatus has a data structure that specifies the position of the transform coefficient to be decoded from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the position of the transform coefficient to be decoded is specified from the position in the transform unit of the transform coefficient decoded immediately before and the relative position.
  • the position of the conversion coefficient can be specified in a chain manner based on the relative position.
  • the conversion unit is a predetermined unit for conversion.
  • the length of the run is counted according to a predetermined scan order, so that the relative position of the two-dimensional coordinate in the conversion unit between the reference non-zero coefficient and the next non-zero coefficient is Even if they are close to each other, as a result, the run may become longer, which may increase the code amount.
  • the code amount can be reduced in such a case.
  • the amount of memory and processing capacity required for the decoding process can be kept low.
  • the first mode decoding process for decoding the length of the continuous non-zero coefficient, the absolute value of the transform coefficient, and the sign of the transform coefficient for the region on the low frequency component side in the transform unit.
  • a decoding means for executing a second mode decoding process for decoding the absolute value of the transform coefficient and the sign of the transform coefficient.
  • the above decoding means performs so-called run-level mode decoding.
  • the region on the low frequency component side in the conversion unit includes, for example, an upper left 8 ⁇ 8 size region including a DC component if the conversion unit has a size of 16 ⁇ 16.
  • the conversion coefficient is not so sparse, so the length of run is also relatively short. Therefore, decoding in the run-level mode can be performed efficiently.
  • the decoding unit changes the size of the region according to the characteristics of the conversion unit to be decoded.
  • the characteristics of the conversion unit include the slice type of the conversion unit, the prediction mode, and the conversion unit size.
  • the above various characteristics may be used in combination or alternatively.
  • the size of the region may be changed according to at least one of the slice type of the conversion unit, the prediction mode, the conversion unit size, and the like.
  • the size of the conversion unit is 16 ⁇ 16
  • it can be configured as follows. That is, if the prediction mode is the intra mode, the size of the region on the low frequency component side is set to 8 ⁇ 8. If the prediction mode is the inter mode, the size of the region on the low frequency component side is set to 4 ⁇ 4.
  • Each block of the above-described moving picture decoding apparatus 1 and moving picture encoding apparatus 2 may be realized in hardware by a logic circuit formed on an integrated circuit (IC chip), or may be a CPU (Central Processing). Unit) may be implemented in software.
  • IC chip integrated circuit
  • CPU Central Processing
  • Unit Central Processing Unit
  • each device includes a CPU that executes instructions of a program that realizes each function, a ROM (Read (Memory) that stores the program, a RAM (Random Memory) that expands the program, the program, and various types
  • a storage device such as a memory for storing data is provided.
  • An object of the present invention is to provide a recording medium in which a program code (execution format program, intermediate code program, source program) of a control program of each of the above devices, which is software that realizes the above-described functions, is recorded so as to be readable by a computer. This can also be achieved by supplying to each of the above devices and reading and executing the program code recorded on the recording medium by the computer (or CPU or MPU).
  • Examples of the recording medium include tapes such as magnetic tape and cassette tape, magnetic disks such as floppy (registered trademark) disks / hard disks, and CD-ROM / MO / MD / DVD / CD-R / Blu-ray disks (registered trademarks). ) And other optical disks, IC cards (including memory cards) / optical cards, semiconductor memories such as mask ROM / EPROM / EEPROM / flash ROM, PLD (Programmable logic device) and FPGA ( Logic circuits such as Field Programmable Gate Array can be used.
  • tapes such as magnetic tape and cassette tape
  • magnetic disks such as floppy (registered trademark) disks / hard disks
  • CD-ROM / MO / MD / DVD / CD-R / Blu-ray disks registered trademarks
  • IC cards including memory cards
  • semiconductor memories such as mask ROM / EPROM / EEPROM / flash ROM, PLD (Programmable logic device) and FPGA ( Logic circuits
  • each of the above devices may be configured to be connectable to a communication network, and the program code may be supplied via the communication network.
  • the communication network is not particularly limited as long as it can transmit the program code.
  • the Internet intranet, extranet, LAN, ISDN, VAN, CATV communication network, virtual private network, telephone line network, mobile communication network, satellite communication network, and the like can be used.
  • the transmission medium constituting the communication network may be any medium that can transmit the program code, and is not limited to a specific configuration or type.
  • wired lines such as IEEE 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, infrared rays such as IrDA and remote control, Bluetooth (registered trademark), IEEE 802.11 wireless, HDR ( It can also be used by radio such as High Data Rate (NFC), Near Field Communication (NFC), Digital Living Network Alliance (DLNA), mobile phone network, satellite line, and digital terrestrial network.
  • the present invention can also be realized in the form of a computer data signal embedded in a carrier wave in which the program code is embodied by electronic transmission.
  • the above-described moving image encoding device 2 and moving image decoding device 1 can be used by being mounted on various devices that perform transmission, reception, recording, and reproduction of moving images.
  • the moving image may be a natural moving image captured by a camera or the like, or may be an artificial moving image (including CG and GUI) generated by a computer or the like.
  • moving image encoding device 2 and moving image decoding device 1 can be used for transmission and reception of moving images.
  • FIG. 37 (a) is a block diagram showing a configuration of a transmission apparatus PROD_A in which the moving picture encoding apparatus 2 is mounted.
  • the transmission device PROD_A modulates a carrier wave with an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, and the encoded data obtained by the encoding unit PROD_A1.
  • a modulation unit PROD_A2 that obtains a modulation signal and a transmission unit PROD_A3 that transmits the modulation signal obtained by the modulation unit PROD_A2 are provided.
  • the moving image encoding apparatus 2 described above is used as the encoding unit PROD_A1.
  • the transmission device PROD_A is a camera PROD_A4 that captures a moving image, a recording medium PROD_A5 that records the moving image, an input terminal PROD_A6 that inputs the moving image from the outside, as a supply source of the moving image input to the encoding unit PROD_A1.
  • An image processing unit A7 that generates or processes an image may be further provided.
  • FIG. 37A illustrates a configuration in which the transmission apparatus PROD_A includes all of these, but a part of the configuration may be omitted.
  • the recording medium PROD_A5 may be a recording of a non-encoded moving image, or a recording of a moving image encoded by a recording encoding scheme different from the transmission encoding scheme. It may be a thing. In the latter case, a decoding unit (not shown) for decoding the encoded data read from the recording medium PROD_A5 according to the recording encoding method may be interposed between the recording medium PROD_A5 and the encoding unit PROD_A1.
  • FIG. 37 (b) is a block diagram showing a configuration of a receiving device PROD_B in which the moving image decoding device 1 is mounted.
  • the receiving device PROD_B includes a receiving unit PROD_B1 that receives the modulated signal, a demodulating unit PROD_B2 that obtains encoded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a demodulator.
  • a decoding unit PROD_B3 that obtains a moving image by decoding the encoded data obtained by the unit PROD_B2.
  • the moving picture decoding apparatus 1 described above is used as the decoding unit PROD_B3.
  • the receiving device PROD_B has a display PROD_B4 for displaying a moving image, a recording medium PROD_B5 for recording the moving image, and an output terminal for outputting the moving image to the outside as a supply destination of the moving image output by the decoding unit PROD_B3.
  • PROD_B6 may be further provided.
  • FIG. 37 (b) a configuration in which all of these are provided in the receiving device PROD_B is illustrated, but a part thereof may be omitted.
  • the recording medium PROD_B5 may be used for recording a non-encoded moving image, or may be encoded using a recording encoding method different from the transmission encoding method. May be. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the recording encoding method may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.
  • the transmission medium for transmitting the modulation signal may be wireless or wired.
  • the transmission mode for transmitting the modulated signal may be broadcasting (here, a transmission mode in which the transmission destination is not specified in advance) or communication (here, transmission in which the transmission destination is specified in advance). Refers to the embodiment). That is, the transmission of the modulation signal may be realized by any of wireless broadcasting, wired broadcasting, wireless communication, and wired communication.
  • a terrestrial digital broadcast broadcasting station (broadcasting equipment or the like) / receiving station (such as a television receiver) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives a modulated signal by wireless broadcasting.
  • a broadcasting station (such as broadcasting equipment) / receiving station (such as a television receiver) of cable television broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives a modulated signal by cable broadcasting.
  • a server workstation etc.
  • Client television receiver, personal computer, smart phone etc.
  • VOD Video On Demand
  • video sharing service using the Internet is a transmitting device for transmitting and receiving modulated signals by communication.
  • PROD_A / reception device PROD_B usually, either a wireless or wired transmission medium is used in a LAN, and a wired transmission medium is used in a WAN.
  • the personal computer includes a desktop PC, a laptop PC, and a tablet PC.
  • the smartphone also includes a multi-function mobile phone terminal.
  • the video sharing service client has a function of encoding a moving image captured by the camera and uploading it to the server. That is, the client of the video sharing service functions as both the transmission device PROD_A and the reception device PROD_B.
  • moving image encoding device 2 and moving image decoding device 1 can be used for recording and reproduction of moving images.
  • FIG. 38 (a) is a block diagram showing a configuration of a recording apparatus PROD_C in which the above-described moving picture encoding apparatus 2 is mounted.
  • the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding a moving image, and the encoded data obtained by the encoding unit PROD_C1 on the recording medium PROD_M.
  • the moving image encoding apparatus 2 described above is used as the encoding unit PROD_C1.
  • the recording medium PROD_M may be of a type built in the recording device PROD_C, such as (1) HDD (Hard Disk Drive) or SSD (Solid State Drive), or (2) SD memory. It may be of the type connected to the recording device PROD_C, such as a card or USB (Universal Serial Bus) flash memory, or (3) DVD (Digital Versatile Disc) or BD (Blu-ray Disc: registration) Or a drive device (not shown) built in the recording device PROD_C.
  • HDD Hard Disk Drive
  • SSD Solid State Drive
  • SD memory such as a card or USB (Universal Serial Bus) flash memory, or (3) DVD (Digital Versatile Disc) or BD (Blu-ray Disc: registration) Or a drive device (not shown) built in the recording device PROD_C.
  • the recording device PROD_C is a camera PROD_C3 that captures moving images as a supply source of moving images to be input to the encoding unit PROD_C1, an input terminal PROD_C4 for inputting moving images from the outside, and reception for receiving moving images.
  • the unit PROD_C5 and an image processing unit C6 that generates or processes an image may be further provided.
  • FIG. 38A illustrates a configuration in which the recording apparatus PROD_C includes all of these, but some of them may be omitted.
  • the receiving unit PROD_C5 may receive a non-encoded moving image, or may receive encoded data encoded by a transmission encoding scheme different from the recording encoding scheme. You may do. In the latter case, a transmission decoding unit (not shown) that decodes encoded data encoded by the transmission encoding method may be interposed between the reception unit PROD_C5 and the encoding unit PROD_C1.
  • Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, and an HDD (Hard Disk Drive) recorder (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 is a main supply source of moving images).
  • a camcorder in this case, the camera PROD_C3 is a main source of moving images
  • a personal computer in this case, the receiving unit PROD_C5 or the image processing unit C6 is a main source of moving images
  • a smartphone in this case In this case, the camera PROD_C3 or the receiving unit PROD_C5 is a main supply source of moving images
  • the camera PROD_C3 or the receiving unit PROD_C5 is a main supply source of moving images
  • FIG. 38 is a block showing a configuration of a playback device PROD_D equipped with the above-described video decoding device 1.
  • the playback device PROD_D reads a moving image by decoding a read unit PROD_D1 that reads encoded data written to the recording medium PROD_M and a coded data read by the read unit PROD_D1. And a decoding unit PROD_D2 to be obtained.
  • the moving picture decoding apparatus 1 described above is used as the decoding unit PROD_D2.
  • the recording medium PROD_M may be of the type built into the playback device PROD_D, such as (1) HDD or SSD, or (2) such as an SD memory card or USB flash memory, It may be of a type connected to the playback device PROD_D, or (3) may be loaded into a drive device (not shown) built in the playback device PROD_D, such as DVD or BD. Good.
  • the playback device PROD_D has a display PROD_D3 that displays a moving image, an output terminal PROD_D4 that outputs the moving image to the outside, and a transmission unit that transmits the moving image as a supply destination of the moving image output by the decoding unit PROD_D2.
  • PROD_D5 may be further provided.
  • FIG. 38B illustrates a configuration in which the playback apparatus PROD_D includes all of these, but some of the configurations may be omitted.
  • the transmission unit PROD_D5 may transmit an unencoded moving image, or transmits encoded data encoded by a transmission encoding method different from the recording encoding method. You may do. In the latter case, it is preferable to interpose an encoding unit (not shown) that encodes a moving image with an encoding method for transmission between the decoding unit PROD_D2 and the transmission unit PROD_D5.
  • Examples of such a playback device PROD_D include a DVD player, a BD player, and an HDD player (in this case, an output terminal PROD_D4 to which a television receiver or the like is connected is a main supply destination of moving images).
  • a television receiver in this case, the display PROD_D3 is a main supply destination of moving images
  • a digital signage also referred to as an electronic signboard or an electronic bulletin board
  • the display PROD_D3 or the transmission unit PROD_D5 is the main supply of moving images.
  • Desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 is the main video image supply destination), laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 is a moving image)
  • a smartphone which is a main image supply destination
  • a smartphone in this case, the display PROD_D3 or the transmission unit PROD_D5 is a main moving image supply destination
  • the like are also examples of such a playback device PROD_D.
  • the present invention can be suitably applied to an image decoding apparatus that decodes encoded data obtained by encoding image data and an image encoding apparatus that generates encoded data obtained by encoding image data. Further, the present invention can be suitably applied to the data structure of encoded data generated by an image encoding device and referenced by the image decoding device.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

 動画像復号装置(1)は、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データのTU情報(TUI)から、該変換係数を復号する動画像復号装置(1)において、上記変換単位である対象ブロックを、複数の復号領域に分割する領域分割部(121)と、TU情報(TUI)から上記変換係数を得るための復号情報であって、上記復号領域ごとに割り当てられているVLCテーブル(TBL11)を参照して、上記復号領域に含まれる変換係数を復号する領域復号部(122)と、を備える。

Description

画像復号装置、画像符号化装置、および符号化データのデータ構造
 本発明は変換係数を復号する画像復号装置、変換係数を符号化する画像符号化装置、および変換係数が符号化された符号化データのデータ構造に関する。
 動画像を効率的に伝送または記録するために、動画像を符号化することによって符号化データを生成する動画像符号化装置、および、当該符号化データを復号することによって復号画像を生成する動画像復号装置が用いられている。
 具体的な動画像符号化方式としては、例えば、H.264/MPEG-4.AVC、VCEG(Video Coding Expert Group)における共同開発用コーデックであるKTAソフトウェアに採用されている方式、TMuC(Test Model under Consideration)ソフトウェアに採用されている方式や、および、その後継コーデックであるHEVC(High-Efficiency Video Coding)にて提案されている方式などが挙げられる(非特許文献1、2)。
 上述の動画像符号化方式では、通常、画像を所定のサイズのブロックに区切り、ブロック毎に画素値を周波数変換することによって変換係数を導出し、導出された変換係数に対して符号化処理を施す。ここで、符号化処理の対象とするブロック(符号化対象ブロック)のサイズとしては、大小様々のものが規定されているが、よりサイズが大きい符号化対象ブロックのほうが、変換係数の符号量がより多くなるという傾向がある。
 HM(HEVC TestModel)2.0では、16×16画素以上のサイズのブロックにおいては、最大でも64個の係数のみを符号化することにより符号量の低減を図る技術が提案されていた。例えば、符号化8×8以上のブロックについては、低周波成分側の最大8×8の領域までしか符号化しないという技術が提案されていた(非特許文献3)。
 しかしながら、上述の技術を採用して、高周波成分側にある係数の符号化を省略した場合、大サイズのブロックの係数符号化において画質が低下する場合があることが知られていた。
 一方で、8×8画素以上のサイズのブロックについて{run,level}の組み合わせを示す値を計算で導出することにより、符号量を低減する技術が提案されている(非特許文献4)。なお、runとは所定のスキャン順に沿って連続するゼロ係数の数(0ラン)であり、levelとは、係数の絶対値を意味する。
 以上のような背景の下、HM2.0の後継であるHM3.0では、符号量を低減するため、非特許文献4の提案を採用するとともに、画質の向上のため、16×16以上のサイズのブロック(変換単位)サイズにおいても、すべての変換係数が符号化されることとなった。
「WD2: Working Draft 2 of High-Efficiency Video Coding (JCTVC-D503) 」, Joint Collaborative Team on Video Coding (JCT-VC)of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 4th Meeting: Daegu, KR, 1/2011(2011年1月公開) 「WD3: Working Draft 3 of High-Efficiency Video Coding (JCTVC-E603)」, Joint Collaborative Team on Video Coding (JCT-VC)of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 5th Meeting: Geneva, CH, 3/2011(2011年3月公開) 「Samsung’s Response to the Call for Proposals on Video Compression Technology (JCTVC-A124)」, Joint Collaborative Team on Video Coding (JCT-VC)of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 1st Meeting: Dresden, DE, 4/2010(2010年4月公開) 「CE5: coefficient coding with LCEC for large blocks (JCTVC-E383)」, Joint Collaborative Team on Video Coding (JCT-VC)of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 5th Meeting: Geneva, CH, 3/2011(2011年3月公開)
 しかしながら、16×16画素以上等のサイズが大きいブロックを符号化処理する場合、係数を符号化するためのテーブルのサイズが大きくなったり、テーブルの値を求めるための演算量が多くなったりするという問題がある。
 例えば、16×16画素のブロックでは、係数の数が最大で256個になる。また、32×32画素のブロックでは、係数の数が最大で1024個になる。
 また、大きいブロックを符号化処理する場合、サイズが大きくなる傾向があるテーブルとしては、スキャン順を指定するスキャンテーブルや、ラン-レベル符号化のVLCテーブル等が挙げられる。
 より具体的には、スキャンテーブルは、係数の個数に比例したサイズが必要となり、ラン-レベル符号化のVLCテーブルは、最大ラン長に応じたサイズが必要となる。
 本発明は、上記の問題点に鑑みてなされたものであり、その目的は、符号化データから変換係数を得るための復号情報の情報量や、復号情報に基づく計算量を低減することができる画像復号装置、画像符号化装置、および符号化データのデータ構造を実現することにある。
 本発明の一側面について説明すると次のとおりである。すなわち、本発明に係る画像復号装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データから、該変換係数を復号する画像復号装置において、上記変換単位を、複数のサブ単位に分割する変換単位分割手段と、上記符号化データから上記変換係数を得るための復号情報であって、上記サブ単位ごとに割り当てられている復号情報を参照して、上記サブ単位に含まれる変換係数を復号する変換係数復号手段と、を備えることを特徴とする。
 また、本発明に係る画像符号化装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する画像符号化装置において、上記変換単位を、複数のサブ単位に分割する変換単位分割手段と、上記変換係数を符号化するための符号化情報であって、上記サブ単位ごとに割り当てられている符号化情報を参照して、上記変換単位に含まれる変換係数を符号化する変換係数符号化手段と、を備えることを特徴とする。
 また、本発明に係る符号化データのデータ構造は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化することにより生成される符号化データのデータ構造において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、これにより、上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定する、ことを特徴とする。
 上記構成によれば、復号処理において、まず、復号の対象となる変換単位を複数のサブ単位に分割する。
 変換単位とは、画素値を周波数領域に変換する単位である。変換単位としては、例えば、64×64画素、32×32画素や16×16画素のサイズ等が挙げられる。
 サブ単位は、変換単位が16×16サイズである場合、例えば、8×8サイズの領域であってもよい。
 また、上記構成によれば、分割により得られた複数のサブ単位を、ひとつずつ処理対象にして、該サブ単位に含まれる変換係数を復号する。サブ単位を復号する順番に特に制限はなく任意の順に復号処理を行うことができる。
 また、上記構成では、変換係数の復号に際し、複数のサブ単位のそれぞれに割り当てられている復号情報を参照する。
 復号情報とは、符号化データのコード(ビット列)から、変換係数の所定のパラメータ値を再現するための情報である。例えば、復号情報は、符号化データのコードから変換係数の所定のパラメータ値を再現するための対応付けを示すテーブルである。また、例えば、復号情報は、符号化データのコードから、変換係数の所定のパラメータ値を導出するための算出式である。
 つまり、上記構成においては、元の変換単位のサイズよりも、小さなサブ単位について規定されている復号情報を用いて、変換係数を復号する。
 このため、元の変換単位のサイズについて規定される復号情報に基づいて復号処理を行うのに比べて、復号情報の情報量や、復号情報に基づく計算量を低減することができるという効果を奏する。
 さらにいえば、復号処理において対象となる変換係数の数を小さくできるので、変換係数のスキャン順序を定義するスキャンテーブルのサイズもより小さくできる。
 また、さらに付言すれば、復号処理において必要となるメモリの量や、処理能力を低く抑えることができる。
 なお、サブ単位は、非特許文献1,2の技術における符号化単位のいずれかと一致していてもよい。この場合、上記符号化単位において予め定義されているVLCテーブル、すなわち復号情報を流用することができる。
 なお、上記のように構成された画像符号化装置または符号化データのデータ構造によれば、本発明に係る画像復号装置と同様の効果を奏する。
 本発明の他の側面について説明すると次のとおりである。すなわち、本発明に係る画像復号装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データから、該変換係数を復号する画像復号装置において、復号対象となる変換係数のひとつ前に復号した変換係数からの相対位置を復号する相対位置復号手段と、ひとつ前に復号した上記変換係数の、上記変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する位置特定手段と、を備えることを特徴とする。
 また、本発明に係る画像符号化装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する画像符号化装置において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を符号化する相対位置符号化手段を備えることを特徴とする。
 また、本発明に係る符号化データのデータ構造は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換に変換して得られた変換係数を符号化することにより生成された符号化データのデータ構造において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、これにより、上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定することを特徴とする。
 上記構成によれば、ひとつ前に復号した上記変換係数の、上記変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する。これにより、相対位置に基づいて、連鎖的に、変換係数の位置を特定していくことができる。なお、変換単位とは、変換を行う所定の単位である。
 上述のrunを符号化する場合、所定のスキャン順序に応じてrunの長さがカウントされるため、基準とする非ゼロ係数と次の非ゼロ係数との変換単位における2次元座標の相対位置が近くても、結果としてrunが長くなるときがあり、これにより、符号量が増大する場合がある。
 この傾向は、変換係数が疎らになりやすい高周波成分の領域において顕著である。また、runが長くなるということは、これに応じた大きなテーブルを用意しなければならないということである。
 これに対して、相対位置で変換係数の位置を特定していけば、このような場合において符号量を削減することができる。
 上記構成によれば、相対位置で変換係数の位置を特定していくので、復号すべき符号量を低減することができる。
 この結果、復号情報の情報量や、復号情報に基づく計算量を低減することができるという効果を奏する。
 また、さらに付言すれば、復号処理において必要となるメモリの量や、処理能力を低く抑えることができる。
 なお、上記のように構成された画像符号化装置または符号化データのデータ構造によれば、本発明に係る画像復号装置と同様の効果を奏する。
 本発明に係る画像復号装置は、変換単位を、複数のサブ単位に分割する変換単位分割手段と、符号化データから上記変換係数を得るための復号情報であって、上記サブ単位ごとに割り当てられている復号情報を参照して、上記サブ単位に含まれる変換係数を復号する変換係数復号手段と、を備える構成である。
 また、本発明に係る画像符号化装置は、変換単位を、複数のサブ単位に分割する変換単位分割手段と、上記変換係数を符号化するための符号化情報であって、上記サブ単位ごとに割り当てられている符号化情報を参照して、上記変換単位に含まれる変換係数を符号化する変換係数符号化手段と、を備える構成である。
 また、本発明に係る符号化データのデータ構造は、符号化対象となる変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、これにより、上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定する、データ構造である。
 本発明に係る画像復号装置は、復号対象となる変換係数のひとつ前に復号した変換係数からの相対位置を復号する相対位置復号手段と、ひとつ前に復号した上記変換係数の、変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する位置特定手段と、を備える構成である。
 また、本発明に係る画像符号化装置は、符号化対象となる変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を符号化する相対位置符号化手段を備える構成である。
 また、本発明に係る符号化データのデータ構造は、符号化対象となる変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、これにより、上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定するデータ構造である。
 上記画像復号装置よれば、符号化データから上記変換係数を得るための復号情報の情報量や、復号情報に基づく計算量を低減することができるという効果を奏する。また、上記画像符号化装置または符号化データのデータ構造によれば、上記画像復号装置と同様の効果を奏する。
本発明の一実施形態に係る動画像復号装置が備えるTU情報復号部の構成例について示す機能ブロック図である。 上記動画像復号装置の概略的構成について示した機能ブロック図である。 本発明の一実施形態に係る動画像符号化装置によって生成され、上記動画像復号装置によって復号される符号化データのデータ構成を示す図であり、(a)~(d)は、それぞれ、ピクチャレイヤ、スライスレイヤ、ツリーブロックレイヤ、およびCUレイヤを示す図である。 対象ブロックを領域分割して係数を符号化/復号する処理の流れについて例示したフローチャートである。 16×16サイズの対象ブロックの分割例を示す図である。 対象ブロックを4つの8×8サイズの領域に分割して復号処理を行う場合の例を示している。 対象ブロックに含まれる4つの復号領域のうち3つに、非ゼロ係数が有り、残りの1つについて非ゼロ係数が無い場合の復号処理について例示する図である。 復号領域の位置に応じてスキャン方法を変更する例について示す図である。 対象ブロックを正方形でない領域に分割する例について示す図である。 本発明の一実施形態に係る動画像符号化装置の構成例について示す機能ブロック図である。 上記動画像符号化装置が備える可変長符号化部の構成例について示すブロック図である。 {run,level}のパラメータの組を、コード番号(ビット列)に変換するためのVLCテーブルの一例を示す図である。 上記VLCテーブルを各復号領域に対応付ける場合の例について示す図である。 動的最適化の例を示す図である。 本発明の他の実施形態に係るTU情報復号部の構成例について示す機能ブロック図である。 相対位置指定により非ゼロ係数を符号化/復号する処理の流れについて例示したフローチャートである。 上記TU情報復号部の復号処理の実行例を示す図である。 本発明の他の実施形態に係るTU情報符号化部の構成例について示す機能ブロック図である。 相対位置指定による非ゼロ係数の符号化の例について示す図である。 上記TU情報復号部が備える相対位置モード用テーブルの構成例を示す図である。 図17や図19に示す対象ブロックにおける高周波成分側の領域を、さらに3つの領域に細分化する例について示す図である。 上記3つの領域に対応付けるVLCテーブルの一例を示している。同図の(a)は、xまたはyの値が、所定以上の正の値にならない場合に参照されるVLCテーブルを示す。また、同図の(b)は、xまたはyの値が、所定以下の負の値にならない場合に参照されるVLCテーブルを示す。 2つの処理を切り替えながら符号化/復号する処理の流れの一例について示すフローチャートである。 対象ブロックにおける低周波成分側の64個の係数のみ符号化/復号する処理について例示する図である。同図の(a)は、インター予測の場合を示しており、(b)は、イントラ予測の場合を示している。 係数符号化データのデータ構造の一例について示す図である。 2階層の分割を行う場合について例示する図である。 図26に示す対象ブロックの分割状況および係数分布状況を表すフラグツリーの表現例(四分木表現)を示している。 再帰的な領域の復号処理の流れの一例について示すフローチャートである。 係数符号化データのデータ構造の他の例について示す図である。 図29に示す係数符号化データの具体例を示す図である。 係数符号化データのデータ構造の別の例について示す図である。 図31に示す係数符号化データの具体例を示す図である。 図26に示す第1階層符号化領域のひとつについて詳細に示す図である。 図26に示す第1階層符号化領域のひとつについて詳細に示す図である。 16×16サイズの対象ブロックにおける分割方式を示す図である。 16×16サイズの対象ブロックにおける分割方式を示す図である。 上記動画像符号化装置を搭載した送信装置、および、上記動画像復号装置を搭載した受信装置の構成について示した図である。(a)は、動画像符号化装置を搭載した送信装置を示しており、(b)は、動画像復号装置を搭載した受信装置を示している。 上記動画像符号化装置を搭載した記録装置、および、上記動画像復号装置を搭載した再生装置の構成について示した図である。(a)は、動画像符号化装置を搭載した記録装置を示しており、(b)は、動画像復号装置を搭載した再生装置を示している。
〔1〕実施形態1
 本発明の一実施形態について図1~図14を参照して説明する。まず、図2を参照しながら、動画像復号装置(画像復号装置)1および動画像符号化装置(画像符号化装置)2の概要について説明する。図2は、動画像復号装置1の概略的構成を示す機能ブロック図である。
 図2に示す動画像復号装置1および動画像符号化装置2は、H.264/MPEG-4 AVC規格に採用されている技術、VCEG(Video Coding Expert Group)における共同開発用コーデックであるKTAソフトウェアに採用されている技術、TMuC(Test Model under Consideration)ソフトウェアに採用されている技術、および、その後継コーデックであるHEVC(High-Efficiency Video Coding)にて提案されている技術を実装している。
 動画像復号装置1には、動画像符号化装置2が動画像を符号化した符号化データ#1が入力される。動画像復号装置1は、入力された符号化データ#1を復号して動画像#2を外部に出力する。動画像復号装置1の詳細な説明に先立ち、符号化データ#1の構成を以下に説明する。
 〔符号化データの構成〕
 図3を用いて、動画像符号化装置2によって生成され、動画像復号装置1によって復号される符号化データ#1の構成例について説明する。符号化データ#1は、例示的に、シーケンス、およびシーケンスを構成する複数のピクチャを含む。
 符号化データ#1におけるピクチャレイヤ以下の階層の構造を図3に示す。図3の(a)~(d)は、それぞれ、ピクチャPICTを規定するピクチャレイヤ、スライスSを規定するスライスレイヤ、ツリーブロック(Tree block)TBLKを規定するツリーブロックレイヤ、ツリーブロックTBLKに含まれる符号化単位(Coding Unit;CU)を規定するCUレイヤを示す図である。
  (ピクチャレイヤ)
 ピクチャレイヤでは、処理対象のピクチャPICT(以下、対象ピクチャとも称する)を復号するために動画像復号装置1が参照するデータの集合が規定されている。ピクチャPICTは、図3の(a)に示すように、ピクチャヘッダPH、及び、スライスS1~SNSを含んでいる(NSはピクチャPICTに含まれるスライスの総数)。
 なお、以下、スライスS1~SNSのそれぞれを区別する必要が無い場合、符号の添え字を省略して記述することがある。また、以下に説明する符号化データ#1に含まれるデータであって、添え字を付している他のデータについても同様である。
 ピクチャヘッダPHには、対象ピクチャの復号方法を決定するために動画像復号装置1が参照する符号化パラメータ群が含まれている。例えば、動画像符号化装置2が符号化の際に用いた可変長符号化のモードを示す符号化モード情報(entropy_coding_mode_flag)は、ピクチャヘッダPHに含まれる符号化パラメータの一例である。
 entropy_coding_mode_flagが0の場合、当該ピクチャPICTは、LCEC(Low Complexity Entropy Coding)またはCAVLC(Context-based Adaptive Variable Length Coding)によって符号化されている。また、entropy_coding_mode_flagが1である場合、当該ピクチャPICTは、CABAC(Context-based Adaptive Binary Arithmetic Coding)によって符号化されている。
 なお、ピクチャヘッダPHは、ピクチャー・パラメーター・セット(PPS:Picture Parameter Set)とも称される。
  (スライスレイヤ)
 スライスレイヤでは、処理対象のスライスS(対象スライスとも称する)を復号するために動画像復号装置1が参照するデータの集合が規定されている。スライスSは、図3の(b)に示すように、スライスヘッダSH、及び、ツリーブロックTBLK1~TBLKNC(NCはスライスSに含まれるツリーブロックの総数)のシーケンスを含んでいる。
 スライスヘッダSHには、対象スライスの復号方法を決定するために動画像復号装置1が参照する符号化パラメータ群が含まれる。スライスタイプを指定するスライスタイプ指定情報(slice_type)は、スライスヘッダSHに含まれる符号化パラメータの一例である。
 スライスタイプ指定情報により指定可能なスライスタイプとしては、(1)符号化の際にイントラ予測のみを用いるIスライス、(2)符号化の際に単方向予測、又は、イントラ予測を用いるPスライス、(3)符号化の際に単方向予測、双方向予測、又は、イントラ予測を用いるBスライスなどが挙げられる。
 また、スライスヘッダSHには、動画像復号装置1の備えるループフィルタ(不図示)によって参照されるフィルタパラメータが含まれていてもよい。
  (ツリーブロックレイヤ)
 ツリーブロックレイヤでは、処理対象のツリーブロックTBLK(以下、対象ツリーブロックとも称する)を復号するために動画像復号装置1が参照するデータの集合が規定されている。
 ツリーブロックTBLKは、ツリーブロックヘッダTBLKHと、符号化単位情報CU~CUNL(NLはツリーブロックTBLKに含まれる符号化単位情報の総数)とを含む。ここで、まず、ツリーブロックTBLKと、符号化単位情報CUとの関係について説明すると次のとおりである。
 ツリーブロックTBLKは、イントラ予測またはインター予測、および、変換の各処理ためのブロックサイズを特定するためのユニットに分割される。
 ツリーブロックTBLKの上記ユニットは、再帰的な4分木分割により分割されている。この再帰的な4分木分割により得られる木構造のことを以下、符号化ツリー(coding tree)と称する。
 以下、符号化ツリーの末端のノードであるリーフ(leaf)に対応するユニットを、符号化ノード(coding node)として参照する。また、符号化ノードは、符号化処理の基本的な単位となるため、以下、符号化ノードのことを、符号化単位(CU)とも称する。
 つまり、符号化単位情報(以下、CU情報と称する)CU~CUNLは、ツリーブロックTBLKを再帰的に4分木分割して得られる各符号化ノード(符号化単位)に対応する情報である。
 また、符号化ツリーのルート(root)は、ツリーブロックTBLKに対応付けられる。換言すれば、ツリーブロックTBLKは、複数の符号化ノードを再帰的に含む4分木分割の木構造の最上位ノードに対応付けられる。
 なお、各符号化ノードのサイズは、当該符号化ノードが直接に属する符号化ノード(すなわち、当該符号化ノードの1階層上位のノードのユニット)のサイズの縦横とも半分である。
 また、各符号化ノードの取り得るサイズは、符号化データ#1のシーケンスパラメータセットSPSに含まれる、符号化ノードのサイズ指定情報および最大階層深度(maximum hierarchical depth)に依存する。例えば、ツリーブロックTBLKのサイズが64×64画素であって、最大階層深度が3である場合には、当該ツリーブロックTBLK以下の階層における符号化ノードは、4種類のサイズ、すなわち、64×64画素、32×32画素、16×16画素、および8×8画素の何れかを取り得る。
  (ツリーブロックヘッダ)
 ツリーブロックヘッダTBLKHには、対象ツリーブロックの復号方法を決定するために動画像復号装置1が参照する符号化パラメータが含まれる。具体的には、図3の(c)に示すように、対象ツリーブロックの各CUへの分割パターンを指定するツリーブロック分割情報SP_TBLK、および、量子化ステップの大きさを指定する量子化パラメータ差分Δqp(qp_delta)が含まれる。
 ツリーブロック分割情報SP_TBLKは、ツリーブロックを分割するための符号化ツリーを表す情報であり、具体的には、対象ツリーブロックに含まれる各CUの形状、サイズ、および、対象ツリーブロック内での位置を指定する情報である。
 なお、ツリーブロック分割情報SP_TBLKは、CUの形状やサイズを明示的に含んでいなくてもよい。例えばツリーブロック分割情報SP_TBLKは、対象ツリーブロック全体またはツリーブロックの部分領域を四分割するか否かを示すフラグ(split_coding_unit_flag)の集合であってもよい。その場合、ツリーブロックの形状やサイズを併用することで各CUの形状やサイズを特定できる。
 また、量子化パラメータ差分Δqpは、対象ツリーブロックにおける量子化パラメータqpと、当該対象ツリーブロックの直前に符号化されたツリーブロックにおける量子化パラメータqp’との差分qp-qp’である。
  (CUレイヤ)
 CUレイヤでは、処理対象のCU(以下、対象CUとも称する)を復号するために動画像復号装置1が参照するデータの集合が規定されている。
 ここで、CU情報CUに含まれるデータの具体的な内容の説明をする前に、CUに含まれるデータの木構造について説明する。符号化ノードは、予測ツリー(prediction tree;PT)および変換ツリー(transform tree;TT)のルートのノードとなる。予測ツリーおよび変換ツリーについて説明すると次のとおりである。
 予測ツリーにおいては、符号化ノードが1または複数の予測ブロックに分割され、各予測ブロックの位置とサイズとが規定される。別の表現でいえば、予測ブロックは、符号化ノードを構成する1または複数の重複しない領域である。また、予測ツリーは、上述の分割により得られた1または複数の予測ブロックを含む。
 予測処理は、この予測ブロックごとに行われる。以下、予測の単位である予測ブロックのことを、予測単位(prediction unit;PU)とも称する。
 予測ツリーにおける分割の種類は、大まかにいえば、イントラ予測の場合と、インター予測の場合との2つがある。
 イントラ予測の場合、分割方法は、2N×2N(符号化ノードと同一サイズ)と、N×Nとがある。
 また、インター予測の場合、分割方法は、2N×2N(符号化ノードと同一サイズ)、2N×N、N×2N、および、N×Nなどがある。
 また、変換ツリーにおいては、符号化ノードが1または複数の変換ブロックに分割され、各変換ブロックの位置とサイズとが規定される。別の表現でいえば、変換ブロックは、符号化ノードを構成する1または複数の重複しない領域のことである。また、変換ツリーは、上述の分割より得られた1または複数の変換ブロックを含む。
 変換処理は、この変換ブロックごとに行われる。以下、変換の単位である変換ブロックのことを、変換単位(transform unit;TU)とも称する。
  (CU情報のデータ構造)
 続いて、図3の(d)を参照しながらCU情報CUに含まれるデータの具体的な内容について説明する。図3の(d)に示すように、CU情報CUは、具体的には、スキップフラグSKIP、PT情報PTI、および、TT情報TTIを含む。
 スキップフラグSKIPは、対象のPUについて、スキップモードが適用されているか否かを示すフラグであり、スキップフラグSKIPの値が1の場合、すなわち、対象CUにスキップモードが適用されている場合、そのCU情報CUにおけるPT情報PTI、および、TT情報TTIは省略される。なお、スキップフラグSKIPは、Iスライスでは省略される。
 PT情報PTIは、CUに含まれるPTに関する情報である。言い換えれば、PT情報PTIは、PTに含まれる1または複数のPUそれぞれに関する情報の集合であり、動画像復号装置1により予測画像を生成する際に参照される。PT情報PTIは、図3の(d)に示すように、予測タイプ情報PType、および、予測情報PInfoを含んでいる。
 予測タイプ情報PTypeは、対象PUについての予測画像生成方法として、イントラ予測を用いるのか、または、インター予測を用いるのかを指定する情報である。
 予測情報PInfoは、予測タイプ情報PTypeが何れの予測方法を指定するのかに応じて、イントラ予測情報、または、インター予測情報より構成される。以下では、イントラ予測が適用されるPUをイントラPUとも呼称し、インター予測が適用されるPUをインターPUとも呼称する。
 また、予測情報PInfoは、対象PUの形状、サイズ、および、位置を指定する情報が含まれる。上述のとおり予測画像の生成は、PUを単位として行われる。予測情報PInfoの詳細については後述する。
 TT情報TTIは、CUに含まれるTTに関する情報である。言い換えれば、TT情報TTIは、TTに含まれる1または複数のTUそれぞれに関する情報の集合であり、動画像復号装置1により残差データを復号する際に参照される。なお、以下、TUのことをブロックと称することもある。
 TT情報TTIは、図3の(d)に示すように、対象CUの各変換ブロックへの分割パターンを指定するTT分割情報SP_TU、および、TU情報TUI1~TUINT(NTは、対象CUに含まれるブロックの総数)んでいる。
 TT分割情報SP_TUは、具体的には、対象CUに含まれる各TUの形状、サイズ、および、対象CU内での位置を決定するための情報である。例えば、TT分割情報SP_TUは、対象となるノードの分割を行うのか否かを示す情報(split_transform_unit_flag)と、その分割の深度を示す情報(trafoDepth)とから実現することができる。
 また、例えば、CUのサイズが、64×64の場合、分割により得られる各TUは、32×32画素から2×2画素までのサイズを取り得る。
 TU情報TUI1~TUINTは、TTに含まれる1または複数のTUそれぞれに関する個別の情報である。例えば、TU情報TUIは、量子化予測残差を含んでいる。
 各量子化予測残差は、動画像符号化装置2が以下の処理1~3を、処理対象のブロックである対象ブロックに施すことによって生成した符号化データである。
 処理1:符号化対象画像から予測画像を減算した予測残差をDCT変換(Discrete Cosine Transform)する;
 処理2:処理1にて得られた変換係数を量子化する;
 処理3:処理2にて量子化された変換係数を可変長符号化する;
 なお、上述した量子化パラメータqpは、動画像符号化装置2が変換係数を量子化する際に用いた量子化ステップQPの大きさを表す(QP=2qp/6)。
  (予測情報PInfo)
 上述のとおり、予測情報PInfoには、インター予測情報およびイントラ予測情報の2種類がある。
 インター予測情報には、動画像復号装置1が、インター予測によってインター予測画像を生成する際に参照される符号化パラメータが含まれる。より具体的には、インター予測情報には、対象CUの各インターPUへの分割パターンを指定するインターPU分割情報、および、各インターPUについてのインター予測パラメータが含まれる。
 インター予測パラメータには、参照画像インデックスと、推定動きベクトルインデックスと、動きベクトル残差とが含まれる。
 一方、イントラ予測情報には、動画像復号装置1が、イントラ予測によってイントラ予測画像を生成する際に参照される符号化パラメータが含まれる。より具体的には、イントラ予測情報には、対象CUの各イントラPUへの分割パターンを指定するイントラPU分割情報、および、各イントラPUについてのイントラ予測パラメータが含まれる。イントラ予測パラメータは、各イントラPUについてのイントラ予測方法(予測モード)を指定するためのパラメータである。
 〔動画像復号装置〕
 以下では、本実施形態に係る動画像復号装置1の構成について、図1~図5を参照して説明する。
  (動画像復号装置の概要)
 動画像復号装置1は、PU毎に予測画像を生成し、生成された予測画像と、符号化データ#1から復号された予測残差とを加算することによって復号画像#2を生成し、生成された復号画像#2を外部に出力する。
 ここで、予測画像の生成は、符号化データ#1を復号することによって得られる符号化パラメータを参照して行われる。符号化パラメータとは、予測画像を生成するために参照されるパラメータのことである。符号化パラメータには、画面間予測において参照される動きベクトルや画面内予測において参照される予測モードなどの予測パラメータに加えて、PUのサイズや形状、ブロックのサイズや形状、および、原画像と予測画像との残差データなどが含まれる。以下では、符号化パラメータに含まれる情報のうち、上記残差データを除く全ての情報の集合を、サイド情報と呼ぶ。
 また、以下では、上記PUおよびTUが、CUと同じサイズである場合を例に挙げ説明を行う。しかしながら、これに限定されるものではなく、PUおよびTUがCUよりも小さい単位である場合に対しても適用することができる。
 また、以下では、復号の対象となるピクチャ(フレーム)、スライス、ツリーブロック、ブロック、および、PUをそれぞれ、対象ピクチャ、対象スライス、対象ツリーブロック、対象ブロック、および、対象PUと呼ぶことにする。
 なお、ツリーブロックのサイズは、例えば64×64画素であり、PUのサイズは、例えば、64×64画素、32×32画素、16×16画素、8×8画素や4×4画素などである。しかしながら、これらのサイズは、単なる例示であり、ツリーブロックおよびPUのサイズは以上に示したサイズ以外のサイズであってもよい。
  (動画像復号装置の構成)
 再び、図2を参照して、動画像復号装置1の概略的構成について説明すると次のとおりである。図2は、動画像復号装置1の概略的構成について示した機能ブロック図である。
 図2に示すように動画像復号装置1は、可変長符号逆多重化部11、TU情報復号部12、逆量子化・逆変換部13、予測画像生成部14、加算器15およびフレームメモリ16を備えている。
   [可変長符号逆多重化部]
 可変長符号逆多重化部11は、動画像復号装置1に入力された1フレーム分の符号化データ#1を、逆多重化することで、図3に示した階層構造に含まれる各種情報に分離する。例えば、可変長符号逆多重化部11は、各種ヘッダに含まれる情報を参照して、符号化データ#1を、スライス、ツリーブロックに順次分離する。
 ここで、各種ヘッダには、(1)対象ピクチャのスライスへの分割方法についての情報、および(2)対象スライスに属するツリーブロックのサイズ、形状および対象スライス内での位置についての情報が含まれる。
 そして、可変長符号逆多重化部11は、ツリーブロックヘッダTBLKHに含まれるツリーブロック分割情報SP_TBLKを参照して、対象ツリーブロックを、CUに分割する。また、可変長符号逆多重化部11は、対象CUについてTT情報TTI、およびPT情報PTIを取得する。
 可変長符号逆多重化部11は、対象CUについて得られたTT情報TTIに含まれるTU情報TUIを所定の順にTU情報復号部12に供給する。また、可変長符号逆多重化部11は、対象CUについて得られたPT情報PTIを予測画像生成部14に供給する。
   [TU情報復号部]
 TU情報復号部12は、可変長符号逆多重化部11から供給されるTU情報TUIの復号を行って、復号済みTU情報TUI’を生成する。
 例えば、TU情報復号部12は、対象ブロックについて、TU情報TUIから量子化予測残差を復号する。ここで、対象ブロックについての量子化予測残差は、量子化された変換係数が2次元の行列に配列された形式で表現できる。以下、量子化された変換係数の2次元の行列表現のことを係数行列と称する。
 TU情報復号部12は、復号済みTU情報TUI’を逆量子化・逆変換部13に供給する。なお、TU情報復号部12の動作詳細については後述する。
   [逆量子化・逆変換部]
 逆量子化・逆変換部13は、TU情報復号部12から供給される復号済み復号済みTU情報TUI’に基づいて対象CUについて、ブロックごとに、量子化予測残差の逆量子化・逆変換を行う。逆量子化・逆変換部13は、復号済みTU情報TUI’に含まれる量子化予測残差を逆量子化および逆DCT変換(Inverse Discrete Cosine Transform)することによって、各対象PUについて、画素毎の予測残差Dを復元する。逆量子化・逆変換部13は、復元した予測残差Dを加算器15に供給する。
   [予測画像生成部]
 予測画像生成部14は、対象CUに含まれる各PUについて、当該PUの周辺の復号済み画像である局所復号画像P’を参照して、イントラ予測またはインター予測により予測画像Predを生成する。予測画像生成部14は、対象CUについて生成した予測画像Predを加算器15に供給する。
   [加算器]
 加算器15は、予測画像生成部14より供給される予測画像Predと、逆量子化・逆変換部13より供給される予測残差Dとを加算することによって、対象CUについての復号画像Pを生成する。
   [フレームメモリ]
 フレームメモリ16には、復号された復号画像Pが順次記録される。フレームメモリ16には、対象ツリーブロックを復号する時点において、当該対象ツリーブロックよりも先に復号された全てのツリーブロック(例えば、ラスタスキャン順で先行する全てのツリーブロック)に対応する復号画像が記録されている。
 なお、動画像復号装置1において、画像内の全てのツリーブロックに対して、ツリーブロック単位の復号画像生成処理が終わった時点で、動画像復号装置1に入力された1フレーム分の符号化データ#1に対応する復号画像#2が外部に出力される。
  (係数のスキャンについて)
 ここで、TU情報復号部12の構成の詳細について説明する前に、係数のスキャンについて説明する。
 係数の符号化においては、動画像符号化装置2において、対象ブロックにおける係数の集合を表す係数行列が所定の順序によりスキャンされる。
 スキャンとは、係数行列における係数の座標を、1次元のスキャン順インデックスに変換する処理である。変換後の係数は、1次元配列に格納されて保持される。スキャンでは従来知られているジグザグスキャン順序等を用いることができる。
 符号化処理では、まず対象ブロックにおいて、左上の位置にあるDC係数から、右下の位置にある最も高周波成分の係数までがジグザグスキャン順序に基づいてスキャンされる。
 なお、高周波成分の係数の値は、ゼロまたはゼロに近くなる傾向がある一方で、低周波成分の係数の値は、ゼロでないまたは値が大きい可能性が高いという特性がある。このためジグザグスキャンは、低周波成分の係数を早くスキャンするスキャン順序となっている。
 また、動画像符号化装置2では、スキャンの後、係数の符号化処理が行われる。係数の符号化処理は、スキャンとは逆の順序にて行われる。
 以下において、この処理順序のことを逆順のジグザグスキャンとも呼ぶ。つまり、符号化処理は、係数行列に対して逆順のジグザグスキャンで行われる。また、符号化処理順序について説明する場合は、便宜的に、係数行列をベースに説明する。
 高周波成分は、値が“0”であることが多いので、本実施形態の符号化処理においては、このようなゼロ係数の符号化を省略し、DC係数からみて最後の非ゼロの係数を基点に行うものとする。
 なお、以上では、説明の便宜上、8×8の係数行列のスキャン例を示したが、4×4の係数行列をスキャンする場合についても同様のことが言える。また、スキャン順序をジグザグスキャン以外の方法、例えばラスタスキャンと同様の順序などに適応的に変更してもよい。
  (係数の符号化について)
 以下において、動画像符号化装置2において符号化される係数符号化データについて説明する。上述のとおり符号化処理においては、まず、最後の非ゼロ係数が符号化される。続いて、所定の条件下において、ランモードによる符号化が行われる。ランモード終了条件が満たされることにより、ランモードが終了すると、残りの係数について、レベルモードによる符号化が行われる。このようにして全ての係数が符号化される。
 以下、最後の非ゼロ係数、ランモード、および、レベルモードで符号化されるデータについて説明する。
   [最後の非ゼロ係数]
 最後の非ゼロ係数は、最後の非ゼロ係数の位置last_pos、最後の非ゼロ係数のレベルlevel、最後の非ゼロ係数の正負符号signが符号化される。ここで係数のレベルとは、係数の絶対値を意味する。last_posは、スキャン順インデックスの形式の値をとる。最後の非ゼロ係数が、スキャン順で10番目(DC係数を1番目とする)であり、その値が“-1”である場合は次のとおりである。
 last_pos:9
 level:1,sign:1
   [ランモード]
 所定の条件下、ランモードを実行する場合、以下の符号化が行われる。ランモードとは、連続するゼロ係数の数(0ラン)を符号化するモードである。どのような条件において、ランモードによる符号化が行われるかについては、後述する。
 ランモードでは、係数ごとに、0ランrun、係数のレベルlevel、および係数の正負符号signが符号化される。次に符号化する非ゼロ係数が、前の非ゼロ係数との間に0を1つ挟んでおり、かつ値が“1”である場合は次のとおりである。
 run:1,level:1,sign:0
 なお、非特許文献1,2では、levelが2以上の係数が出現した場合を、ランモードの終了条件のひとつとして用いている。
   [レベルモード]
 レベルモードとは、ゼロ係数であっても、1つずつ符号化するモードである。レベルモードでは、係数ごとに、係数のレベルlevel、および係数の正負符号signが符号化される。符号化する係数の値が“-6”である場合、および符号化する係数がゼロ係数(値が“0”)である場合はそれぞれ次のとおりである。
 level:6,sign:1
 level:0
   [備考]
 なお、以上に示した符号化データは、単なる例示であり、シンタクス要素としては上記と異なる名称または定義により符号化されていてもよい。例えば、ランモードの係数のレベルlevelは、isLevelOne(levelは1)、および、これに追加してlevel_magnitude_minus2(係数値が2以上の場合であり、符号化するのは2を引いたレベル)のシンタクス要素として符号化されていてもよい(非特許文献1,2の「level_magnitude_minus2」を参照)。
 なお、levelは、最後の非ゼロ係数については、last_pos_levelのシンタクス要素として符号化されていてもよい。
 また、非ランモードのlevelは、level_magnitudeのシンタクス要素として符号化されていてもよい。
  (TU情報復号部)
 次に、図1を用いてTU情報復号部12の構成についてさらに詳しく説明する。図1は、TU情報復号部12の構成例について示す機能ブロック図である。
 なお、以下では、TU情報復号部12が、TU情報TUIに含まれる符号化データのうち、係数に関するデータを復号するための構成について説明する。言い換えれば、以下では、TU情報復号部12が、符号化された量子化予測残差、すなわち係数符号化データを復号するための構成について説明する。
 しかしながら、これに限られずTT情報復号部12は、符号化データに含まれる係数符号化データ以外のデータ、例えば、サイド情報等を復号することができる。
 図1に示すように、TU情報復号部12は、VLCテーブルTBL11と、領域分割部(変換単位分割手段)121と、領域復号部(変換係数復号手段)122とを備える。
 また、以下では、TU情報復号部12が、16×16サイズの対象ブロックについてTU情報TUIを復号して、復号済みTU情報TUI’を出力する例について説明する。しかしながら、これに限られず、TU情報復号部12が復号する対象ブロックのサイズは、64×64、32×32等であってもかまわない。
 VLCテーブルTBL11は、ビット列(コード)と相互に変換可能なコード番号と、復号すべきパラメータとが対応付けられたテーブルである。VLCテーブルTBL11は、領域復号部122の復号処理において参照される。
 VLCテーブルTBL11には、例示的に、8×8サイズの符号化単位について規定されているテーブルを用いている。すなわち、VLCテーブルTBL11には、入力となる変換単位(あるいは符号化単位)16×16サイズよりも小さな8×8サイズのものを用いている。
 また、VLCテーブルTBL11は、適応的な復号処理ができるよう、復号処理におけるコンテキストに応じて複数規定されている。コンテキストには、例えば、処理中の係数の位置や、対象ブロックの属性(輝度/色差といった画素の種類や予測方法)などを用いることができる。
 領域分割部121は、対象ブロックを複数の領域に分割する。以下において、領域分割部121の分割により得られる各領域のことを復号領域と称する。
 また、以下では、領域分割部121が、16×16サイズの対象ブロックを、4つの8×8サイズの復号領域に分割する例について説明する。しかしながら、これに限られず、領域分割部121の分割の手法には、種々のものを採用することが可能である。その変形例については、後に詳しく説明する。
 領域復号部122は、領域分割部121が対象ブロックを分割することで得られた復号領域のそれぞれについて、当該復号領域のサイズに応じて規定されているVLCテーブルTBL11を参照しながら、復号処理を行う。
 以下では、領域復号部122が、8×8サイズの符号化単位について規定されているVLCテーブルTBL11をコンテキストに応じて参照しながら復号処理を行う例について説明する。
 領域復号部122の具体的構成は次のとおりである。すなわち、領域復号部122は、最後の非ゼロ係数復号部101と、ランモード復号部102と、レベルモード復号部103とを備える。
 最後の非ゼロ係数復号部101は、復号の対象となる復号領域(以下、対象復号領域と称する)について最後の非ゼロ係数を復号する。より具体的には、最後の非ゼロ係数復号部101は、対象復号領域に対応する係数符号化データに含まれるlast_pos、level、およびsignを復号する。
 ランモード復号部102は、対象復号領域について、係数符号化データから、ランモードで符号化された係数を復号する。すなわち、ランモード復号部102は、コンテキストに応じたVLCテーブルTBL11を参照しながら、係数符号化データから、ランモードで符号化されたrun、levelおよびsignを復号する。以下、ランモード復号部102が行う復号処理のことを、ランモード復号処理と称する。
 また、ランモード復号部102は、ランモード終了条件が満たされるまで、ランモード復号処理を繰り返して行う。ランモード終了条件の条件としては、例えば、復号した係数の値が閾値を超えている場合や、所定の数だけ係数を復号した場合などが挙げられる。
 そして、ランモード終了条件が満たされた場合、ランモード復号部102は、レベルモード復号部103にレベルモードによる復号を開始させる。
 レベルモード復号部103は、対象復号領域について、係数符号化データから、レベルモードで符号化された係数を復号する。すなわち、レベルモード復号部103は、コンテキストに応じたVLCテーブルTBL11を参照しながら、係数符号化データから、レベルモードで符号化されたlevelおよびsignを復号する。
 以下、レベルモード復号部103が行う復号処理のことをレベルモード復号処理と称する。また、レベルモード復号部103は、直流成分の係数を復号するまで、レベルモード復号処理を繰り返して行う。
 なお、領域復号部122は、対象ブロックの各対象復号領域について復号処理が完了すると、当該復号処理により得られた係数を含む復号済みTU情報TUI’を出力する。
  (処理の流れ)
 次に、図4を用いて、TU情報復号部12における復号処理について説明する。図4は、対象ブロックを領域分割して係数を符号化/復号する処理S10の流れについて例示したフローチャートである。
 なお、動画像符号化装置2の符号化処理と動画像復号装置1の復号処理との相違点は、符号化処理を行うか復号処理を行うかであり、それ以外の流れは符号化処理および復号処理において概ね同じである。よって、図4では、動画像符号化装置2の符号化処理と動画像復号装置1の復号処理とをまとめて示している。
 図4に示すように、まず、動画像復号装置1では、領域分割部121が、対象ブロックを分割する(S11)。
 続いて、分割した各復号領域についてのループLP1に入る(S12)。ループLP1では、領域復号部122が、対象復号領域について係数符号化データに含まれる係数を復号する(S13)。
 具体的には、まず、最後の非ゼロ係数復号部101が、対象復号領域における最後の非ゼロ係数を復号する。
 続いて、ランモード終了条件が満たされるまで、ランモード復号部102が、ランモード復号処理を行う。
 ランモードが終了した後、レベルモード復号部103が、レベルモード復号処理を行う。
このようにして、対象復号領域についての復号処理が完了すると、ループLP1の先頭に戻り(S14からS12へ)、次の対象復号領域についてさらに復号処理が行われる。
 ブロックの各対象復号領域の復号処理が完了するとループLP1が終了する。その後、領域分割をして係数を復号する処理S10が終了する。
  (具体例)
 次に、図5および図6に加えて、さらに図4に示すフローチャートを参照しながら、TU情報復号部12における復号処理の具体例を示す。図5は、16×16サイズの対象ブロック(変換単位)BLKを示している。また、図6は、対象ブロックBLKを4つの8×8サイズの領域に分割して復号処理を行う場合の例を示している。
 S11では、領域分割部121が、図5に示すように、16×16サイズの対象ブロックBLKを、4つの8×8サイズの復号領域(サブ単位)R11~R14に分割する。
 そして、ループLP1のS13では、領域復号部122が、図6に示す各復号領域を、復号領域R11、R12、R13およびR14の順で処理する。
 なお、図6において復号領域R11~R14内に示す矢印はスキャン順序を示している。すなわち、復号領域R11~R14のスキャン順序は、ジグザグスキャンである。
 S13では、まず、最後の非ゼロ係数復号部101が、スキャン順序上、最後の非ゼロ係数を復号する。そして、ランモード復号処理およびレベルモード復号処理により、最後の非ゼロ係数から、直流成分の係数までが、逆順のジグザグスキャンにより復号される。
 復号領域R11の復号処理が完了すると、以下、復号領域R12、R13およびR14について同様に復号処理が行われる。
 このように、復号領域R11~R14では、8×8サイズの符号化単位に対する従来の復号方式で復号処理が行われる。すなわち、復号領域R11~R14の復号処理では、最後の非ゼロ係数を復号し、ランモード復号処理およびレベルモード復号処理を行う従来の復号方式を用いることができる。
  (変形例)
 以下において、動画像復号装置1の好ましいいくつかの変形例について説明する。
  変形例1-1:[非ゼロ係数の有無の判定]
 係数符号化データにおいて、各復号領域について非ゼロ係数の有無を示す非ゼロ係数フラグが符号化されている場合、領域復号部122は非ゼロ係数の有無に応じて復号処理を行ってもよい。
 非ゼロ係数の有無に応じて復号処理を行う場合、領域復号部122を次のように構成すればよい。
 すなわち、まず、領域復号部122は非ゼロ係数フラグを復号して、非ゼロ係数の有無を判定する。そして、非ゼロ係数フラグが対象復号領域において非ゼロ係数が無いことを示す場合、領域復号部122は、当該対象領域における復号処理をスキップする。
 図7を用いて、具体的に説明する。図7は、復号領域R11、R13およびR13に、非ゼロ係数が有り、復号領域R12について非ゼロ係数が無い場合の例を示している。すなわち、復号領域R12に示している「0」は、非ゼロ係数がないことを示している。
 非ゼロ係数フラグでは、例えば、非ゼロ係数が無い場合、“0”が、非ゼロ係数が有る場合、“1”が符号化されている。
 図7に示す例では、領域R11、R13、およびR14には非ゼロ係数が有り、領域R12には非ゼロ係数が無い。このため、非ゼロ係数フラグは、“1011”が符号化されている。なお、非ゼロ係数フラグの“1011”等のパターンをそのまま4ビットで固定長符号化せず、パターンの出現頻度に応じた可変長符号化を行ってもよい。
 図7に示す例の復号処理は次のようにして行われる。まず、復号領域R11については、非ゼロ係数フラグ“1”が符号化されているので、領域復号部122は、復号領域R11に対して復号処理を行う。
 続いて、復号領域R12については、非ゼロ係数フラグ“0”が符号化されているので、領域復号部122は、復号領域R12に対する復号処理をスキップする。
 残りの復号領域R13およびR14については、復号領域R11と同様、非ゼロ係数が有ることを示す非ゼロ係数フラグ“1”が符号化されているので、領域復号部122は、復号領域R13およびR14のそれぞれに対して復号処理を行う。
 以上の構成によれば、復号領域単位での不要な復号処理を省略することができるので、処理効率の向上を図ることができる。
  変形例1-2:[復号領域の位置に応じて処理方法を変更する]
  [1]スキャン方法の変更
 以上の説明では、各復号領域のスキャン方式にジグザグスキャンを採用していたが、復号領域の位置に応じてスキャン方法を変更してもよい。
 例えば、図8に示すように、対象ブロックBLKにおいて右上に位置する復号領域R12では水平スキャンを採用してもよい。また、図8に示すように、対象ブロックBLKにおいて左下に位置する復号領域R13では、垂直スキャンを採用してもよい。また、図8に示すように、復号領域R11およびR14では、ジグザグスキャンを採用してもよい。
  [2]モード限定
 以上の説明では、各復号領域の復号処理において、ランモード復号処理およびレベルモード復号処理を行っていたが、復号領域の位置に応じてランモード復号処理かレベルモード復号処理のいずれかを行う構成としてもよい。
 例えば、対象ブロックBLKの右下に位置する復号領域R14は、高周波成分であるためゼロ係数の数が多い傾向がある。このため、復号領域R14では、復号処理をランモード復号処理のみで行ってもよい。
  [3]VLCテーブルやコード番号の計算方法を変更する
 ランモード復号処理において、復号領域の位置に応じて、参照するVLCテーブルや、コード番号の計算方法を変更してもよい。例えば、コード番号を示すビット列を、{run,level}のパラメータの組に変換するためのVLCテーブルを各復号領域の位置に応じて変更してもよい。
 例えば、VLCテーブルを次のように構成する。すなわち、高周波成分側の復号領域(例えば、図6における復号領域R12~R14)で参照されるVLCテーブルでは、短いビット列には、level=0である組を対応付けておく。また、低周波成分側の復号領域(例えば、図6における復号領域R11)で参照されるVLCテーブルでは、短いビット列には、runの短いものを対応付けておく。
 高周波成分側の復号領域では、ゼロ係数が多くなる傾向があり、低周波成分側の復号領域では、非ゼロ係数が多くなりrunが短くなる傾向がある。
 よって、以上のように構成すれば、各復号領域における非ゼロ係数の分布状況に応じた効率のよい復号処理を行うことができる。すなわち、復号すべき符号量を低減することができる。VLCテーブルの定義の詳細について、図12および図13を用いて、説明すれば以下の通りである。
 図12は、{run,level}のパラメータの組を、コード番号に変換するためのVLCテーブルTBL11の一例を示している。また、同図は、runの最大値が4のときにおける2種類のVLCテーブルを示しており、(a)は、level=0の優先度が高いVLCテーブルT1を示しており、(b)は、level=0の優先度が低いVLCテーブルT2を示している。
 なお、VLCテーブルT1およびVLCテーブルT2では、次の非ゼロ係数の絶対値が“1”のときlevel=0であり、また、次の非ゼロ係数の絶対値が“1”より大きいときlevel=1である。
 図13は、上記VLCテーブルT1およびT2を各復号領域に対応付ける場合の例について示している。図13に示すように、復号領域R11は、VLCテーブルT2に対応付けられており、復号領域R12~R14は、VLCテーブルT1に対応付けられている。
 ランモード復号部102の復号処理の流れについて説明すると次のとおりである。まず、ランモード復号部102は、図13に示す対応付けに従い、復号領域R11のランモード復号処理においてVLCテーブルT2を参照する。
 図12の(b)を用いて、より詳細に説明すると次のとおりである。すなわち、VLCテーブルT2では、runが短い{run,level}の組み合わせに、より小さいコード番号(短いコード)が割り当てられている。
 例えば、VLCテーブルT1では、{run,level}=(0,0)の組み合わせに最も小さいコード番号“0”が割り当てられている。また、出現頻度を考慮して、{run,level}=(4,0)の組み合わせに次に小さいコード番号“1”が割り当てられている。
 それ以外の組み合わせについては、levelが同じであれば、runが長いほうに、より大きなコード番号が割り当てられている。
 level=0となる{run,level}の組み合わせについては、run=3のとき、コード番号“7”が割り当てられている。また、level=0となる{run,level}の組み合わせについては、run=3のとき、コード番号“8”が割り当てられている。
 低周波成分側の復号領域R11では、runが短くなる傾向があるため、runが短い{run,level}の組み合わせに、より小さいコード番号を割り当てることで、符号化効率を向上させることができる。
 また、ランモード復号部102は、図13に示す対応付けに従い、復号領域R12~14のランモード復号処理においてVLCテーブルT1を参照する。
 図12の(a)を用いて、より詳細に説明すると次のとおりである。すなわち、VLCテーブルT1では、level=0となる{run,level}の組み合わせに、より小さいコード番号(より短いコード)が割り当てられている。
 VLCテーブルT1では、{run,level}=(0,0)の組み合わせに最も小さいコード番号“0”が割り当てられている。また、出現頻度を考慮して、{run,level}=(4,0)の組み合わせに次に小さいコード番号“1”が割り当てられている。
 level=0となる{run,level}の組み合わせには、上記の組み合わせよりも大きなコード番号が割り当てられている。すなわち、runの長さに応じて、{run,level}=(0,1)~(0,3)には、それぞれ、コード番号“5”~“8”が割り当てられている。
 高周波成分側の復号領域R12~14では、level=0を復号することになる可能性が高い傾向があるため、level=0となる{run,level}の組み合わせに、より小さいコード番号を割り当てることで、符号化効率を向上させることができる。
 なお、ランモード復号処理におけるパラメータ-コードの変換をVLCテーブルT1およびT2を用いて実行する例について示したが、これに限られず、VLCテーブルT1およびT2と同等の変換処理が演算により実現されていてもかまわない。
 また、以上では、VLCテーブルT1およびT2の2種類のテーブルを用いてランモード復号処理を実行する例について説明したが、これに限られず、2種類より多くの種類のテーブルを定義しておいてもかまわない。例えば、ランモード復号部102は、復号領域R11~R14において、相互に異なるVLCテーブルを参照するように構成されていてもかまわない。
  [4]動的最適化
 領域復号部122による、復号状況に応じた動的最適化の例について説明する。
  [4-1]VLCテーブルの更新
 領域復号部122は、パラメータの値の出現頻度をカウントし、出現頻度に応じてVLCテーブルのコード番号を書き換えてもよい。
 すなわち、領域復号部122は、出現頻度が高いパラメータの値のコード番号をより小さな(よりコードが短い)ものに書き換え、出現頻度が低いパラメータの値のコード番号をより大きな(よりコードが長い)ものに書き換えてもよい。
 図14を用いて説明すると次のとおりである。図14に例示するように、最適化前のVLCテーブルT3では、あるパラメータの値yにコード番号CN-1が割り当てられており、別のパラメータの値xにコード番号CNが割り当てられている。
 領域復号部122が、最適化前のVLCテーブルT3のxと、これに対応するコード番号CNを参照したとする。
 このとき、領域復号部122は、コード番号CNを1デクリメントして、xに対応するコード番号をCN-1に繰り上げる。その一方で、領域復号部122は、元々CN-1に対応していたyのコード番号を1インクリメントしてCNに対応付ける。ここでは、このようなコード番号の繰り上げ処理のことを最適化と称している。図14に示すVLCテーブルT4が、上記最適化後に得られるテーブルである。
 さらに、ここで、領域復号部122がxを参照した場合、領域復号部122は、xに対応するコード番号をCN-2とする。
 このようにして、領域復号部122は、復号するパラメータの値の出現頻度に応じて、短いコード番号が割り当てられるように、復号処理中に動的にVLCテーブルを更新する。
  [4-2]適応速度について
 動的最適化の適応速度について説明すると次のとおりである。図14では、パラメータの値を1回参照するたびに、コード番号を1デクリメントしていたが、これに限られず、コード番号を2以上デクリメントしてもよい。
 適応速度の速さは、例えば、コード番号をデクリメントする大きさとして表現することができる。つまり、コード番号のデクリメントの量が“1”である場合よりも、“2”である場合のほうが、“適応速度が速い”といえる。
 また、{run,level}のVLCテーブルを動的最適化する場合、高周波成分側の復号領域では、最適化時のインクリメントの量を“0”とし、低周波成分側では、“1”としてもよい。
 すなわち、高周波成分側の復号領域では、最適化を行わない構成としてもよい。この構成を採用する理由としては次のとおりである。すなわち、高周波成分側ではランが長くなりやすく、特定の{run,level}が高頻度で発生する可能性が低い。このため、上記構成によれば、パラメータの値が出現するたびに最適化を行うことによる計算量の増加を防ぐことができる。
 なお、動画像復号装置1の領域復号部122において上記最適化を行ってもかまわない。
  [5]ランモード終了条件の変更
 復号領域の位置に応じてランモード終了条件を変更してもよい。例えば、高周波成分側の復号領域では、ランモード終了条件を厳しくし、低周波成分側の復号領域では、ランモード終了条件を緩和してもよい。つまり、対象ブロックにおいて左上に位置する復号領域ほど、ランモードを終了しやすくすればよい。また、高周波成分側の復号領域では、レベルモード復号処理を行わないようにしてもよい。
 高周波成分側の復号領域では、係数の絶対値が小さくなる傾向がある。このため、高周波成分側の復号領域ではランが長くなる可能性が高い。よって、なるべくランモード復号処理が実行されるようにすることが好ましい。
 また、ランモード復号部102において、ランモード復号処理を開始するか否かを判定してもよい。ランモード復号部102は、例えば、最後の非ゼロ係数が復号された時点で参照可能なデータに基づいて、「係数の絶対値が全体的に小さいと判断できる」場合、ランモード復号処理を開始すると判定してもよい。
 「係数の絶対値が全体的に小さいと判断できる」場合、ランが長くなる可能性が高いので、ランモード復号部102は、ランモード復号処理を実行する。
 なお、最後の非ゼロ係数が復号された時点で参照可能なデータとは、例えば、対象ブロックの予測情報、最後の非ゼロ係数についての情報、その他エンコーダ側で符号化されていて既に復号されているフラグ等が挙げられる。
 具体的には、最後の係数の値が1より大きい場合、ランモード復号部102は、ランモード復号処理をスキップしてもよい。
 このほか、復号領域の位置に応じて係数復号処理における各種判定の閾値を変更してもよい。
  変形例1-3:[対象ブロックの予測モードに応じて分割の有無を決定する]
 領域分割部121は、対象ブロックの予測モードに応じて分割の有無を決定する構成であってもよい。
 例えば、領域分割部121は、対象ブロックの予測モードが、イントラ予測である場合に分割を行い、インター予測である場合に分割を行わないようにしてもよい。
 また、対象ブロックの予測モードがインター予測であり、領域分割部121が分割を行わなかった場合、領域復号部122は、対象ブロックにおいて、スキャン順序上、1番目から64番目までの係数のみを復号するようにしてもよい。すなわち、この場合、領域復号部122は、対象ブロックの左上側にある64個の係数のみを復号してもよい。
 この場合、分割の有無を示すフラグの復号を省略することができるという効果を得ることができる。
  変形例1-4:[領域の数およびサイズ]
 図5および図6に示す例では、領域分割部121が、対象ブロックを4つの復号領域に分割することについて説明した。しかしながら、これに限られず、領域分割部121は、対象ブロックを4つより多い領域に分割してもかまわない。
 また、対象ブロックは、16×16サイズのものに限られない。例えば、対象ブロックは、32×32サイズであっても、64×64サイズのものであってもかまわない。
 例えば、対象ブロックが64×64サイズである場合、領域分割部121は、対象ブロックを、以下のとおり4つ以上の復号領域に分割してもかまわない。
 すなわち、領域分割部121は、対象ブロックを64個の8×8サイズの復号領域に分割してもよい。また、領域分割部121は、対象ブロックを16個の16×16サイズの復号領域に分割してもよい。また、領域分割部121は、対象ブロックを4個の32×32サイズの復号領域に分割してもよい。
 また、領域分割部121が対象ブロックを分割して得られる復号領域は、正方形のものには限られない。例えば、復号領域は、矩形であってもよい。また、対象ブロックを、左上の8×8サイズの復号領域(図5における復号領域R11に相当)と、それ以外の復号領域(図5における復号領域R12~R14に相当)との2つに分割してもよい。
 また、各復号領域の形状は同じでなくてもよい。例えば、対象ブロックを、図9に示すように、ジグザグスキャン順序上、64個の係数ごとに分割して復号領域を得てもよい。
 すなわち、図9に示すように、対象ブロックを、復号領域R21~R24に分割してもよい。復号領域R21~R24の領域内に示す数字“64”は、領域内に係数が64個含まれることを示している。
 復号領域R21は、DC係数を含む領域であり、同図では、その形状を直角三角形にて表している。また、復号領域R22およびR23の形状は、それぞれ台形にて表している。また、復号領域R24は、各領域のうち最も高周波成分側に有る領域であり、その形状を直角三角形にて表している。
 なお、図9では、説明の便宜のため、領域R21およびR24の形状を直角三角形にて示し、領域R22およびR23の形状を台形にて表しているが、実際のところ厳密に言えば、直角三角形、台形にならない点には留意されたい。
 また、各復号領域の形状および各復号領域に含まれる係数の数は、同じでなくてもよい。
 このように、領域分割部121は、対象ブロックを、当該対象ブロックのサイズより小さいサイズの複数の復号領域に任意に分割することができる。
  変形例1-5:[復号領域の処理順序]
 図5および図6に示す例では、領域復号部122が、復号領域R11、R12、R13およびR14の順(いわゆるラスタスキャン順)で復号処理を行うと説明した。しかしながら、これに限られず、領域復号部122は、これ以外の順序で復号処理を行ってもよい。
 例えば、領域復号部122は、復号領域R14、R13、R12およびR11の順、あるいは別の例として、復号領域R11、R13、R12およびR14の順で復号処理を行ってもかまわない。また、復号領域の数が4つより多い場合、例えば、16である場合には、ジグザグスキャン順に処理してもよい。
 〔動画像符号化装置〕
 まず、以下では、本実施形態に係る動画像符号化装置2について、図10~図14を参照して説明する。
  (動画像符号化装置の概要)
 動画像符号化装置2は、概略的に言えば、入力画像#10を符号化することによって符号化データ#1を生成し、出力する装置である。
  (動画像符号化装置の構成)
 まず、図10を用いて、動画像符号化装置2の構成例について説明する。図10は、動画像符号化装置2の構成について示す機能ブロック図である。図10に示すように、動画像符号化装置2は、符号化設定部21、逆量子化・逆変換部22、予測画像生成部23、加算器24、フレームメモリ25、減算器26、変換・量子化部27、および可変長符号化部28を備えている。
 符号化設定部21は、入力画像#10に基づいて、符号化に関する画像データおよび各種の設定情報を生成する。
 具体的には、符号化設定部21は、次の画像データおよび設定情報を生成する。
 まず、符号化設定部21は、入力画像#10を、スライス単位、ツリーブロック単位に順次分割することにより、対象CUについてのCU画像#100を生成する。
 また、符号化設定部21は、分割処理の結果に基づいて、ヘッダ情報H’を生成する。ヘッダ情報H’は、(1)対象スライスに属するツリーブロックのサイズ、形状および対象スライス内での位置についての情報、並びに、(2)各ツリーブロックに属するCUのサイズ、形状および対象ツリーブロック内での位置についてのCU情報CU’を含んでいる。
 さらに、符号化設定部21は、CU画像#100、および、CU情報CU’を参照して、PT設定情報PTI’を生成する。PT設定情報PTI’には、(1)対象CUの各PUへの可能な分割パターン、および、(2)各PUに割り付ける可能な予測モード、の全ての組み合わせに関する情報が含まれる。
 符号化設定部21は、CU画像#100を減算器26に供給する。また、符号化設定部21は、ヘッダ情報H’を可変長符号化部28に供給する。また、符号化設定部21は、PT設定情報PTI’を予測画像生成部23に供給する。
 逆量子化・逆変換部22は、変換・量子化部27より供給される、ブロック毎の量子化予測残差を、逆量子化、および、逆DCT変換(Inverse Discrete Cosine Transform)することによって、ブロック毎の予測残差を復元する。
 また、逆量子化・逆変換部22は、ブロック毎の予測残差を、TT分割情報(後述)により指定される分割パターンに従って統合し、対象CUについての予測残差Dを生成する。逆量子化・逆変換部22は、生成した対象CUについての予測残差Dを、加算器24に供給する。
 予測画像生成部23は、フレームメモリ25に記録されている局所復号画像P’、および、PT設定情報PTI’を参照して、対象CUについての予測画像Predを生成する。予測画像生成部23は、予測画像生成処理により得られた予測パラメータを、PT設定情報PTI’に設定し、設定後のPT設定情報PTI’を可変長符号化部28に転送する。なお、予測画像生成部23による予測画像生成処理は、動画像復号装置1の備える予測画像生成部14と同様であるので、ここでは説明を省略する。
 加算器24は、予測画像生成部23より供給される予測画像Predと、逆量子化・逆変換部22より供給される予測残差Dとを加算することによって、対象CUについての復号画像Pを生成する。
 フレームメモリ25には、復号された復号画像Pが順次記録される。フレームメモリ25には、対象ツリーブロックを復号する時点において、当該対象ツリーブロックよりも先に復号された全てのツリーブロック(例えば、ラスタスキャン順で先行する全てのツリーブロック)に対応する復号画像が記録されている。
 減算器26は、CU画像#100から予測画像Predを減算することによって、対象CUについての予測残差Dを生成する。減算器26は、生成した予測残差Dを、変換・量子化部27に供給する。
 変換・量子化部27は、予測残差Dに対して、DCT変換(Discrete Cosine Transform)および量子化を行うことで量子化予測残差を生成する。
 具体的には、変換・量子化部27は、CU画像#100、および、CU情報CU’を参照し、対象CUの1または複数のブロックへの分割パターンを決定する。また、決定された分割パターンに従って、予測残差Dを、各ブロックについての予測残差に分割する。
 また、変換・量子化部27は、各ブロックについての予測残差をDCT変換(Discrete Cosine Transform)することによって周波数領域における予測残差を生成した後、当該周波数領域における予測残差を量子化することによってブロック毎の量子化予測残差を生成する。
 また、変換・量子化部27は、生成したブロック毎の量子化予測残差と、対象CUの分割パターンを指定するTT分割情報と、対象CUの各ブロックへの可能な全分割パターンに関する情報とを含むTT設定情報TTI’を生成する。変換・量子化部27は、生成したTT設定情報TTI’を逆量子化・逆変換部22に供給する。
 また、変換・量子化部27は、対象ブロックの量子化予測残差を含むTU設定情報TUI’を生成し、可変長符号化部28に供給する。
 可変長符号化部28は、TU設定情報TUI’、PT設定情報PTI’、およびヘッダ情報H’に基づいて符号化データ#1を生成し、出力する。なお、可変長符号化部28の詳細については以下に説明する。
  (可変長符号化部)
 次に、図11を用いて可変長符号化部28の構成についてさらに詳しく説明する。図11は、可変長符号化部28の構成例について示すブロック図である。
 図11に示すように、可変長符号化部28は、TU情報符号化部280、ヘッダ情報符号化部40、PTI情報符号化部41、および符号化データ多重化部42を備える。
 TU情報符号化部280は、TU設定情報TUI’を符号化して、符号化データ多重化部42に供給する。また、ヘッダ情報符号化部40は、ヘッダ情報H’を符号化して、符号化データ多重化部42に供給する。また、PTI情報符号化部41は、PTI情報PTI’を符号化して符号化データ多重化部42に供給する。
 符号化データ多重化部42は、TU設定情報TUI’、ヘッダ情報H’およびPTI情報PTI’を多重化して符号化データ#1を生成し、出力する。
 ここで、TU情報符号化部280の構成についてさらに詳しく説明すると次のとおりである。なお、以下では、TU設定情報TUI’に含まれる量子化予測残差、すなわち係数行列を符号化して係数符号化データを得るための構成について説明する。
 また、以下では、TU情報符号化部280が、16×16サイズの対象ブロックについてTU情報TUIを符号化して、符号化済みTU情報TUI’を出力する例について説明する。しかしながら、これに限られず、TU情報符号化部280が符号化する対象ブロックのサイズは、64×64、32×32等であってもかまわない。
 図11に示すように、TU情報符号化部280は、VLCテーブルTBL21と、領域分割部(変換単位分割手段)281と、領域符号化部(変換係数符号化手段)282とを備える。
 VLCテーブルTBL21は、各パラメータと、符号化データのビット列であるコードとの対応づけが定義されたテーブルである。
 VLCテーブルTBL21には、例示的に、8×8サイズの符号化単位について規定されているテーブルを用いている。すなわち、VLCテーブルTBL21には、入力となる変換単位(あるいは符号化単位)16×16サイズよりも小さな8×8サイズのものを用いている。
 また、VLCテーブルTBL21は、適応的な符号化処理ができるよう、符号化処理におけるコンテキストに応じて複数規定されている。
 領域分割部281は、対象ブロックを複数の領域に分割する。以下において、領域分割部281の分割により得られる各領域のことを符号化領域と称する。
 また、以下では、領域分割部281が、16×16サイズの対象ブロックを、4つの符号化領域に分割する例について説明する。すなわち、領域分割部281は16×16サイズの対象ブロックを4つの8×8サイズの符号化領域に分割する。
 領域分割部281は、例えば、図5を用いて示したように対象ブロックを分割することができる。なお、以下の説明では、図5において示した復号領域R11~R14を、符号化領域R11~R14と読み替える。
 すなわち、図5に示すように、領域分割部281は、対象ブロックBLKを、符号化領域R11~R14に分割する。
 しかしながら、これに限られず領域分割部281の分割の手法には、種々のものを適用することが可能である。その変形例については、後に詳しく説明する。
 領域符号化部282は、領域分割部281が対象ブロックを分割することで得られた符号化領域のそれぞれについて、当該符号化領域のサイズに応じて規定されているVLCテーブルTBL21を参照しながら符号化処理を行う。すなわち、以下では、領域符号化部282が、8×8サイズについて規定されているVLCテーブルTBL21をコンテキストに応じて参照しながら符号化処理を行う例について説明する。
 領域符号化部282の具体的構成は次のとおりである。すなわち、領域符号化部282は、最後の非ゼロ係数符号化部201と、ランモード符号化部202と、レベルモード符号化部203とを備える。
 最後の非ゼロ係数符号化部201は、符号化の対象となる符号化領域(以下、対象符号化領域と称する)について最後の非ゼロ係数を符号化する。より具体的には、最後の非ゼロ係数符号化部201は、対象符号化領域に対応する係数行列に含まれる最後の非ゼロ係数についてlast_pos、level、およびsignを符号化する。
 ランモード符号化部202は、対象符号化領域についての係数行列に対してランモード符号化処理を施す。すなわち、ランモード符号化部202は、コンテキストに応じたVLCテーブルTBL21を参照しながら、ランモードにより、係数行列に含まれる非ゼロ係数について、順次、run、levelおよびsignを符号化する。以下、ランモード符号化部202が行う符号化処理のことを、ランモード符号化処理と称する。
 また、ランモード符号化部202は、ランモード終了条件が満たされるまで、ランモード符号化処理を繰り返して行う。ランモード終了条件の条件としては、例えば、符号化した係数の値が閾値を超えている場合や、所定の数だけ係数を符号化した場合などが挙げられる。
 そして、ランモード終了条件が満たされた場合、ランモード符号化部202は、レベルモード符号化部203にレベルモードによる符号化を開始させる。
 レベルモード符号化部203は、対象符号化領域について、係数行列に含まれる係数を、レベルモードにより符号化する。すなわち、レベルモード符号化部203は、コンテキストに応じたVLCテーブルTBL21を参照しながら、レベルモードにより、係数行列に含まれる係数のlevelおよびsignを順次符号化する。
 以下、レベルモード符号化部203が行う符号化処理のことをレベルモード符号化処理と称する。また、レベルモード符号化部203は、対象符号化領域のスキャン順で最初の係数を符号化するまで、レベルモード符号化処理を繰り返して行う。
 なお、領域符号化部282は、対象ブロックの各対象符号化領域について符号化処理が完了すると、当該符号化処理により得られた係数を含む符号化済みTU設定情報TUI’を出力する。
 なお、領域符号化部282は、図6に示したとおりに符号化処理を行うことができる。以下の説明では、図6において示した復号領域R11~R14を、符号化領域R11~R14と読み替える。
 すなわち、領域符号化部282は、図6に示す符号化領域R11、R12、R13およびR14の順で符号化処理を行う。また、符号化領域R11~R14のそれぞれでは、まず、最後の非ゼロ係数復号部101が、スキャン順序上、最後の非ゼロ係数を復号する。そして、ランモード復号処理およびレベルモード復号処理により、最後の非ゼロ係数から、対象符号化領域のスキャン順で最初の係数までが、スキャン順序の逆順に復号される。
  (処理の流れ)
 上述のとおり、動画像符号化装置2の符号化処理の流れは、図4を用いて示した動画像復号装置1の復号処理の流れと概ね同じであるので、ここではその詳細な説明については省略する。
  (変形例)
 以下において、動画像符号化装置2の好ましいいくつかの変形例について説明する。
  変形例1-1’:[非ゼロ係数の有無の判定]
 動画像符号化装置2の備える領域符号化部282は、各符号化領域について非ゼロ係数の有無を判定すると共に、判定した結果を示す非ゼロ係数フラグを符号化し、係数符号化データに含める構成としてもよい。また、このような構成において、領域符号化部282は、非ゼロ係数が存在しない符号化領域についての係数の符号化を省略することができる。
 符号化領域及び、非ゼロ係数フラグの具体例については、動画像復号装置1の変形例1-1において既に述べたものと略同様であるので、ここでは説明を省略する。ただし、変形例1-1の説明における「復号領域」、「復号処理」、及び、「領域復号部122」をそれぞれ「符号化領域」、「符号化処理」、及び、「領域符号化部282」と読み替えるものとする。
  変形例1-2’:[符号化領域の位置に応じて処理方法を変更する]
  [1]スキャン方法の変更
 以上の説明では、各符号化領域のスキャン方式にジグザグスキャンを採用していたが、符号化領域の位置に応じてスキャン方法を変更してもよい。具体的なスキャン方法については、例えば、動画像復号装置1の変形例1-2[1]において説明したものと同様であるので、ここでは説明を省略する。ただし、変形例1-2[1]の説明における「復号領域」、「復号処理」、及び、「領域復号部122」をそれぞれ「符号化領域」、「符号化処理」、及び、「領域符号化部282」と読み替えるものとする。
  [2]モード限定
 以上の説明では、各符号化領域の符号化処理において、ランモード符号化処理およびレベルモード符号化処理を行っていたが、符号化領域の位置に応じてランモード符号化処理を行う構成としてもよい。具体的な処理例は、例えば、動画像復号装置1の変形例1-2[2]において説明したものと同様であるので、ここでは説明を省略する。ただし、変形例1-2[2]の説明における「復号領域」、「復号処理」、及び、「領域復号部122」をそれぞれ「符号化領域」、「符号化処理」、及び、「領域符号化部282」と読み替えるものとする。
  [3]VLCテーブルやコード番号の計算方法を変更する
 ランモード符号化処理において、符号化領域の位置に応じて、参照するVLCテーブルや、コード番号の計算方法を変更してもよい。例えば、コード番号を示すビット列を、{run,level}のパラメータの組に変換するためのVLCテーブルを各符号化領域の位置に応じて変更してもよい。具体的な処理例は、例えば、動画像復号装置1の変形例1-2[3]において説明したものと同様であるので、ここでは説明を省略する。ただし、変形例1-2[3]の説明における「復号領域」、「復号処理」、及び、「領域復号部122」をそれぞれ「符号化領域」、「符号化処理」、及び、「領域符号化部282」と読み替えるものとする。
  [4]動的最適化
 領域符号化部282による、符号化状況に応じた動的最適化の例について説明する。
  [4-1]VLCテーブルの更新
 領域符号化部282は、パラメータの値の出現頻度をカウントし、出現頻度に応じてVLCテーブルのコード番号を書き換えてもよい。
 すなわち、領域符号化部282は、出現頻度が高いパラメータの値のコード番号をより小さな(よりコードが短い)ものに書き換え、出現頻度が低いパラメータの値のコード番号をより大きな(よりコードが長い)ものに書き換えてもよい。具体的な処理例は、例えば、動画像復号装置1の変形例1-2[4-1]において説明したものと同様であるので、ここでは説明を省略する。ただし、変形例1-2[4-1]の説明における「復号領域」、「復号処理」、及び、「領域復号部122」をそれぞれ「符号化領域」、「符号化処理」、及び、「領域符号化部282」と読み替えるものとする。
  [4-2]適応速度について
 動画像符号化装置2における動的最適化の適応速度については、動画像復号装置1の変形例1-2[4-2]において既に説明したものと同様であるのでここでは説明を省略する。
  [5]ランモード終了条件の変更
 符号化領域の位置に応じてランモード終了条件を変更してもよい。また、符号化領域の位置に応じてランモード開始条件を判定してもよい。その詳細については、動画像復号装置1の変形例1-2[5]の説明において説明したとおりであるので、ここでは、その詳細な説明を省略する。
  変形例1-3’:[対象ブロックの予測モードに応じて分割の有無を決定する]
 領域分割部281は、対象ブロックの予測モードに応じて分割の有無を決定する構成であってもよい。具体的な処理は、例えば、動画像復号装置1の変形例1-3において説明したものと同様であるので、ここでは説明を省略する。
  変形例1-4’:[領域の数およびサイズ]
 領域分割部281は、対象ブロックを4つより多い領域に分割してもかまわない。具体的な処理は、例えば、動画像復号装置1の変形例1-4において説明したものと同様であるので、ここでは説明を省略する。
  変形例1-5’:[符号化領域の処理順序]
 領域符号化部282による符号化処理のスキャン順序は、上述した例に限定されるものではない。領域符号化部282は、例えば、動画像復号装置1の変形例1-5において説明した処理と同様の処理を行う構成としてもよい。
 なお、上述のように、動画像復号装置1の[変形例]は、動画像符号化装置2に適用することが可能である。逆に、ここで示した動画像符号化装置2の[変形例]についても、符号化処理を復号処理に変更することで動画像復号装置1に適応することができる。
  (作用・効果)
 以上に説明したように、動画像復号装置1は、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データのTU情報TUIから、該変換係数を復号する動画像復号装置1において、上記変換単位である対象ブロックを、複数の復号領域に分割する領域分割部121と、TU情報TUIから上記変換係数を得るための復号情報であって、上記復号領域ごとに割り当てられているVLCテーブルTBL11を参照して、上記復号領域に含まれる変換係数を復号する領域復号部122と、を備える構成である。
 また、動画像符号化装置2は、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する動画像符号化装置2において、上記変換単位である対象ブロックを、複数の符号化領域に分割する領域分割部281と、上記変換係数を符号化するためのVLCテーブルTBL21であって、上記符号化領域ごとに割り当てられているVLCテーブルTBL21を参照して、上記変換単位に含まれる変換係数を符号化する領域符号化部282と、を備える構成である。
 動画像復号装置1の上記構成によれば、8×8サイズの復号領域について規定されているVLCテーブル11に基づいて復号処理を行うので、元の対象ブロックのサイズ(16×16)について規定されているVLCテーブルに基づいて復号処理を行うのに比べて、VLCテーブルのサイズを低減することができる。また、スキャン順を表すテーブルについても同様に、8×8サイズの復号領域について規定されているテーブルに基づいて復号処理を行うので、テーブルのサイズを低減することができる。
 なお、動画像符号化装置2についても同様の作用・効果を得ることができる。
〔2〕実施形態2
 本発明の他の実施形態について図15~図22に基づいて説明すると、以下の通りである。なお、説明の便宜上、前記実施形態1にて説明した図面と同じ機能を有する部材については、同じ符号を付記し、その説明を省略する。
 以下では、相対位置指定による係数の符号化および復号について示す。また、以下の説明では対象ブロックのサイズは、一例として16×16であるとする。
 〔動画像復号装置〕
 まず、図15を参照しながら、動画像復号装置1の構成について説明すると以下のとおりである。本実施形態に係る動画像復号装置1では、図2に示した動画像復号装置1において、TU情報復号部12を、図15に示すTU情報復号部12Aに変更する。
 以下、図15に示すTU情報復号部12Aについて説明すると次のとおりである。すなわち、図15に示すように、TU情報復号部12Aは、VLCテーブルTBL30、ラン-レベルモード復号部310、相対位置モード復号部320、および処理モード制御部330を備える。
 VLCテーブルTBL30は、ビット列(コード)と相互に変換可能なコード番号と、復号すべきパラメータとが対応付けられたテーブルである。VLCテーブルTBL30は、後述するラン-レベルモード復号部310によって参照されるラン-レベルモード用テーブルTBL31と、相対位置モード復号部320によって参照される相対位置モード用テーブルTBL32とを備える。
 ラン-レベルモード用テーブルTBL31は、図1に示したTU情報復号部12のVLCテーブルTBL11と同様のものを用いることができる。よって、ここではその説明は省略する。また、相対位置モード用テーブルTBL32の定義については、後述する。
 ラン-レベルモード復号部310は、処理モード制御部330の制御により、ランモード復号処理およびレベルモード復号処理を行う。以下、ラン-レベルモード復号部310による復号処理のことを、ラン-レベルモード復号処理と称する。なお、ランモード復号処理およびレベルモード復号処理については、前記実施形態1において既に説明済みであるので、ここではその説明を省略する。
 ラン-レベルモード復号部(復号手段)310は、最後の非ゼロ係数復号部311、ランモード復号部312、およびレベルモード復号部313を備える。
 最後の非ゼロ係数復号部311、ランモード復号部312、およびレベルモード復号部313は、それぞれ、図1を用いて示した、最後の非ゼロ係数復号部101、ランモード復号部102、およびレベルモード復号部103と同様の機能を有している。よって、その機能については既に説明済みであるので、ここではその説明を省略する。
 なお、ランモード復号部102、およびレベルモード復号部103は、ラン-レベルモード復号処理において、ラン-レベルモード用テーブルTBL31を参照するように構成されている。
 相対位置モード復号部320は、非ゼロ係数の相対位置が符号化された係数符号化データの復号処理を行う。相対位置モード復号部320は、具体的には、最後の非ゼロ係数復号部321、相対位置復号部(相対位置復号手段)322、および係数位置決定部(位置特定手段)323を備える。
 最後の非ゼロ係数復号部321は、対象ブロックにおける最後の非ゼロ係数を復号する。最後の非ゼロ係数復号部321は、対象ブロックについての係数符号化データに含まれるlast_pos、level、およびsignを復号する。また、最後の非ゼロ係数復号部321は、last_posを、DC係数を原点(0,0)とする座標表示(lastx,lasty)に変換する。以下、このDC係数を原点(0,0)とする座標表示で示される位置のことを絶対位置と称する。
 相対位置復号部322は、対象ブロックについて、非ゼロ係数の相対位置が符号化された係数符号化データから、復号対象となる非ゼロ係数の相対位置(dx,dy)と、非ゼロ係数の値(levelおよびsign)とを復号する。ここで非ゼロ係数の相対位置とは、ひとつ前に復号した非ゼロ係数の絶対位置から見た、復号対象となる非ゼロ係数の相対的な位置のことをいう。また、以下において、このような相対位置による非ゼロ係数の位置の表現を相対位置指定と称する。
 係数位置決定部323は、相対位置復号部322が復号した非ゼロ係数の相対位置(dx,dy)、および、ひとつ前に復号した非ゼロ係数の絶対位置(x,y)から復号対象となる非ゼロ係数の絶対位置を決定する。
 最後の非ゼロ係数Cからn+1番目の非ゼロ係数Cn+1までの相対位置は、例えば、次の関係式(1-1)~(1-3)にて表現することができる。
 C :x=lastx, y=lasty … (1-1)
 C :x=x+dx, y=y+dy … (1-2)
 …
 Cn+1:xn+1=x+dx, yn+1=y+dy … (1-3)
 係数位置決定部323は、上記関係式(1-1)~(1-3)を用いて復号対象となる非ゼロ係数の絶対位置を決定する。
 処理モード制御部330は、復号対象となる非ゼロ係数が、所定の領域内に位置するか否かを判定し、当該判定結果に応じて、ラン-レベルモード復号部310によって復号処理を行うか、相対位置モード復号部320によって復号処理を行うかを制御する。
 具体的には、処理モード制御部330は、復号対象となる非ゼロ係数の位置が低周波成分側の8×8サイズの領域内か否かを判定する。
 すなわち、上記関係式(1-1)~(1-3)の例に従えば、復号対象となる非ゼロ係数Cの位置(x,y)が、“x<8&&y<8”であるか否かを判定する。なお、“&&”は、論理積を示す演算子である。
 復号対象となる非ゼロ係数の位置が低周波成分側の8×8サイズの領域内であれば、処理モード制御部330は、ラン-レベルモード復号部310に次の非ゼロ係数Cn+1からの復号処理を実行させる。
 復号対象となる非ゼロ係数の位置が低周波成分側の8×8サイズの領域内でなければ、処理モード制御部330は、相対位置モード復号部320に復号処理を実行させる。
  (処理の流れ)
 次に、図16を用いて、TU情報復号部12Aにおける復号処理について説明する。図16は、相対位置指定により非ゼロ係数を符号化/復号する処理S20の流れについて例示したフローチャートである。
 なお、図16では、動画像符号化装置2の符号化処理と動画像復号装置1の復号処理とをまとめて示している。
 図16に示すように、相対位置指定により非ゼロ係数を復号する処理S20が開始されると、まず、動画像復号装置1では、処理モード制御部330が相対位置モード復号部320に復号処理を実行させる。これに応じて、最後の非ゼロ係数復号部321が、対象ブロックの最後の非ゼロ係数を復号する(S21)。
 処理モード制御部330は、S21において復号された最後の非ゼロ係数の位置が、相対位置指定された係数を復号する領域内であるか否かを判定する(S22)。
 相対位置指定された係数を復号する領域内である場合(S22においてYES)、処理モード制御部330は、相対位置モード復号部320に復号処理を実行させる。これに応じて、相対位置復号部322が、相対位置指定された係数の位置を復号し、係数位置決定部323が、当該係数の絶対位置を決定することで復号対象となる係数を復号する(S23)。
 以下、復号された係数の位置が、相対位置指定された係数を復号する領域内である場合、相対位置指定による係数の復号処理を継続して行う(S22、S23)。
 復号された係数の位置が、相対位置指定された係数を復号する領域内でなくなった場合(S22においてNO)、処理モード制御部330は、ラン-レベルモード復号部310に復号処理を実行させる。これに応じて、ランモード復号部312がランモード復号処理を実行し、その後、レベルモード復号部313がレベルモード復号処理を実行する(S24)。このようにして、対象ブロックのすべての係数が復号され相対位置指定により係数を復号する処理が終了する。
  (具体例)
 図17を用いて、TU情報復号部12Aにおける復号処理の具体例について説明する。図17は、TU情報復号部12Aの復号処理の実行例を示す図である。
 図17に示すように、対象ブロックBLKは、2つの領域R1およびR2からなる。ここで、対象ブロックBLKは、以上の説明と同様、16×16サイズのブロックであるとし、領域R1は、低周波成分側の8×8サイズの領域であるとする。
 なお、領域R1に示す矢印は、領域R1におけるスキャン順序を示しており、そのスキャン順序は、ジグザグスキャンである。また、領域R2は、相対位置指定により係数を復号する領域である。
 相対位置指定により係数を復号する処理が開始されると、まず、最後の非ゼロ係数復号部321が、対象ブロックBLKにおける最後の非ゼロ係数Cを復号する(S21)。
 最後の非ゼロ係数Cは、領域R2に位置するので(S22においてYES)、処理モード制御部330は、相対位置モード復号部320に復号処理を実行させる。
 そして、相対位置復号部322が、非ゼロ係数CN-1の相対位置を復号し、係数位置決定部323が、復号された当該相対位置と、最後の非ゼロ係数Cの絶対位置とから、非ゼロ係数CN-1の絶対位置を決定する。さらに、level、signが復号されることで非ゼロ係数CN-1が復号される(S23)。
 非ゼロ係数CN-1以下、非ゼロ係数Cまでは、領域R2に位置するので、上記と同様に、相対位置モード復号部320が復号処理を実行する(S22、S23)。
 さらに、非ゼロ係数Cまで、相対位置モード復号部320が復号処理を完了すると、非ゼロ係数Cは、領域R1に位置するので、(S22においてNO)、処理モード制御部330は、ラン-レベルモード復号部310に復号処理を実行させる。
 ラン-レベルモード復号部310では、ランモード復号部312が、非ゼロ係数Cを基点として、領域R1においてランモード復号処理を実行する。また、ランモード復号処理が終了すると、レベルモード復号部313が、レベルモード復号処理を実行する。
 ここで、ラン-レベルモード復号部310の復号処理は、8×8サイズの符号化単位を(逆順の)ジグザグスキャンにより復号する処理と同様のものが採用可能である。また、領域R1において最初に復号される非ゼロ係数Cは、8×8サイズの領域R1における最後の非ゼロ係数であることが好ましい。
  (実施例)
 以下において、TU情報復号部12Aのより具体的な実施例について説明する。相対位置(dx,dy)の復号に用いる相対位置モード用テーブルTBL32は、次のように構成することができる。
 すなわち、処理対象の非ゼロ係数とひとつ前の非ゼロ係数との相対距離が小さいほど、相対位置モード用テーブルTBL32において短いコード(ビット列)が対応付けられていてもよい。処理対象の非ゼロ係数とひとつ前の非ゼロ係数との相対距離は、例えば、処理対象の非ゼロ係数の相対位置(dx、dy)から導出することが可能である。
 あるいは、出現頻度がより高い相対距離を示す(dx、dy)の組に、相対位置モード用テーブルTBL32おいて、より短いコードを対応付けるようにしてもよい。
 上記構成によれば、相対距離や、出現頻度に応じて適応的に短いコードを割り当てることができるので復号する符号量の低減を図ることができる。
 なお、dxおよびdyの絶対値の最大値は、符号化対象ブロックの辺の長さ-1である。つまり、符号化対象ブロックが、16×16である場合、dxおよびdyの絶対値の最大値は、15である。
  (変形例)
  [所定の領域の変更]
 処理モード制御部330が、対象ブロックのサイズを判定し、当該判定に応じて、ラン-レベルモード復号部310によって復号処理を行うか、相対位置モード復号部320によって復号処理を行うかを制御してもよい。
 例えば、処理モード制御部330は、対象ブロックのサイズが8×8以下であれば、ラン-レベルモード復号部310に復号処理を実行させる。また、処理モード制御部330は、対象ブロックのサイズが16×16以上であれば、相対位置モード復号部320に復号処理を実行させる。
 また、処理モード制御部330は、対象ブロックにおいて、復号する非ゼロ係数の数が所定の個数、例えば、64個以下であれば、ラン-レベルモード復号部310に復号処理を実行させてもよい。
 また、係数符号化データにおいて、すべての非ゼロ係数が相対位置の指定により表現されていてもよい。この場合、すべての非ゼロ係数を相対位置モード復号部330が復号する。
 また、処理モード制御部330は、所定の領域を、スライスタイプや、予測モード、対象ブロックのサイズに応じて変更してもよい。
 例えば、処理モード制御部330は、対象ブロックにおいて、イントラ予測が符号化されていた場合、低周波成分側の8×8サイズの領域を所定の領域として設定する。また、例えば、処理モード制御部330は、対象ブロックにおいて、インター予測が符号化されていた場合、低周波成分側の4×4サイズの領域を所定の領域として設定する。
 [相対位置モード用テーブルの変更]
 相対位置モード用テーブルTBL32を、非ゼロ係数の絶対位置(x,y)に応じて複数用意してもよい。この相対位置モード用テーブルTBL32は、それぞれの非ゼロ係数の絶対位置において、dx、dyが取り得る値の範囲に基づいて最適化しておくことが好ましい。
そして、相対位置モード復号部420は、非ゼロ係数の絶対位置(x,y)に応じて参照する相対位置モード用テーブルTBL32を変更してもよい。
 上記構成によれば、適応的に参照する相対位置モード用テーブルTBL42を切り替えることができるので、復号する符号量を低減することができる。
 なお、本変形例の詳細については、後の動画像符号化装置2の説明において行う。
  [非ゼロ係数の相対位置の表現]
 以上では、非ゼロ係数の相対位置を(dx,dy)の形式で表現したが、これに限られない。例えば、相対位置は、方向と距離により表現されていてもよい。
   [VLCテーブルについて]
 図20に示すように相対位置モード用テーブルTBL32を構成してもよい。図20に示す相対位置モード用テーブルTBL32では、dxまたはdyの値が小さいほど、小さいコード番号が対応付けられている。
 例えば、dxまたはdyが0のとき、コード番号は“0”である。つまり、(dx,dy)=(0,1)であれば、コード番号“0,1”を割り当てる。
 さらに図20を参照すれば、dx,dy“-1”に“1”が対応付けられており、dx,dy“1”に“2”が対応付けられている。このように、相対位置モード用テーブルTBL32では、絶対値が同じとなる値であれば符号が負になるほうが、小さいコード番号が割り当てられている。
 なお、図20に示す相対位置モード用テーブルTBL32では、コード番号が小さいほど短いコードが対応付けられていることが好ましい。
  (変形例)
   [dx、dyが取り得る値の範囲に基づいてVLCテーブルを最適化する手法]
 相対位置モード用テーブルTBL32を、係数の絶対位置(x,y)に応じて複数用意してもよい。この相対位置モード用テーブルTBL32は、それぞれの係数の絶対位置において、dx、dyが取り得る値の範囲に基づいて最適化しておくことが好ましい。
 図21および図22を用いて、その具体的構成例について以下に示す。図21は、図17や図19(後に説明する)に示す対象ブロックBLKにおける領域R2を、さらに領域R2a、領域R2b、および領域R2cの3つの領域に細分化する例について示している。
 また、図22は、領域R2a、領域R2b、および領域R2cに対応付けるVLCテーブルの一例を示している。
 同図の(a)は、xまたはyの値が、所定以上の正の値にならない場合に相対位置モード復号部320の相対位置決定部323が参照する相対位置モード用テーブルTd1を示している。
 また、同図の(b)は、xまたはyの値が、所定以下の負の値にならない場合に、相対位置モード復号部320の相対位置決定部323が参照する相対位置モード用テーブルTd2を示している。
 図21に示す対象ブロックBLKは、領域R1、領域R2a~R2cから構成されている。図21では、対象ブロックBLKのDC係数の位置を、(0,0)の座標表示で表している。また、対象ブロックBLK内における、ジグザグスキャン順上最後の係数の位置、すなわち最も高周波の成分の位置を(15,15)で表している。
 図21では、例示的に、領域R1は、(0,0)を左上頂点とする正方形の内側領域としている。また、領域R2aは、(8,0)を左上頂点とする正方形の内側領域としている。また、領域R2bは、(0,8)を左上頂点とする正方形の内側領域とし、領域R2cは、(8,8)を左上頂点とする正方形の内側領域としている。
 図22の(a)に示す相対位置モード用テーブルTd1は、係数の値“0”にコード番号“0”が割り当てられている。以下、係数の値“-1”、“1”…“-7”、“7”に、それぞれ、コード番号“1”、“2”、…、“13”、“14”が割り当てられている。
 つまり、係数の絶対値が小さいほうが、より小さいコード番号が割り当てられている。また、係数の絶対値が同じで正負符号が異なる場合、負の値のほうが、正の値よりも小さなコード番号が割り当てられている。例えば、係数の値“-1”、“1”…“-7”、“7”には、それぞれ、コード番号“1”、“2”、…、“13”、“14”が割り当てられている。
 なお、係数の絶対値が“8”~“15”の範囲においては、負の値“-8”~“-15”のみが定義されている。
 一方、図22の(b)に示す相対位置モード用テーブルTd2は、係数の絶対値が“0”~“7”の範囲においては、相対位置モード用テーブルTd1と同じ定義となっている。
 なお、係数の絶対値が“8”~“15”の範囲においては、正の値“8”~“15”のみが定義されている。
 相対位置決定部323は、(dx,dy)の位置の取り得る範囲に応じて、下記のとおり参照するVLCテーブルを切り替える。
 領域R2aについては、dxの取り得る範囲は、-15≦dx≦7であるので、相対位置決定部323は、相対位置モード用テーブルTd1を参照する。また、dyの取り得る範囲は、-7≦dy≦15であるので、相対位置決定部323は、相対位置モード用テーブルTd2を参照する。
 また、領域R2bについては、dxの取り得る範囲は、-7≦dx≦15であるので、相対位置決定部323は、相対位置モード用テーブルTd2を参照する。また、dyの取り得る範囲は、-15≦dy≦7であるので、相対位置決定部323は、相対位置モード用テーブルTd1を参照する。
 また、領域R2cについては、dxの取り得る範囲は、-15≦dx≦7であるので、相対位置決定部323は、相対位置モード用テーブルTd1を参照する。また、dyの取り得る範囲は、-15≦dy≦7であるので、相対位置決定部323は、相対位置モード用テーブルTd1を参照する。
 以上のように、dx,dyの値の取り得る範囲に基づいてVLCテーブルを最適化することで、符号化効率を向上するこができる。
 〔動画像符号化装置〕
 まず、図18を参照しながら、動画像符号化装置2の構成について説明すると以下のとおりである。本実施形態に係る動画像符号化装置2では、図11に示した動画像符号化装置2の可変長符号化部11において、TU情報符号化部280を、図18に示すTU情報符号化部280Aに変更する。
 以下、図18に示すTU情報符号化部280Aについて説明すると次のとおりである。すなわち、図18に示すように、TU情報符号化部280Aは、VLCテーブルTBL40、ラン-レベルモード符号化部410、相対位置モード符号化部420、および処理モード制御部430を備える。
 VLCテーブルTBL40は、各パラメータと、符号化データのビット列であるコードとの対応づけが定義されたテーブルである。VLCテーブルTBL40は、後述するラン-レベルモード符号化部410によって参照されるラン-レベルモード用テーブルTBL41と、相対位置モード符号化部420によって参照される相対位置モード用テーブルTBL42とを備える。
 ラン-レベルモード用テーブルTBL41は、図11に示したTU情報符号化部280のVLCテーブルTBL21と同様のものを用いることができる。よって、ここではその説明は省略する。また、相対位置モード用テーブルTBL42の定義については、後述する。
 ラン-レベルモード符号化部410は、処理モード制御部330の制御により、ランモード符号化処理およびレベルモード符号化処理を行う(以下、ラン-レベルモード符号化処理と称する)。なお、ランモード符号化処理およびレベルモード符号化処理については、前記実施形態1において既に説明済みであるので、ここではその説明を省略する。
 ラン-レベルモード符号化部410は、最後の非ゼロ係数符号化部411、ランモード符号化部412、およびレベルモード符号化部413を備える。
 最後の非ゼロ係数符号化部411、ランモード符号化部412、およびレベルモード符号化部413は、それぞれ、図11を用いて示した、最後の非ゼロ係数符号化部201、ランモード符号化部202、およびレベルモード符号化部203と同様の機能を有している。よって、その機能については既に説明済みであるので、ここではその説明を省略する。
 なお、ランモード符号化部202、およびレベルモード符号化部203は、ラン-レベルモード用テーブルTBL41を参照するように構成されている。
 相対位置モード符号化部420は、符号化対象となる非ゼロ係数の対象ブロックにおける相対位置を符号化して係数符号化データを生成する。具体的には、相対位置モード符号化部420は、最後の非ゼロ係数符号化部421、相対位置算出部(相対位置符号化手段)422、および相対位置符号化部(相対位置符号化手段)423を備える。
 最後の非ゼロ係数符号化部421は、対象ブロックにおける最後の非ゼロ係数を符号化する。最後の非ゼロ係数符号化部421は、対象ブロックにおいて、逆順のジグザグスキャンに従って最後の非ゼロ係数のlast_pos、level、およびsignを符号化する。また、最後の非ゼロ係数符号化部421は、last_posを、DC係数を原点(0,0)とする座標表示(lastx,lasty)に変換する。以下、このDC係数を原点(0,0)とする座標表示で示される位置のことを絶対位置と称する。
 相対位置算出部422は、符号化対象となる係数の相対位置を、当該係数の絶対位置と、ひとつ前に符号化した非ゼロ係数の絶対位置とから算出する。
 相対位置算出部422は、符号化対象となる非ゼロ係数の相対位置(dx,dy)を、前述の関係式(1-1)~(1-3)に基づいて算出することができる。
 相対位置符号化部423は、相対位置モード用テーブルTBL42を参照しながら、所定の順番で、相対位置算出部422が算出した非ゼロ係数の相対位置(dx,dy)と、非ゼロ係数のlevelおよびsignとを符号化することにより係数符号化データを生成する。
 相対位置符号化部423が対象ブロックにおいて、どのような順番で非ゼロ係数を相対位置指定により符号化するかについては、後の実施例において詳しく説明する。
 処理モード制御部430は、符号化対象となる非ゼロ係数が、所定の領域内に位置するか否かを判定し、当該判定結果に応じて、ラン-レベルモード符号化部410によって符号化処理を行うか、相対位置モード符号化部420によって符号化処理を行うかを制御する。
 処理モード制御部430の制御の手法については、TU情報復号部12Aの処理モード制御部330について説明したものに準ずるので、ここではその説明を省略する。
  (処理の流れ)
 上述のとおり、動画像符号化装置2の符号化処理の流れは、図16を用いて示した動画像復号装置1の復号処理の流れと概ね同じであるので、ここではその詳細な説明については省略する。
  (実施例)
   [符号化順序について]
 以下、図19を用いて、相対位置算出部422が対象ブロックにおいてどのような順番で非ゼロ係数を相対位置指定により符号化するかについての実施例を示す。図19は、相対位置指定による非ゼロ係数の符号化の例について示す図である。
 図19に示す対象ブロックBLKは、例示的に、16×16サイズであるとし、領域R1は、当該対象ブロックにおける低周波成分側の8×8サイズの領域であるとする。また、領域R2は、対象ブロックにおける領域R1以外の領域である。
 また、処理モード制御部430は、領域R1については、ラン-レベルモード符号化部410に符号化処理を実行させ、領域R2については、相対位置モード符号化部420に符号化処理を実行させる。
 なお、以下では、例示的に、予め領域R2におけるN個の非ゼロ係数C(1)~C(N)が検出されており、かつ領域R1における最後の非ゼロ係数Cが検出されているものとする。
 相対位置符号化部423は、以下の工程を経て連鎖的に非ゼロ係数を符号化する(或いは、非ゼロ係数の相対位置を数珠繋ぎで符号化すると表現してもよい)。
 工程[1] まず、相対位置符号化部423は、領域R1における最後の非ゼロ係数Cを基点とし、所定の選択基準により、領域R2における非ゼロ係数C(1)~C(N)の中から、非ゼロ係数Cに連鎖させる次の非ゼロ係数Cを選択する。
 所定の選択基準としては、例えば、相対位置算出部422によって算出されたCと非ゼロ係数C(1)~C(N)との相対位置(dx,dy)の符号化をそれぞれ試行して得られた符号量である。この場合、相対位置符号化部423は、符号量が小さい非ゼロ係数を次の非ゼロ係数Cとして選択する。
 工程[2] 相対位置符号化部423は、選択した非ゼロ係数を基点に、上記選択基準による選択工程を、領域R2における未選択の非ゼロ係数が無くなるまで順に繰り返す。
 上記選択基準は、より一般的には、CとCn+1との相対位置(dx,dy)の符号量であるといえる。
 なお、上記選択基準は、単なる例示に過ぎずこれに限られない。例えば、CとCn+1とのマンハッタン距離(|dx|+|dy|)や、CとCn+1とのユークリッド距離(dx+dy)などでもよい。
 上記工程[2]の選択工程を繰り返すことにより、非ゼロ係数C~Cが連鎖される。結果として、図19に例示するような非ゼロ係数C~Cの連鎖が得られる。
 工程[3] 続いて、相対位置符号化部423は、非ゼロ係数Cの位置、絶対値(level)、正負の符号(sign)を符号化する。上記位置は、例えば、符号化対象ブロックの右下からの相対位置(dx,dy)で指定してもよいし、スキャン順序によるLastPosで指定してもよい。なお、非ゼロ係数Cは、必ずしも符号化対象ブロックを何らかのスキャン順序でスキャンした場合における最後の非ゼロ係数でなくてもかまわない。
 工程[4] 相対位置符号化部423は、上記工程[2]における選択工程と逆順に、非ゼロ係数CN-1からCまで、それぞれの相対位置(dx,dy)、絶対値(level)、正負の符号(sign)を符号化する。
 非ゼロ係数Cを符号化が完了すると、処理モード制御部430が、ラン-レベルモード符号化部410に符号化処理を実行させる。ラン-レベルモード符号化部410が、ラン-レベルモード符号化処理により、非ゼロ係数C以外の領域R1の係数を符号化する。
   [VLCテーブルについて]
 相対位置モード用テーブルTBL42は、図20に示すように構成してもよい。図20に示す相対位置モード用テーブルについては、既に説明したため、ここでは説明を省略する。
  (変形例)
   [dx、dyが取り得る値の範囲に基づいてVLCテーブルを最適化する手法]
 相対位置モード用テーブルTBL42を、係数の絶対位置(x,y)に応じて複数用意してもよい。この相対位置モード用テーブルTBL42は、それぞれの係数の絶対位置において、dx、dyが取り得る値の範囲に基づいて最適化しておくことが好ましい。
 本変形例の具体的な処理は、動画像復号装置の説明において図21及び図22を用いて説明したものと同様であるので、ここではその説明を省略する。ただし、「相対位置決定部323」を「相対位置符号化部423」と読み替えるものとする。
   [その他]
 なお、本実施形態に係る動画像復号装置1の[変形例]についても動画像符号化装置2に適用することが可能である。
  (作用・効果)
 以上に説明したように、動画像復号装置1は、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数が符号化された符号化データのTU情報TUIから、該変換係数を復号する動画像復号装置1において、復号対象となる変換係数のひとつ前に復号した変換係数からの相対位置を復号する相対位置復号部320と、ひとつ前に復号した上記変換係数の、上記変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する係数位置決定部323と、を備える構成である。
 また、動画像符号化装置2は、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する動画像符号化装置2において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を符号化する相対位置符号化部423を備える構成である。
 高周波成分側の領域では、係数が疎らになる傾向がある。このため、スキャン順序に従ってランを符号化した場合、runの長さが非常に長くなる傾向がある。このため、符号化および復号に大きなテーブルを用いなければならなくなったり、符号量が増大したりする傾向がある。また、対象ブロックのサイズが大きいほどこれらの傾向がより顕著に現れる。
 また、{run,level}の組み合わせを表すVLCテーブルのサイズは、基本的にはスキャン順におけるrunの長さの最大値、すなわち、対象ブロックの面積に比例する。
 動画像復号装置1の上記構成によれば、runの長さに基づく復号処理ではなく、相対位置に基づく復号処理を行うので、{run,level}の組み合わせを表すVLCテーブルを用いなくても済む。
 この結果、VLCテーブルのサイズを低減することができる。なお、動画像符号化装置2についても同様の作用・効果を得ることができる。また、大きな符号化対象ブロック全体のスキャン順序を表すテーブルも不要である。
〔3〕実施形態3
 本発明のさらに他の実施形態について図23~図25に基づいて説明すると、以下の通りである。なお、説明の便宜上、前記実施形態1にて説明した図面と同じ機能を有する部材については、同じ符号を付記し、その説明を省略する。
 以下では、処理A:『対象ブロックにおける低周波成分側のn個の係数(例えば16×16以上のサイズではn=64)のみ符号化/復号する処理』と、処理B:『領域分割して係数を符号化/復号する処理(図4、S10参照)』とを切り替えながら符号化処理/復号処理を行う方法について説明する。
 また、符号化データには、処理Aと処理Bとを切り替えるための符号化方法識別子が符号化されているものとする。すなわち、以下の例では、符号化方法識別子に、処理Aおよび処理Bのいずれかが指定されている。
 TU情報復号部12が上記処理Bを行うことについては既に説明したが、以下の例では、TU情報復号部12が上記処理Bに加えて上記処理Aも実行することとする。
 また、以下の説明では対象ブロックのサイズは、一例として16×16であるとする。
  (処理の流れ)
 図23を用いて、処理Aと処理Bとを切り替えながら符号化/復号する処理の流れについて説明する。図23は、処理Aと処理Bとを切り替えながら符号化/復号する処理の流れの一例について示すフローチャートである。
 図23では、動画像符号化装置2の符号化処理と動画像復号装置1の復号処理とをまとめて示している。なお、以下では、動画像復号装置1側の動作について説明するが、動画像符号化装置2側の動作も概ね同じである。
 また、以下の例では、動画像復号装置1のTU情報復号部12が符号化方法識別子を判定することとする。
 処理が開始すると、図23に示すように、まず、TU情報復号部12が符号化方法識別子を判定する(S101)。
 符号化方法識別子が処理Aを示している場合(S101において処理A)、TU情報復号部12は、処理Aを実行する(S102)。すなわち、TU情報復号部12は、対象ブロックの低周波成分側に位置する左上64個の係数のみを復号する。
 ここで、図24を用いて、処理Aにおいて復号対象とする係数について詳しく説明すると次のとおりである。対象ブロックが、イントラ予測を行うブロックであるか、インター予測を行うブロックかに応じて処理Aにおいて復号対象にする係数を変更してもよい。
 具体例を示すと次のとおりである。まず、インター予測を行うブロックについては、図24の(a)に示すように、TU情報復号部12は、低周波成分側の左上8×8サイズの領域RInterに位置する係数を復号対象とする。
 また、イントラ予測を行うブロックについては、TU情報復号部12は、図24の(b)に示す領域RIntraに位置する係数を復号対象とする。すなわち、TU情報復号部12は、ジグザグスキャン順序上、1番目のDC係数から、64番目の係数までを復号対象とする。領域RIntra内に示す数字“64”は、領域内に含まれる係数の数を示している。
 なお、図24の(b)では、説明の便宜のため、領域RIntraの形状を直角三角形にて示しているが、実際のところ厳密にいえば、領域RIntraは、直角三角形にならない点には留意されたい。
 一方、符号化方法識別子が処理Bを示している場合(S101において処理B)、TU情報復号部12は、処理Bを実行する(S10,図4参照)。
  (データ構造)
 図25を用いて、係数符号化データのデータ構造について例示する。図25は、係数符号化データのデータ構造について示す図である。
 符号化方法識別子FLGは、上述したとおり、処理Aまたは処理Bが指定されているフラグである。
 符号化方法識別子FLGに処理Aが指定される場合、係数符号化データは、DATA1に示すデータ構造を採用することができる(以下、係数符号化データDATA1と表記する)。符号化データDATA1には、対象ブロックの低周波成分側に位置する左上64個分の係数データ(run,level,sign)が含まれる。
 一方、符号化方法識別子FLGに処理Bが指定される場合、係数符号化データは、DATA2に示すデータ構造を採用することができる(以下、係数符号化データDATA2と表記する)。
 係数符号化データDATA2は、非ゼロ係数フラグ×n(n=領域数)、係数データ[領域1]~[領域n]を含む。
 非ゼロ係数フラグとは、上述のとおり、復号領域における非ゼロ係数の有無を示す。非ゼロ係数フラグは、復号領域の数だけ符号化されているため“×n”で示している。
 例えば、非ゼロ係数フラグが“1(真)”であれば、その復号領域に非ゼロ係数が有ることを示し、非ゼロ係数フラグが“0(偽)”であれば、その復号領域に非ゼロ係数が無いことを示す。
 係数データ[領域x](x=1~n)は、各復号領域における係数データを含む。なお、非ゼロ係数フラグが領域xにおける非ゼロ係数が無いことを示す場合、係数データ[領域x]は省略される。
  (作用・効果)
 対象ブロックのサイズが大きい場合、符号化する係数の数が多くなるため、処理Aおよび処理Bの間における効率の差は大きく表れる。さらにいえば、処理Aおよび処理Bの間で、得意とする映像特性が異なる。よって、一方の処理方式でのみ符号化を行うと、大きく符号化効率を低下させるおそれがある。
 上記処理Aと処理Bとを切り替えながら符号化/復号することにより、処理方式の多様化による符号化効率の低下の抑止、または向上を図ることができる。効率化により、処理Aおよび処理Bを切り替えるための符号化方式識別子の符号量を相殺することも可能である。
  (変形例)
 また、符号化方式識別子を符号化する単位は、任意である。例えば、LCU単位で符号化することができる。
 また、符号化の状況に応じて条件判定により、処理Aおよび処理Bを切り替えることも可能である。この場合、符号化方法識別子によって処理Aおよび処理Bを明示しなくても、暗黙的に実行する処理方式が決まるため符号化方法識別子を省略することもできる。
 条件判定の基準としては、ブロックの属性や状態、所定のパラメータとすることができる。
 より具体的には次のとおりである。対象ブロック(変換単位)のサイズに基づいて判定してもよい。例えば、16×16であれば、処理Aを実行し、32×32であれば、処理Bを実行するように構成することが可能である。
 また、対象ブロックの予測モードで判定してもよい。例えば、イントラ予測モードである場合、処理Aを実行し、インター予測モードであれば、処理Bを実行するように構成することが可能である。
 また、対象ブロックの隣接ブロックの変換単位サイズに基づいて判定してもよい。例えば、対象ブロックの左隣接ブロックの変換単位が所定サイズ(例えば、32×32サイズ)以上であれば、当該対象ブロックでは、処理Aを実行し、そうでなければ、処理Bを実行するように構成してもよい。
 また、“所定サイズ以上”のほかにも、“対象ブロックのサイズ以上”、“所定サイズ未満”、“所定サイズに等しい”といった判定条件を適用することも可能である。
 変換単位のサイズが小さい場合、対象ブロックに含まれるエッジが多くなり、これにより対象ブロックに含まれる非ゼロ係数の数が多くなる傾向がある。よって、変換単位のサイズが小さい場合、処理Bを実行するように構成することがより好ましい。
 また、符号化方法識別子以外の所定のパラメータを処理方式の判定に用いてもよい。ヘッダ等に付加されているプロファイル識別子に基づいて処理方式を切り替えてもかまわない。また、ヘッダ等に付加されているデコーダの能力やビット・ストリームの複雑さを規定するレベルに基づいて処理方式を切り替えてもかまわない。
 また、処理Bは、図4に示した、領域分割して係数を符号化/復号する処理としたが、これに限られない。処理Bは、図16に示した、相対位置指定により係数を符号化/復号する処理S20であってもよい。
 また、例えば、動画像符号化装置2は、処理S10および処理S20の符号化を試行し、あるいは符号量を推定し、符号化効率が良い処理を符号化方法識別子に指定してもよい。
〔4〕実施形態4
 本発明のさらに他の実施形態について図26~図36に基づいて説明すると、以下の通りである。なお、説明の便宜上、前記実施形態1にて説明した図面と同じ機能を有する部材については、同じ符号を付記し、その説明を省略する。
  (階層的領域分割)
 以下では、対象ブロックを階層的に領域分割して符号化/復号する方法について説明する。
 動画像符号化装置2は、対象ブロックBLKの符号化において、対象ブロックBLKを階層的に分割してもよい。この場合、対象ブロックBLKの符号化において、動画像符号化装置2のTU情報符号化部280は、次の処理ENCおよびDIVを繰り返す。
 処理ENC:処理対象の領域を分割せずに符号化
 処理DIV:処理対象の領域を分割し、分割により得られた各領域を次の処理対象の領域とする。
 なお、対象ブロックが、一番初めに処理対象とする領域となる。また、TU情報符号化部280は、処理PおよびQのうち、全体として符号化効率が良くなるほうを選択する。
 例えば、TU情報符号化部280は、処理対象について、処理ENCのみを実行する場合と、処理DIVを試行してから処理ENCを実行する場合との間で、符号量を比較し、符号量が少なくなるよう処理手順を決定することができる。
 図26を用いて、その具体的な例について説明する。図26は、2階層の分割を行う場合について例示している。
 図26に示すように、対象ブロックBLKは、1階層の分割を行った場合の第1階層符号化領域R10、R20、R30、R40を含む。
 また、第1階層符号化領域R30、およびR40については、2階層まで分割されている。すなわち、第1階層符号化領域R30は、第2階層符号化領域R31~R34を含み、また、第1階層符号化領域R40は、第2階層符号化領域R41~R44を含む。
  (分割の決定)
 TU情報符号化部280は、次の工程により、図26に示す対象ブロックBLKを分割する。
 工程[1] TU情報符号化部280は、対象ブロックBLKを処理対象の領域とし、処理ENCおよび処理DIVを実行する。すなわち、領域符号化部282が処理ENCを実行するとともに、領域符号化部282は、対象ブロックBLKに対する処理ENCによる符号量を記憶する。また、領域分割部281が、対象ブロックBLKに対して処理DIVを実行し、第1階層符号化領域R10~R40を得る。
 工程[2] 領域符号化部282は、第1階層符号化領域R10~R40のそれぞれについて処理ENCを実行するとともに、第1階層符号化領域R10~R40全体に対する処理ENCの符号量を記憶する。
 工程[3] TU情報符号化部280は、工程[1]の符号量と、工程[2]の符号量とを比較する。ここでは、工程[2]の符号量の方が小さかったとする。
 工程[4] 領域分割部281が、さらに第1階層符号化領域R10~R40に対して処理DIVを実行し、第2階層の符号化領域を得る。
 工程[5] TU情報符号化部280は、第1階層までの符号化による符号量と、第2階層までの符号化による符号量とを比較し、符号量が小さいほうの処理手順を採用する。
 工程[6] 工程[5]の比較結果により、第1階層符号化領域R10、R20については、第2階層まで符号化したほうが、符号量が多く、また、第1階層符号化領域R30、R40については、第2階層まで符号化したほうが、符号量が少なかったとする。
 これにより最終的に、第1階層符号化領域R10、R20、第2階層符号化領域R31~R34、R41~R44が確定する。
  (分割フラグおよび非ゼロ係数フラグ)
 領域符号化部282は、領域分割部281による分割の状況を示す分割フラグを符号化する。
 具体的に例示すると次のとおりである。領域符号化部282は、領域分割部281による分割が確定した領域については、分割フラグ“1”を符号化する。また、領域符号化部282は、領域分割部281に分割されないことが確定した領域については、分割フラグ“0”を符号化する。
 また、図26において、網掛けをした領域は、非ゼロ係数が1つでもある領域であることを示しており、網掛けをしなかった領域は、非ゼロ係数が1つもない領域であることを示している。
 そこで、領域符号化部282は、処理対象の領域における非ゼロ係数の有無を示す非ゼロ係数フラグを符号化してもよい。例えば、非ゼロ係数がある場合、非ゼロ係数フラグは、“1”であり、非ゼロ係数がない場合、非ゼロ係数フラグは、“0”である。
  (フラグツリー)
 図27を用いて、分割フラグおよび非ゼロ係数フラグをフラグツリーFTの形式で符号化する例について説明する。図27は、図26に示す対象ブロックBLKの分割状況および係数分布状況を表すフラグツリーFTの表現例(四分木表現)を示している。
 フラグツリーFTのROOTレベル、LEVEL1、および、LEVEL2の階層構造からなる。フラグツリーFTのROOTレベルおよびLEVEL1は、分割フラグに対応している。フラグツリーFTのLEVEL2は、非ゼロ係数フラグに対応している。
 フラグツリーFTのROOTレベルでは、対象ブロックBLKの分割フラグFRootが符号化される。また、フラグツリーFTのLEVEL1では、第1階層符号化領域R10、R20、R30、およびR40の分割フラグF10、F20、F30、F40が符号化される。
 そして、フラグツリーFTのリーフ(末端ノード)となるLEVEL2では、非ゼロ係数フラグが符号化される。
 図27において、丸で囲ったリーフは、左から順に、第2階層符号化領域R41、R42、R43、およびR44の非ゼロ係数フラグを示しており、それぞれ、“1”、“0”、“0”、および“1”が符号化されている。
 なお、第2階層符号化領域R41、R42、R43、およびR44について、計4個の非ゼロ係数フラグが符号化されているが、このフラグ4個のパターンを1組にして、当該パターンの出現頻度に応じた可変長符号化を行ってもよい。当該可変長符号化によれば、平均的には、四分木の符号量を削減することができる。
 非ゼロ係数が多い領域については、runの長さが平均的に短くなるため、分割しないほうが、符号化効率が向上する傾向がある。
 これに対して、ゼロ係数が多い、すなわち係数の分布密度が低い場合、runの長さが平均的に長くなるため、領域分割による符号化を採用することで、符号化効率を向上させることができる。
  (領域の復号処理)
 次に、図28を用いて、上述の手法により符号化した係数符号化データを動画像復号装置1において復号する処理について説明する。図28は、動画像復号装置1における領域の復号処理S200の流れの一例について示すフローチャートである。
 なお、領域の復号処理S200は、再帰的処理であるため、最上位の呼び出しでは、対象ブロックを処理対象の領域として領域の復号処理S200に含まれる各手順が実行される。
 図28に示すように、まず、動画像復号装置1では、領域分割部121が、分割フラグを参照することで、処理対象の領域が分割されているか否かを判定する(S201)。
 処理対象の領域が分割されていない場合(S201においてNO)、係数復号処理S210を実行する。
 すなわち、まず、領域復号部122が、非ゼロ係数フラグを参照することで、領域内に非ゼロ係数があるか否かを判定する(S211)。
 領域内に非ゼロ係数がある場合(S211においてYES)、領域復号部122は、処理対象の領域に含まれる係数全体を、まとめてラン-レベルモード復号処理により復号する(S212)。
 これに対し、領域内に非ゼロ係数がない場合(S211においてNO)、領域復号部122は、処理対象の領域の復号処理をスキップする。
 係数復号処理S210が終了し、処理対象の領域の復号処理が終了して、次の処理対象の領域に対する領域の復号処理S200が実行される。
 一方、処理対象の領域が分割されている場合(S201においてYES)、分割領域復号処理S220を実行する。
 すなわち、まず、領域復号部122が、各領域について符号化されている分割フラグを復号し(S221)、分割した各領域についてのループLP200に入る。
 ループLP200では、領域の復号処理S200が再帰的に実行される(S223)、そして、ループLP200の先頭にもどって(S224からS223へ)、以後、分割した各領域について順にループ内の処理が実行される。
 分割した領域すべてについて処理終了すると、ループLP200を抜けて、分割領域復号処理S220が終了する。その後、実行中の領域の復号処理S200は終了する。再帰的に領域の復号処理S200が呼び出されている場合は、呼び出し元に制御を返す。
 なお、以上において、不要なフラグの符号化を避けるため、分割フラグは、これ以上分割しないサイズのブロックでは、常に偽と判定してもよい。また、非ゼロ係数フラグは、処理対象の領域が、対象ブロックそのものである場合、常に真と判定してもよい。
  (データ構造)
 以下において、図29~図32を用いて、領域の復号処理S200において復号される係数符号化データのデータ構造について例示する。
  (1) フラグを分散して格納する例
 図29を用いて、フラグを分散して格納する場合のデータ構造について説明する。図29に示すように、係数符号化データの先頭には、分割フラグFRootが格納される。分割フラグFRootが「0」であれば、係数符号化データは、一例として、DATA11に示すデータ構造が採用される(以下、係数符号化データDATA11と表記する)。符号化データDATA11には、対象ブロックの16×16個分の係数データ(run,level,sign)が含まれる。
 一方、分割フラグFrootが「1」であれば、係数符号化データは、一例として、DATA12に示すデータ構造が採用される(以下、係数符号化データDATA12と表記する)。
 係数符号化データDATA12は、領域情報[領域1]F1~領域情報[領域n]Fnを含む。なお、上記“領域n”とは、第1階層の“領域”のことを示している。図26に示す対象ブロックBLKを用いて例示すると、領域R10~R40が上記“領域”に該当する。
 ここで、図29に従って、領域情報[領域1]F1の詳細なデータ構造について説明する。領域情報[領域1]F1は、分割フラグ[領域1]F10を含む。
 ここで、非ゼロ係数フラグおよび係数データは、分割フラグ[領域1]F10が、「0」の場合は、係数情報F12を含み、また、分割フラグ[領域1]F10が、「1」の場合は、係数情報F11を含む。
 分割フラグ[領域1]F10が、「0」の場合、さらなる領域分割はないので、係数情報には、非ゼロ係数フラグ[1]と、係数データ[1]とが含まれるのみである。
 分割フラグ[領域1]F10が、「1」の場合において、領域1がn個に分割されるとする。この場合、係数情報F11には、非ゼロ係数フラグ[1-1]、係数データ[1-1]~非ゼロ係数フラグ[1-n]、係数データ[1-n]が含まれる。
 その他の領域情報については、領域情報[領域1]F1と同様であるので、その説明を省略する。
 次に、図30を用いて、図29に示す係数符号化データDATA12の具体例について説明する。図30は、図26を用いて説明した対象ブロックBLKを示す係数符号化データの例を示す図である。
 図26に示す対象ブロックBLKは、ROOTレベルで4分木分割が指定されるので、分割フラグFrootは、“1”である。
 また、係数符号化データDATA12は、対象ブロックBLKに含まれる復号領域R10~R40に対応する領域情報F1~F4を含む。係数符号化データDATA12では、
領域情報F1、F2、F3およびF4の順でデータが格納されている。以下、領域情報F1~F4に含まれるデータについて順に説明する。
 領域情報F1に含まれるデータについて説明すると次のとおりである。復号領域R10は、それ以上分割されないので、領域情報F1では、分割フラグ[1]=0である。また、復号領域R10は、非ゼロ係数を含むので、領域情報F1では、非ゼロ係数フラグ[1]=1である。また、領域情報F1は、係数データ[1]を含む。
 復号領域R20は、分割されず、また非ゼロ係数を1つも含まないので、領域情報F2には、分割フラグ[2]=0、非ゼロ係数フラグ[2]=0が含まれる。
 領域情報F3に含まれるデータについて説明すると次のとおりである。復号領域R30は、復号領域R31~R34に分割される。よって、領域情報F3は、分割フラグ[3]=1を含む。
 また、復号領域R31、R33には、非ゼロ係数が含まれ、復号領域R32、R34には、非ゼロ係数が1つも含まれない。
 よって、復号領域R31について、非ゼロ係数フラグ[3-1]=1と、係数データ[3-1]とが含まれる。復号領域R33についても同様である。
 また、復号領域R32およびR34については、それぞれ、非ゼロ係数フラグ[3-2]=0、および、非ゼロ係数フラグ[3-4]=0が含まれる。
 領域情報F4に含まれるデータについて説明すると次のとおりである。復号領域R40は、復号領域R41~R44に分割される。よって、領域情報F4は、分割フラグ[4]=1を含む。
 また、復号領域R41、R44には、非ゼロ係数が含まれ、復号領域R42、R43には、非ゼロ係数が1つも含まれない。
 よって、復号領域R41について、非ゼロ係数フラグ[4-1]=1と、係数データ[4-1]とが含まれる。復号領域R44についても同様である。
 また、復号領域R42およびR43については、それぞれ、非ゼロ係数フラグ[4-2]=0、および、非ゼロ係数フラグ[4-3]=0が含まれる。
  (2) フラグツリーをデータ先頭にまとめて格納する例
 図31を用いて、フラグツリーをデータ先頭にまとめて格納する場合のデータ構造について説明する。図31に示すように、係数符号化データの先頭には、フラグツリーFTが格納される。
 ここで、フラグツリーFTが、対象ブロックの分割がないことを示している場合、DATA21に示すデータ構造(係数符号化データDATA21)が採用される。符号化データDATA21には、対象ブロックの16×16個分の係数データ(run,level,sign)が含まれる。
 一方、フラグツリーFTが、対象ブロックを1回以上分割することを示している場合、DATA22に示すデータ構造(係数符号化データDATA22)が採用される。
 係数符号化データDATA22は、係数データ[領域1]~係数データ[領域n]を含む。なお、上記“領域n”とは、それ以上分割されない“領域”のことを示している。図26に示す対象ブロックBLKを用いて例示すると、領域R10や、領域R31等が上記“領域”に該当する。
 次に、図32を用いて、図31に示す係数符号化データDATA22の具体例について説明する。図32は、図26を用いて説明した対象ブロックBLKを示す係数符号化データの例を示す図である。
 係数符号化データDATA22には、まず、フラグツリーFTが格納される。フラグツリーFTでは、まず分割フラグFRootが格納される。続いて、復号領域R10~R40に関するフラグが順に格納される。
 係数符号化データDATA22では、フラグツリーFTの後に、各復号領域の係数データが格納される。係数データが1つ以上ある復号領域R10、R31、R33、R41、およびR44に対応する係数データ[1]、係数データ[3-1]、係数データ[3-3]、係数データ[4-1]、および係数データ[4-4]が順に格納される。
  (変形例)
   [runのカウントの省略]
 以下において、図33および図34を用いて、非ゼロ係数フラグが“0(偽)”の領域はrunのカウントに含めない例について説明する。図33および図34は、図26の第1階層符号化領域R30について示している。
 なお、以下の説明では、例示的に、対象ブロックBLKは、16×16サイズであるとする。第1階層符号化領域R30は、8×8サイズであるとし、第2階層符号化領域R31~R34は、4×4サイズであるとする。
 このとき、TU情報符号化部280は、処理対象の符号化領域が、所定サイズ(例えば、8×8サイズ)以下である場合、領域分割および非ゼロ係数の有無の判定は行うが、スキャンおよび符号化は8×8単位で行う。
 図33は、図26の第1階層符号化領域R30について詳細に示す図である。図33に示すとおり、領域分割部281は、第1階層符号化領域R30を、第2階層符号化領域R31~R34に分割する。
 領域符号化部282は、第1階層符号化領域R30について、分割フラグ“1”を符号化し、さらに非ゼロ係数フラグ“1001”を符号化する。
 また、領域符号化部282は、第1階層符号化領域R30を、対象符号化領域として符号化処理を行う。第1階層符号化領域R30の全体にわたって示す矢印は、スキャン順序を示している。
 ここで、ランモード符号化部202は、非ゼロ係数フラグ“0(偽)”である符号化対象領域R32およびR34では、runをカウントしない。
 すなわち、係数の符号化処理において、ランモード符号化部202は、次のようにランモード符号化処理を行う。
 まず、図34に示す非ゼロ係数A1の符号化が完了しているものとする。そして、次の非ゼロ係数の符号化処理において、ランモード符号化部202は、図34において矢印で示す逆順のジグザグスキャン順序にて、係数を読み取っていく。ここで、逆順のジグザグスキャン順序、次の非ゼロ係数は、A3である(以下、非ゼロ係数A3にて参照する)。非ゼロ係数A1と、非ゼロ係数A3との間には、9つのゼロ係数が存在する。
 ランモード符号化部202は、第2階層符号化領域R32およびR34におけるゼロ係数を、runのカウントに含めない。このため、次の非ゼロ係数A3の符号化処理において、runとしては、第2階層符号化領域R33におけるゼロ係数A2のみがカウントされる。よって、ランモード符号化部202は、run=1を符号化する。
 なお、ランモード符号化部202は、各係数の2次元座標値から非ゼロ係数フラグが“0”の領域か否かを判定してもよい。
 なお、動画像復号装置1側では次のようにして復号処理を行えばよい。すなわち、各係数の復号後にスキャン順に配列に格納された係数を、再び係数行列に格納する処理を行う際に、非ゼロ係数フラグが“0”である領域では、格納処理をスキップすればよい。
   [規定による領域分割]
 以下において、図35および図36を用いて、規定による領域分割について説明する。
 TU情報復号部12は、所定の復号対象領域については、予め定められたとおりに分割を行って、係数を復号してもよい。
 図35および図36は、16×16サイズの対象ブロックBLKにおける分割方式を示している。また、図35および図36では、対象ブロックBLKは、第1階層復号領域R101~R104に分割されている。また、第2階層復号領域まで分割可能な第1階層復号領域については、点線を付している(例えば、図35において、第1階層復号領域R102~R104)。
 すなわち、TU情報復号部12の領域分割部121は、所定の復号対象領域を、常に分割するよう設定されていてもよいし(分割可能)、常に分割しないように設定されていてもよい(分割不可)。
 例えば、図35に示すように、対象ブロックBLKにおいて、左上8×8の第1階層復号領域R101は、分割不可としてもよい。第1階層復号領域R102~R104については、領域分割部121が分割可能である。
 なお、DC成分に近い領域では、非ゼロ係数が存在する可能性が高い。よって、第1階層復号領域R101では、分割しない場合のほうが、分割する場合よりも、動画像符号化装置2側での符号化効率が向上する傾向がある。
 また、例えば、図36に示すように、対象ブロックBLKにおいて、右下8×8の第1階層復号領域R104のみ、分割可能としてもよい。第1階層復号領域R104については、領域分割部121が分割可能である。
 高周波成分側の領域は、ゼロ係数が存在する可能性が高いため分割するほうが、動画像符号化装置2側での符号化効率が向上する傾向がある。
 また、領域分割部121は、所定サイズより大きな復号領域を常に分割するようにしてもよい。また、領域分割部281は、所定サイズより小さなサイズの復号領域は、常に分割しないようにしてもよい。例えば、以下のように構成することができる。
 領域分割部121は、8×8サイズより大きなサイズの復号領域を常に分割してもよい。また、領域分割部121は、最上位(つまり、対象ブロック)分割フラグを常に“1”と設定してもよい。
 また、領域分割部121は、4×4サイズ以上のサイズの対象ブロックに含まれる4×4サイズの領域は、それ以上分割しないようにしてもよい。
 また、領域分割部121は、3階層以上の分割は行わないようにしてもよい。
 なお、本変形例は、動画像符号化装置2のTU情報符号化部280について適用することも可能である。
〔その他の変形例〕
 実施形態3に示した符号化方式の選択や、実施形態4に示した規定による領域分割を、特定サイズ・形状の対象ブロックや、特定のスライスタイプにのみ適用してもよい。
 例えば、Bスライスでは、ゼロ係数が比較的多い傾向がある。よって、Bスライスでは、分割を行わず、対象ブロックの左上64個の係数のみを符号化してもよい。
 これらの変形によれば、フラグ符号量の増加を抑制することが可能である。
 また、実施形態4に示した分割手法により得られた領域において、実施形態2に示す相対位置指定による係数の符号化を行ってもよい。
 また、実施形態1で示した[領域の数およびサイズ]の変形例は、実施形態2~4においても適用可能である。
 また、本発明は以下のように表現することもできる。
 すなわち、本発明に係る画像復号装置では、変換単位分割手段が、変換単位を、複数のサブ単位に分割し、変換係数復号手段が、符号化データから上記変換係数を得るための復号情報であって、上記サブ単位ごとに割り当てられている復号情報を参照して、上記サブ単位に含まれる変換係数を復号する。
 また、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データでは、所定のスキャン順序に沿って連続するゼロ係数の長さが符号化されており、低周波側に位置する領域について定義されている上記復号情報では、上記連続するゼロ係数の長さが短いほど、短い符号が割り当てられている。
 低周波成分側の領域では、非ゼロ係数が出現する頻度が高くため、連続するゼロ係数の長さは短くなる傾向がある。
 上記構成によれば、連続するゼロ係数の長さは短いほど、短い符号を割り当てるので、領域の位置に応じて、上記傾向を考慮した効率のよい復号処理を行うことができる。これにより、復号する符号量を低減することができる。
 また、高周波側に位置する領域について定義されている上記復号情報では、上記変換係数の絶対値を含むパラメータの組について、上記絶対値が1であるパラメータの組に、より短い符号が割り当てられている。
 高周波成分側の領域では、変換係数の絶対値が全体的に小さくなる傾向がある。このため、変換係数がゼロ係数でなければ、絶対値が1となる傾向がある。
 上記構成によれば、高周波成分側の位置において出現頻度が高くなるようなパラメータの組に、より短い符号を割り当てることができる。
 なお、パラメータの組とは、例えば、ランモードにおける{run,level}の組のことであり、ここでの絶対値とは、levelに対応する。
 よって、領域の位置に応じて、上記傾向を考慮した効率のよい復号処理を行うことができる。これにより、復号する符号量を低減することができる。
 また、上記復号情報では、符号の長さに応じて順序が指定されており、
 上記復号情報更新手段は、上記更新において、上記変換単位における上記サブ単位の位置に応じて、上記順序を繰り上げる。
 上記構成によれば、上記変換単位における上記サブ単位の位置に応じて、パラメータに割り当てられている符号の長さに応じて定められている順序を繰り上げる。
 上記符号の長さに応じて定められている順序とは、例えば、コード番号のことをいう。すなわち、上記構成では、コード番号の繰り上げを行うことにより、パラメータに割り当てられている符号の長さをより短いものに更新する。
 上記構成によれば、コード番号の繰り上げという比較的簡便な手順により符号の長さの更新を実現することができる。
 また、上記変換単位分割手段は、分割対象の領域の位置およびサイズの少なくとも一方に応じて、上記再帰的な分割を行う。
 上記構成によれば、分割を行うことを示すフラグを復号しなくても済むため、復号する符号量を低減することができる。
〔付記的情報〕
 本発明の一側面について説明すると次のとおりである。すなわち、本発明に係る画像復号装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データから、該変換係数を復号する画像復号装置において、上記変換単位を、複数のサブ単位に分割する変換単位分割手段と、上記符号化データから上記変換係数を得るための復号情報であって、上記サブ単位ごとに割り当てられている復号情報を参照して、上記サブ単位に含まれる変換係数を復号する変換係数復号手段と、を備える構成である。
 また、本発明に係る画像符号化装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する画像符号化装置において、上記変換単位を、複数のサブ単位に分割する変換単位分割手段と、上記変換係数を符号化するための符号化情報であって、上記サブ単位ごとに割り当てられている符号化情報を参照して、上記変換単位に含まれる変換係数を符号化する変換係数符号化手段と、を備える構成である。
 また、本発明に係る符号化データのデータ構造は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化することにより生成される符号化データのデータ構造において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、これにより、上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定するデータ構造である。
 上記構成によれば、復号処理において、まず、復号の対象となる変換単位を複数のサブ単位に分割する。
 変換単位とは、画素値を周波数領域に変換する単位である。変換単位としては、例えば、64×64画素、32×32画素や16×16画素のサイズ等が挙げられる。
 サブ単位は、変換単位が16×16サイズである場合、例えば、8×8サイズの領域であってもよい。
 また、上記構成によれば、分割により得られた複数のサブ単位を、ひとつずつ処理対象にして、該サブ単位に含まれる変換係数を復号する。サブ単位を復号する順番に特に制限はなく任意の順に復号処理を行うことができる。
 また、上記構成では、変換係数の復号に際し、複数のサブ単位のそれぞれに割り当てられている復号情報を参照する。
 復号情報とは、符号化データのコード(ビット列)から、変換係数の所定のパラメータ値を再現するための情報である。例えば、復号情報は、符号化データのコードから変換係数の所定のパラメータ値を再現するための対応付けを示すテーブルである。また、例えば、復号情報は、符号化データのコードから、変換係数の所定のパラメータ値を導出するための算出式である。
 つまり、上記構成においては、元の変換単位のサイズよりも、小さなサブ単位について規定されている復号情報を用いて、変換係数を復号する。
 このため、元の変換単位のサイズについて規定される復号情報に基づいて復号処理を行うのに比べて、復号情報の情報量や、復号情報に基づく計算量を低減することができるという効果を奏する。
 さらにいえば、復号処理において対象となる変換係数の数を小さくできるので、変換係数のスキャン順序を定義するスキャンテーブルのサイズもより小さくできる。
 また、さらに付言すれば、復号処理において必要となるメモリの量や、処理能力を低く抑えることができる。
 なお、サブ単位は、非特許文献1,2の技術における符号化単位のいずれかと一致していてもよい。この場合、上記符号化単位において予め定義されているVLCテーブル、すなわち復号情報を流用することができる。
 なお、上記のように構成された画像符号化装置または符号化データのデータ構造によれば、本発明に係る画像復号装置と同様の効果を奏する。
 また、本発明に係る画像復号装置では、上記変換係数復号手段は、上記サブ単位における非ゼロの変換係数の有無を示す非ゼロ情報を参照し、該非ゼロ情報が、上記サブ単位における非ゼロの変換係数が無いことを示すとき、上記サブ単位の復号処理を省略することが好ましい。
 上記構成によれば、非ゼロ係数の有無に応じて、領域単位で不要な復号処理を行うことを避けることができる。
 また、本発明に係る画像復号装置では、上記復号情報は、上記変換単位における、上記サブ単位の位置に応じて適応的に定義されていることが好ましい。
 変換単位では、DC成分を含む低周波成分側と、高周波成分側とで、変換係数の値の出現傾向が異なる。例えば、DC成分付近の低周波成分側では、非ゼロ係数が出現する可能性が高い。また、高周波成分側では、ゼロ係数が出現する可能性が高い。
 位置に応じて適応的に定義するとは、例えば、上記サブ単位の位置が、低周波成分側であるのか、低周波成分側であるのかに応じて適応的に復号情報を定義するということである。
 また適応的とは、上記の出現傾向に応じて、コードを割り当てることである。例えば、高周波成分側では、ゼロ係数に、または変換係数の絶対値が小さいものに、より短いコードを割り当てるということである。
 また、例えば、低周波成分側では、非ゼロ係数に、または変換係数の絶対値が大きいものに、より短いコードを割り当てるということである。
 上記構成によれば、領域の位置に応じた効率のよい復号処理を行うことができ、復号する符号量を低減することができる。
 また、本発明に係る画像復号装置では、上記変換係数を示すパラメータの出現頻度に応じて、上記復号情報において該パラメータに割り当てられている符号をより短いものに更新する復号情報更新手段を備えることが好ましい。
 上記構成によれば、出現頻度が高いパラメータについて、より短い符号が割り当てることができる。すなわち、パラメータの出現頻度を、符号の短さに動的に反映させることができる。
 これにより、出現頻度が高いパラメータについて復号する符号量を低減することができる。
 また、本発明に係る画像復号装置では、上記変換係数復号手段は、連続する非ゼロ係数の長さと、変換係数の絶対値と、変換係数の符号とを復号する第1モード復号手順を所定条件下において実行した後、変換係数の絶対値と、変換係数の符号とを復号する第2モード復号手順を実行する復号処理を行うことが好ましい。
 上記構成によれば、復号処理において、連続する非ゼロ係数の長さ(run)と、変換係数の絶対値(level)と、変換係数の符号(sign)とを復号する第1モード復号手順を所定条件下において実行した後、変換係数の絶対値(level)と、変換係数の符号(sign)とを復号する第2モード復号手順を実行する。第1モード復号手順とは、いわゆるランモードであり、第2モード復号手順とは、いわゆるレベルモードである。
 上記所定条件とは、例えば、復号した変換係数の数、変換係数の絶対値などが挙げられる。他にも、所定条件は、サブ単位の変換単位における位置や、係数の出現傾向に応じた条件であってもよい。
 ランモードおよびレベルモードは、例えば、非特許文献1、2に採用されている技術である。各領域における復号処理では、このような従来構成を流用することができる。これにより、高い符号化効率を実現することができる。
 また、本発明に係る画像復号装置では、上記変換係数復号手段は、上記変換単位における、上記サブ単位の位置に応じて、上記所定条件を、変更することが好ましい。
 上記変換単位における、上記サブ単位の位置に応じて、上記所定条件を、変更するとは、例えば、低周波成分側の領域では、第1モード復号手順を終了しにくくし、高周波成分側の領域では、第1モード復号手順を終了しやすくするということである。
 低周波成分側の領域では、連続する非ゼロ係数の長さが比較的短く、高周波成分側の領域では、連続する非ゼロ係数の長さが比較的長くなる傾向がある。このため、連続する非ゼロ係数の長さが長くなる場合には第1モード復号手順を優先的に用いるようにする。
 なお、所定条件の変更には、第1モード復号手順のみを実行し、第2モード復号手順を実行しないことも含まれる。
 上記構成によれば、ラン-レベルモードによる効率的な復号処理を実現することができる。
 また、本発明に係る画像復号装置では、上記変換単位における低周波成分側の所定領域に限り変換係数を復号する限定領域復号手段と、上記変換単位分割手段および上記変換係数復号手段による復号処理と、上記限定領域復号手段による復号処理とを切り替える切り替え手段とを備えることが好ましい。
 上記構成によれば、上記変換単位分割手段および上記変換係数復号手段による復号処理方式と、上記変換単位における低周波成分側の所定領域に限り変換係数を復号する復号処理方式とのいずれか符号化効率がよい復号処理方式に適宜切り替えることができる。
 また、本発明に係る画像復号装置では、上記変換単位分割手段は、分割した上記複数の領域を再帰的に分割することが好ましい。
 より小さいサイズの領域について復号処理を行ったほうが復号する符号量が少なくて済む場合がある。上記構成によれば、より小さいサイズの領域について復号処理を行ったほうが復号する符号量が少なくて済む場合、効率的に復号処理を行うことができる。
 さらに付言しておくと、非ゼロ係数が多いときは、分割しないほうが効率的な場合がある。また、非ゼロ係数の有無が符号化されていると、より小さいサイズの領域について、細かく制御を行うことができるので、さらに効果的である。
 本発明の他の側面について説明すると次のとおりである。すなわち、本発明に係る画像復号装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データから、該変換係数を復号する画像復号装置において、復号対象となる変換係数のひとつ前に復号した変換係数からの相対位置を復号する相対位置復号手段と、ひとつ前に復号した上記変換係数の、上記変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する位置特定手段と、を備える構成である。
 また、本発明に係る画像符号化装置は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する画像符号化装置において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を符号化する相対位置符号化手段を備える構成である。
 また、本発明に係る符号化データのデータ構造は、上記課題を解決するために、対象画像の画素値を変換単位ごとに周波数変換に変換して得られた変換係数を符号化することにより生成された符号化データのデータ構造において、符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、これにより、上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定するデータ構造である。
 上記構成によれば、ひとつ前に復号した上記変換係数の、上記変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する。これにより、相対位置に基づいて、連鎖的に、変換係数の位置を特定していくことができる。なお、変換単位とは、変換を行う所定の単位である。
 上述のrunを符号化する場合、所定のスキャン順序に応じてrunの長さがカウントされるため、基準とする非ゼロ係数と次の非ゼロ係数との変換単位における2次元座標の相対位置が近くても、結果としてrunが長くなるときがあり、これにより、符号量が増大する場合がある。
 この傾向は、変換係数が疎らになりやすい高周波成分の領域において顕著である。また、runが長くなるということは、これに応じた大きなテーブルを用意しなければならないということである。
 これに対して、相対位置で変換係数の位置を特定していけば、このような場合において符号量を削減することができる。
 上記構成によれば、相対位置で変換係数の位置を特定していくので、復号すべき符号量を低減することができる。
 この結果、復号情報の情報量や、復号情報に基づく計算量を低減することができるという効果を奏する。
 また、さらに付言すれば、復号処理において必要となるメモリの量や、処理能力を低く抑えることができる。
 なお、上記のように構成された画像符号化装置または符号化データのデータ構造によれば、本発明に係る画像復号装置と同様の効果を奏する。
 また、本発明に係る画像復号装置では、上記変換単位における低周波成分側の領域について、連続する非ゼロ係数の長さと変換係数の絶対値と変換係数の符号とを復号する第1モード復号処理、および、変換係数の絶対値と変換係数の符号とを復号する第2モード復号処理を実行する復号手段を備えることが好ましい。
 上記復号手段は、いわゆるラン-レベルモードの復号を実行するものである。変換単位における低周波成分側の領域とは、例えば、当該変換単位が16×16サイズのものであれば、DC成分を含む左上の8×8サイズの領域が挙げられる。
 低周波成分側の領域では、変換係数がそれほど疎らにならないため、runの長さも比較的短くなる。よって、ラン-レベルモードによる復号を効率的に行うことができる。
 上記構成によれば、変換係数がそれほど疎らにならないような領域について、ラン-レベルモードによる復号を効率的に行うことができる。
 また、本発明に係る画像復号装置では、上記復号手段は、復号対象となる変換単位の特性に応じて上記領域のサイズを変更することが好ましい。
 変換単位の特性とは、変換単位のスライスタイプ、予測モード、変換単位サイズなどである。上記各種特性を組み合わせて用いてもよいし、択一的に用いてもよい。例えば、変換単位のスライスタイプ、予測モード、変換単位サイズなどの少なくとも1つに応じて上記領域のサイズを変更すればよい。
 変換単位のサイズが16×16であるとき、例えば、次のように構成することができる。すなわち、予測モードが、イントラモードであれば、上記低周波成分側の領域のサイズを8×8とする。また、予測モードが、インターモードであれば、上記低周波成分側の領域のサイズを4×4とする。
 これにより、変換単位の特性に応じて、ラン-レベルモードにより復号する領域を変更することができる。
 この結果、ラン-レベルモードによる復号を効率的に行うことができる。
 また、上述した動画像復号装置1および動画像符号化装置2の各ブロックは、集積回路(ICチップ)上に形成された論理回路によってハードウェア的に実現してもよいし、CPU(Central Processing Unit)を用いてソフトウェア的に実現してもよい。
 後者の場合、上記各装置は、各機能を実現するプログラムの命令を実行するCPU、上記プログラムを格納したROM(Read Only Memory)、上記プログラムを展開するRAM(Random Access Memory)、上記プログラムおよび各種データを格納するメモリ等の記憶装置(記録媒体)などを備えている。そして、本発明の目的は、上述した機能を実現するソフトウェアである上記各装置の制御プログラムのプログラムコード(実行形式プログラム、中間コードプログラム、ソースプログラム)をコンピュータで読み取り可能に記録した記録媒体を、上記各装置に供給し、そのコンピュータ(またはCPUやMPU)が記録媒体に記録されているプログラムコードを読み出し実行することによっても、達成可能である。
 上記記録媒体としては、例えば、磁気テープやカセットテープ等のテープ類、フロッピー(登録商標)ディスク/ハードディスク等の磁気ディスクやCD-ROM/MO/MD/DVD/CD-R/ブルーレイディスク(登録商標)等の光ディスクを含むディスク類、ICカード(メモリカードを含む)/光カード等のカード類、マスクROM/EPROM/EEPROM/フラッシュROM等の半導体メモリ類、あるいはPLD(Programmable logic device)やFPGA(Field Programmable Gate Array)等の論理回路類などを用いることができる。
 また、上記各装置を通信ネットワークと接続可能に構成し、上記プログラムコードを通信ネットワークを介して供給してもよい。この通信ネットワークは、プログラムコードを伝送可能であればよく、特に限定されない。例えば、インターネット、イントラネット、エキストラネット、LAN、ISDN、VAN、CATV通信網、仮想専用網(Virtual Private Network)、電話回線網、移動体通信網、衛星通信網等が利用可能である。また、この通信ネットワークを構成する伝送媒体も、プログラムコードを伝送可能な媒体であればよく、特定の構成または種類のものに限定されない。例えば、IEEE1394、USB、電力線搬送、ケーブルTV回線、電話線、ADSL(Asymmetric Digital Subscriber Line)回線等の有線でも、IrDAやリモコンのような赤外線、Bluetooth(登録商標)、IEEE802.11無線、HDR(High Data Rate)、NFC(Near Field Communication)、DLNA(Digital Living Network Alliance)、携帯電話網、衛星回線、地上波デジタル網等の無線でも利用可能である。なお、本発明は、上記プログラムコードが電子的な伝送で具現化された、搬送波に埋め込まれたコンピュータデータ信号の形態でも実現され得る。
<<応用例>>
 上述した動画像符号化装置2及び動画像復号装置1は、動画像の送信、受信、記録、再生を行う各種装置に搭載して利用することができる。なお、動画像は、カメラ等により撮像された自然動画像であってもよいし、コンピュータ等により生成された人工動画像(CGおよびGUIを含む)であってもよい。
 まず、上述した動画像符号化装置2及び動画像復号装置1を、動画像の送信及び受信に利用できることを、図37を参照して説明する。
 図37の(a)は、動画像符号化装置2を搭載した送信装置PROD_Aの構成を示したブロック図である。図37の(a)に示すように、送信装置PROD_Aは、動画像を符号化することによって符号化データを得る符号化部PROD_A1と、符号化部PROD_A1が得た符号化データで搬送波を変調することによって変調信号を得る変調部PROD_A2と、変調部PROD_A2が得た変調信号を送信する送信部PROD_A3と、を備えている。上述した動画像符号化装置2は、この符号化部PROD_A1として利用される。
 送信装置PROD_Aは、符号化部PROD_A1に入力する動画像の供給源として、動画像を撮像するカメラPROD_A4、動画像を記録した記録媒体PROD_A5、動画像を外部から入力するための入力端子PROD_A6、及び、画像を生成または加工する画像処理部A7を更に備えていてもよい。図37の(a)においては、これら全てを送信装置PROD_Aが備えた構成を例示しているが、一部を省略しても構わない。
 なお、記録媒体PROD_A5は、符号化されていない動画像を記録したものであってもよいし、伝送用の符号化方式とは異なる記録用の符号化方式で符号化された動画像を記録したものであってもよい。後者の場合、記録媒体PROD_A5と符号化部PROD_A1との間に、記録媒体PROD_A5から読み出した符号化データを記録用の符号化方式に従って復号する復号部(不図示)を介在させるとよい。
 図37の(b)は、動画像復号装置1を搭載した受信装置PROD_Bの構成を示したブロック図である。図37の(b)に示すように、受信装置PROD_Bは、変調信号を受信する受信部PROD_B1と、受信部PROD_B1が受信した変調信号を復調することによって符号化データを得る復調部PROD_B2と、復調部PROD_B2が得た符号化データを復号することによって動画像を得る復号部PROD_B3と、を備えている。上述した動画像復号装置1は、この復号部PROD_B3として利用される。
 受信装置PROD_Bは、復号部PROD_B3が出力する動画像の供給先として、動画像を表示するディスプレイPROD_B4、動画像を記録するための記録媒体PROD_B5、及び、動画像を外部に出力するための出力端子PROD_B6を更に備えていてもよい。図37の(b)においては、これら全てを受信装置PROD_Bが備えた構成を例示しているが、一部を省略しても構わない。
 なお、記録媒体PROD_B5は、符号化されていない動画像を記録するためのものであってもよいし、伝送用の符号化方式とは異なる記録用の符号化方式で符号化されたものであってもよい。後者の場合、復号部PROD_B3と記録媒体PROD_B5との間に、復号部PROD_B3から取得した動画像を記録用の符号化方式に従って符号化する符号化部(不図示)を介在させるとよい。
 なお、変調信号を伝送する伝送媒体は、無線であってもよいし、有線であってもよい。また、変調信号を伝送する伝送態様は、放送(ここでは、送信先が予め特定されていない送信態様を指す)であってもよいし、通信(ここでは、送信先が予め特定されている送信態様を指す)であってもよい。すなわち、変調信号の伝送は、無線放送、有線放送、無線通信、及び有線通信の何れによって実現してもよい。
 例えば、地上デジタル放送の放送局(放送設備など)/受信局(テレビジョン受像機など)は、変調信号を無線放送で送受信する送信装置PROD_A/受信装置PROD_Bの一例である。また、ケーブルテレビ放送の放送局(放送設備など)/受信局(テレビジョン受像機など)は、変調信号を有線放送で送受信する送信装置PROD_A/受信装置PROD_Bの一例である。
 また、インターネットを用いたVOD(Video On Demand)サービスや動画共有サービスなどのサーバ(ワークステーションなど)/クライアント(テレビジョン受像機、パーソナルコンピュータ、スマートフォンなど)は、変調信号を通信で送受信する送信装置PROD_A/受信装置PROD_Bの一例である(通常、LANにおいては伝送媒体として無線又は有線の何れかが用いられ、WANにおいては伝送媒体として有線が用いられる)。ここで、パーソナルコンピュータには、デスクトップ型PC、ラップトップ型PC、及びタブレット型PCが含まれる。また、スマートフォンには、多機能携帯電話端末も含まれる。
 なお、動画共有サービスのクライアントは、サーバからダウンロードした符号化データを復号してディスプレイに表示する機能に加え、カメラで撮像した動画像を符号化してサーバにアップロードする機能を有している。すなわち、動画共有サービスのクライアントは、送信装置PROD_A及び受信装置PROD_Bの双方として機能する。
 次に、上述した動画像符号化装置2及び動画像復号装置1を、動画像の記録及び再生に利用できることを、図38を参照して説明する。
 図38の(a)は、上述した動画像符号化装置2を搭載した記録装置PROD_Cの構成を示したブロック図である。図38の(a)に示すように、記録装置PROD_Cは、動画像を符号化することによって符号化データを得る符号化部PROD_C1と、符号化部PROD_C1が得た符号化データを記録媒体PROD_Mに書き込む書込部PROD_C2と、を備えている。上述した動画像符号化装置2は、この符号化部PROD_C1として利用される。
 なお、記録媒体PROD_Mは、(1)HDD(Hard Disk Drive)やSSD(Solid State Drive)などのように、記録装置PROD_Cに内蔵されるタイプのものであってもよいし、(2)SDメモリカードやUSB(Universal Serial Bus)フラッシュメモリなどのように、記録装置PROD_Cに接続されるタイプのものであってもよいし、(3)DVD(Digital Versatile Disc)やBD(Blu-ray Disc:登録商標)などのように、記録装置PROD_Cに内蔵されたドライブ装置(不図示)に装填されるものであってもよい。
 また、記録装置PROD_Cは、符号化部PROD_C1に入力する動画像の供給源として、動画像を撮像するカメラPROD_C3、動画像を外部から入力するための入力端子PROD_C4、動画像を受信するための受信部PROD_C5、及び、画像を生成または加工する画像処理部C6を更に備えていてもよい。図38の(a)においては、これら全てを記録装置PROD_Cが備えた構成を例示しているが、一部を省略しても構わない。
 なお、受信部PROD_C5は、符号化されていない動画像を受信するものであってもよいし、記録用の符号化方式とは異なる伝送用の符号化方式で符号化された符号化データを受信するものであってもよい。後者の場合、受信部PROD_C5と符号化部PROD_C1との間に、伝送用の符号化方式で符号化された符号化データを復号する伝送用復号部(不図示)を介在させるとよい。
 このような記録装置PROD_Cとしては、例えば、DVDレコーダ、BDレコーダ、HDD(Hard Disk Drive)レコーダなどが挙げられる(この場合、入力端子PROD_C4又は受信部PROD_C5が動画像の主な供給源となる)。また、カムコーダ(この場合、カメラPROD_C3が動画像の主な供給源となる)、パーソナルコンピュータ(この場合、受信部PROD_C5又は画像処理部C6が動画像の主な供給源となる)、スマートフォン(この場合、カメラPROD_C3又は受信部PROD_C5が動画像の主な供給源となる)なども、このような記録装置PROD_Cの一例である。
 図38の(b)は、上述した動画像復号装置1を搭載した再生装置PROD_Dの構成を示したブロックである。図38の(b)に示すように、再生装置PROD_Dは、記録媒体PROD_Mに書き込まれた符号化データを読み出す読出部PROD_D1と、読出部PROD_D1が読み出した符号化データを復号することによって動画像を得る復号部PROD_D2と、を備えている。上述した動画像復号装置1は、この復号部PROD_D2として利用される。
 なお、記録媒体PROD_Mは、(1)HDDやSSDなどのように、再生装置PROD_Dに内蔵されるタイプのものであってもよいし、(2)SDメモリカードやUSBフラッシュメモリなどのように、再生装置PROD_Dに接続されるタイプのものであってもよいし、(3)DVDやBDなどのように、再生装置PROD_Dに内蔵されたドライブ装置(不図示)に装填されるものであってもよい。
 また、再生装置PROD_Dは、復号部PROD_D2が出力する動画像の供給先として、動画像を表示するディスプレイPROD_D3、動画像を外部に出力するための出力端子PROD_D4、及び、動画像を送信する送信部PROD_D5を更に備えていてもよい。図38の(b)においては、これら全てを再生装置PROD_Dが備えた構成を例示しているが、一部を省略しても構わない。
  なお、送信部PROD_D5は、符号化されていない動画像を送信するものであってもよいし、記録用の符号化方式とは異なる伝送用の符号化方式で符号化された符号化データを送信するものであってもよい。後者の場合、復号部PROD_D2と送信部PROD_D5との間に、動画像を伝送用の符号化方式で符号化する符号化部(不図示)を介在させるとよい。
 このような再生装置PROD_Dとしては、例えば、DVDプレイヤ、BDプレイヤ、HDDプレイヤなどが挙げられる(この場合、テレビジョン受像機等が接続される出力端子PROD_D4が動画像の主な供給先となる)。また、テレビジョン受像機(この場合、ディスプレイPROD_D3が動画像の主な供給先となる)、デジタルサイネージ(電子看板や電子掲示板等とも称され、ディスプレイPROD_D3又は送信部PROD_D5が動画像の主な供給先となる)、デスクトップ型PC(この場合、出力端子PROD_D4又は送信部PROD_D5が動画像の主な供給先となる)、ラップトップ型又はタブレット型PC(この場合、ディスプレイPROD_D3又は送信部PROD_D5が動画像の主な供給先となる)、スマートフォン(この場合、ディスプレイPROD_D3又は送信部PROD_D5が動画像の主な供給先となる)なども、このような再生装置PROD_Dの一例である。
 本発明は、画像データが符号化された符号化データを復号する画像復号装置、および、画像データが符号化された符号化データを生成する画像符号化装置に好適に適用することができる。また、画像符号化装置によって生成され、画像復号装置によって参照される符号化データのデータ構造に好適に適用することができる。
   1  動画像復号装置(画像復号装置)
   2  動画像符号化装置(画像符号化装置)
  12、12A  TU情報復号部
 121  領域分割部(変換単位分割手段)
 122  領域復号部(変換係数復号手段)
 280、280A  TU情報符号化部
 281  領域分割部(変換単位分割手段)
 282  領域符号化部(変換係数符号化手段)
 320  相対位置モード復号部
 321  最後の非ゼロ係数復号部
 322  相対位置復号部(相対位置復号手段)
 323  係数位置決定部(位置特定手段)
 310  ラン-レベルモード復号部(復号手段)
 420  相対位置モード符号化部
 421  最後の非ゼロ係数符号化部
 422  相対位置算出部(相対位置符号化手段)
 423  相対位置符号化部(相対位置符号化手段)
 BLK  対象ブロック(変換単位)
 R11~R14 復号領域(サブ単位)
 TBL11、TBL30 VLCテーブル(復号情報)
 TBL21、TBL40 VLCテーブル(符号化情報)
 TUI  TU情報(符号化データ)

Claims (15)

  1.  対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データから、該変換係数を復号する画像復号装置において、
     上記変換単位を、複数のサブ単位に分割する変換単位分割手段と、
     上記符号化データから上記変換係数を得るための復号情報であって、上記サブ単位ごとに割り当てられている復号情報を参照して、上記サブ単位に含まれる変換係数を復号する変換係数復号手段と、を備えることを特徴とする画像復号装置。
  2.  上記変換係数復号手段は、上記サブ単位における非ゼロの変換係数の有無を示す非ゼロ情報を参照し、該非ゼロ情報が、上記サブ単位における非ゼロの変換係数が無いことを示すとき、上記サブ単位の復号処理を省略することを特徴とする請求項1に記載の画像復号装置。
  3.  上記復号情報は、上記変換単位における、上記サブ単位の位置に応じて適応的に定義されていることを特徴とする請求項1または2に記載の画像復号装置。
  4.  上記変換係数を示すパラメータの出現頻度に応じて、上記復号情報において該パラメータに割り当てられている符号をより短いものに更新する復号情報更新手段を備えることを特徴とする請求項1から3のいずれか1項に記載の画像復号装置。
  5.  上記変換係数復号手段は、連続する非ゼロ係数の長さと、変換係数の絶対値と、変換係数の符号とを復号する第1モード復号手順を所定条件下において実行した後、変換係数の絶対値と、変換係数の符号とを復号する第2モード復号手順を実行する復号処理を行うことを特徴とする請求項1から4のいずれか1項に記載の画像復号装置。
  6.  上記変換係数復号手段は、上記変換単位における、上記サブ単位の位置に応じて、上記所定条件を、変更することを特徴とする請求項5に記載の画像復号装置。
  7.  上記変換単位における低周波成分側の所定領域に限り変換係数を復号する限定領域復号手段と、
     上記変換単位分割手段および上記変換係数復号手段による復号処理と、上記限定領域復号手段による復号処理とを切り替える切り替え手段とを備えることを特徴とする請求項1から6のいずれか1項に記載の画像復号装置。
  8.  上記変換単位分割手段は、分割した上記複数のサブ単位を再帰的に分割することを特徴とする請求項1から7のいずれか1項に記載の画像復号装置。
  9.  対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する画像符号化装置において、
     上記変換単位を、複数のサブ単位に分割する変換単位分割手段と、
     上記変換係数を符号化するための符号化情報であって、上記サブ単位ごとに割り当てられている符号化情報を参照して、上記変換単位に含まれる変換係数を符号化する変換係数符号化手段と、を備えることを特徴とする画像符号化装置。
  10.  対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化することにより生成される符号化データのデータ構造において、
     上記変換単位を複数のサブ単位に分割させるか否かを示す分割フラグを含み、
     上記符号化データを復号する画像復号装置は、上記分割フラグが上記変換単位を複数のサブ単位に分割させることを示しているときに、上記変換単位を複数のサブ単位に分割すると共に、サブ単位毎に上記変換係数を復号する、
    ことを特徴とする符号化データのデータ構造。
  11.  対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化して得られる符号化データから、該変換係数を復号する画像復号装置において、
     復号対象となる変換係数のひとつ前に復号した変換係数からの相対位置を復号する相対位置復号手段と、
     ひとつ前に復号した上記変換係数の、上記変換単位における位置と、上記相対位置とから上記復号対象となる上記変換係数の位置を特定する位置特定手段と、を備えることを特徴とする画像復号装置。
  12.  上記変換単位における低周波成分側の領域について、連続する非ゼロ係数の長さと変換係数の絶対値と変換係数の符号とを復号する第1モード復号処理、および、変換係数の絶対値と変換係数の符号とを復号する第2モード復号処理を実行する復号手段を備えることを特徴とする請求項11に記載の画像復号装置。
  13.  上記復号手段は、復号対象となる変換単位の特性に応じて上記領域のサイズを変更することを特徴とする請求項12に記載の画像復号装置。
  14.  対象画像の画素値を変換単位ごとに周波数変換して得られた変換係数を符号化する画像符号化装置において、
     符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を符号化する相対位置符号化手段を備えることを特徴とする画像復号装置。
  15.  対象画像の画素値を変換単位ごとに周波数変換に変換して得られた変換係数を符号化することにより生成された符号化データのデータ構造において、
     符号化対象となる上記変換係数の位置のひとつ前に符号化した上記変換係数の位置に対する相対位置を含み、
     上記符号化データを復号する画像復号装置は、ひとつ前に復号した上記変換係数の上記変換単位における位置と、上記相対位置とから、復号対象となる上記変換係数の位置を特定することを特徴とする符号化データのデータ構造。
PCT/JP2012/061478 2011-04-27 2012-04-27 画像復号装置、画像符号化装置、および符号化データのデータ構造 Ceased WO2012147966A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
JP2013512497A JP6051156B2 (ja) 2011-04-27 2012-04-27 画像復号装置および画像符号化装置
RU2013152171A RU2609096C2 (ru) 2011-04-27 2012-04-27 Устройство декодирования изображений, устройство кодирования изображений и структура данных кодированных данных
CN201280020067.0A CN103493494A (zh) 2011-04-27 2012-04-27 图像解码装置、图像编码装置以及编码数据的数据结构

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2011100081 2011-04-27
JP2011-100081 2011-04-27

Publications (1)

Publication Number Publication Date
WO2012147966A1 true WO2012147966A1 (ja) 2012-11-01

Family

ID=47072477

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2012/061478 Ceased WO2012147966A1 (ja) 2011-04-27 2012-04-27 画像復号装置、画像符号化装置、および符号化データのデータ構造

Country Status (4)

Country Link
JP (2) JP6051156B2 (ja)
CN (1) CN103493494A (ja)
RU (2) RU2609096C2 (ja)
WO (1) WO2012147966A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110456983A (zh) * 2019-04-17 2019-11-15 上海酷芯微电子有限公司 面向深度学习芯片稀疏计算的数据存储结构和方法
CN115379216A (zh) * 2018-06-03 2022-11-22 Lg电子株式会社 视频信号的解码、编码和发送设备及存储视频信号的介质

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP6051156B2 (ja) * 2011-04-27 2016-12-27 シャープ株式会社 画像復号装置および画像符号化装置
WO2020007489A1 (en) * 2018-07-06 2020-01-09 Huawei Technologies Co., Ltd. A picture encoder, a picture decoder and corresponding methods
CN109840471B (zh) * 2018-12-14 2023-04-14 天津大学 一种基于改进Unet网络模型的可行道路分割方法
US11184622B2 (en) * 2019-09-18 2021-11-23 Sharp Kabushiki Kaisha Video decoding apparatus and video coding apparatus
KR20230133891A (ko) 2021-02-04 2023-09-19 베이징 다지아 인터넷 인포메이션 테크놀로지 컴퍼니 리미티드 비디오 코딩을 위한 잔차 및 계수 코딩
CN116600130B (zh) * 2022-01-19 2024-10-29 杭州海康威视数字技术股份有限公司 一种系数解码方法、装置、图像解码器及电子设备

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2004134896A (ja) * 2002-10-08 2004-04-30 Ntt Docomo Inc 画像符号化方法、画像復号方法、画像符号化装置、画像復号装置、画像処理システム、画像符号化プログラム、画像復号プログラム。
JP2006033508A (ja) * 2004-07-16 2006-02-02 Olympus Corp 適応型可変長符号化装置、適応型可変長復号化装置、適応型可変長符号化・復号化方法、及び適応型可変長符号化・復号化プログラム
JP2011501535A (ja) * 2007-10-12 2011-01-06 クゥアルコム・インコーポレイテッド ビデオブロックのインターリーブされたサブブロックのエントロピーコード化

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP4931034B2 (ja) * 2004-06-10 2012-05-16 株式会社ソニー・コンピュータエンタテインメント 復号装置および復号方法、並びに、プログラムおよびプログラム記録媒体
CN1779716A (zh) * 2005-05-26 2006-05-31 智多微电子(上海)有限公司 一种快速游程长度编解码电路的实现方法
WO2010063883A1 (en) * 2008-12-03 2010-06-10 Nokia Corporation Switching between dct coefficient coding modes
US8406546B2 (en) * 2009-06-09 2013-03-26 Sony Corporation Adaptive entropy coding for images and videos using set partitioning in generalized hierarchical trees
JP2011049816A (ja) * 2009-08-27 2011-03-10 Kddi R & D Laboratories Inc 動画像符号化装置、動画像復号装置、動画像符号化方法、動画像復号方法、およびプログラム
KR101457894B1 (ko) * 2009-10-28 2014-11-05 삼성전자주식회사 영상 부호화 방법 및 장치, 복호화 방법 및 장치
JP6051156B2 (ja) * 2011-04-27 2016-12-27 シャープ株式会社 画像復号装置および画像符号化装置

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2004134896A (ja) * 2002-10-08 2004-04-30 Ntt Docomo Inc 画像符号化方法、画像復号方法、画像符号化装置、画像復号装置、画像処理システム、画像符号化プログラム、画像復号プログラム。
JP2006033508A (ja) * 2004-07-16 2006-02-02 Olympus Corp 適応型可変長符号化装置、適応型可変長復号化装置、適応型可変長符号化・復号化方法、及び適応型可変長符号化・復号化プログラム
JP2011501535A (ja) * 2007-10-12 2011-01-06 クゥアルコム・インコーポレイテッド ビデオブロックのインターリーブされたサブブロックのエントロピーコード化

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
MARTA KARCZEWICZ ET AL.: "Variable Length Coding for Coded Block Flag and Large Transform", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC) OF ITU-T SG16 WP3 AND ISO/IEC JTC1/SC29/WG11 2ND MEETING, 21 July 2010 (2010-07-21), GENEVA, CH *
SUNIL LEE ET AL.: "Efficient coefficient coding method for large transform in VLC mode", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC) OF ITU-T SG16 WP3 AND ISO/IEC JTC1/SC29/WG11 3RD MEETING, vol. 7-15, October 2010 (2010-10-01), GUANGZHOU, CN *

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115379216A (zh) * 2018-06-03 2022-11-22 Lg电子株式会社 视频信号的解码、编码和发送设备及存储视频信号的介质
CN115379215A (zh) * 2018-06-03 2022-11-22 Lg电子株式会社 视频信号的解码、编码和传输方法及存储视频信号的介质
US12132922B2 (en) 2018-06-03 2024-10-29 Lg Electronics Inc. Method and apparatus for processing video signals using reduced transform
CN115379215B (zh) * 2018-06-03 2024-12-10 Lg电子株式会社 视频信号的解码、编码和传输方法及存储视频信号的介质
CN110456983A (zh) * 2019-04-17 2019-11-15 上海酷芯微电子有限公司 面向深度学习芯片稀疏计算的数据存储结构和方法

Also Published As

Publication number Publication date
RU2653319C1 (ru) 2018-05-07
CN103493494A (zh) 2014-01-01
JPWO2012147966A1 (ja) 2014-07-28
RU2609096C2 (ru) 2017-01-30
JP6407944B2 (ja) 2018-10-17
JP2017085586A (ja) 2017-05-18
JP6051156B2 (ja) 2016-12-27
RU2013152171A (ru) 2015-06-10

Similar Documents

Publication Publication Date Title
US12149735B2 (en) Image decoding device, image encoding device, and image decoding method
JP7200320B2 (ja) 画像フィルタ装置、フィルタ方法および動画像復号装置
US10547861B2 (en) Image decoding device
JP6407944B2 (ja) 画像復号装置、画像符号化装置、画像復号方法、画像符号化方法、制御プログラム、および記録媒体
JP5972888B2 (ja) 画像復号装置、画像復号方法および画像符号化装置
JP2013034161A (ja) 画像復号装置、画像符号化装置、および符号化データのデータ構造
JP2013141094A (ja) 画像復号装置、画像符号化装置、画像フィルタ装置、および符号化データのデータ構造
WO2012090962A1 (ja) 画像復号装置、画像符号化装置、および符号化データのデータ構造、ならびに、算術復号装置、算術符号化装置
AU2015264943B2 (en) Image decoding device, image encoding device, and data structure of encoded data
HK1191161A (en) Image decoding apparatus, image encoding apparatus, and data structure of encoded data
WO2012081706A1 (ja) 画像フィルタ装置、フィルタ装置、復号装置、符号化装置、および、データ構造
JP2012182753A (ja) 画像復号装置、画像符号化装置、および符号化データのデータ構造
WO2012043676A1 (ja) 復号装置、符号化装置、および、データ構造
HK1200255B (zh) 图像解码装置、图像编码装置
HK1191484B (en) Image decoding apparatus, image encoding apparatus, and data structure of encoded data

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 12777577

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2013512497

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2013152171

Country of ref document: RU

Kind code of ref document: A

122 Ep: pct application non-entry in european phase

Ref document number: 12777577

Country of ref document: EP

Kind code of ref document: A1