WO2005020587A1 - Adaptive interframe wavelet video coding method, computer readable recording medium and system therefor - Google Patents

Adaptive interframe wavelet video coding method, computer readable recording medium and system therefor Download PDF

Info

Publication number
WO2005020587A1
WO2005020587A1 PCT/KR2004/002050 KR2004002050W WO2005020587A1 WO 2005020587 A1 WO2005020587 A1 WO 2005020587A1 KR 2004002050 W KR2004002050 W KR 2004002050W WO 2005020587 A1 WO2005020587 A1 WO 2005020587A1
Authority
WO
WIPO (PCT)
Prior art keywords
frames
motion vectors
pixels
group
boundary
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2004/002050
Other languages
French (fr)
Inventor
Ho-Jin Ha
Chang-Hoon Yim
Bae-Keun Lee
Woo-Jin Han
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Samsung Electronics Co Ltd
Original Assignee
Samsung Electronics Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020030065863A external-priority patent/KR100577364B1/en
Application filed by Samsung Electronics Co Ltd filed Critical Samsung Electronics Co Ltd
Priority to JP2006524561A priority Critical patent/JP2007503750A/en
Publication of WO2005020587A1 publication Critical patent/WO2005020587A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
    • H04N19/615Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding using motion compensated temporal filtering [MCTF]
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/12Selection from among a plurality of transforms or standards, e.g. selection between discrete cosine transform [DCT] and sub-band transform or selection between H.263 and H.264
    • H04N19/122Selection of transform size, e.g. 8x8 or 2x4x8 DCT; Selection of sub-band transforms of varying structure or type
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • H04N19/137Motion inside a coding unit, e.g. average field, frame or block difference
    • H04N19/139Analysis of motion vectors, e.g. their magnitude, direction, variance or reliability
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/63Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding using sub-band based transform, e.g. wavelets
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/63Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding using sub-band based transform, e.g. wavelets
    • H04N19/635Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding using sub-band based transform, e.g. wavelets characterised by filter definition or implementation details
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/13Adaptive entropy coding, e.g. adaptive variable length coding [AVLC] or context adaptive binary arithmetic coding [CABAC]

Definitions

  • the present invention relates to a wavelet video coding method, a computer readable recording medium and system therefor, and more particularly, to an interframe wavelet video coding (IWVC) method which decreases an average temporal distance by changing a temporal filtering direction.
  • IWVC interframe wavelet video coding
  • Multimedia data requires large capacity storage mediums and wide bandwidths for transmission since the amount of multimedia data is usually large.
  • a 24-bit true color image having a resolution of 640 * 480 needs a capacity of 640 * 480 * 24 bits, i.e., data of about 7.37 Mbits, per frame.
  • a bandwidth of 221 Mbits/sec is required.
  • a 90-minute movie based on such an image is stored, a storage space of about 1200 Gbits is required.
  • a compression coding method is a requisite for transmitting multimedia data including text, video, and audio.
  • a basic principle of data compression is removing data redundancy.
  • Data can be compressed by removing spatial redundancy in which the same color or object is repeated in an image, temporal redundancy in which there is little change between adjacent frames in a moving image or the same sound is repeated in audio, or mental visual redundancy taking into account human eyesight and limited perception of high frequency.
  • Data compression can be classified into lossy/lossless compression according to whether source data is lost, intraframe/interframe compression according to whether individual frames are compressed independently, and symmetric/ asymmetric compression according to whether time required for compression is the same as time required for recovery.
  • data compression is defined as realtime compression when a compression/recovery time delay does not exceed 50 ms and as scalable compression when frames have different resolutions.
  • lossless compression is usually used.
  • For multimedia data lossy compression is usually used.
  • intraframe compression is usually used to remove spatial redundancy
  • interframe compression is usually used to remove temporal redundancy.
  • FIG. 1 is a flowchart of a conventional three-dimensional IWVC method.
  • an image is received in group-of-frames (GOF) units in step SI.
  • the GOF includes a plurality of frames, e.g., 16 frames.
  • various operations are performed in GOF units.
  • HVSBM hierarchical variable size block matching
  • FIG. 2 which illustrates a motion estimation using HVSBM
  • an original image has a size of N * N
  • images of level 0 (N * N), of level 1 (W2 * N2), and of level 2 (W4 * NA) are obtained using wavelet transform.
  • a motion estimation block size is changed from 16 * 16 to 8 * 8 and 4 * 4
  • a motion estimation (ME) and a Magnitude of Absolute Distortion (MAD) are obtained with respect to each block.
  • the motion estimation block size is changed from 32 * 32 to 16 * 16, 8 * 8, and 4 * 4, and an ME and a MAD are obtained with respect to each block.
  • the motion estimation block size is changed from 64 * 64 to 32 * 32, 16 * 16, 8 * 8, and 4 * 4, and an ME and a MAD are obtained with respect to each block.
  • Motion compensated temporal filtering is performed using a pruned optimal ME in step S4.
  • MCTF is performed forward with respect to 16 image frames, thereby obtaining 8 low-frequency frames and 8 high-frequency frames.
  • MCTF is performed forward with respect to the 8 low- frequency frames, thereby obtaining 4 low- frequency frames and 4 high-frequency frames.
  • MCTF is performed forward with respect to the 4 low- frequency frames obtained at temporal level 1, thereby obtaining 2 low- frequency frames and 2 high-frequency frames.
  • MCTF is performed forward with respect to the 2 low- frequency frames obtained at temporal level 2, thereby obtaining a single low-frequency frame and a single high-frequency frame. Accordingly, as a result of MCTF, a total of 16 subbands HI, H3, H5, H7, H9, Hl l, H13, H15, LH2, LH6, LH10, LH14, LLH4, LLH12, LLLH8, and LLLL16 including 15 high-frequency frames and a single low- frequency frame at the last level are obtained.
  • step S5 After obtaining the 16 subbands, spatial transform and quantization are performed on the 16 subbands in step S5. Thereafter, a bitstream including data obtained by performing spatial transform and quantization on the 16 subbands, motion estimation data, and a header is generated in step S6.
  • FIG. 4 is diagram comparing performances of conventional MCTF with respect to a boundary condition.
  • FIG. 4 illustrates a best case of forward MCTF where an external image comes into a frame and a worst case of forward MCTF where an internal image goes out of the frame.
  • MCTF is performed forward, a temporally preceding image is replaced with a filtered high-frequency image, and a temporally succeeding image is replaced with a filtered low- frequency image.
  • high-frequency frames and a single low- frequency frame at a highest level are used. In other words, performance of video coding depends on whether a component of a high-frequency frame is large or small.
  • the present invention provides an adaptive interframe wavelet video coding (IWVC) method allowing a direction of temporal filtering to be changed according to a boundary condition.
  • IWVC adaptive interframe wavelet video coding
  • the present invention also provides a computer readable recording medium and a system which can perform the adaptive IWVC method.
  • an IWVC method comprising, (a) receiving a group-of-frames including a plurality of frames and determining a mode flag according to a predetermined procedure using motion vectors of boundary pixels; (b) temporally decomposing the frames included in the group- of-frames in predetermined directions in accordance with the determined mode flag; and (c) performing spatial transform and quantization on the frames obtained by performing step (b), thereby generating a bitstream.
  • the group-of-frames comprises 16 frames.
  • Step (a) may comprise determining the mode flag according to the predetermined procedure using motion vectors obtained at a boundary having a predetermined thickness among motion vectors of pixels obtained through motion estimation using hierarchical variable size block matching (HVSBM).
  • the motion vectors used to determine the mode flag may be motion vectors of pixels at left and right boundaries, or motion vectors of pixels at left, right, upper and lower boundaries.
  • the mode flag F is preferably determined using the following algorithm:
  • Programs executing the adaptive IWVC method may be recorded onto a computer readable recording medium to be used in a computer.
  • an IWVC system which receives a group-of-frames including a plurality of frames and generates a bitstream.
  • the IWVC system comprises a motion estimation/mode determination block which receives the group-of-frames, obtains motion vectors of pixels in each of the frames using a predetermined procedure, and determines a mode flag using motion vectors of boundary pixels among the obtained motion vectors; and a motion compensation temporal filtering block which decomposes the frames into low- and high- frequency frames in a predetermined temporal direction in accordance with the mode flag determined by the motion estimation/mode determination block using the motion vectors.
  • the interframe wavelet video coding system may further comprise a spatial transform block which wavelet-decomposes the low- and high-frequency frames generated by the motion compensation temporal filtering block into spatial low- and high-frequency components.
  • FIG. 1 is a flowchart of a conventional three-dimensional interframe wavelet video coding (IWVC) method
  • FIG. 2 illustrates conventional motion estimation using hierarchical variable size block matching (HVSBM);
  • FIG. 3 illustrates conventional motion compensated temporal filtering (MCTF);
  • FIG. 4 is diagram comparing performances of conventional MCTF with respect to a boundary condition
  • FIG. 5 is a flowchart of an adaptive IWVC method according to an embodiment of the present invention.
  • FIG. 6 illustrates a reference for determining an MCTF direction according to a boundary condition
  • FIGS. 7 and 8 illustrate boundary pixels used to determine a mode flag
  • FIG. 9 illustrates MCTF directions according to a mode flag representing a boundary condition
  • FIG. 10 is a functional block diagram of a system for adaptive IWVC according to an embodiment of the present invention. Mode for Invention
  • FIG. 5 is a flowchart of an adaptive interframe wavelet video coding (IWVC) method according to an embodiment of the present invention.
  • IWVC adaptive interframe wavelet video coding
  • a single GOF includes a plurality of frames and preferably includes 2 frames (where 'n' is a natural number), e.g., 2, 4, 8, 16, or 32 frames, to facilitate computation and management.
  • 2 frames where 'n' is a natural number
  • video coding efficiency increases while buffering time and coding time also increases unfavorably.
  • video coding efficiency decreases.
  • a single GOF includes 16 frames.
  • step S20 After receiving the image, motion estimation is performed and a mode flag is set in step S20.
  • the motion estimation is performed using hierarchical variable size block matching (HVSBM) as described with reference to FIG. 1.
  • HVSBM hierarchical variable size block matching
  • the mode flag is used to determine a direction of temporal filtering according to a boundary condition. A reference for determining a mode flag will be described with reference to FIGS. 6, 7 and 8.
  • MCTF motion compensated temporal filtering
  • step S50 16 subbands resulting from the MCTF are subjected to spatial transform and quantization in step S50. Thereafter, a bitstream including data resulting from the spatial transform and quantization, motion vector data, and the mode flag is generated in step S60.
  • FIG. 6 illustrates a reference for determining an MCTF direction according to a boundary condition
  • FIGS. 7 and 8 illustrate boundary pixels used to determine a mode flag.
  • FIGS. 6 illustrates cases where an internal image goes out of the frame.
  • FIG. 6 illustrates forward MCTF and backward MCTF.
  • image blocks B and N flow out of the frame when a T-l frame is converted into a T frame.
  • the image blocks B and N in the T-l frame do not have their matches in the T frame.
  • the image blocks B and N in the T-l frame are compared with image blocks C and M, respectively, in the T frame.
  • a difference between the image blocks B and C and a difference between the image blocks N and M are large, which increases the amount of information of the T-l frame to be replaced with a high-frequency frame.
  • each image block in the T frame to be replaced with a high-frequency frame has its matches in the T-l frame, and therefore, the amount of information of the high-frequency frame, i.e., the T frame, may be decreased.
  • forward MCTF is more efficient in a case where a new image comes into a frame through a boundary while backward MCTF is more efficient in a case where an image goes out of the frame through a boundary.
  • video coding efficiency and performance can be increased by properly selecting either forward or backward MCTF according to a boundary condition of an input GOF.
  • a mode flag a basic principle is made that forward MCTF is used when a new image comes into a frame, backward MCTF is used when an image goes out of a frame, and forward MCTF and backward MCTF are properly combined in other cases.
  • the mode flag can be determined using a motion vector for pixels at a boundary of a frame. As shown in FIG. 7, pixels at right and left boundaries of a frame may be used in a first embodiment. Alternatively, as shown in FIG. 8, pixels at right, left, upper, and lower boundaries of a frame may be used in a second embodiment. Video coding performance depends on a thickness of a boundary used to determine the mode flag. Where the boundary is too thin, information regarding output/input of a particular image may be missed. Conversely, where the boundary is too thick, a boundary condition may not be sharply identified. Accordingly, the thickness of the boundary needs to be appropriately determined. In embodiments of the present invention, the boundary has a thickness of 32 pixels.
  • determining the mode flag motion vectors of pixels in each frames are obtained using HVSBM.
  • a mode flag is determined based on the motion vectors of pixels in the frames.
  • the mode flag may be different according to a temporal level, but it is preferable to determine the mode flag at temporal level 0.
  • a mode flag is determined using motion vectors at left and right boundaries of each frame because a new image usually comes into or goes out of a frame of a moving picture in an X direction.
  • An average of motion vectors of pixels at the left boundary of each of all frames included in a single GOF is obtained.
  • An X component of the average motion vector at the left boundary is denoted by 'L.'
  • an average of motion vectors of pixels at the right boundary of each of all frames included in a single GOF is obtained.
  • An X component of the average motion vector at the right boundary is denoted by 'R.'
  • an L value less than 0 indicates that an image comes into the frame through the left boundary
  • an R value less than 0 indicates that an image goes out of the frame through the right boundary.
  • the L value greater than 0 and the R value greater than 0 account for the opposite cases, respectively.
  • the L or R value may not be 0 even if an image does not come in or go out of the frame. Accordingly, it is preferable that the L and R values not exceeding a predetermined threshold are determined as 0.
  • the L value is less than 0 and the R value is equal to or greater than 0, or the L value is less than 0 and the R value is greater than 0. In this case, it is preferable to use forward MCTF. Conversely, when an image goes out of the frame through the left or right boundary, the L value is greater than 0 and the R value is equal to or less than 0, or the L value is greater than 0 and the R value is less than 0. In this case, it is preferable to use backward MCTF. When an image comes into the frame through the left boundary and an image goes out of the frame through the right boundary, it is preferable to appropriately combine forward MCTF and backward MCTF.
  • a mode flag F can be determined by the following algorithm:
  • a mode flag F can be determined by the following algorithm:
  • the first and second embodiments are exemplary, and the spirit of the present invention is not restricted thereto.
  • a direction of MCTF is appropriately determined using information regarding image input/output at a boundary. Accordingly, the present invention will be considered as including a case where a mode flag is determined to be different among two or some frames more than two in a GOF in addition to the first and second embodiments where a mode flag is determined using average motion vectors obtained with respect to all of the frames in a GOF.
  • FIG. 9 illustrates MCTF directions according to a mode flag representing a boundary condition.
  • MCTF directions are depicted as + + + + + + + + +.
  • MCTF directions are depicted as .
  • MCTF directions may be depicted in various ways, but FIG. 9 illustrates an example where MCTF directions are depicted as + - + - + - + - at temporal level 0.
  • '+' indicates a forward direction
  • '-' indicates a backward direction.
  • MCTF is performed in the same direction.
  • video coding performance changes depending on a combination of forward and backward directions.
  • a sequence of forward and backward directions may be determined in various ways. Representative examples of a sequence of MCTF directions in the forward, backward, and bi-directional modes are shown in Table 1.
  • the cases V and 'd' are characterized in that a low- frequency frame (hereinafter, referred to as a reference frame) at a last level is positioned at a center (i.e., an 8th frame) among 1st through 16th frames.
  • the reference frame is a most essential frame in video coding.
  • the other frames are recovered based on the reference frame. As a temporal distance between a frame and the reference frame increases, recovery performance decreases.
  • a combination of forward MCTF and backward MCTF is made such that the reference frame is positioned at the center, i.e., the 8th frame, to minimize a temporal distance between the reference frame and each of the other frames.
  • an average temporal distance is minimized.
  • ATD average temporal distance
  • temporal distances are calculated.
  • a temporal distance is defined as a positional difference between two frames. Referring to FIG. 3, a temporal distance between a first frame and a second frame is defined as 1, and a temporal distance between the frame L2 and the frame L4 is defined as 2.
  • An ATD is obtained by dividing the sum of temporal distances between frames subjected to an operation for motion estimation in pairs by the number of pairs of frames defined for the motion estimation.
  • FIG. 10 is a functional block diagram of a system for adaptive IWVC according to an embodiment of the present invention.
  • the system for adaptive IWVC includes a motion estimation/mode determination block 10 which obtains a motion vector and determines a mode using the motion vector, a motion compensation temporal filtering block 40 which removes temporal redundancy using the motion vector and the determined mode, a spatial transform block 50 which removes spatial redundancy, a motion vector encoding block 20 which encodes the motion vector using a predetermined algorithm, a quantization block 60 which quantizes wavelet coefficients of respective components generated by the spatial transform block 50, and a buffer 30 which temporarily stores an encoded bitstream received from the quantization block 60.
  • the motion estimation/mode determination block 10 obtains a motion vector used by the motion compensation temporal filtering block 40 using a hierarchical method such as HVSBM. In addition, the motion estimation/mode determination block 10 determines a mode flag for determining temporal filtering directions.
  • the motion compensation temporal filtering block 40 decomposes frames into low- and high-frequency frames in a temporal direction using the motion vector obtained by the motion estimation/mode determination block 10. A direction of the decomposition is determined according to the mode flag. Frames are decomposed in GOF units. Through such decomposition, temporal redundancy is removed.
  • the spatial transform block 50 wavelet-decomposes frames that have been decomposed in the temporal direction by the motion compensation temporal filtering block 40 into spatial low- and high-frequency components, thereby removing spatial redundancy.
  • the motion vector encoding block 20 encodes the motion vector and the mode flag hierarchically obtained by the motion estimation/mode determination block 10 and then transmits the encoded motion vector and the encoded mode flag to the buffer 30.
  • the quantization block 60 quantizes and encodes wavelet coefficients of components generated by the spatial transform block 50.
  • the buffer 30 stores a bitstream including encoded data, the encoded motion vector, and the encoded mode flag before transmission and is controlled by a rate control algorithm.
  • IWVC can be adaptively performed in accordance with a boundary condition.
  • a PSNR is increased in the present invention.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Discrete Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

An adaptive interframe wavelet video coding method, a computer readable recording medium and system therefor are provided. The interframe wavelet video coding method includes (a) receiving a group-of-frames including a plurality of frames and determining a mode flag according to a predetermined procedure using motion vectors of boundary pixels, (b) temporally decomposing the frames included in the group-of-frames in predetermined directions in accordance with the determined mode flag, and (c) performing spatial transform and quantization on the frames obtained by performing step (b), thereby generating a bitstream. Since an appropriate temporal filtering is performed in accordance with a boundary condition, efficiency of interframe wavelet video coding is increased.

Description

Description ADAPTIVE INTERFRAME WAVELET VIDEO CODING METHOD, COMPUTER READABLE RECORDING MEDIUM AND SYSTEM THEREFOR Technical Field
[1] The present invention relates to a wavelet video coding method, a computer readable recording medium and system therefor, and more particularly, to an interframe wavelet video coding (IWVC) method which decreases an average temporal distance by changing a temporal filtering direction. Background Art
[2] With the development of information communication technology including the Internet, video communication as well as text and voice communication has increased. Conventional text communication cannot satisfy the various demands of users, and thus multimedia services that can provide various types of information such as text, pictures, and music have increased. Multimedia data requires large capacity storage mediums and wide bandwidths for transmission since the amount of multimedia data is usually large. For example, a 24-bit true color image having a resolution of 640 * 480 needs a capacity of 640 * 480 * 24 bits, i.e., data of about 7.37 Mbits, per frame. When this image is transmitted at a speed of 30 frames per second, a bandwidth of 221 Mbits/sec is required. When a 90-minute movie based on such an image is stored, a storage space of about 1200 Gbits is required. Accordingly, a compression coding method is a requisite for transmitting multimedia data including text, video, and audio.
[3] A basic principle of data compression is removing data redundancy. Data can be compressed by removing spatial redundancy in which the same color or object is repeated in an image, temporal redundancy in which there is little change between adjacent frames in a moving image or the same sound is repeated in audio, or mental visual redundancy taking into account human eyesight and limited perception of high frequency. Data compression can be classified into lossy/lossless compression according to whether source data is lost, intraframe/interframe compression according to whether individual frames are compressed independently, and symmetric/ asymmetric compression according to whether time required for compression is the same as time required for recovery. In addition, data compression is defined as realtime compression when a compression/recovery time delay does not exceed 50 ms and as scalable compression when frames have different resolutions. For text or medical data, lossless compression is usually used. For multimedia data, lossy compression is usually used. Meanwhile, intraframe compression is usually used to remove spatial redundancy, and interframe compression is usually used to remove temporal redundancy.
[4] FIG. 1 is a flowchart of a conventional three-dimensional IWVC method.
[5] First, an image is received in group-of-frames (GOF) units in step SI. The GOF includes a plurality of frames, e.g., 16 frames. In IWVC, various operations are performed in GOF units.
[6] Next, motion estimation is performed using hierarchical variable size block matching (HVSBM) in step S2. Referring to FIG. 2, which illustrates a motion estimation using HVSBM, an original image has a size of N * N, images of level 0 (N * N), of level 1 (W2 * N2), and of level 2 (W4 * NA) are obtained using wavelet transform. For the image of level 2, a motion estimation block size is changed from 16 * 16 to 8 * 8 and 4 * 4, and a motion estimation (ME) and a Magnitude of Absolute Distortion (MAD) are obtained with respect to each block. Similarly, for the image of level 1, the motion estimation block size is changed from 32 * 32 to 16 * 16, 8 * 8, and 4 * 4, and an ME and a MAD are obtained with respect to each block. For the image of level 0, the motion estimation block size is changed from 64 * 64 to 32 * 32, 16 * 16, 8 * 8, and 4 * 4, and an ME and a MAD are obtained with respect to each block.
[7] Next, as shown in FIG. 1, an ME tree is pruned to minimize the MAD in step S3.
[8] Motion compensated temporal filtering (MCTF) is performed using a pruned optimal ME in step S4. Referring to FIG. 3, at temporal level 0, MCTF is performed forward with respect to 16 image frames, thereby obtaining 8 low-frequency frames and 8 high-frequency frames. At temporal level 1, MCTF is performed forward with respect to the 8 low- frequency frames, thereby obtaining 4 low- frequency frames and 4 high-frequency frames. At temporal level 2, MCTF is performed forward with respect to the 4 low- frequency frames obtained at temporal level 1, thereby obtaining 2 low- frequency frames and 2 high-frequency frames. Lastly, at temporal level 3, MCTF is performed forward with respect to the 2 low- frequency frames obtained at temporal level 2, thereby obtaining a single low-frequency frame and a single high-frequency frame. Accordingly, as a result of MCTF, a total of 16 subbands HI, H3, H5, H7, H9, Hl l, H13, H15, LH2, LH6, LH10, LH14, LLH4, LLH12, LLLH8, and LLLL16 including 15 high-frequency frames and a single low- frequency frame at the last level are obtained.
[9] After obtaining the 16 subbands, spatial transform and quantization are performed on the 16 subbands in step S5. Thereafter, a bitstream including data obtained by performing spatial transform and quantization on the 16 subbands, motion estimation data, and a header is generated in step S6.
[10] Although such conventional IWVC has excellent scalability, it does not have satisfactory performance as compared to other conventional video coding methods. An example of IWVC performance depending upon a boundary condition will be described with reference to FIG. 4.
[11] FIG. 4 is diagram comparing performances of conventional MCTF with respect to a boundary condition.
[12] FIG. 4 illustrates a best case of forward MCTF where an external image comes into a frame and a worst case of forward MCTF where an internal image goes out of the frame. Where MCTF is performed forward, a temporally preceding image is replaced with a filtered high-frequency image, and a temporally succeeding image is replaced with a filtered low- frequency image. For video coding, high-frequency frames and a single low- frequency frame at a highest level are used. In other words, performance of video coding depends on whether a component of a high-frequency frame is large or small.
[13] In a case where the external image comes into the frame, a T-l frame is replaced with a high-frequency image, and a T frame is replaced with a low-frequency image. All image blocks in the T-l frame can be exactly matched with image blocks, respectively, in the T-frame, and thus a magnitude of a high frequency component proportional to a difference between two image blocks is less compared to a case where the image blocks are not matched exactly. In other words, a size of the T-l frame to be replaced with a high-frequency image is small. Disclosure of Invention Technical Problem
[14] Conversely, in the worst case where the internal image goes out of the frame, all of the image blocks in the T-l frame are not exactly matched with the image blocks in the T-frame. Here, image blocks A and N that do not have their matches are coupled with image blocks B and M, respectively, giving a least difference therebetween. Since a difference between the image blocks A and B and a difference between the image blocks N and M are needed to be expressed, the size of the T-l frame is increased. Technical Solution
[15] As described above, performance of MCTF greatly changes depending on a boundary condition such as an incoming image or an outgoing image. Therefore, a video coding method allowing a filtering direction to be adaptively changed according to a boundary condition during MCTF is desired.
[16] The present invention provides an adaptive interframe wavelet video coding (IWVC) method allowing a direction of temporal filtering to be changed according to a boundary condition.
[17] The present invention also provides a computer readable recording medium and a system which can perform the adaptive IWVC method.
[18] According to an aspect of the present invention, there is provided an IWVC method comprising, (a) receiving a group-of-frames including a plurality of frames and determining a mode flag according to a predetermined procedure using motion vectors of boundary pixels; (b) temporally decomposing the frames included in the group- of-frames in predetermined directions in accordance with the determined mode flag; and (c) performing spatial transform and quantization on the frames obtained by performing step (b), thereby generating a bitstream.
[19] Preferably, in step (a), the group-of-frames comprises 16 frames. Step (a) may comprise determining the mode flag according to the predetermined procedure using motion vectors obtained at a boundary having a predetermined thickness among motion vectors of pixels obtained through motion estimation using hierarchical variable size block matching (HVSBM). Meanwhile, the motion vectors used to determine the mode flag may be motion vectors of pixels at left and right boundaries, or motion vectors of pixels at left, right, upper and lower boundaries. In the first case, the mode flag F is preferably determined using the following algorithm:
[20] if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))then F=0 else if((L > 0 and R==0)or (L==0 and R < 0)or (L > 0 and R < 0))then F=l else F=2
[21] where L denotes an average of X components of motion vectors of pixels at the left boundary having the predetermined thickness, and R denotes an average of X components of motion vectors of pixels at the right boundary having the predetermined thickness, wherein step (b) comprises temporally decomposing the frames included in the group-of-frames in a forward direction when F=0, temporally decomposing the frames included in the group-of-frames in a backward direction when F=l, and temporally decomposing the frames included in the group-of-frames in forward and backward directions combined in a predetermined sequence when F=2. In the latter case, the mode flag F is preferably determined using the following algorithm: [22] if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if (abs(U)< Threshold)then U=0 if (abs(D)< Threshold)then D=0 if(((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))and ((D < 0 and U==0)or (D==0 and U > 0)or (D < 0 and U > 0)or (D —0 and U=0)))then F=0 else if(((L > 0 and R==0)or (L=0 and R < 0)or (L > 0 and R < 0))and ((D > 0 and U=0) or (D==0 and U < 0)or (D > 0 and U < 0)or (D ==0 and U=0)))then F=l else F=2
[23] where L denotes an average of X components of motion vectors of pixels at the left boundary having the predetermined thickness, R denotes an average of X components of motion vectors of pixels at the right boundary having the predetermined thickness, U denotes an average of Y components of motion vectors of pixels at the upper boundary having the predetermined thickness, and D denotes an average of Y components of motion vectors of pixels at the lower boundary having the predetermined thickness, wherein step (b) comprises temporally decomposing the frames included in the group-of-frames in a forward direction when F=0, temporally decomposing the frames included in the group-of-frames in a backward direction when F=l, and temporally decomposing the frames included in the group-of-frames in forward and backward directions combined in a predetermined sequence when F=2.
[24] In either case, when F=2 in step (b), the frames are preferably decomposed such that an average temporal distance between frames is minimized.
[25] Programs executing the adaptive IWVC method may be recorded onto a computer readable recording medium to be used in a computer.
[26] According to another aspect of the present invention, there is provided an IWVC system which receives a group-of-frames including a plurality of frames and generates a bitstream. The IWVC system comprises a motion estimation/mode determination block which receives the group-of-frames, obtains motion vectors of pixels in each of the frames using a predetermined procedure, and determines a mode flag using motion vectors of boundary pixels among the obtained motion vectors; and a motion compensation temporal filtering block which decomposes the frames into low- and high- frequency frames in a predetermined temporal direction in accordance with the mode flag determined by the motion estimation/mode determination block using the motion vectors.
[27] The interframe wavelet video coding system may further comprise a spatial transform block which wavelet-decomposes the low- and high-frequency frames generated by the motion compensation temporal filtering block into spatial low- and high-frequency components. Description of Drawings
[28] The above and other features and advantages of the present invention will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[29] FIG. 1 is a flowchart of a conventional three-dimensional interframe wavelet video coding (IWVC) method;
[30] FIG. 2 illustrates conventional motion estimation using hierarchical variable size block matching (HVSBM);
[31] FIG. 3 illustrates conventional motion compensated temporal filtering (MCTF);
[32] FIG. 4 is diagram comparing performances of conventional MCTF with respect to a boundary condition;
[33] FIG. 5 is a flowchart of an adaptive IWVC method according to an embodiment of the present invention;
[34] FIG. 6 illustrates a reference for determining an MCTF direction according to a boundary condition;
[35] FIGS. 7 and 8 illustrate boundary pixels used to determine a mode flag;
[36] FIG. 9 illustrates MCTF directions according to a mode flag representing a boundary condition; and [37] FIG. 10 is a functional block diagram of a system for adaptive IWVC according to an embodiment of the present invention. Mode for Invention
[38] Exemplary, non-limiting, embodiments of the present invention will now be described with reference to the accompanying drawings.
[39] FIG. 5 is a flowchart of an adaptive interframe wavelet video coding (IWVC) method according to an embodiment of the present invention.
[40] An image is received in group-of-frames (GOF) units in step S10. A single GOF includes a plurality of frames and preferably includes 2 frames (where 'n' is a natural number), e.g., 2, 4, 8, 16, or 32 frames, to facilitate computation and management. When the number of frames included in a GOF increases, video coding efficiency increases while buffering time and coding time also increases unfavorably. As the number of frames included in a GOF decreases, video coding efficiency decreases. In the embodiment of the present invention, a single GOF includes 16 frames.
[41] After receiving the image, motion estimation is performed and a mode flag is set in step S20. Preferably, the motion estimation is performed using hierarchical variable size block matching (HVSBM) as described with reference to FIG. 1. The mode flag is used to determine a direction of temporal filtering according to a boundary condition. A reference for determining a mode flag will be described with reference to FIGS. 6, 7 and 8.
[42] After the motion estimation and mode flag setup, pruning is performed in the same manner as in conventional technology in step S30.
[43] Next, motion compensated temporal filtering (MCTF) is performed using a pruned motion vector in step S40. An MCTF direction in accordance with the mode flag will be described with reference to FIG. 9.
[44] After completing the MCTF, 16 subbands resulting from the MCTF are subjected to spatial transform and quantization in step S50. Thereafter, a bitstream including data resulting from the spatial transform and quantization, motion vector data, and the mode flag is generated in step S60.
[45] FIG. 6 illustrates a reference for determining an MCTF direction according to a boundary condition, and FIGS. 7 and 8 illustrate boundary pixels used to determine a mode flag.
[46] FIGS. 6 illustrates cases where an internal image goes out of the frame. FIG. 6 illustrates forward MCTF and backward MCTF. In other words, image blocks B and N flow out of the frame when a T-l frame is converted into a T frame. In the worst case of forward MCTF shown in FIG. 6, the image blocks B and N in the T-l frame do not have their matches in the T frame. Thus, the image blocks B and N in the T-l frame are compared with image blocks C and M, respectively, in the T frame. In this situation, a difference between the image blocks B and C and a difference between the image blocks N and M are large, which increases the amount of information of the T-l frame to be replaced with a high-frequency frame. Conversely, in the best case of backward MCTF shown in FIG. 6, each image block in the T frame to be replaced with a high-frequency frame has its matches in the T-l frame, and therefore, the amount of information of the high-frequency frame, i.e., the T frame, may be decreased.
[47] In a comprehensive conception, forward MCTF is more efficient in a case where a new image comes into a frame through a boundary while backward MCTF is more efficient in a case where an image goes out of the frame through a boundary. In other cases, it is efficient to properly combine forward MCTF and backward MCTF. In other words, video coding efficiency and performance can be increased by properly selecting either forward or backward MCTF according to a boundary condition of an input GOF. In setting a mode flag, a basic principle is made that forward MCTF is used when a new image comes into a frame, backward MCTF is used when an image goes out of a frame, and forward MCTF and backward MCTF are properly combined in other cases.
[48] The mode flag can be determined using a motion vector for pixels at a boundary of a frame. As shown in FIG. 7, pixels at right and left boundaries of a frame may be used in a first embodiment. Alternatively, as shown in FIG. 8, pixels at right, left, upper, and lower boundaries of a frame may be used in a second embodiment. Video coding performance depends on a thickness of a boundary used to determine the mode flag. Where the boundary is too thin, information regarding output/input of a particular image may be missed. Conversely, where the boundary is too thick, a boundary condition may not be sharply identified. Accordingly, the thickness of the boundary needs to be appropriately determined. In embodiments of the present invention, the boundary has a thickness of 32 pixels.
[49] In determining the mode flag, motion vectors of pixels in each frames are obtained using HVSBM. A mode flag is determined based on the motion vectors of pixels in the frames. The mode flag may be different according to a temporal level, but it is preferable to determine the mode flag at temporal level 0.
[50] In the first embodiment shown in FIG. 7, a mode flag is determined using motion vectors at left and right boundaries of each frame because a new image usually comes into or goes out of a frame of a moving picture in an X direction. An average of motion vectors of pixels at the left boundary of each of all frames included in a single GOF is obtained. An X component of the average motion vector at the left boundary is denoted by 'L.' Similarly, an average of motion vectors of pixels at the right boundary of each of all frames included in a single GOF is obtained. An X component of the average motion vector at the right boundary is denoted by 'R.' Where an L value less than 0 indicates that an image comes into the frame through the left boundary, an R value less than 0 indicates that an image goes out of the frame through the right boundary. Similarly, the L value greater than 0 and the R value greater than 0 account for the opposite cases, respectively. Actually, the L or R value may not be 0 even if an image does not come in or go out of the frame. Accordingly, it is preferable that the L and R values not exceeding a predetermined threshold are determined as 0. When an image comes into the frame through the left or right boundary, the L value is less than 0 and the R value is equal to or greater than 0, or the L value is less than 0 and the R value is greater than 0. In this case, it is preferable to use forward MCTF. Conversely, when an image goes out of the frame through the left or right boundary, the L value is greater than 0 and the R value is equal to or less than 0, or the L value is greater than 0 and the R value is less than 0. In this case, it is preferable to use backward MCTF. When an image comes into the frame through the left boundary and an image goes out of the frame through the right boundary, it is preferable to appropriately combine forward MCTF and backward MCTF.
[51] As such, a mode flag F can be determined by the following algorithm:
[52] if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))then F=0 else if((L > 0 and R=0)or (L=0 and R < 0)or (L > 0 and R < 0))then F=l else F=2 ,
[53] Here, F=0 indicates a forward mode, F=l indicates a backward mode, and F=2 indicates a bi-directional mode.
[54] In a second embodiment shown in FIG. 8, left, right, upper, and lower boundaries are used. L and R values are obtained in the same manner as described in the first embodiment, and U and D values are obtained using averages of Y components of motion vectors. Like the first embodiment, where an image comes into a frame through at least one boundary and an image does not go out of the frame through any of the boundaries, it is preferable to use forward MCTF. Where an image goes out of the frame through at least one boundary and an image does not come into the frame through any of the boundaries, it is preferable to use backward MCTF. In other cases, it is preferable to appropriately combine forward MCTF and backward MCTF.
[55] As such, a mode flag F can be determined by the following algorithm:
[56] if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if (abs(U)< Threshold)then U=0 if (abs(D)< Threshold)then D=0 if(((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))and ((D < 0 and U==0)or (D==0 and U > 0)or (D < 0 and U > 0)or (D =0 and U==0)))then F=0 else if(((L > 0 and R==0)or (L==0 and R < 0)or (L > 0 and R < 0))and ((D > 0 and U==0) or (D=0 and U < 0)or (D > 0 and U < 0)or (D ==0 and U==0)))then F=l else F=2,
[57] Here, F=0 indicates a forward mode, F=l indicates a backward mode, and F=2 indicates a bi-directional mode. The first and second embodiments are exemplary, and the spirit of the present invention is not restricted thereto. In other words, a direction of MCTF is appropriately determined using information regarding image input/output at a boundary. Accordingly, the present invention will be considered as including a case where a mode flag is determined to be different among two or some frames more than two in a GOF in addition to the first and second embodiments where a mode flag is determined using average motion vectors obtained with respect to all of the frames in a GOF.
[58] FIG. 9 illustrates MCTF directions according to a mode flag representing a boundary condition.
[59] In a forward mode, MCTF directions are depicted as + + + + + + + +. In a backward mode, MCTF directions are depicted as . In a bi-directional mode, MCTF directions may be depicted in various ways, but FIG. 9 illustrates an example where MCTF directions are depicted as + - + - + - + - at temporal level 0. Here, '+' indicates a forward direction, and '-' indicates a backward direction.
[60] In each of the forward and backward modes, MCTF is performed in the same direction. However, in the bi-directional mode, video coding performance changes depending on a combination of forward and backward directions. In other words, in the bi-directional mode, a sequence of forward and backward directions may be determined in various ways. Representative examples of a sequence of MCTF directions in the forward, backward, and bi-directional modes are shown in Table 1.
[61] Table 1
Figure imgf000013_0001
[62] Various combinations of forward and backward directions may be made in the bi¬ directional mode, but four cases 'a', 'b', V, and 'd' are shown as examples. The cases V and 'd' are characterized in that a low- frequency frame (hereinafter, referred to as a reference frame) at a last level is positioned at a center (i.e., an 8th frame) among 1st through 16th frames. The reference frame is a most essential frame in video coding. The other frames are recovered based on the reference frame. As a temporal distance between a frame and the reference frame increases, recovery performance decreases. Accordingly, in the cases 'c' and 'd', a combination of forward MCTF and backward MCTF is made such that the reference frame is positioned at the center, i.e., the 8th frame, to minimize a temporal distance between the reference frame and each of the other frames.
[63] In the cases 'a' and 'b', an average temporal distance (ATD) is minimized. To calculate an ATD, temporal distances are calculated. A temporal distance is defined as a positional difference between two frames. Referring to FIG. 3, a temporal distance between a first frame and a second frame is defined as 1, and a temporal distance between the frame L2 and the frame L4 is defined as 2. An ATD is obtained by dividing the sum of temporal distances between frames subjected to an operation for motion estimation in pairs by the number of pairs of frames defined for the motion estimation. In the case 'a', ATD= 8xl +4xl+2x4+l x3 = l s3 1 5 In the case 'b', 8 x l + 4 x l + 2 x 3 ÷ l x 3 15 In the forward mode and the backward mode shown in Table 1,
Figure imgf000014_0001
In the case 'c', . __ 8x 1 + 4x 2 + 2 x 4 + 1x 2 ATD= _________ — = 1.73. 15 In the case 'd', 8 * 1 + 4x 1 + 2x 4 + 1 x 1 15 In actual simulations, as an ATD was decreased, a PSNR value was increased so that performance of video coding was increased.
[64] FIG. 10 is a functional block diagram of a system for adaptive IWVC according to an embodiment of the present invention.
[65] The system for adaptive IWVC includes a motion estimation/mode determination block 10 which obtains a motion vector and determines a mode using the motion vector, a motion compensation temporal filtering block 40 which removes temporal redundancy using the motion vector and the determined mode, a spatial transform block 50 which removes spatial redundancy, a motion vector encoding block 20 which encodes the motion vector using a predetermined algorithm, a quantization block 60 which quantizes wavelet coefficients of respective components generated by the spatial transform block 50, and a buffer 30 which temporarily stores an encoded bitstream received from the quantization block 60. [66] The motion estimation/mode determination block 10 obtains a motion vector used by the motion compensation temporal filtering block 40 using a hierarchical method such as HVSBM. In addition, the motion estimation/mode determination block 10 determines a mode flag for determining temporal filtering directions.
[67] The motion compensation temporal filtering block 40 decomposes frames into low- and high-frequency frames in a temporal direction using the motion vector obtained by the motion estimation/mode determination block 10. A direction of the decomposition is determined according to the mode flag. Frames are decomposed in GOF units. Through such decomposition, temporal redundancy is removed.
[68] The spatial transform block 50 wavelet-decomposes frames that have been decomposed in the temporal direction by the motion compensation temporal filtering block 40 into spatial low- and high-frequency components, thereby removing spatial redundancy.
[69] The motion vector encoding block 20 encodes the motion vector and the mode flag hierarchically obtained by the motion estimation/mode determination block 10 and then transmits the encoded motion vector and the encoded mode flag to the buffer 30.
[70] The quantization block 60 quantizes and encodes wavelet coefficients of components generated by the spatial transform block 50. [71] The buffer 30 stores a bitstream including encoded data, the encoded motion vector, and the encoded mode flag before transmission and is controlled by a rate control algorithm.
[72] In experiments, performance was increased by about 0.8 dB. In the experiments, Mobile , Tempete, Canoa, and Bus were used, and results of the experiments are shown in Tables 2 through 5.
[73] Table 2: Mobile, CIF, Frames: 0-299
[74] Table 3: Tempete CIF, Frames: 0-259
Figure imgf000016_0001
[75] Table 4: Canoa, CIF, Frames: 0-208
Figure imgf000016_0002
[76] Table 5: Bus, CIF, Frames: 0-150
Figure imgf000016_0003
Industrial Applicability
[77] According to the present invention, IWVC can be adaptively performed in accordance with a boundary condition. In other words, as compared to conventional methods, a PSNR is increased in the present invention.
[78] It will be understood by those of ordinary skill in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present invention as defined by the following claims. Therefore, it is to be appreciated that the above described embodiment is for purposes of illustration only and not to be construed as a limitation of the invention. The scope of the invention is given by the appended claims, rather than the preceding description, and all variations and equivalents which fall within the range of the claims are intended to be embraced therein.

Claims

Claims
[1] An interframe wavelet video coding method comprising: (a) receiving a group-of-frames including a plurality of frames and determining a mode flag according to a predetermined procedure using motion vectors of boundary pixels; (b) temporally decomposing the frames included in the group-of-frames in predetermined directions in accordance with the determined mode flag; and (c) performing spatial transform and quantization on the frames obtained by performing step (b), thereby generating a bitstream.
[2] The interframe wavelet video coding method of claim 1, wherein in step (a), the group-of-frames comprises 16 frames.
[3] The interframe wavelet video coding method of claim 1, wherein step (a) comprises determining the mode flag according to the predetermined procedure using motion vectors obtained at a boundary having a predetermined thickness among motion vectors of pixels obtained through motion estimation using hierarchical variable size block matching (HVSBM).
[4] The interframe wavelet video coding method of claim 3, wherein the motion vectors used to determine the mode flag are motion vectors of pixels at left and right boundaries.
[5] The interframe wavelet video coding method of claim 4, wherein the mode flag F is determined using the following algorithm: if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))then F=0 else if((L > 0 and R==0)or (L==0 and R < 0)or (L > 0 and R < 0))then F=l else F=2 where L denotes an average of X components of motion vectors of pixels at the left boundary having the predetermined thickness, and R denotes an average of X components of motion vectors of pixels at the right boundary having the predetermined thickness, wherein step (b) comprises temporally decomposing the frames included in the group-of-frames in a forward direction when F=0, temporally decomposing the frames included in the group-of-frames in a backward direction when F=l, and temporally decomposing the frames included in the group-of-frames in forward and backward directions combined in a predetermined sequence when F=2.
[6] The interframe wavelet video coding method of claim 5, wherein when F=2 in step (b), the frames are decomposed such that an average temporal distance between frames is minimized.
[7] The interframe wavelet video coding method of claim 3, wherein the motion vectors used to determine the mode flag are motion vectors of pixels at left, right, upper, and lower boundaries.
[8] The interframe wavelet video coding method of claim 7, wherein the mode flag F is determined using the following algorithm: if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if (abs(U)< Threshold)then U=0 if (abs(D)< Threshold)then D=0 if(((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))and ((D < 0 and U==0)or (D==0 and U > 0)or (D < 0 and U > 0)or (D ==0 and U==0)))then F=0 else if(((L > 0 and R==0)or (L==0 and R < 0)or (L > 0 and R < 0))and ((D > 0 and U==0) or (D==0 and U < 0)or (D > 0 and U < 0)or (D ==0 and U==0)))then F=l else F=2 where L denotes an average of X components of motion vectors of pixels at the left boundary having the predetermined thickness, R denotes an average of X components of motion vectors of pixels at the right boundary having the predetermined thickness, U denotes an average of Y components of motion vectors of pixels at the upper boundary having the predetermined thickness, and D denotes an average of Y components of motion vectors of pixels at the lower boundary having the predetermined thickness, wherein step (b) comprises temporally decomposing the frames included in the group-of-frames in a forward direction when F=0, temporally decomposing the frames included in the group-of-frames in a backward direction when F=l, and temporally decomposing the frames included in the group-of-frames in forward and backward directions combined in a predetermined sequence when F=2.
[9] The interframe wavelet video coding method of claim 8, wherein when F=2 in step (b), the frames are decomposed such that an average temporal distance between frames is minimized.
[10] A recording medium comprising commands which can be executed in a computer, the commands executing: (a) receiving a group-of-frames including a plurality of frames and determining a mode flag according to a predetermined procedure using motion vectors of boundary pixels; (b) temporally decomposing the frames included in the group-of-frames in predetermined directions in accordance with the determined mode flag; and (c) performing spatial transform and quantization on the frames obtained by performing step (b), thereby generating a bitstream.
[11] The recording medium of claim 10, wherein in step (a), the group-of-frames comprises 16 frames.
[12] The recording medium of claim 9, wherein step (a) comprises determining the mode flag according to the predetermined procedure using motion vectors obtained at a boundary having a predetermined thickness among motion vectors of pixels obtained through motion estimation using hierarchical variable size block matching (HVSBM).
[13] The recording medium of claim 12, wherein the motion vectors used to determine the mode flag are motion vectors of pixels at left and right boundaries.
[14] The recording medium of claim 13, wherein the mode flag F is determined using the following algorithm: if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))then F=0 else if((L > 0 and R==0)or (L=0 and R < 0)or (L > 0 and R < 0))then F=l else F=2 where L denotes an average of X components of motion vectors of pixels at the left boundary having the predetermined thickness, and R denotes an average of X components of motion vectors of pixels at the right boundary having the predetermined thickness, wherein step (b) comprises temporally decomposing the frames included in the group-of-frames in a forward direction when F=0, temporally decomposing the frames included in the group-of-frames in a backward direction when F=l, and temporally decomposing the frames included in the group-of-frames in forward and backward directions combined in a predetermined sequence when F=2.
[15] The recording medium of claim 14, wherein when F=2 in step (b), the frames are decomposed such that an average temporal distance between frames is minimized.
[16] The recording medium of claim 12, wherein the motion vectors used to determine the mode flag are motion vectors of pixels at left, right, upper, and lower boundaries.
[17] The recording medium of claim 16, wherein the mode flag F is determined using the following algorithm: if (abs(L)< Threshold)then L=0 if (abs(R)< Threshold)then R=0 if (abs(U)< Threshold)then U=0 if (abs(D)< Threshold)then D=0 if(((L < 0 and R==0)or (L==0 and R > 0)or (L < 0 and R > 0))and ((D < 0 and U==0)or (D==0 and U > 0)or (D < 0 and U > 0)or (D -=0 and l =0)))then F=0 else if(((L > 0 and R==0)or (L==0 and R < 0)or (L > 0 and R < 0))and ((D > 0 and U==0) or (D==0 and U < 0)or (D > 0 and U < 0)or (D ==0 and U=0)))then F-l else F=2 where L denotes an average of X components of motion vectors of pixels at the left boundary having the predetermined thickness, R denotes an average of X components of motion vectors of pixels at the right boundary having the predetermined thickness, U denotes an average of Y components of motion vectors of pixels at the upper boundary having the predetermined thickness, and D denotes an average of Y components of motion vectors of pixels at the lower boundary having the predetermined thickness, wherein step (b) comprises temporally decomposing the frames included in the group-of-frames in a forward direction when F=0, temporally decomposing the frames included in the group-of-frames in a backward direction when F=l, and temporally decomposing the frames included in the group-of-frames in forward and backward directions combined in a predetermined sequence when F=2.
[18] The recording medium of claim 17, wherein when F=2 in step (b), the frames are decomposed such that an average temporal distance between frames is minimized.
[19] An interframe wavelet video coding system which receives a group-of-frames including a plurality of frames and generates a bitstream, the interframe wavelet video coding system comprising: a motion estimation/mode determination block which receives the group- of-frames, obtains motion vectors of pixels in each of the frames using a predetermined procedure, and determines a mode flag using motion vectors of boundary pixels among the obtained motion vectors; and a motion compensation temporal filtering block which decomposes the frames into low- and high-frequency frames in a predetermined temporal direction in accordance with the mode flag determined by the motion estimation/mode determination block using the motion vectors.
[20] The interframe wavelet video coding system of claim 19, further comprising a spatial transform block which wavelet-decomposes the low- and high-frequency frames generated by the motion compensation temporal filtering block into spatial low- and high-frequency components.
PCT/KR2004/002050 2003-08-26 2004-08-16 Adaptive interframe wavelet video coding method, computer readable recording medium and system therefor Ceased WO2005020587A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2006524561A JP2007503750A (en) 2003-08-26 2004-08-16 Adaptive interframe wavelet video coding method, computer-readable recording medium and apparatus for the method

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US49756703P 2003-08-26 2003-08-26
US60/497,567 2003-08-26
KR10-2003-0065863 2003-09-23
KR1020030065863A KR100577364B1 (en) 2003-09-23 2003-09-23 Adaptive interframe video coding method, computer readable recording medium for the method, and apparatus

Publications (1)

Publication Number Publication Date
WO2005020587A1 true WO2005020587A1 (en) 2005-03-03

Family

ID=34220840

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2004/002050 Ceased WO2005020587A1 (en) 2003-08-26 2004-08-16 Adaptive interframe wavelet video coding method, computer readable recording medium and system therefor

Country Status (3)

Country Link
US (1) US20050047508A1 (en)
JP (1) JP2007503750A (en)
WO (1) WO2005020587A1 (en)

Families Citing this family (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100668345B1 (en) * 2004-10-05 2007-01-12 삼성전자주식회사 Motion compensated layer generator and method
US7634148B2 (en) * 2005-01-07 2009-12-15 Ntt Docomo, Inc. Image signal transforming and inverse-transforming method and computer program product with pre-encoding filtering features
CN100512439C (en) * 2005-10-27 2009-07-08 中国科学院研究生院 Small wave region motion estimation scheme possessing frame like small wave structure
US8358693B2 (en) * 2006-07-14 2013-01-22 Microsoft Corporation Encoding visual data with computation scheduling and allocation
US8311102B2 (en) * 2006-07-26 2012-11-13 Microsoft Corporation Bitstream switching in multiple bit-rate video streaming environments
US8340193B2 (en) * 2006-08-04 2012-12-25 Microsoft Corporation Wyner-Ziv and wavelet video coding
US7388521B2 (en) * 2006-10-02 2008-06-17 Microsoft Corporation Request bits estimation for a Wyner-Ziv codec
US8340192B2 (en) * 2007-05-25 2012-12-25 Microsoft Corporation Wyner-Ziv coding with multiple side information
WO2010005691A1 (en) * 2008-06-16 2010-01-14 Dolby Laboratories Licensing Corporation Rate control model adaptation based on slice dependencies for video coding
US9635385B2 (en) 2011-04-14 2017-04-25 Texas Instruments Incorporated Methods and systems for estimating motion in multimedia pictures
US9747255B2 (en) * 2011-05-13 2017-08-29 Texas Instruments Incorporated Inverse transformation using pruning for video coding
CN110543095A (en) * 2019-09-17 2019-12-06 南京工业大学 A Design Method of Control System of CNC Gear Chamfering Machine Based on Quantum Framework
EP4274207A4 (en) 2021-04-13 2024-07-10 Samsung Electronics Co., Ltd. ELECTRONIC DEVICE AND CONTROL METHOD THEREFOR

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5646997A (en) * 1994-12-14 1997-07-08 Barton; James M. Method and apparatus for embedding authentication information within digital data
GB2301970B (en) * 1995-06-06 2000-03-01 Sony Uk Ltd Motion compensated video processing
WO1997017797A2 (en) * 1995-10-25 1997-05-15 Sarnoff Corporation Apparatus and method for quadtree based variable block size motion estimation
US5956026A (en) * 1997-12-19 1999-09-21 Sharp Laboratories Of America, Inc. Method for hierarchical summarization and browsing of digital video
US6480615B1 (en) * 1999-06-15 2002-11-12 University Of Washington Motion estimation within a sequence of data frames using optical flow with adaptive gradients
JP2004514351A (en) * 2000-11-17 2004-05-13 コーニンクレッカ フィリップス エレクトロニクス エヌ ヴィ Video coding method using block matching processing

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
BOTTEREAU V. ET AL.: "A fully scalable 3D subband video codec", INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, vol. 2, 7 October 2001 (2001-10-07) - 10 October 2001 (2001-10-10), pages 1017 - 1020 *
PESQUET-POPESCU B. ET AL.: "Three-dimensional lifting schemes for motion compensated video compression", IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING, vol. 3, 7 May 2001 (2001-05-07) - 11 May 2001 (2001-05-11), pages 1793 - 1796 *
SHIH-TA HSIANG ET AL.: "Embedded video coding using invertible motion compensated 3-D subband/wavelet filter bank", SIGNAL PROCESSING: IMAGE COMMUNICATION, vol. 16, no. 8, May 2001 (2001-05-01), pages 705 - 724 *

Also Published As

Publication number Publication date
US20050047508A1 (en) 2005-03-03
JP2007503750A (en) 2007-02-22

Similar Documents

Publication Publication Date Title
CN1906945B (en) Method and apparatus for scalable video encoding and decoding
US20050047509A1 (en) Scalable video coding and decoding methods, and scalable video encoder and decoder
US20060013313A1 (en) Scalable video coding method and apparatus using base-layer
US20060013309A1 (en) Video encoding and decoding methods and video encoder and decoder
US20050226334A1 (en) Method and apparatus for implementing motion scalability
US20050169379A1 (en) Apparatus and method for scalable video coding providing scalability in encoder part
WO2005074277A1 (en) Method and device for transmitting scalable video bitstream
US20050157793A1 (en) Video coding/decoding method and apparatus
KR20040106417A (en) Scalable wavelet based coding using motion compensated temporal filtering based on multiple reference frames
AU2004302413B2 (en) Scalable video coding method and apparatus using pre-decoder
WO2005020587A1 (en) Adaptive interframe wavelet video coding method, computer readable recording medium and system therefor
US20060159173A1 (en) Video coding in an overcomplete wavelet domain
US20050163224A1 (en) Device and method for playing back scalable video streams
EP0892557A1 (en) Image compression
US20050158026A1 (en) Method and apparatus for reproducing scalable video streams
US20050163217A1 (en) Method and apparatus for coding and decoding video bitstream
US7292635B2 (en) Interframe wavelet video coding method
CA2547628C (en) Method and apparatus for scalable video encoding and decoding
US20060013312A1 (en) Method and apparatus for scalable video coding and decoding
WO2006006764A1 (en) Video decoding method using smoothing filter and video decoder therefor
KR100664930B1 (en) Video coding method and apparatus supporting temporal scalability
KR100577364B1 (en) Adaptive interframe video coding method, computer readable recording medium for the method, and apparatus
WO2005009046A1 (en) Interframe wavelet video coding method
WO2006006793A1 (en) Video encoding and decoding methods and video encoder and decoder
WO2006043750A1 (en) Video coding method and apparatus

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A1

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KP KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): BW GH GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 2006524561

Country of ref document: JP

122 Ep: pct application non-entry in european phase