EP4627794A1 - Delayed frame filtering in in-loop filtering - Google Patents
Delayed frame filtering in in-loop filteringInfo
- Publication number
- EP4627794A1 EP4627794A1 EP23837514.1A EP23837514A EP4627794A1 EP 4627794 A1 EP4627794 A1 EP 4627794A1 EP 23837514 A EP23837514 A EP 23837514A EP 4627794 A1 EP4627794 A1 EP 4627794A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- frame
- restored
- filtering
- task
- degraded
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/172—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/59—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial sub-sampling or interpolation, e.g. alteration of picture size or resolution
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
- H04N19/82—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation involving filtering within a prediction loop
Definitions
- Digital video streams may represent video using a sequence of frames or still images.
- Digital video can be used for various applications including, for example, video conferencing, high-definition video entertainment, video advertisements, or sharing of usergenerated videos.
- a digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data.
- Various approaches have been proposed to reduce the amount of data in video streams, including compression and other coding techniques. These techniques may include both lossy and lossless coding techniques.
- This disclosure relates generally to encoding and decoding video data and more particularly relates to motion vector coding candidate signaling.
- a system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions.
- One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
- One general aspect includes a method.
- the method includes obtaining a degraded frame from a reconstructed frame of a current frame; adding the degraded frame to a reference frame store: coding a next frame based on the degraded frame; obtaining, using a restoration task, a restored frame from the degraded frame; and coding a frame subsequent to the next frame based on the restored frame.
- Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
- the method may include coding in a header of the current frame a syntax element indicating that the restoration task is to be performed with respect to the degraded frame to obtain the restored frame.
- the method may include initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame; and in response to a command to output the next frame: outputting the next restored frame; and discarding the next restored frame.
- the method may include initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame; and in response to a command to output the next frame before the next restored frame is available: waiting until the next restored frame to become available; and outputting the next restored frame.
- Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
- One general aspect includes a method.
- the method may include decoding a first frame using a reference frame; applying a filtering task to the reference frame to obtain a restored reference frame; and decoding a second frame using the restored reference frame.
- Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
- FIG. 1 is a schematic of a video encoding and decoding system.
- FIG. 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
- FIG. 7 is a process flow of coding using delayed frame filtering in in-loop filtering.
- FIG. 8 is a flow chart of a technique for delayed frame filtering in in-loop filtering.
- FIG. 9 is a flow chart of a technique for delayed frame filtering in in-loop filtering.
- FIG. 10 is a flowchart of a technique for delayed frame filtering in in-loop filtering.
- compression schemes related to coding video streams may include breaking images into blocks and generating a digital video output bitstream (i.e., an encoded bitstream) using one or more techniques to limit the information included in the output bitstream.
- a received bitstream can be decoded to re-create the blocks and the source images from the limited information.
- Encoding a video stream, or a portion thereof, such as a frame or a block can include using temporal similarities in the video stream to improve coding efficiency. For example, a current block of a video stream may be encoded based on identifying a difference (residual) between the previously coded pixel values, or between a combination of previously coded pixel values, and those in the current block.
- Inter prediction can attempt to predict the pixel values of a block using a possibly displaced block or blocks from a temporally nearby frame (i.e., reference frame) or frames.
- a temporally nearby frame is a frame that appears earlier or later in time in the video stream than the frame of the block being encoded.
- a prediction block resulting from inter prediction is referred to herein as inter predictor.
- a motion vector used to generate a prediction block refers to a frame other than a current frame, i.e., a reference frame.
- Reference frames can be located before or after the current frame in the sequence of the video stream.
- Some codecs may use up to eight reference frames, which can be stored in frame buffers of a reference frame store.
- the motion vector can refer to (i.e., use) one of the reference frames stored in the reference frame store.
- Reference frames stored in the reference frame store can also be used to generate motion fields.
- Residuals i.e.. differences
- Decoding i.e., reconstructing
- the encoder and decoder may perform operations that improve the qualify of the reconstructed data, such as described below with respect to a loop filtering stage 416 of FIG. 4 and a loop filtering stage 512 of FIG. 5 (e.g., within the reconstruction loop at the encoder or prior the outputting at the decoder).
- Loop restoration may be performed to process a reconstructed video frame for use as a reference frame.
- a filtering stage may include more than one restoration tasks (i.e., operations or tools) that are to be applied to a reconstructed frame of a current frame.
- the restorations tasks are expected to complete within the time allocated for decoding the current frame.
- a codec may be designed for a certain throughput.
- the throughput of a codec may be specified in terms of frames per second.
- a decoder may be designed to decode fames at a rate of 60 frames per second. As such, the decoder allocates, on average, l/60 th of a second to decode a frame.
- restoration tasks may be too complex to complete without significant added complexity to a codec.
- more complex restoration tasks may result in additional codec complexity'.
- a hardware-implemented codec may require additional circuitry (i.e., larger silicon footprint), and a software-implemented coded may require additional processing cores and threads, to complete complex restoration tasks within the codec throughput requirements.
- additional circuitry i.e., larger silicon footprint
- a software-implemented coded may require additional processing cores and threads, to complete complex restoration tasks within the codec throughput requirements.
- CNN convolutional neural network
- ML machine-learning
- CNNs may not be suitable for widespread deployment in video coding due to their complexity requirements (in terms of, for example, model size and requisite inferencing time) at the throughput desired for video coding.
- CNNs (and more generally, ML models) may be too complex to implement in a cost-effective manner with a small enough silicon footprint to meet stringent throughput requirements.
- Other examples of complex restoration tasks other than CNNs are possible and the disclosure herein is not limited to complex restoration tasks that are CNNs.
- an inter-predicted frame can use any of the reference frames stored in a reference frame store through prediction, motion field generation, or for some other purpose related to motion determination.
- a new frame e.g., output by a filtering stage
- it updates the reference frame store by replacing one of the existing slots so that it can be used to decode subsequent frames.
- any filtering or restoration operation conducted on the frame before producing the final output is characterized as in-loop.
- some in-loop filtering tasks e.g., operations
- Implementations according to this disclosure provide a solution to the problem of complex in-loop filtering tasks by relaxing the in-loop-ness of the decoding mechanism by allowing the complex filtering task to have more time to work on a frame just decoded (up to the usage of the complex filtering task) while retaining the low complexity of a codec. Rather.
- the throughput required by complex restoration tasks (such as for ML inference in a case that the complex restoration task is implemented using an ML model) is reduced at a modest loss in efficiency and increase in delay, which can be acceptable in applications such as video-on-demand applications. Reduced throughput translates directly to reduced gatecount and silicon area for implementation of the hardware.
- a reconstructed frame can be processed by one or more filtering tasks to obtain a degraded frame, which is further processed by a complex filtering task to generate a restored frame.
- relaxing the in-loop-ness of the coding mechanism can mean that the coding of some subsequent frames need not wait until the restored frame is available. Rather, coding of a certain number subsequent frames proceeds based on the degraded frame and, when the restored frame becomes available, then coding of later subsequence frames proceeds based on the restored frame. As the use of restored frame is delayed until the restore frame is available and, until then, coding proceeds based on the degraded frame.
- FIG. 1 is a schematic of a video encoding and decoding system 100.
- a transmitting station 102 can be, for example, a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.
- a network 104 can connect the transmitting station 102 and a receiving station 106 for encoding and decoding of the video stream.
- the video stream can be encoded in the transmitting station 102 and the encoded video stream can be decoded in the receiving station 106.
- the network 104 can be, for example, the Internet.
- the network 104 can also be a local area network (LAN), wide area network (WAN), virtual private network (VPN), cellular telephone netw ork or any other means of transferring the video stream from the transmitting station 102 to. in this example, the receiving station 106.
- the receiving station 106 in one example, can be a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
- an implementation can omit the network 104.
- a video stream can be encoded and then stored for transmission at a later time to the receiving station 106 or any other device having memory.
- the receiving station 106 receives (e.g., via the network 104, a computer bus, and/or some communication pathway) the encoded video stream and stores the video stream for later decoding.
- a real-time transport protocol RTP
- a transport protocol other than RTP may be used, e.g., a video streaming protocol based on the Hypertext Transfer Protocol (HTTP).
- the transmitting station 102 and/or the receiving station 106 may include the ability to both encode and decode a video stream as described below.
- the receiving station 106 could be a video conference participant who receives an encoded video bitstream from a video conference sen- er (e.g., the transmitting station 102) to decode and view and further encodes and transmits its own video bitstream to the video conference sen- er for decoding and viewing by other participants.
- FIG. 2 is a block diagram of an example of a computing device 200 (e.g., an apparatus) that can implement a transmitting station or a receiving station.
- the computing device 200 can implement one or both of the transmitting station 102 and the receiving station 106 of FIG. 1.
- the computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of one computing device, for example, a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.
- a CPU 202 in the computing device 200 can be a conventional central processing unit.
- the CPU 202 can be any other type of device, or multiple devices, capable of manipulating or processing information now existing or hereafter developed.
- the disclosed implementations can be practiced with one processor as shown, e.g., the CPU 202, advantages in speed and efficiency can be achieved using more than one processor.
- a memory 204 in computing device 200 can be a read only memory (ROM) device or a random-access memory (RAM) device in an implementation. Any other suitable type of storage device can be used as the memory 204.
- the memory 204 can include code and data 206 that is accessed by the CPU 202 using a bus 212.
- the memory 204 can further include an operating system 208 and application programs 210, the application programs 210 including at least one program that permits the CPU 202 to perform the methods described here.
- the application programs 210 can include applications 1 through N. which further include a video coding application that performs the techniques described here, such as those for delayed frame filtering in in-loop filtering.
- Computing device 200 can also include a secondary' storage 214. which can, for example, be a memory card used with a mobile computing device. Because the video communication sessions may contain a significant amount of information, they can be stored in whole or in part in the secondary storage 214 and loaded into the memoiy 204 as needed for processing.
- the computing device 200 can also include one or more output devices, such as a display 218.
- the display 218 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs.
- the display 218 can be coupled to the CPU 202 via the bus 212.
- Other output devices that permit a user to program or otherwise use the computing device 200 can be provided in addition to or as an alternative to the display 218.
- the output device is or includes a display
- the display can be implemented in various ways, including by a liquid cry stal display (LCD), a cathode-ray tube (CRT) display or light emitting diode (LED) display, such as an organic LED (OLED) display.
- LCD liquid cry stal display
- CRT cathode-ray tube
- LED light emitting diode
- OLED organic LED
- the computing device 200 can also include or be in communication with a soundsensing device 222, for example a microphone, or any other sound-sensing device nowexisting or hereafter developed that can sense sounds near the computing device 200.
- the sound-sensing device 222 can be positioned such that it is directed toward the user operating the computing device 200 and can be configured to receive sounds, for example, speech or other utterances, made by the user while the user operates the computing device 200.
- FIG. 2 depicts the CPU 202 and the memory 204 of the computing device 200 as being integrated into one unit, other configurations can be utilized.
- the operations of the CPU 202 can be distributed across multiple machines (wherein individual machines can have one or more of processors) that can be coupled directly or across a local area or other network.
- the memory 204 can be distributed across multiple machines such as a network-based memory or memory' in multiple machines performing the operations of the computing device 200.
- the bus 212 of the computing device 200 can be composed of multiple buses.
- the secondary storage 214 can be directly coupled to the other components of the computing device 200 or can be accessed via a netw ork and can comprise an integrated unit such as a memory' card or multiple units such as multiple memory cards.
- the computing device 200 can thus be implemented in a wide variety of configurations.
- FIG. 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded.
- the video stream 300 includes a video sequence 302.
- the video sequence 302 includes a number of adjacent frames 304. While three frames are depicted as the adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304.
- the adjacent frames 304 can then be further subdivided into individual frames, e.g., a frame 306.
- the frame 306 can be divided into a series of planes or segments 308.
- the segments 308 can be subsets of frames that permit parallel processing, for example.
- the segments 308 can also be subsets of frames that can separate the video data into separate colors.
- a frame 306 of color video data can include a luminance plane and two chrominance planes.
- the segments 308 may be sampled at different resolutions.
- the frame 306 may be further subdivided into blocks 310, which can contain data corresponding to. for example, 16x16 pixels in the frame 306.
- the blocks 310 can also be arranged to include data from one or more segments 308 of pixel data.
- the blocks 310 can also be of any other suitable size such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger. Unless otherwise noted, the terms block and macro-block are used interchangeably herein.
- FIG. 4 is a block diagram of an encoder 400.
- the encoder 400 can be implemented, as described above, in the transmitting station 102 such as by providing a computer software program stored in memory, for example, the memory 204.
- the computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the transmitting station 102 to encode video data in the manner described in FIG. 4.
- the encoder 400 can also be implemented as specialized hardware included in, for example, the transmitting station 102. In one particularly desirable implementation, the encoder 400 is a hardware encoder.
- the encoder 400 has the following stages to perform the various functions in a forward path (shown by the solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra/inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408.
- the encoder 400 may also include a reconstruction path (shown by the dotted connection lines) to reconstruct a frame for encoding of future blocks.
- the encoder 400 has the following stages to perform the various functions in the reconstruction path: a dequantization stage 410, an inverse transform stage 412. a reconstruction stage 414. and a loop filtering stage 416.
- Other structural variations of the encoder 400 can be used to encode the video stream 300.
- respective frames 304 can be processed in units of blocks.
- respective blocks can be encoded using intra-frame prediction (also called intra-prediction) or inter-frame prediction (also called inter-prediction).
- intra-frame prediction also called intra-prediction
- inter-frame prediction also called inter-prediction
- a prediction block can be formed.
- intra-prediction a prediction block may be formed from samples in the current frame that have been previously encoded and reconstructed.
- interprediction a prediction block may be formed from samples in one or more previously constructed reference frames.
- the prediction block can be subtracted from the current block at the intra/inter prediction stage 402 to produce a residual block (also called a residual).
- the transform stage 404 transforms the residual into transform coefficients in, for example, the frequency domain using block-based transforms.
- the quantization stage 406 converts the transform coefficients into discrete quantum values, which are referred to as quantized transform coefficients, using a quantizer value or a quantization level. For example, the transform coefficients may be divided by the quantizer value and truncated.
- the quantized transform coefficients are then entropy encoded by the entropy encoding stage 408.
- the entropy-encoded coefficients, together with other information used to decode the block, which may include for example the type of prediction used, transform type, motion vectors and quantizer value, are then output to the compressed bitstream 420.
- the compressed bitstream 420 can be formatted using various techniques, such as variable length coding (VLC) or arithmetic coding.
- VLC variable length coding
- the compressed bitstream 420 can also be referred to as an encoded video stream or encoded video bitstream, and the terms will be used interchangeably herein.
- the reconstruction path in FIG. 4 can be used to ensure that the encoder 400 and a decoder 500 (described below-) use the same reference frames to decode the compressed bitstream 420.
- the reconstruction path performs functions that are similar to functions that take place dunng the decoding process that are discussed in more detail below, including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (also called a derivative residual).
- the prediction block that was predicted at the intra/inter prediction stage 402 can be added to the derivative residual to create a reconstructed block.
- the loop filtering stage 416 can be applied to the reconstructed block to reduce distortion such as blocking artifacts.
- encoder 400 can be used to encode the compressed bitstream 420.
- a non-transform-based encoder can quantize the residual signal directly without the transform stage 404 for certain blocks or frames.
- an encoder can have the quantization stage 406 and the dequantization stage 410 combined in a common stage.
- FIG. 5 is a block diagram of a decoder 500.
- the decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204.
- the computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described in FIG. 5.
- the decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.
- the decoder 500 similar to the reconstruction path of the encoder 400 discussed above, includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra/inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512 and a deblocking filtering stage 514.
- stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420 includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra/inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512 and a deblocking filtering stage 514.
- Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.
- the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients.
- the dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by the quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derivative residual that can be identical to that created by the inverse transform stage 412 in the encoder 400.
- the decoder 500 can use the intra/inter prediction stage 508 to create the same prediction block as was created in the encoder 400, e.g., at the intra/inter prediction stage 402.
- the prediction block can be added to the derivative residual to create a reconstructed block.
- the loop filtering stage 512 can be applied to the reconstructed block to reduce blocking artifacts.
- Other filtering can be applied to the reconstructed block.
- the deblocking filtering stage 514 is applied to the reconstructed block to reduce blocking distortion, and the result is output as the output video stream 516.
- the output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein.
- Other variations of the decoder 500 can be used to decode the compressed bitstream 420.
- the decoder 500 can produce the output video stream 516 without the deblocking filtering stage 514.
- FIG. 6 is a diagram of an example of a reference frame store 600.
- a motion vector used to generate a prediction block refers to (i.e., uses) a reference frame.
- Reference frames can be stored in a reference frame buffers of the reference frame store 600.
- a current frame, or a current block or region of a frame can be coded (encoded or decoded) using a reference frame such as a “last frame,’ 7 which is the adjacent frame immediately before the current frame in the video sequence.
- a reference frame such as a “last frame,’ 7 which is the adjacent frame immediately before the current frame in the video sequence.
- motion information from video frames in the past or future can be included in the candidate reference motion vectors. Coding video frames can occur, for example, using so-called “alternate reference frames” that are not temporally neighboring to the frames coded immediately before or after them.
- An alternate reference frame can be a synthesized frame that does not occur in the input video stream or is a duplicate frame to one in the input video stream that is used for prediction.
- the reference frame store 600 can include a last frame LAST FRAME 602, a golden frame GOLDEN FRAME 604, and an alternative reference frame ALTREF FRAME 606.
- a reference frame buffer can include more, fewer, or other reference frames.
- the reference frame store 600 is shown as including eight reference frames. However, a reference frame buffer can include more or fewer than eight reference frames. Any particular semantics (e.g., labels) associated with reference frames stored in the reference frame store 600 are not necessary to the understanding or use of this disclosure.
- the reference frames stored in the reference frame store 600 can be used to identify motion vectors for predicting blocks of frames to be encoded or decoded. Different reference frames may be used depending on the type of prediction used to predict a current block of a current frame. For example, when compound prediction is used, multiple frames, such as one for forward prediction (e.g.. LAST FRAME 602 or GOLDEN FRAME 604) and one for backward prediction (e.g., ALTREF FRAME 606) can be used for predicting the current block.
- forward prediction e.g.. LAST FRAME 602 or GOLDEN FRAME 604
- ALTREF FRAME 606 backward prediction
- reference frame store 600 There may be a finite number of reference frames that can be stored within the reference frame store 600. As shown in FIG. 6, the reference frame store 600 can store up to eight reference frames. Although three of the eight spaces in the reference frame store 600 are used by the LAST FRAME 602, the GOLDEN FRAME 604, and the ALTREF FRAME 606, five spaces remain available to store other reference frames.
- one or more available spaces in the reference frame store 600 may be used to store a second last frame LAST2_FRAME and/or a third last frame
- LAST3 FRAME as additional forward reference frames, in addition to the LAST FRAME 602.
- a backward frame BWDREF FRAME 608 can be stored as an additional backward prediction reference frame, in addition to ALTREF FRAME 606.
- the BWDREF FRAME can be closer in relative distance to the current frame than the ALTREF_FRAME 606, for example.
- the pair of ⁇ LAST FRAME, BWDREF FRAME ⁇ can be used to generate a compound predictor for coding the current block.
- LAST FRAME is a “nearest” forward reference frame for forward prediction
- BWDREF_FRAME is a “nearest” backward reference frame for backward prediction.
- a current block is predicted based on a prediction mode.
- the prediction mode may be selected from one of multiple intra-prediction modes.
- the prediction mode may be selected from one of multiple inter-prediction modes using one or more reference frames of the reference frame store 600 including, for example, the LAST FRAME 602, the GOLDEN FRAME 604, the ALTREF FRAME 606. or any other reference frame.
- the prediction mode of the current block can be transmitted from an encoder, such as the encoder 400 of FIG. 4, to a decoder, such as the decoder 500 of FIG. 5, in an encoded bitstream, such as the compressed bitstream 420 of FIGS. 4-5.
- a bitstream syntax can support three categories of inter prediction modes in an example. These inter prediction modes can include a mode (referred to herein as the ZERO_MV mode) in which a block from the same location within a reference frame as the current block is used as the prediction block, a mode (referred to herein as the NEW_MV mode) in which a motion vector is transmitted to indicate the location of a block within a reference frame to be used as the prediction block relative to the current block, or a mode (referred to herein as the REF_MV mode and comprising aNEAR_MV or NEAREST MV mode) in which no motion vector is transmitted and the current block uses the last or second- to-last non-zero motion vector used by neighboring, previously coded blocks to generate the prediction block.
- ZERO_MV mode a mode in which a block from the same location within a reference frame as the current block is used as the prediction block
- the NEW_MV mode a mode in which a motion vector is transmitted to indicate the location of a block within a reference frame to
- the previously coded blocks may be those coded in the scan order, e.g., a raster or other scan order, before the current block.
- Inter-prediction modes may be used with any of the available reference frames.
- NEAREST_MV and NEAR_MV can refer to the most and second most likely motion vectors for the current block obtained by a survey of motion vectors in the context for a reference.
- the reference can be a causal neighborhood in current frame.
- the reference can be co-located motion vectors in the previous frame.
- FIG. 7 is a process flow 700 of coding using delayed frame filtering in in-loop filtering.
- the process flow 700 can be implemented by a decoder, such as the decoder 500 of FIG 5.
- the process flow 700 can be implemented by an encoder, such as in the reconstruction phase of the encoder 400 of FIG. 4.
- the process flow 700 includes a filtering stage 702, upstream stages 704, and a reference frame store 706.
- the filtering stage 702 can be the loop filtering stage 416 of FIG. 4 or the loop filtering stage 512 of FIG. 5.
- the upstream stages 704 can be or include a prediction stage, an dequantization stage, an inverse transform stage, and a reconstruction phase, such as described with respect to FIG. 4 or FIG. 5.
- the reference frame store 706 can be the reference frame store 600 of FIG. 6.
- the filtering stage 702 receives reconstructed frames (such as a reconstructed frame 701) from the upstream stages 704.
- the filtering stage applies one or more filtering tasks (i.e., tools), in a pipeline, to the reconstructed frames to obtain filtered frames.
- the filtered frames obtained from the filtering stage 702 are added to a reference frame store 706.
- a filtered frame can be a degraded frame (such as a degraded frame 703) or a restored frame (such as a restored frame 705).
- a reference frame may be removed from the reference frame store 706 when it is no longer needed, such as when it is output.
- a degraded frame that is stored in the reference frame store 706 may become no longer needed and is, thus, removed from the reference frame store 706 before the restored frame therefrom is obtained from the complex restoration task 710.
- the restored frame is discarded. That is, the restored frame is not added to the reference frame store 706.
- a codec may include a number of additional buffers corresponding to the restoration frame interval parameter. That is, the number of additional buffers may equal to the restoration frame interval parameter.
- Each of the additional buffers is used to hold an in-process restored frame.
- An in-process restored frame may be initialized to a copy of a degraded frame (such as the degraded frame 703) and which undergoes (incremental) updates by the complex restoration task 710 until a restored frame (such as the restored frame 705) is obtained and that is to replace the degraded frame in the reference frame store 706. Additionally, D instances of the complex restoration task 710 may need to be executing (e.g., running, available, etc.) simultaneously.
- D the more complex (and therefore more expensive) the codec hardware implementation needed to implement delayed frame filtering in in-loop filtering.
- the added complexity is due to the additional buffers and the additional executing instances of the complex restoration task 710.
- D can be set to 1.
- a syntax element in a picture parameter set may indicate a restoration frame interval parameter, 1). for the group of frames corresponding to the PPS.
- a PPS can contain parameters common to all frames of the group of frames.
- different groups of pictures may be associated with difference values of the syntax element.
- FIG. 8 is a flowchart of a technique 800 for delayed frame filtering in in-loop filtering.
- the technique 800 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106.
- the software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary' storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 800.
- the technique 800 may be implemented in whole or in part in the loop filtering stage 416 of the encoder 400 of FIG. 4 and/or the loop filtering stage 512 of the decoder 500 of FIG. 5.
- the technique 800 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.
- the technique 800 can be executed for at least some (e.g., all) of the reconstructed frames obtained from a reconstruction phase, such as one of the reconstruction stage 414 of FIG. 4 or the reconstruction stage 510 of FIG. 5.
- a reconstructed frame is received.
- the reconstructed frame can be the reconstructed frame 701 that is generated by the upstream stages 704 and received by the filtering stage 702 of FIG. 7.
- a degraded frame is obtained.
- the degraded frame can be obtained from one or more filtering tasks (e.g., tools) of a the filtering stage.
- the one or more filtering tasks can be the set filtering tasks 707 of FIG. 7. From 802, the technique proceeds to 804 and 806, which, in an example, may be performed in parallel.
- the degraded frame is added to a reference frame store, which can be the reference frame store 706 of FIG. 7.
- a next reconstructed frame is received.
- the next reconstructed frame is a frame that is coded based on the reference frames currently stored in the reference frame store.
- the next reconstructed frame can be received from the upstream coding stage (e.g., the upstream stages 704 of Fig. 7). It is noted that the next reconstructed frame is not necessarily a frame that immediately follows the degraded frame in display order.
- the temporal relationship of the next reconstructed frame obtained at 808 to the degraded frame obtained at 802 can be based on a determined coding structure (i.e., the order in which frames of a video sequence are coded). Additionally, the next reconstructed frame may temporally precede or follow the degraded frame in display order. Receiving the next reconstructed frame at 808 can be the same at receiving the reconstructed frame at 801. Thus, the technique 800 can also proceed to 802 for the next reconstructed frame received at 808. [0085] At 810, the technique 800 determines (e.g., tests, checks, etc.) whether a restored frame is available. That is, the technique 800 determines whether the complex filtering task has completed generating a restored frame from a degraded frame.
- the technique 800 determines whether the complex filtering task that was commenced a number of cycles prior has completed (which is as described below with respect to 818-820). The number of cycles prior can be given by a restoration frame interval parameter, D. which is described above. If the restored frame is not available, then the technique 800 proceeds back to 808 to receive a new reconstructed frame. If the restored frame is available, at 810, then the technique 800 proceeds to 812.
- the technique 800 determines whether to update the reference frame store with the restored frame or whether to discard the restored frame. In an example, if the degraded frame corresponding to restored frame is removed from the reference frame store (such as because it is no longer to be used as a reference frame or because it has been output), then the restored frame is discarded, at 814. Otherwise, if the degraded frame is still in the reference frame store, then the degraded frame is replaced by the restored frame in the reference frame buffer, at 816. As such, subsequent frames are reconstructed based on the reference frame store, which at this point includes the restored frame, as described above. From 816, the technique 800 proceeds to 808 (or, equivalently, to 801).
- determining whether a complex filtering task is available can mean determining whether to perform the complex filtering task with respect to the degraded frame.
- a syntax element (which may be coded in a header associated with the current frame), can indicate whether the complex filtering task is to be performed on the degraded frame to obtain the restored frame. If the syntax element associated with the current indicates that a restored frame is not to be obtained from the restored frame, then technique 800 does not proceed to 818.
- the technique 800 proceeds to 818.
- the complex filtering task is initiated (e.g., executed).
- the restored frame is obtained from (e.g. generated by) the complex filtering task.
- updating the reference frame store at 816 may be performed after the restored frame is obtained at 820.
- the technique 800 may first determine whether, generally, the reference frame store is being used (e.g., whether any frame is current being reconstructed), or more specifically, whether a currently being reconstructed frame is using the degraded frame as a reference frame.
- FIG. 9 is a flowchart of a technique 900 for delayed frame filtering in in-loop filtering.
- the technique 900 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106.
- the software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary' storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 900.
- the technique 900 may be implemented in whole or in part in the loop filtering stage 416 of the encoder 400 of FIG. 4 and/or the loop filtering stage 512 of the decoder 500 of FIG. 5.
- the technique 900 can be implemented using specialized hardware or firmware.
- the technique 900 can be executed for at least some(e.g., all) of the reconstructed frames obtained from a reconstruction phase, such as one of the reconstruction stage 414 of FIG. 4 or the reconstruction stage 510 of FIG. 5.
- a degraded frame is obtained from a reconstructed frame.
- the reconstructed frame can be generated by upstream stages of a coding process.
- the reconstructed frame can be the reconstructed frame 701 generated by upstream stages 704 of FIG. 7.
- the degraded frame can be obtained from (e.g., generated by) one or more filtering tasks (e.g., tools, processes, etc.) of a filtering stage, such as the filtering stage 702 of FIG. 7.
- the one or more filtering tasks can be one or more of the filtering tasks described with respect to the set filtering tasks 707 of FIG. 7 or some other one or more filtering tasks.
- the degraded frame is added to a reference frame store.
- the reference frame store can be the reference frame store 706 of FIG. 7.
- frames that are inter-predicted or more generally, predicted based on motion information with respect to other frames
- a next frame is coded based on the degraded frame. More generally, the next frame is coded based on the current state (i.e., the contents) of the reference frame buffer.
- a restored frame is obtained from the degraded frame using a restoration task (i.e., a complex restoration task).
- the restored frame can be stored in the reference frame store. That is. the restored frame can replace the degraded frame in the reference frame store.
- the technique 900 can include initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame.
- a command which may be included in a compressed bitstream
- the next restored frame is output and discarded. That is. the next restored frame is not added to the reference frame store.
- the technique 900 waits until the next restored frame becomes available (i.e., after the complex filtering task completes) and then outputs the next restored frame.
- FIG. 10 is a flowchart of a technique 1000 for delayed frame filtering in in-loop filtering.
- the technique 1000 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106.
- the software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1000.
- the technique 1000 may be implemented in whole or in part in the loop filtering stage 416 of the encoder 400 of FIG. 4 and/or the loop filtering stage 512 of the decoder 500 of FIG. 5.
- the technique 900 can be implemented using specialized hardware or firmware.
- the technique 1000 can be executed for at least some(e.g., all) of the reconstructed frames obtained from a reconstruction phase, such as one of the reconstruction stage 414 of FIG. 4 or the reconstruction stage 510 of FIG. 5.
- a first frame is decoded using a reference frame.
- the reference frame may be stored in a reference frame store, which can be the reference frame store 706 of FIG. 7.
- the reference frame can be a reconstructed frame, such as the reconstructed frame 701 of FIG. 7, to which zero or more filters of a filtering stage have been applied.
- at least one filtering task may be applied to the reconstructed frame to obtain the reference frame.
- constrained directional enhancement filters may be applied to the reconstructed frame.
- the reconstructed frame can be upscaled to an original resolution to obtain the reference frame.
- the first frame can be a frame of a set of frames where the number of frames in the set can be based on a parameter associated with a filtering task, which can be the complex restoration task 710 of FIG. 7.
- the technique 1000 can decode all of the frames of the set using the reference frame.
- the filtering task is applied to the reference frame to obtain a restored reference frame, which can be the restored frame 705 of FIG. 5.
- a second frame is decoded using the restored reference frame.
- the techniques described herein such as the techniques 800, 900, and 1000 of FIGS. 8, 9, and 10, respectively, are each depicted and described as respective series of steps or operations. However, the steps or operations in accordance with this disclosure can occur in various orders and/or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, not all illustrated steps or operations may be required to implement a method in accordance with the disclosed subject matter. Furthermore, the term “coding 7 ’ (or variations thereof), when used in conjunction with an encoder means “encoding” and when used in conjunction with a decoder means “decoding.”
- example is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the word “example” is intended to present concepts in a concrete fashion.
- the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X includes A or B” is intended to mean any of the natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then “X includes A or B” is satisfied under any of the foregoing instances.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
A degraded frame is obtained from a reconstructed frame of a current frame. The degraded frame is added to a reference frame store. A next frame is decoded based on the degraded frame. Using a restoration task, a restored frame is obtained from the degraded frame. A frame subsequent to the next frame is then coded based on the restored frame.
Description
DELAYED FRAME FILTERING IN IN-LOOP FILTERING
CROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63/436.525, filed December 31, 2022, the entire disclosure of which is hereby incorporated by reference.
BACKGROUND
[0002] Digital video streams may represent video using a sequence of frames or still images. Digital video can be used for various applications including, for example, video conferencing, high-definition video entertainment, video advertisements, or sharing of usergenerated videos. A digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data. Various approaches have been proposed to reduce the amount of data in video streams, including compression and other coding techniques. These techniques may include both lossy and lossless coding techniques.
SUMMARY
[0003] This disclosure relates generally to encoding and decoding video data and more particularly relates to motion vector coding candidate signaling.
[0004] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0005] One general aspect includes a method. The method includes obtaining a degraded frame from a reconstructed frame of a current frame; adding the degraded frame to a reference frame store: coding a next frame based on the degraded frame; obtaining, using a restoration task, a restored frame from the degraded frame; and coding a frame subsequent to the next frame based on the restored frame. Other embodiments of this aspect include
corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0006] Implementations may include one or more of the following features. The method may include storing the restored frame in the reference frame store. The method may include replacing the degraded frame by the restored frame in the reference frame store. Coding the frame subsequent to the next frame may include determining, based on a parameter associated with the restoration task, that the frame subsequent to the next frame is to be coded based on the restored frame. The parameter associated with the restoration task identifies a number of frames that are coded using the degraded frame before coding other frames using the restored frame. The parameter associated with the restoration task can be included in a sequence parameter set of a compressed bitstream. The parameter associated with the restoration task can be included in a picture parameter set of a compressed bitstream.
[0007] The method may include coding in a header of the current frame a syntax element indicating that the restoration task is to be performed with respect to the degraded frame to obtain the restored frame. The method may include initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame; and in response to a command to output the next frame: outputting the next restored frame; and discarding the next restored frame.
[0008] The method may include initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame; and in response to a command to output the next frame before the next restored frame is available: waiting until the next restored frame to become available; and outputting the next restored frame. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
[0009] One general aspect includes a method. The method may include decoding a first frame using a reference frame; applying a filtering task to the reference frame to obtain a restored reference frame; and decoding a second frame using the restored reference frame. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0010] Implementations may include one or more of the following features. The method may include: applying at least one filtering task to a reconstructed frame to obtain the reference frame. Applying the at least one filtering task may include applying constrained
directional enhancement filters to the reconstructed frame. Applying the at least one filtering task may include upscaling the reconstructed frame to an original resolution to obtain the reference frame. Decoding the first frame using the reference frame may include decoding a number of frames that includes the first frame using the reference frame, where the number of the frames is based on a parameter associated with the filtering task. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
[0011] It will be appreciated that aspects can be implemented in any convenient form. For example, aspects may be implemented by appropriate computer programs which may be carried on appropriate carrier media which may be tangible carrier media (e.g. disks) or intangible carrier media (e.g. communications signals). Aspects may also be implemented using suitable apparatus which may take the form of programmable computers running computer programs arranged to implement the methods and/or techniques disclosed herein. Aspects can be combined such that features described in the context of one aspect may be implemented in another aspect.
[0012] These and other aspects of the present disclosure are disclosed in the follow ing detailed description of the embodiments, the appended claims, and the accompanying figures.
BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The description herein refers to the accompanying drawings described below wherein like reference numerals refer to like parts throughout the several views.
[0014] FIG. 1 is a schematic of a video encoding and decoding system.
[0015] FIG. 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
[0016] FIG. 3 is a diagram of an example of a video stream to be encoded and subsequently decoded.
[0017] FIG. 4 is a block diagram of an encoder.
[0018] FIG. 5 is a block diagram of a decoder.
[0019] FIG. 6 is a diagram of an example of a reference frame store.
[0020] FIG. 7 is a process flow of coding using delayed frame filtering in in-loop filtering.
[0021] FIG. 8 is a flow chart of a technique for delayed frame filtering in in-loop filtering.
[0022] FIG. 9 is a flow chart of a technique for delayed frame filtering in in-loop filtering.
[0023] FIG. 10 is a flowchart of a technique for delayed frame filtering in in-loop filtering.
DETAILED DESCRIPTION
[0024] As mentioned, compression schemes related to coding video streams may include breaking images into blocks and generating a digital video output bitstream (i.e., an encoded bitstream) using one or more techniques to limit the information included in the output bitstream. A received bitstream can be decoded to re-create the blocks and the source images from the limited information. Encoding a video stream, or a portion thereof, such as a frame or a block, can include using temporal similarities in the video stream to improve coding efficiency. For example, a current block of a video stream may be encoded based on identifying a difference (residual) between the previously coded pixel values, or between a combination of previously coded pixel values, and those in the current block.
[0025] Encoding using temporal similarities can be known as inter prediction. Inter prediction can attempt to predict the pixel values of a block using a possibly displaced block or blocks from a temporally nearby frame (i.e., reference frame) or frames. A temporally nearby frame is a frame that appears earlier or later in time in the video stream than the frame of the block being encoded. A prediction block resulting from inter prediction is referred to herein as inter predictor.
[0026] Inter prediction is performed using a motion vector (MV). A motion vector used to generate a prediction block refers to a frame other than a current frame, i.e., a reference frame. Reference frames can be located before or after the current frame in the sequence of the video stream. Some codecs may use up to eight reference frames, which can be stored in frame buffers of a reference frame store. The motion vector can refer to (i.e., use) one of the reference frames stored in the reference frame store. Reference frames stored in the reference frame store can also be used to generate motion fields.
[0027] Residuals (i.e.. differences) can be encoded using a lossy quantization step. Decoding (i.e., reconstructing) from such residuals often results in distortion or artefacts (e.g., ringing artefacts or blockiness artefacts) in the reconstructed data. The encoder and decoder may perform operations that improve the qualify of the reconstructed data, such as described below with respect to a loop filtering stage 416 of FIG. 4 and a loop filtering stage 512 of FIG. 5 (e.g., within the reconstruction loop at the encoder or prior the outputting at the
decoder). Loop restoration may be performed to process a reconstructed video frame for use as a reference frame.
[0028] A filtering stage may include more than one restoration tasks (i.e., operations or tools) that are to be applied to a reconstructed frame of a current frame. The restorations tasks are expected to complete within the time allocated for decoding the current frame. For example, a codec may be designed for a certain throughput. The throughput of a codec may be specified in terms of frames per second. For example, a decoder may be designed to decode fames at a rate of 60 frames per second. As such, the decoder allocates, on average, l/60th of a second to decode a frame.
[0029] However, some restoration tasks may be too complex to complete without significant added complexity to a codec. To meet the codec throughput constraints, more complex restoration tasks may result in additional codec complexity'. For example, a hardware-implemented codec may require additional circuitry (i.e., larger silicon footprint), and a software-implemented coded may require additional processing cores and threads, to complete complex restoration tasks within the codec throughput requirements. However, these are not desirable solutions.
[0030] An example of a complex restoration task is a convolutional neural network (CNN), which is a type of machine-learning (ML) models, that is trained to perform restoration. However, CNNs may not be suitable for widespread deployment in video coding due to their complexity requirements (in terms of, for example, model size and requisite inferencing time) at the throughput desired for video coding. CNNs (and more generally, ML models) may be too complex to implement in a cost-effective manner with a small enough silicon footprint to meet stringent throughput requirements. Other examples of complex restoration tasks other than CNNs are possible and the disclosure herein is not limited to complex restoration tasks that are CNNs.
[0031] To summarize, an inter-predicted frame, or inter-predicted blocks therein, can use any of the reference frames stored in a reference frame store through prediction, motion field generation, or for some other purpose related to motion determination. Once a new frame has been decoded (e.g., output by a filtering stage), it updates the reference frame store by replacing one of the existing slots so that it can be used to decode subsequent frames. For this reason, any filtering or restoration operation conducted on the frame before producing the final output is characterized as in-loop. However, some in-loop filtering tasks (e.g., operations) may be or involve complex processing.
[0032] Implementations according to this disclosure provide a solution to the problem of complex in-loop filtering tasks by relaxing the in-loop-ness of the decoding mechanism by allowing the complex filtering task to have more time to work on a frame just decoded (up to the usage of the complex filtering task) while retaining the low complexity of a codec. Rather. The throughput required by complex restoration tasks (such as for ML inference in a case that the complex restoration task is implemented using an ML model) is reduced at a modest loss in efficiency and increase in delay, which can be acceptable in applications such as video-on-demand applications. Reduced throughput translates directly to reduced gatecount and silicon area for implementation of the hardware.
[0033] Herein, a reconstructed frame can be processed by one or more filtering tasks to obtain a degraded frame, which is further processed by a complex filtering task to generate a restored frame. As such, relaxing the in-loop-ness of the coding mechanism can mean that the coding of some subsequent frames need not wait until the restored frame is available. Rather, coding of a certain number subsequent frames proceeds based on the degraded frame and, when the restored frame becomes available, then coding of later subsequence frames proceeds based on the restored frame. As the use of restored frame is delayed until the restore frame is available and, until then, coding proceeds based on the degraded frame.
[0034] Further details of delayed frame filtering in in-loop filtering are described herein with initial reference to a system in which it can be implemented.
[0035] FIG. 1 is a schematic of a video encoding and decoding system 100. A transmitting station 102 can be, for example, a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.
[0036] A network 104 can connect the transmitting station 102 and a receiving station 106 for encoding and decoding of the video stream. Specifically, the video stream can be encoded in the transmitting station 102 and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), wide area network (WAN), virtual private network (VPN), cellular telephone netw ork or any other means of transferring the video stream from the transmitting station 102 to. in this example, the receiving station 106.
[0037] The receiving station 106, in one example, can be a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable
implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
[0038] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit the network 104. In another implementation, a video stream can be encoded and then stored for transmission at a later time to the receiving station 106 or any other device having memory. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and/or some communication pathway) the encoded video stream and stores the video stream for later decoding. In an example implementation, a real-time transport protocol (RTP) is used for transmission of the encoded video over the network 104. In another implementation, a transport protocol other than RTP may be used, e.g., a video streaming protocol based on the Hypertext Transfer Protocol (HTTP).
[0039] When used in a video conferencing system, for example, the transmitting station 102 and/or the receiving station 106 may include the ability to both encode and decode a video stream as described below. For example, the receiving station 106 could be a video conference participant who receives an encoded video bitstream from a video conference sen- er (e.g., the transmitting station 102) to decode and view and further encodes and transmits its own video bitstream to the video conference sen- er for decoding and viewing by other participants.
[0040] FIG. 2 is a block diagram of an example of a computing device 200 (e.g., an apparatus) that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement one or both of the transmitting station 102 and the receiving station 106 of FIG. 1. The computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of one computing device, for example, a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.
[0041] A CPU 202 in the computing device 200 can be a conventional central processing unit. Alternatively, the CPU 202 can be any other type of device, or multiple devices, capable of manipulating or processing information now existing or hereafter developed. Although the disclosed implementations can be practiced with one processor as shown, e.g., the CPU 202, advantages in speed and efficiency can be achieved using more than one processor.
[0042] A memory 204 in computing device 200 can be a read only memory (ROM) device or a random-access memory (RAM) device in an implementation. Any other suitable
type of storage device can be used as the memory 204. The memory 204 can include code and data 206 that is accessed by the CPU 202 using a bus 212. The memory 204 can further include an operating system 208 and application programs 210, the application programs 210 including at least one program that permits the CPU 202 to perform the methods described here. For example, the application programs 210 can include applications 1 through N. which further include a video coding application that performs the techniques described here, such as those for delayed frame filtering in in-loop filtering. Computing device 200 can also include a secondary' storage 214. which can, for example, be a memory card used with a mobile computing device. Because the video communication sessions may contain a significant amount of information, they can be stored in whole or in part in the secondary storage 214 and loaded into the memoiy 204 as needed for processing.
[0043] The computing device 200 can also include one or more output devices, such as a display 218. The display 218 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 218 can be coupled to the CPU 202 via the bus 212. Other output devices that permit a user to program or otherwise use the computing device 200 can be provided in addition to or as an alternative to the display 218. When the output device is or includes a display, the display can be implemented in various ways, including by a liquid cry stal display (LCD), a cathode-ray tube (CRT) display or light emitting diode (LED) display, such as an organic LED (OLED) display.
[0044] The computing device 200 can also include or be in communication with an image-sensing device 220, for example a camera, or any other image-sensing device 220 now existing or hereafter developed that can sense an image such as the image of a user operating the computing device 200. The image-sensing device 220 can be positioned such that it is directed toward the user operating the computing device 200. In an example, the position and optical axis of the image-sensing device 220 can be configured such that the field of vision includes an area that is directly adjacent to the display 218 and from which the display 218 is visible.
[0045] The computing device 200 can also include or be in communication with a soundsensing device 222, for example a microphone, or any other sound-sensing device nowexisting or hereafter developed that can sense sounds near the computing device 200. The sound-sensing device 222 can be positioned such that it is directed toward the user operating
the computing device 200 and can be configured to receive sounds, for example, speech or other utterances, made by the user while the user operates the computing device 200.
[0046] Although FIG. 2 depicts the CPU 202 and the memory 204 of the computing device 200 as being integrated into one unit, other configurations can be utilized. The operations of the CPU 202 can be distributed across multiple machines (wherein individual machines can have one or more of processors) that can be coupled directly or across a local area or other network. The memory 204 can be distributed across multiple machines such as a network-based memory or memory' in multiple machines performing the operations of the computing device 200. Although depicted here as one bus, the bus 212 of the computing device 200 can be composed of multiple buses. Further, the secondary storage 214 can be directly coupled to the other components of the computing device 200 or can be accessed via a netw ork and can comprise an integrated unit such as a memory' card or multiple units such as multiple memory cards. The computing device 200 can thus be implemented in a wide variety of configurations.
[0047] FIG. 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a number of adjacent frames 304. While three frames are depicted as the adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. The adjacent frames 304 can then be further subdivided into individual frames, e.g., a frame 306. At the next level, the frame 306 can be divided into a series of planes or segments 308. The segments 308 can be subsets of frames that permit parallel processing, for example. The segments 308 can also be subsets of frames that can separate the video data into separate colors. For example, a frame 306 of color video data can include a luminance plane and two chrominance planes. The segments 308 may be sampled at different resolutions.
[0048] Whether or not the frame 306 is divided into segments 308, the frame 306 may be further subdivided into blocks 310, which can contain data corresponding to. for example, 16x16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more segments 308 of pixel data. The blocks 310 can also be of any other suitable size such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger. Unless otherwise noted, the terms block and macro-block are used interchangeably herein.
[0049] FIG. 4 is a block diagram of an encoder 400. The encoder 400 can be implemented, as described above, in the transmitting station 102 such as by providing a
computer software program stored in memory, for example, the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the transmitting station 102 to encode video data in the manner described in FIG. 4. The encoder 400 can also be implemented as specialized hardware included in, for example, the transmitting station 102. In one particularly desirable implementation, the encoder 400 is a hardware encoder.
[0050] The encoder 400 has the following stages to perform the various functions in a forward path (shown by the solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra/inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408. The encoder 400 may also include a reconstruction path (shown by the dotted connection lines) to reconstruct a frame for encoding of future blocks. In FIG. 4, the encoder 400 has the following stages to perform the various functions in the reconstruction path: a dequantization stage 410, an inverse transform stage 412. a reconstruction stage 414. and a loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.
[0051] When the video stream 300 is presented for encoding, respective frames 304, such as the frame 306, can be processed in units of blocks. At the intra/inter prediction stage 402, respective blocks can be encoded using intra-frame prediction (also called intra-prediction) or inter-frame prediction (also called inter-prediction). In any case, a prediction block can be formed. In the case of intra-prediction, a prediction block may be formed from samples in the current frame that have been previously encoded and reconstructed. In the case of interprediction. a prediction block may be formed from samples in one or more previously constructed reference frames.
[0052] Next, still referring to FIG. 4, the prediction block can be subtracted from the current block at the intra/inter prediction stage 402 to produce a residual block (also called a residual). The transform stage 404 transforms the residual into transform coefficients in, for example, the frequency domain using block-based transforms. The quantization stage 406 converts the transform coefficients into discrete quantum values, which are referred to as quantized transform coefficients, using a quantizer value or a quantization level. For example, the transform coefficients may be divided by the quantizer value and truncated. The quantized transform coefficients are then entropy encoded by the entropy encoding stage 408. The entropy-encoded coefficients, together with other information used to decode the
block, which may include for example the type of prediction used, transform type, motion vectors and quantizer value, are then output to the compressed bitstream 420. The compressed bitstream 420 can be formatted using various techniques, such as variable length coding (VLC) or arithmetic coding. The compressed bitstream 420 can also be referred to as an encoded video stream or encoded video bitstream, and the terms will be used interchangeably herein.
[0053] The reconstruction path in FIG. 4 (show n by the dotted connection lines) can be used to ensure that the encoder 400 and a decoder 500 (described below-) use the same reference frames to decode the compressed bitstream 420. The reconstruction path performs functions that are similar to functions that take place dunng the decoding process that are discussed in more detail below, including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (also called a derivative residual). At the reconstruction stage 414, the prediction block that was predicted at the intra/inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filtering stage 416 can be applied to the reconstructed block to reduce distortion such as blocking artifacts.
[0054] Other variations of the encoder 400 can be used to encode the compressed bitstream 420. For example, a non-transform-based encoder can quantize the residual signal directly without the transform stage 404 for certain blocks or frames. In another implementation, an encoder can have the quantization stage 406 and the dequantization stage 410 combined in a common stage.
[0055] FIG. 5 is a block diagram of a decoder 500. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described in FIG. 5. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106. [0056] The decoder 500, similar to the reconstruction path of the encoder 400 discussed above, includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra/inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512 and a deblocking filtering
stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.
[0057] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by the quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derivative residual that can be identical to that created by the inverse transform stage 412 in the encoder 400. Using header information decoded from the compressed bitstream 420, the decoder 500 can use the intra/inter prediction stage 508 to create the same prediction block as was created in the encoder 400, e.g., at the intra/inter prediction stage 402. At the reconstruction stage 510, the prediction block can be added to the derivative residual to create a reconstructed block. The loop filtering stage 512 can be applied to the reconstructed block to reduce blocking artifacts. [0058] Other filtering can be applied to the reconstructed block. In this example, the deblocking filtering stage 514 is applied to the reconstructed block to reduce blocking distortion, and the result is output as the output video stream 516. The output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein. Other variations of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can produce the output video stream 516 without the deblocking filtering stage 514.
[0059] FIG. 6 is a diagram of an example of a reference frame store 600. As described above with respect to inter prediction, a motion vector used to generate a prediction block refers to (i.e., uses) a reference frame. Reference frames can be stored in a reference frame buffers of the reference frame store 600.
[0060] A current frame, or a current block or region of a frame, can be coded (encoded or decoded) using a reference frame such as a “last frame,’7 which is the adjacent frame immediately before the current frame in the video sequence. When video frames are coded out of order (i.e., not in the sequence that they appear in the video stream), motion information from video frames in the past or future can be included in the candidate reference motion vectors. Coding video frames can occur, for example, using so-called “alternate reference frames” that are not temporally neighboring to the frames coded immediately before or after them. An alternate reference frame can be a synthesized frame that does not
occur in the input video stream or is a duplicate frame to one in the input video stream that is used for prediction. An alternate frame may not be displayed following decoding. Such a frame can resemble a video frame in the non-adjacent future. Another example in which out of order encoding may occur is through the use of a so-called “golden frame,” which is a reconstructed video frame that may or may not be neighboring to a cunent frame and is stored in memory for use as a reference frame until replaced, e.g., by a new golden frame. [0061] The reference frame store 600 stores reference frames used to encode or decode blocks of frames of a video sequence. The reference frame buffer can include reference frames such as those described above. For example, the reference frame store 600 can include a last frame LAST FRAME 602, a golden frame GOLDEN FRAME 604, and an alternative reference frame ALTREF FRAME 606. A reference frame buffer can include more, fewer, or other reference frames. The reference frame store 600 is shown as including eight reference frames. However, a reference frame buffer can include more or fewer than eight reference frames. Any particular semantics (e.g., labels) associated with reference frames stored in the reference frame store 600 are not necessary to the understanding or use of this disclosure.
[0062] The last frame LAST_FRAME 602 can be, for example, the adjacent frame immediately before the current frame in the video sequence. The golden frame GOLDEN_FRAME 604 can be, for example, a reconstructed video frame for use as a reference frame that may or may not be adjacent to the current frame. The alternative reference frame ALTREF FRAME 606 can be, for example, a video frame in the non- adjacent future, which is a backward reference frame.
[0063] The reference frames stored in the reference frame store 600 can be used to identify motion vectors for predicting blocks of frames to be encoded or decoded. Different reference frames may be used depending on the type of prediction used to predict a current block of a current frame. For example, when compound prediction is used, multiple frames, such as one for forward prediction (e.g.. LAST FRAME 602 or GOLDEN FRAME 604) and one for backward prediction (e.g., ALTREF FRAME 606) can be used for predicting the current block.
[0064] There may be a finite number of reference frames that can be stored within the reference frame store 600. As shown in FIG. 6, the reference frame store 600 can store up to eight reference frames. Although three of the eight spaces in the reference frame store 600 are
used by the LAST FRAME 602, the GOLDEN FRAME 604, and the ALTREF FRAME 606, five spaces remain available to store other reference frames.
[0065] In particular, one or more available spaces in the reference frame store 600 may be used to store a second last frame LAST2_FRAME and/or a third last frame
LAST3 FRAME as additional forward reference frames, in addition to the LAST FRAME 602. A backward frame BWDREF FRAME 608 can be stored as an additional backward prediction reference frame, in addition to ALTREF FRAME 606. The BWDREF FRAME can be closer in relative distance to the current frame than the ALTREF_FRAME 606, for example.
[0066] In one example, the pair of {LAST FRAME, BWDREF FRAME} can be used to generate a compound predictor for coding the current block. In this example, LAST FRAME is a “nearest” forward reference frame for forward prediction, and BWDREF_FRAME is a “nearest” backward reference frame for backward prediction.
[0067] A current block is predicted based on a prediction mode. The prediction mode may be selected from one of multiple intra-prediction modes. In the case of inter prediction, the prediction mode may be selected from one of multiple inter-prediction modes using one or more reference frames of the reference frame store 600 including, for example, the LAST FRAME 602, the GOLDEN FRAME 604, the ALTREF FRAME 606. or any other reference frame. The prediction mode of the current block can be transmitted from an encoder, such as the encoder 400 of FIG. 4, to a decoder, such as the decoder 500 of FIG. 5, in an encoded bitstream, such as the compressed bitstream 420 of FIGS. 4-5.
[0068] A bitstream syntax can support three categories of inter prediction modes in an example. These inter prediction modes can include a mode (referred to herein as the ZERO_MV mode) in which a block from the same location within a reference frame as the current block is used as the prediction block, a mode (referred to herein as the NEW_MV mode) in which a motion vector is transmitted to indicate the location of a block within a reference frame to be used as the prediction block relative to the current block, or a mode (referred to herein as the REF_MV mode and comprising aNEAR_MV or NEAREST MV mode) in which no motion vector is transmitted and the current block uses the last or second- to-last non-zero motion vector used by neighboring, previously coded blocks to generate the prediction block. The previously coded blocks may be those coded in the scan order, e.g., a raster or other scan order, before the current block. Inter-prediction modes may be used with any of the available reference frames. NEAREST_MV and NEAR_MV can refer to the most
and second most likely motion vectors for the current block obtained by a survey of motion vectors in the context for a reference. The reference can be a causal neighborhood in current frame. The reference can be co-located motion vectors in the previous frame.
[0069] FIG. 7 is a process flow 700 of coding using delayed frame filtering in in-loop filtering. The process flow 700 can be implemented by a decoder, such as the decoder 500 of FIG 5. The process flow 700 can be implemented by an encoder, such as in the reconstruction phase of the encoder 400 of FIG. 4. The process flow 700 includes a filtering stage 702, upstream stages 704, and a reference frame store 706. The filtering stage 702 can be the loop filtering stage 416 of FIG. 4 or the loop filtering stage 512 of FIG. 5. The upstream stages 704 can be or include a prediction stage, an dequantization stage, an inverse transform stage, and a reconstruction phase, such as described with respect to FIG. 4 or FIG. 5. The reference frame store 706 can be the reference frame store 600 of FIG. 6.
[0070] The filtering stage 702 receives reconstructed frames (such as a reconstructed frame 701) from the upstream stages 704. The filtering stage applies one or more filtering tasks (i.e., tools), in a pipeline, to the reconstructed frames to obtain filtered frames. The filtered frames obtained from the filtering stage 702 are added to a reference frame store 706. A filtered frame can be a degraded frame (such as a degraded frame 703) or a restored frame (such as a restored frame 705).
[0071] The filtering stage 702 is illustrated as including a set of filtering tasks 707 and a complex restoration task 710. The set of filtering tasks 707 includes filtering tasks 708A to 708N. The filtering tasks 708A to 708N are considered to be simple because they do not require a significant amount of time to complete. The disclosure herein is not limited by the number or the specific details (semantics) of the filtering tasks 708A to 708N. While, as already mentioned, the disclosure is not limited by the number of filtering tasks of the set filtering tasks 707 and their semantics, a brief description of the filtering tasks 708A to 708N is provided herein. The filtering tasks 708A to 708N can be or includes filtering tasks applied by the AV 1 codec; however, other filtering tasks are possible.
[0072] The filtering task 708A may be used to reduce blocking artifacts at transform block boundaries caused by quantization. In an example, the selection of filter length can be determined by the minimum transform block sizes that are applied on both sides of the transform block boundary. The filtering task 708B may apply constrained directional enhancement filters (CDEFs) that perform edge direction searching at blocks of a certain size (e.g., 8x8). CDEFs identify edge directions within blocks and apply filters accordingly. The
CDEF filtering process may consist of a primary filter and a secondary filter. The primary filter processes reconstruction samples along an identified edge direction, while the secondary filter processes reconstruction samples along a direction 45-degrees from the edge direction.
[0073] The filtering task 708C can be used to upscale a downscaled frame back to its original resolution. The filtering task 708N may apply filters at what is referred to as loop restoration units (LRUs), which may be 64x64, 128x 128, or 256x256 blocks. A loop restoration filter may be applied on an LRU subject to one of three options: 1) applying a Wiener filter, 2) applying a self-guided filter (SGF), and 3) not applying any filter. The Wiener filter may sequentially apply a 7-tap vertical filter followed by a 7-tap horizontal filter. Self-guided filters (SGFs) obtain two initial filtered frames, Xi and X2. The final restoration Xr can be obtained as a combination of the degraded samples and the difference between the degraded sample and the coarse restoration.
[0074] The complex restoration task 710 is deemed or considered to be complex at least because it requires more time to complete than is available to complete within an allotted time (which may be a frame processing time). The filtering stage 702 is such that the most complex filtering task (i.e., the complex restoration task 710) is the last filtering task in the in-loop filtering pipeline. The filtering stage 702 is further characterized by a design parameter D > 0 that determines or is related to a time that the complex restoration task 710 can take to obtain a restored frame from a degraded frame. The parameter D is referred to herein as a “restoration frame interval parameter.” The parameter D may be given in units of the frame interval.
[0075] When D = C (e.g., C=l). the complex restoration task 710 is allowed to take up to C (e g., 1) frame-intervals of time to complete the generation of a restored frame from a degraded frame. When
= 0 (i.e., a no-delay case), frame restoration needs to be completed immediately and before the next frame can be decoded. That is. in the case that D=0. the complex restoration task 710 is not performed with respect to a degraded frame.
[0076] The restoration frame interval parameter, D, is the number of frames that are to be coded using a degraded frame obtained from a reconstructed frame before coding subsequence frames using the restored frame. To illustrate, assume that frames are coded in an order N. N+L N+2. N+3, etc. and that restoration frame interval parameter D=2. Assume further that a degraded frame N is now available and that the complex filtering task is started on the degraded on the degraded frame N to obtain a restored frame N. As such, frames N+l
and N+2 are coded using the degrade frame N and then frames starting with frame N+3 can be coded using the restored frame N (if available).
[0077] The degraded frame 703 output by a last filtering task (e.g., the filtering task 708N) of the set filtering tasks 707 is used to update the reference frame store 706. The next D frames are coded using the degraded frame 703 as a reference frame. The degraded frame 703 is also provided (i.e., input) to the complex restoration task 710, which produces a restored frame 705. The restored frame 705 may replace (e.g., overwrite) the degraded frame 703 in the reference frame store 706 and may then be used (as reference frame) for coding subsequent frames.
[0078] In some situations, a reference frame may be removed from the reference frame store 706 when it is no longer needed, such as when it is output. As such, in some situations, a degraded frame that is stored in the reference frame store 706 may become no longer needed and is, thus, removed from the reference frame store 706 before the restored frame therefrom is obtained from the complex restoration task 710. In such a case, the restored frame is discarded. That is, the restored frame is not added to the reference frame store 706. [0079] To implement delayed frame filtering in in-loop filtering, a codec may include a number of additional buffers corresponding to the restoration frame interval parameter. That is, the number of additional buffers may equal to the restoration frame interval parameter. Each of the additional buffers is used to hold an in-process restored frame. An in-process restored frame may be initialized to a copy of a degraded frame (such as the degraded frame 703) and which undergoes (incremental) updates by the complex restoration task 710 until a restored frame (such as the restored frame 705) is obtained and that is to replace the degraded frame in the reference frame store 706. Additionally, D instances of the complex restoration task 710 may need to be executing (e.g., running, available, etc.) simultaneously.
[0080] As such, the larger the value of D, the more complex (and therefore more expensive) the codec hardware implementation needed to implement delayed frame filtering in in-loop filtering. The added complexity is due to the additional buffers and the additional executing instances of the complex restoration task 710. Thus, in an implementation, to balance the complexity (e.g., cost) of delayed frame filtering in in-loop filtering with improved quality of reconstructed frames, D can be set to 1.
[0081] In an example, the restoration frame interval parameter, I). may be a design parameter of a codec. For example, a codec may be designed to meet certain throughput requirements, cost, or other criteria, and, as such, the hardware implementation may be
designed based on a value of D that meets these criteria. In another example, the restoration frame interval parameter, />. may be transmitted in a compressed bitstream from an encoder to a decoder. For example, a syntax element in a sequence parameter set (SPS) may indicate the restoration frame interval parameter. D. As is known, an SPS can contain parameters common to an entire video sequence (i.e., to each of the frames of the video sequence). In an example, the restoration frame interval parameter, I). may be indicated for a group frames. For example, a syntax element in a picture parameter set (PPS) may indicate a restoration frame interval parameter, 1). for the group of frames corresponding to the PPS. As is known, a PPS can contain parameters common to all frames of the group of frames. As such, different groups of pictures may be associated with difference values of the syntax element. [0082] FIG. 8 is a flowchart of a technique 800 for delayed frame filtering in in-loop filtering. The technique 800 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary' storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 800. The technique 800 may be implemented in whole or in part in the loop filtering stage 416 of the encoder 400 of FIG. 4 and/or the loop filtering stage 512 of the decoder 500 of FIG. 5. The technique 800 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used. The technique 800 can be executed for at least some (e.g., all) of the reconstructed frames obtained from a reconstruction phase, such as one of the reconstruction stage 414 of FIG. 4 or the reconstruction stage 510 of FIG. 5.
[0083] At 801. a reconstructed frame is received. The reconstructed frame can be the reconstructed frame 701 that is generated by the upstream stages 704 and received by the filtering stage 702 of FIG. 7. At 802, a degraded frame is obtained. The degraded frame can be obtained from one or more filtering tasks (e.g., tools) of a the filtering stage. The one or more filtering tasks can be the set filtering tasks 707 of FIG. 7. From 802, the technique proceeds to 804 and 806, which, in an example, may be performed in parallel.
[0084] At 804, the degraded frame is added to a reference frame store, which can be the reference frame store 706 of FIG. 7. At 808, a next reconstructed frame is received. The next reconstructed frame is a frame that is coded based on the reference frames currently stored in the reference frame store. The next reconstructed frame can be received from the upstream coding stage (e.g., the upstream stages 704 of Fig. 7). It is noted that the next reconstructed
frame is not necessarily a frame that immediately follows the degraded frame in display order. The temporal relationship of the next reconstructed frame obtained at 808 to the degraded frame obtained at 802 can be based on a determined coding structure (i.e., the order in which frames of a video sequence are coded). Additionally, the next reconstructed frame may temporally precede or follow the degraded frame in display order. Receiving the next reconstructed frame at 808 can be the same at receiving the reconstructed frame at 801. Thus, the technique 800 can also proceed to 802 for the next reconstructed frame received at 808. [0085] At 810, the technique 800 determines (e.g., tests, checks, etc.) whether a restored frame is available. That is, the technique 800 determines whether the complex filtering task has completed generating a restored frame from a degraded frame. More specifically, the technique 800 determines whether the complex filtering task that was commenced a number of cycles prior has completed (which is as described below with respect to 818-820). The number of cycles prior can be given by a restoration frame interval parameter, D. which is described above. If the restored frame is not available, then the technique 800 proceeds back to 808 to receive a new reconstructed frame. If the restored frame is available, at 810, then the technique 800 proceeds to 812.
[0086] At 812, the technique 800 determines whether to update the reference frame store with the restored frame or whether to discard the restored frame. In an example, if the degraded frame corresponding to restored frame is removed from the reference frame store (such as because it is no longer to be used as a reference frame or because it has been output), then the restored frame is discarded, at 814. Otherwise, if the degraded frame is still in the reference frame store, then the degraded frame is replaced by the restored frame in the reference frame buffer, at 816. As such, subsequent frames are reconstructed based on the reference frame store, which at this point includes the restored frame, as described above. From 816, the technique 800 proceeds to 808 (or, equivalently, to 801).
[0087] At 806, the technique 800 determines whether a complex filtering task is available to obtain a restored frame from the degraded frame. In an example, the complex filtering task may not be available if the complex filtering task is already being executed with respect to (e.g., for) another degraded task. In an example, a codec may be configured to execute a certain number (e.g., instances) of the complex filtering task according to the restoration frame interval parameter, D. As such, if an instance of the complex filtering task is not available (not shown in FIG. 8), then a restored frame is not obtained from the degraded frame.
[0088] In an example, determining whether a complex filtering task is available can mean determining whether to perform the complex filtering task with respect to the degraded frame. For example, a syntax element (which may be coded in a header associated with the current frame), can indicate whether the complex filtering task is to be performed on the degraded frame to obtain the restored frame. If the syntax element associated with the current indicates that a restored frame is not to be obtained from the restored frame, then technique 800 does not proceed to 818.
[0089] If the complex filtering task is available, then the technique 800 proceeds to 818. At 818. the complex filtering task is initiated (e.g., executed). At 820, the restored frame is obtained from (e.g. generated by) the complex filtering task.
[0090] While one particular arrangement of blocks (e.g., steps) is shown and described with respect to FIG. 8, other arrangements are possible. In an example, updating the reference frame store at 816 may be performed after the restored frame is obtained at 820. However, to avoid a situation where the degraded frame is replaced by the restored frame while the degraded frame is being used, the technique 800 may first determine whether, generally, the reference frame store is being used (e.g., whether any frame is current being reconstructed), or more specifically, whether a currently being reconstructed frame is using the degraded frame as a reference frame.
[0091] If the frame is to be output (e.g., displayed) from the reference frame store (such in response to a "show frame” instruction received in a compressed bitstream), then two alternative implementations are possible. In a first implementation, the degraded frame that is currently in the reference frame store can be output and a corresponding restore frame is then discarded. In a second implementation, outputting the frame can be delayed until the restored frame is obtained and the restored frame is then output.
[0092] FIG. 9 is a flowchart of a technique 900 for delayed frame filtering in in-loop filtering. The technique 900 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary' storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 900. The technique 900 may be implemented in whole or in part in the loop filtering stage 416 of the encoder 400 of FIG. 4 and/or the loop filtering stage 512 of the decoder 500 of FIG. 5. The technique 900 can be implemented using specialized hardware or firmware. Multiple
processors. memories, or both, may be used. The technique 900 can be executed for at least some(e.g., all) of the reconstructed frames obtained from a reconstruction phase, such as one of the reconstruction stage 414 of FIG. 4 or the reconstruction stage 510 of FIG. 5.
[0093] At 902, a degraded frame is obtained from a reconstructed frame. The reconstructed frame can be generated by upstream stages of a coding process. For example, the reconstructed frame can be the reconstructed frame 701 generated by upstream stages 704 of FIG. 7. The degraded frame can be obtained from (e.g., generated by) one or more filtering tasks (e.g., tools, processes, etc.) of a filtering stage, such as the filtering stage 702 of FIG. 7. The one or more filtering tasks can be one or more of the filtering tasks described with respect to the set filtering tasks 707 of FIG. 7 or some other one or more filtering tasks.
[0094] At 904, the degraded frame is added to a reference frame store. The reference frame store can be the reference frame store 706 of FIG. 7. As described above, frames that are inter-predicted (or more generally, predicted based on motion information with respect to other frames) can use, as reference frames, frames stored in the reference frame store. At 906, a next frame is coded based on the degraded frame. More generally, the next frame is coded based on the current state (i.e., the contents) of the reference frame buffer. At 908, a restored frame is obtained from the degraded frame using a restoration task (i.e., a complex restoration task). In an example, the restored frame can be stored in the reference frame store. That is. the restored frame can replace the degraded frame in the reference frame store.
[0095] At 910, a frame that is subsequent to the next frame is coded based on the restored frame. The subsequent frame may not necessarily be a frame that immediately follows the next frame in display order. Furthermore, the subsequent frame may not necessarily be a frame that immediately follows the next frame in coding order. Rather, and as described above, the distance between the degraded frame and the subsequent frame is based on the restoration frame interval parameter, I). As such, coding the frame subsequent to the next frame can include determining, based on a parameter (i.e., the restoration frame interval parameter, /)) associated with the restoration task, that the frame subsequent to the next frame is to be coded based on the restored frame. The parameter associated with the restoration task identifies a number of frames that are coded using the degraded frame before coding frames using the restored frame.
[0096] In an example, the technique 900 can include initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame. In response to a command (which may be included in a compressed bitstream) to output the next frame, the
next restored frame is output and discarded. That is. the next restored frame is not added to the reference frame store. In an example, in response to a command to output the next frame before the next restored frame is available, the technique 900 waits until the next restored frame becomes available (i.e., after the complex filtering task completes) and then outputs the next restored frame.
[0097] FIG. 10 is a flowchart of a technique 1000 for delayed frame filtering in in-loop filtering. The technique 1000 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1000. The technique 1000 may be implemented in whole or in part in the loop filtering stage 416 of the encoder 400 of FIG. 4 and/or the loop filtering stage 512 of the decoder 500 of FIG. 5. The technique 900 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used. The technique 1000 can be executed for at least some(e.g., all) of the reconstructed frames obtained from a reconstruction phase, such as one of the reconstruction stage 414 of FIG. 4 or the reconstruction stage 510 of FIG. 5.
[0098] At 1002, a first frame is decoded using a reference frame. The reference frame may be stored in a reference frame store, which can be the reference frame store 706 of FIG. 7. The reference frame can be a reconstructed frame, such as the reconstructed frame 701 of FIG. 7, to which zero or more filters of a filtering stage have been applied. As such, in an example, at least one filtering task may be applied to the reconstructed frame to obtain the reference frame. In an example, constrained directional enhancement filters may be applied to the reconstructed frame. In an example, the reconstructed frame can be upscaled to an original resolution to obtain the reference frame. The first frame can be a frame of a set of frames where the number of frames in the set can be based on a parameter associated with a filtering task, which can be the complex restoration task 710 of FIG. 7. As such, the technique 1000 can decode all of the frames of the set using the reference frame.
[0099] At 1004, the filtering task is applied to the reference frame to obtain a restored reference frame, which can be the restored frame 705 of FIG. 5. At 1006, a second frame is decoded using the restored reference frame.
[00100] For simplicity of explanation, the techniques described herein, such as the techniques 800, 900, and 1000 of FIGS. 8, 9, and 10, respectively, are each depicted and
described as respective series of steps or operations. However, the steps or operations in accordance with this disclosure can occur in various orders and/or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, not all illustrated steps or operations may be required to implement a method in accordance with the disclosed subject matter. Furthermore, the term “coding7’ (or variations thereof), when used in conjunction with an encoder means “encoding” and when used in conjunction with a decoder means “decoding.”
[00101] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it is to be understood that encoding and decoding, as those terms are used in the claims, could mean compression, decompression, transformation, or any other processing or change of data.
[00102] The word “example” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the word “example” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X includes A or B” is intended to mean any of the natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then “X includes A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form. Moreover, use of the term “an implementation” or “one implementation” throughout is not intended to mean the same embodiment or implementation unless described as such.
[00103] Implementations of the transmitting station 102 and/or the receiving station 106 (and the algorithms, methods, instructions, etc., stored thereon and/or executed thereby, including by the encoder 400 and the decoder 500) can be realized in hardware, software, or any combination thereof. The hardware can include, for example, computers, intellectual property (IP) cores, application-specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, microcontrollers, servers, microprocessors, digital signal processors or any other suitable circuit. In the claims, the term “processor” should be understood as encompassing any of the foregoing hardware, either singly or in combination. The terms “signal” and “data” are used interchangeably.
Further, portions of the transmitting station 102 and the receiving station 106 do not necessarily have to be implemented in the same manner.
[00104] Further, in one aspect, for example, the transmitting station 102 or the receiving station 106 can be implemented using a general-purpose computer or general-purpose processor with a computer program that, when executed, carries out any of the respective methods, algorithms and/or instructions described herein. In addition, or alternatively, for example, a special purpose computer/processor can be utilized which can contain other hardware for carrying out any of the methods, algorithms, or instructions described herein. [00105] The transmitting station 102 and the receiving station 106 can, for example, be implemented on computers in a video conferencing system. Alternatively, the transmitting station 102 can be implemented on a server and the receiving station 106 can be implemented on a device separate from the server, such as a hand-held communications device. In this instance, the transmitting station 102 can encode content using an encoder 400 into an encoded video signal and transmit the encoded video signal to the communications device. In turn, the communications device can then decode the encoded video signal using a decoder 500. Alternatively, the communications device can decode content stored locally on the communications device, for example, content that was not transmitted by the transmitting station 102. Other suitable transmitting and receiving implementation schemes are available. For example, the receiving station 106 can be a generally stationary personal computer rather than a portable communications device and/or a device including an encoder 400 may also include a decoder 500.
[00106] Further, all or a portion of implementations of the present disclosure can take the form of a computer program product accessible from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can, for example, tangibly contain, store, communicate, or transport the program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or a semiconductor device. Other suitable mediums are also available.
[00107] The above-described embodiments, implementations and aspects have been described in order to allow easy understanding of the present invention and do not limit the present invention. On the contrary, the invention is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope
is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structure as is permitted under the law.
Claims
1. A method, comprising: obtaining a degraded frame from a reconstructed frame of a current frame; adding the degraded frame to a reference frame store; coding a next frame based on the degraded frame; obtaining, using a restoration task, a restored frame from the degraded frame; and coding a frame subsequent to the next frame based on the restored frame.
2. The method of claim 1, further comprising: storing the restored frame in the reference frame store.
3. The method of claim 2. further comprising: replacing the degraded frame by the restored frame in the reference frame store.
4. The method of any one of claims 1 to 3, wherein coding the frame subsequent to the next frame comprises: determining, based on a parameter associated with the restoration task, that the frame subsequent to the next frame is to be coded based on the restored frame.
5. The method of claim 4, wherein the parameter associated with the restoration task identifies a number of frames that are coded using the degraded frame before coding other frames using the restored frame.
6. The method of claim 4, wherein the parameter associated with the restoration task is included in a sequence parameter set of a compressed bitstream.
7. The method of claim 4, wherein the parameter associated with the restoration task is included in a picture parameter set of a compressed bitstream.
8. The method of any one of claims 1 to 7, further comprising: coding in a header of the current frame a syntax element indicating that the restoration task is to be performed with respect to the degraded frame to obtain the restored frame.
9. The method of any one of claims 1 to 8, further comprising: initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame; and in response to a command to output the next frame: outputting the next restored frame; and discarding the next restored frame.
10. The method of any one of claims 1 to 8, further comprising: initiating the restoration task on a next reconstructed frame of the next frame to obtain a next restored frame; and in response to a command to output the next frame before the next restored frame is available: waiting until the next restored frame to become available; and outputting the next restored frame.
11. A method, comprising: decoding a first frame using a reference frame; applying a filtering task to the reference frame to obtain a restored reference frame; and decoding a second frame using the restored reference frame.
12. The method of claim 11, further comprising: applying at least one filtering task to a reconstructed frame to obtain the reference frame.
13. The method of claim 12, wherein applying the at least one filtering task comprises: applying constrained directional enhancement filters to the reconstructed frame.
14. The method of any one of claims 12 to 13, wherein applying the at least one filtering task comprises:
upscaling the reconstructed frame to an original resolution to obtain the reference frame.
15. The method of any one of claims 11 to 14, wherein decoding the first frame using the reference frame comprises: decoding a number of frames that includes the first frame using the reference frame, wherein the number of the frames is based on a parameter associated with the filtering task.
16. A device, comprising: a processor that is configured to perform the method of any one of claims 1-10 or claims 11-15.
17. A device, comprising: a memory; and a processor, the processor configured to execute instructions stored in the memory to perform the method of any one of claims 1-10 or claims 11-15.
18. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising operations that perform the method of any one of claims 1-10 or claims 11-15.
19. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is configured for decoding by the method of any one of claims 1-10 or claims 11-15.
20. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is generated by an encoder performing the method of any one of claims 1-10 or claims 11-15.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263436525P | 2022-12-31 | 2022-12-31 | |
| PCT/US2023/082520 WO2024144978A1 (en) | 2022-12-31 | 2023-12-05 | Delayed frame filtering in in-loop filtering |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4627794A1 true EP4627794A1 (en) | 2025-10-08 |
Family
ID=89508853
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23837514.1A Pending EP4627794A1 (en) | 2022-12-31 | 2023-12-05 | Delayed frame filtering in in-loop filtering |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4627794A1 (en) |
| WO (1) | WO2024144978A1 (en) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021025080A1 (en) * | 2019-08-07 | 2021-02-11 | パナソニック インテレクチュアル プロパティ コーポレーション オブ アメリカ | Encoding device, decoding device, encoding method, and decoding method |
| WO2021193775A1 (en) * | 2020-03-27 | 2021-09-30 | パナソニック インテレクチュアル プロパティ コーポレーション オブ アメリカ | Coding device, decoding device, coding method, and decoding method |
-
2023
- 2023-12-05 EP EP23837514.1A patent/EP4627794A1/en active Pending
- 2023-12-05 WO PCT/US2023/082520 patent/WO2024144978A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024144978A1 (en) | 2024-07-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109983770B (en) | Multi-level compound forecast | |
| US10798408B2 (en) | Last frame motion vector partitioning | |
| US12425636B2 (en) | Segmentation-based parameterized motion models | |
| US10798402B2 (en) | Same frame motion estimation and compensation | |
| US10506240B2 (en) | Smart reordering in recursive block partitioning for advanced intra prediction in video coding | |
| US10009622B1 (en) | Video coding with degradation of residuals | |
| US10382767B2 (en) | Video coding using frame rotation | |
| US10225573B1 (en) | Video coding using parameterized motion models | |
| WO2019036080A1 (en) | Constrained motion field estimation for inter prediction | |
| US11095890B2 (en) | Memory-efficient filtering approach for image and video coding | |
| US10491923B2 (en) | Directional deblocking filter | |
| US12470732B2 (en) | Palette mode coding with designated bit depth precision | |
| EP4627794A1 (en) | Delayed frame filtering in in-loop filtering | |
| US20250142050A1 (en) | Overlapped Filtering For Temporally Interpolated Prediction Blocks | |
| WO2024072438A1 (en) | Motion vector candidate signaling | |
| CN121533014A (en) | Frame-level nonlinear motion offset in video bitmap processing | |
| EP4584956A1 (en) | Inter-prediction with filtering | |
| WO2025049689A1 (en) | Projected motion field hole filling for motion vector referencing |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250630 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |