EP4732539A2 - Coding techniques and metadata for video communications using generative face video - Google Patents
Coding techniques and metadata for video communications using generative face videoInfo
- Publication number
- EP4732539A2 EP4732539A2 EP24743096.0A EP24743096A EP4732539A2 EP 4732539 A2 EP4732539 A2 EP 4732539A2 EP 24743096 A EP24743096 A EP 24743096A EP 4732539 A2 EP4732539 A2 EP 4732539A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- picture
- gfv
- face
- video
- features
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/46—Embedding additional information in the video signal during the compression process
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/587—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal sub-sampling or interpolation, e.g. decimation or subsequent interpolation of pictures in a video sequence
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/59—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial sub-sampling or interpolation, e.g. alteration of picture size or resolution
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/478—Supplemental services, e.g. displaying phone caller identification, shopping application
- H04N21/4788—Supplemental services, e.g. displaying phone caller identification, shopping application communicating with other users, e.g. chatting
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/816—Monomedia components thereof involving special video data, e.g 3D video
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/84—Generation or processing of descriptive data, e.g. content descriptors
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- General Engineering & Computer Science (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Methods, systems, and metadata for video communications using generative face video (GFV) are described. In encoder, a GFV bitstream is generating comprising multiplexed coded face video pictures and GFV metadata. Using supplemental enhancement information (SEI), a GFV SEI message comprises syntax elements describing face features and at least one or more of: presence of a single or multiple faces, spatial sampling, temporal sampling, primary code picture characteristics and driving-picture handling, background handling, persistence of SEI, and compression parameters for face features. In a decoder, the decoder combines information extracted from the GFV metadata and the decoded face video pictures to generate a reconstructed output video.
Description
CODING TECHNIQUES AND METADATA FOR VIDEO COMMUNICATIONS USING GENERATIVE FACE VIDEO CROSS-REFERENCE TO RELATED APPLICATIONS [0001] This application claims the benefit of priority from U.S. Provisional Patent Application Ser. No.63/509,119, filed on June 20, 2023, U.S. Provisional Patent Application Ser. No.63/511,827, filed on July 3, 2023, U.S. Provisional Patent Application Ser. No.63/587,703, filed on Oct 3, 2023, and U.S. Provisional Patent Application Ser. No.63/572,783, filed on April 1, 2024. TECHNOLOGY [0002] The present invention relates generally to video communications. More particularly, embodiments of the present invention relate to coding techniques and metadata for video communications using generative video, including generative face video (GFV). BACKGROUND [0003] In recent years, video conferencing has seen a tremendous growth due to constraints in face-to-face meetings and increased demand for remote working. Recent developments on generative face video (GFV) allow for a more efficient digital communication among parties by replacing the traditional pixel-based representation of images and video by a feature set, which when received by a suitable decoder, can generate facial images at ultra-low bit rates. [0004] To improve existing communication techniques using generative face video, as appreciated by the inventors here, improved coding techniques and metadata are developed. [0005] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, issues identified with respect to one or more approaches should not assume to have been recognized in any prior art on the basis of this section, unless otherwise indicated.
BRIEF DESCRIPTION OF THE DRAWINGS [0006] An embodiment of the present invention is illustrated by way of example, and not in way by limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which: [0007] FIG.1 depicts an example process for generating generative feature video (GFV) according to prior art; [0008] FIG.2 depicts an encoding and decoding process for video communications using GFV according to a first example embodiment of the present invention; [0009] FIG.3 depicts an encoding and decoding process for video communications using GFV according to a second example embodiment of the present invention; [00010] FIG.4 depicts an example process for feature compression in GFV communications according to an example embodiment of the present invention; [00011] FIG.5 depicts an example process using GFV-related messaging to facilitate machine analysis according to an example embodiment of the present invention; and [00012] FIG.6 depicts an example process using GFV-related messaging to facilitate video improvements or enhancements according to an example embodiment of the present invention. DESCRIPTION OF EXAMPLE EMBODIMENTS [00013] Methods and metadata for video communications using generative face video techniques are described herein. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well- known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily occluding, obscuring, or obfuscating the present invention.
SUMMARY [00014] Example embodiments described herein relate to methods and metadata for video communications using generative face video (GFV) techniques. In an embodiment, a processor receives a sequence of pictures in a video sequence. Then, for a current picture in the sequence of video pictures, the processor generates a GFV bitstream comprising multiplexed coded face video pictures and GFV metadata. Using supplemental enhancement information (SEI), a GFV SEI message comprises syntax elements describing face features and at least one or more of: presence of a single or multiple faces, spatial sampling, temporal sampling, primary code picture characteristics and driving-picture handling, background handling, persistence of SEI, and compression parameters for the face features. In a decoder, the decoder combines information extracted from the GFV metadata and the decoded face video pictures to generate a reconstructed output video. REVIEW OF GENERATIVE FACE VIDEO METHODS [00015] Example embodiments described herein focus on generative face video methods based on the first order motion model for image animation (FOMM) described in Ref.[1] and its derivatives (Refs.[2-5]). In all these methods, a key problem is how to estimate the temporal motion in video coding using facial features. As depicted in FIG.1, the inputs of this process are a key frame (a reference frame) (102) and a driving frame (104). The output is a generative driving frame (target frame) (127). The algorithm can be summarized in four modules as depicted in FIG. 1: 1) feature extraction (105): The feature extraction module is applied to both the key frame and the driving frame. The features can be face landmarks, 2D or 3D key-points, region matrices, compact features, face semantics, etc., to be discussed in more details in a later section. 2) motion estimation (110): The motion estimation module is to estimate (coarse) motion based on key-frame features and driving frame-features, and/or the key frame.
3) input preparation for the Generator (115): The estimated motion combined with the key frame (102) can be used to predict a dense motion field, a warp key frame or other features, create occlusion masks, and the like, which are the input to the Generator (120). 4) Generator (image generation) (120): The Generator is used to output the generative driving frame (127), e.g., using a GAN (Generative Adversarial Network). In essence, the facial features serve as a compact motion description of the face along the time domain. [00016] The process of FIG.1 implies that the driving frame (104) can be generated by the key frame and the features of the key frame and the driving frame. Therefore, when using a generative face video framework for an ultra-low bit rate face video communication, what needs to be sent in the bitstream are: 1) a key frame 2) features for the driving frame. The features for the key frame can be derived at the decoder. The compactness of the extracted features impacts the bit rate for both the metadata and the video key bitstream. One can potentially compress the features of the driving frame to further improve the efficiency. It also affects the decoding complexity of the decoder reconstruction algorithm. In general, the facial features and the reconstruction algorithm are tightly coupled and affect the final R-D performance. [00017] Before examining specific embodiments of this invention, the next section provides a short overview of the most common generative facial features. Commonly used facial reconstruction algorithms using those facial features are discussed as well. Description of Facial features [00018] Landmarks are used to identify and represent key parts of a human body part. For example, for the face region, there are landmarks for the nose, eyebrow, mouth, or eye corners. The landmarks are usually the 2D or 3D coordinates of their positions in the image plane. [00019] Facial landmarks are used to identify and represent key parts of a human face, such as the mouth, right eyebrow, left eyebrow, right eye, left eye, nose and jaw. Standard facial datasets provide annotations of 68 x and y coordinates that indicate
the most important points on a person’s face (e.g., as supported by the Dlib library in dlib.net). Newer datasets and algorithms leverage a dense “face mesh” with over 468 3D face landmarks (e.g., as supported by the MediaPipe library in https://developers.google.com/mediapipe). [00020] In some use cases for video conferencing, hand gestures are also important to convey information, thus a variety of hand landmark areas are also defined (e.g., wrist, index_finger_tip, etc.). [00021] In some use cases for video conferencing, pose landmarks (such as: nose, left_elbow, and the like) can be used to exchange the body language. Without limitation, example embodiments presented here will focus mostly on face landmarks. [00022] Facial keypoints are the vital areas in the face from which a person's facial expressions, and, therefore, emotions can be evaluated. The term “keypoint” is used to denote the location of key object parts in an image. For example, it defines spatial locations or points that stand out in an image, like key parts of our faces (nose tip, eyebrow, lips) or key points of our body (joints, hips, elbow). Additionally, detection of keypoints might be task-dependent and can potentially be learned in a self- supervised manner. Keypoints learned in a self-supervised manner might also lie in parts of a face that are not necessarily key parts (Ref.[1]). In recent developments, detecting of keypoints is often done by a neural network. The keypoint representation acts as a bottleneck resulting in a compact motion representation. One can represent the location of keypoints via 2D coordinates (Ref.[1]) or 3D coordinates (Ref.[3]). [00023] One example of 2D keypoints is defined in Kaggle as follows: - Each predicted keypoint is specified by an (x, y) real-valued pair in the space of pixel indices. There are 15 keypoints, which represent the following elements of the face: left_eye_center, right_eye_center, left_eye_inner_corner, left_eye_outer_corner, right_eye_inner_corner, right_eye_outer_corner, left_eyebrow_inner_end, left_eyebrow_outer_end, right_eyebrow_inner_end, right_eyebrow_outer_end, nose_tip, mouth_left_corner, mouth_right_corner, mouth_center_top_lip, mouth_center_bottom_lip [00024] Using 3D keypoints (or point clouds) has potential to achieve better accuracy than its 2D counterpart when handling larger head poses, such as rotation and translation. Sometimes it is beneficial to fuse 2D keypoints and 3D keypoints to have a better accuracy and complexity tradeoff (Ref.[21]). Landmarks can be considered as a special type of keypoints with semantic meaning.
[00025] An affine transform in each local neighborhood is used to describe the local motion. A simple 2x2 affine transform (Jacobian matrix) can be used for representation (Ref.[1]). The affine transform can be applied globally (i.e., one affine transform per image) or can be associated with each of the keypoints (Ref.[1]). [00026] In some cases, instead of using keypoints, one may use region-based features. Multiple types of matrices, such as a shift matrix (1x2), covariance matrix (2x2), and affine transform matrix (2x2) can be used as region matrices (Ref.[2]). For face generation, the number of regions can be pre-defined, e.g., 10 (Ref.[2]). [00027] Neural network (such as U-Net in Ref.[4]) learned features are generally learned in an end-to-end manner and can represent temporal motion very sparsely. The features can be further compressed/compacted at the corresponding size (such as 4x4 in Ref.[4]) using a learned-based compression scheme (Auto Encoder) (Ref.[22]). The compact features can be packed in a matrix (such as 4x4 in Ref.[4]). [00028] Intrinsic visual representations or facial semantics have an advantage to represent facial expressions explicitly. In Ref[5], facial semantics are represented using 14 parameters, as shown in Table 1. Table 1. Definition of facial semantics in Ref.[5]
[00029] Ref.[5] relies on a 3D Morphable Model (3DMM) model for 3D face reconstruction. The deep learning-based methods have greatly improved the 3D face mesh reconstruction performance by directly regressing 3DMM coefficients from the input and transforming the face template to reconstruct the corresponding face meshes in a supervised or unsupervised manner, which significantly advances generative face video development. [00030] The Facial Action Coding System (FACS) is a system to taxonomize human facial movements by their appearance on the face, based on a system
originally developed by Swedish anatomist Carl-Herman Hjortsjö. Using the FACS, human coders can manually code nearly any anatomically possible facial expression, deconstructing it into the specific "action units" (AU) and their temporal segments that produced the expression. As AUs are independent of any interpretation, they can be used for any higher order decision making process, including recognition of basic emotions or pre-programmed commands for an ambient intelligent environment. [00031] For clarification, the FACS is an index of facial expressions, but does not actually provide any bio-mechanical information about the degree of muscle activation. Though muscle activation is not part of the FACS, the main muscles involved in the facial expression have been added here. [00032] Action units (AUs) are the fundamental actions of individual muscles or groups of muscles. Action descriptors (ADs) are unitary movements that may involve the actions of several muscle groups (e.g., a forward‐thrusting movement of the jaw). The muscular basis for these actions has not been specified and specific behaviors have not been distinguished as precisely as for the AUs. [00033] Intensities of the FACS are annotated by appending letters A–E (for minimal to maximal intensity) to the action unit number (e.g., AU 1A is the weakest trace of AU 1, and AU 1E is the maximum intensity possible for the individual person). A- Trace B- Slight C- Marked or pronounced D- Severe or extreme E- Maximum [00034] There are other modifiers present in FACS codes for emotional expressions, such as "R," which represents an action that occurs on the right side of the face, and "L," for actions which occur on the left side. An action which is unilateral (occurs on only one side of the face), but has no specific side, is indicated with a "U," and an action which is bilateral, but has a stronger side, is indicated with an "A" for asymmetric. List of AUs and ADs (~100 codes)
• Main codes • Head movement codes • Eye movement codes • Visibility codes • Gross behavior codes [00035] Segmentation maps could be generated for well-defined regions of the face. For instance, maps could be made that are labeled for 15 categories (eyes, hairs, ears, etc.) [00036] MRAA (Ref.[2]) proposes to use meaningful object parts such as torso, upper arms, lower arms and so on. This might have variable number of pixels. A potential option might be to normalize the number of pixels for each region so that it can be deterministic. [00037] Ref.[3] includes 12-dimensional head transformation parameters – a 3x3 rotation matrix, and 3x1 translation matrix. The head pose could be defined using 3 values pertaining to the parameters that define the yaw, roll and pitch of the 3d head pose. [00038] In addition to image/video features, some techniques use the speech signal for generating talking head. It is noted that some of the above features can be possibly merged into one category. Summary of Reconstruction Algorithms [00039] Almost all the techniques use a generator network to reconstruct the final output image. The FOMM method (Ref.[1]) uses a dense motion network that takes as input the facial landmarks from the source/reference frame and the driving frames (other frames), along with the source frame to produce a dense motion field and an occlusion map. These are then provided to a generation module for additional processing. For the motion module, Ref.[1] employs an architecture based on U-Net (Ref.[18]) with five conv3x3 – bn – relu – avg - pool2x2 blocks in the encoders and five upsample2x2 - conv3x3 – bn – relu blocks in the decoders. [00040] The generation module takes as input the source frame, the dense motion field, and an occlusion map. The dense motion field is used to align/warp the feature
maps computed from the source frame with the object pose in the driving frame. The feature map is obtained from the source frame after two down-sampling convolutional blocks. The warped feature map is then multiplied with the occlusion map to mask out the feature map regions that cannot be recovered using image warping, and thus should be in-painted. For the generator network, FOMM uses the Johnson architecture (Ref.[19]) with two down-sampling blocks, six residual-blocks and two up-sampling blocks. [00041] Refs.[2,9,10] use a similar reconstruction mechanism as Ref.[1]. Ref.[8] uses a slightly different mechanism in the dense motion network module. It does not use information about the affine transform matrix of the keypoints during predicting the flow and the occlusion map. Instead, it relies just on the key points. [00042] Like Ref.[8], Ref.[7] also does not use the affine transform matrix during the flow estimation. Furthermore, it modifies the FOMM architecture by adding SPADE (spatially adaptive) (Ref.[20]) normalization layers in the up-sampling blocks of the decoder network. They use facial keypoints and draw polygons for eyes, eyebrows, lips and inner mouth and use those as semantic maps for SPADE. The generator network consists of a stack of five residual blocks and three up-sampling blocks that apply the SPADE normalization. [00043] Similar to Ref.[1], Ref.[3] also uses the source and driving keypoints to estimate optical flows which are used to warp the source feature. This is fed to the motion estimation network to produce a flow composition mask. This is then used to warp the 3D source feature. The generator then converts the warped feature to the output image. [00044] In addition to the motion feature map (similar to Ref.[1]), Ref.[12] uses an attention map and speech as additional features. This is passed to a generator module in addition to the source frame to create the output. The generator module consists of an identity-encoder (where the source frame is fed to the encoder), which is different from Ref.[1]. VIDEO CODING BASED ON GENERATIVE FACE VIDEO [00045] In an embodiment, a general video compression scheme using generative face video has at least the following requirements:
1) The key frame needs to be compressed using an image or video codec. For a video codec, the key frame is generally coded as an intra frame. The face features can be derived at the decoder to save some bits because the decoded image is available at the decoder. To save the face reconstruction complexity, the key frame face features can also be extracted at the encoder and sent to the decoder in the bitstream. It is noted that the key frame can be coded as an Inter frame. When coded as inter, the reference picture needs to be carefully handled to make sure it exists in the decoder picture buffer (DPB) when inter prediction is applied. For example, the reference picture could be set as a long-term reference picture. 2) For the driving frame, the basic idea is to only send face features into the coded bitstream. Then, the decoder decodes the face features and generates the driving frame via the Generator block while using information from the Input Generation for the Generator block. The face features are carried in High Level Syntax (HLS) in video codecs, such as the video parameter set (VPS), the sequence parameter set (SPS), the picture parameter set (PPS), a picture header (PH), supplemental enhancement information (SEI), or other means of metadata. In a proposed example embodiment, without limitation, the face features are carried in a novel SEI message. In some cases, to retain face details (such as wrinkles, which is different from the key frame), or to cover the background occlusion part due to head pose movement or camera motion, or to cover the appearance of new objects, it is proposed to code a driving frame with much lower bitrate using inter frame coding, or just code the missing details. For example, one can apply some preprocessing scheme: e.g., blur the part which can be generated by the Generator and leave only the part which cannot be covered by Generator for compression. In another method, one can simply code the residue between the original image and the generated image. At the decoder, the generative driving frame can be fused together with the decoded driving frame, using simple addition and filtering or some more advance neural-network (NN) approach. This has another advantage in that if the SEI message is lost or not understood by the decoder, the decoder can still output some meaningful video. It is noted that in a video codec standard, such as AVC, HEVC, VVC, and the like, for any access unit, an SEI message must be bundled together with the primary coded picture (e.g., see structure of an access unit in AVC in Figure 7.1 in Ref.[25]). Therefore, in case it is decided not to code the driving frame, one still needs to encode a dummy picture as a primary
coded picture. The decoder needs to be informed via SEI messaging to neither display nor use the dummy picture for any processing. [00046] To further reduce the bitrate, in an embodiment, it is proposed to use the spatial and/or temporal downsampling scheme in the overall coding framework. In addition, it can reduce the complexity at the decoder for generative face video reconstruction. In a first example embodiment, one can downsample (spatially and/or temporally) the video first, then use the proposed generative face video compression scheme. In this case, both the coded video and extracted face features are downsampled spatially and/or temporally. At the decoder side, one can upsample the reconstructed video to the original spatial resolution and/or framerate. The downsampling and upsampling processes can use a neural network method and possibly can be absorbed in the neural network for feature generation and generator reconstruction. [00047] In a second example embodiment, one only spatially downsamples the video for video compression. The face feature is extracted in its original resolution. Temporal downsampling can still be applied to both video and the face features. At the decoder, the decoded video is upsampled spatially and temporally (if the driving frame is coded and temporally downsampled). The temporal downsampled face features are interpolated for full frame rate. Then the driving frame is reconstructed using the generator at full resolution and/or full frame rate. Alternatively, instead of first temporally interpolating features and then generating a full framerate driving frame, one could also first generate the driving frame and then interpolate a full framerate driving frame using a frame interpolation method. [00048] FIG.2 depicts an example framework for the first example embodiment, comprising an encoder and a decoder, where both video compression and feature extraction are at the same resolution and frame rate. As depicted in FIG.2, at the encoder side, there are two modules to process the input video. • Spatial/temporal down-sampling (205): this module is to perform down- sampling to reduce the spatial resolution and/or temporal frame rate. As explained earlier, the purpose is to reduce the bitrate and possibly decoder complexity. This module is optional. • Face Feature Extractions (210): this module is to extract the features of the face, and/or body/hands. Those extracted features will be compressed and
signaled in the SEI message (222), generated by the GFV SEI generation block (220). [00049] A multiplexer (muxer 225) will multiplex the SEI messaging (222) to the output of the video compression block (215) to generate the coded video bitstream (227). In FIG.2, the term “face video” denotes any of a key frame, a driving frame, or a dummy frame. [00050] At the decoder side, the compressed bit stream and SEI messaging (227) will be demultiplexed by demuxer (230). Three main modules are needed: • Spatial/temporal up-sampling (245): This module will perform up-sampling (e.g., using super-resolution) in both the spatial and/or temporal domain to restore the video frame back to the original resolution and/or frame rate. This module is optional. • GFV SEI Reconstruction (235): This module decompresses the compressed features stored in the SEI message to generate reconstructed face features (237). • Face Reconstruction (240): This module takes the input from the decompressed SEI features (237) and the decoded face video to reconstruct a high quality face video, which, if needed, can be temporally and spatially upsampled to generate the final reconstructed face video (250) at full spatial resolution and/or frame rate. [00051] FIG.3 depicts an example framework for the second example embodiment, wherein the feature extraction (210) is performed in the full resolution video, and the video compression (215) happens in a downsampled resolution after spatial downsampling (207). Temporal downsampling (206) can be applied before both the video compression and feature extractions. [00052] In the decoder, spatial upsampling (247) follows video decoding, while temporal upsampling (246) (if any) follows the reconstruction of the face. Syntax and Semantics of Proposed Metadata [00053] In JVET, two proposals have been made related to generative face video (Refs. [23-24]). These proposed SEI messages only covers the face features. The supported face features are listed in Table 2.
Table 2. Summary of facial representations for generative face video compression algorithms in Refs. [23,24]
[00054] In example embodiments, following the processing pipelines depicted in FIG.2 and FIG.3, a series of novel SEI messages is proposed. Besides the face features proposed by Refs. [23, 24], the new SEI message covers the following aspects: 1) Single face or multiple faces 2) Spatial sampling 3) Temporal sampling 4) Primary coded picture characteristics and driving-picture handling 5) Background handling 6) Persistence of the SEI 7) Face features and compression
8) How it works with a NN post-filter representation (e.g., NNPFC and NNPFA SEI messaging (Ref.[26])) Multiple Face IDs [00055] In a video chat or conference, there might be more than one face in the same video. In Refs. [23,24], a syntax gfv_id is proposed, where gfv_id is used to identify a generative face video filter. gfv_id contains an identifying number that may be used to identify a generative face video filter. The value of gfv_id shall be in the range of 0 to 232 − 2, inclusive. [00056] In example embodiments, two alternative designs are proposed. In the first design, the gfv_id is repurposed to represent both face_id and generative face video filters. In one embodiment, one may use bits from 0 to 7 to identify the generative face video filters and bits from 9 to 15 to identify face IDs. In the second design, one explicitly signals the number of faces using syntax gfv_num_faces_minus1 and loops over the number of faces using face_id inside one generative face video SEI message. For example, in a frame with two faces (say, person1 and person2) gfv_num_faces_minus_1 is set to 1 and one can assign to person1 face_id equal to 0 and to person2 face_id equal to 1. One also needs to specify the order the people in the picture. In one embodiment, one can follow a “Z” scanning order: from left to right, and then top to bottom. Thus: gfv_num_faces_minus1 plus 1 specifies the number of faces described in this SEI message. Spatial sampling (gfv_spatial_sampling_flag) [00057] At the encoder, spatial downsampling (205 or 207) of the input signal may be used. In general, a generative face video network is trained in an end-to-end fashion (e.g., from the input of the input video in the encoder to the output of the reconstructed video in the decoder). It is beneficial to specify the input resolution to the generative face video network and the output resolution after the generative face video network. For example, in Ref.[4], the input resolution of the network is
256x256, and the output resolution of the network is 256x256. It is proposed to use a presence flag to signal the spatial sampling, and a presence-resolution flag of input picture resolution and output picture resolution. The resolution can be explicitly signaled by width and height, or implicitly signaled by a scaling ratio for width and height. An example syntax is shown in Table 3. Table 3. Example syntax of spatial sampling parameters
gfv_spatial_sampling_flag equal to 1 indicate that spatial sampling information corresponding generative face video is present. gfv_spatial_sampling_flag equal to 0 indicate that spatial sampling information corresponding generative face video is not present. gfv_input_resolution_present_flag equal to 1 specifies that the syntax gfv_input_pic_width_in_luma_samples and gfv_input_pic_height_in_luma_samples are present. gfv_input_resolution_present_flag equal to 0 specifies that the syntax gfv_input_pic_width_in_luma_samples and gfv_input_pic_height_in_luma_samples are not present. gfv_input_pic_width_in_luma_samples and gfv_input_pic_height_in_luma_samples specify the width and height, respectively,
of the luma sample array of the input picture to the generative face network in the SEI message. When not present, they are inferred to be equal to CroppedWidth and CroppedHeight, which are width and height respectively, of the decoded cropped picture in units of luma samples. The same semantics may be defined for the output resolution parameters. Temporal sampling (gfv_temporal_sampling_flag) [00058] To further reduce bitrate, at the encoder, temporal subsampling (205, 206) (frame rate reduction) can be used to reduce the bitrate to transmit the face features. At the decoder, frame-interpolation can be applied to interpolate the missing pictures. An alternative way is to interpolate the face features of the missing pictures (such as keypoints) and then use GFV to generate the missing picture. [00059] Indication of temporal rate reduction can be implicit or explicit. For implicit signaling, in an example, one may just not code the primary coded picture, so no SEI is available for the missing frame. The decoder knows the missing frame from system level information or other means. For explicit signaling, there are two example methods: 1) when a primary picture is coded as intra random access picture (IRAP), one can use a temporal sub layer to indicate whether GFV SEI is sent for that sub layer. This will signal explicitly if GFV SEI is intentionally missing for a particular picture and not dropped.2) signal a new flag (e.g., gfv_temporal_sampling_flag) in the SEI to tell if feature information is signaled in this SEI. This method will only be explicit if the primary coded picture is present. An example for the explicit signaling method is shown in Table 4. Table 4. Example syntax for signaling temporal sampling
gfv_IRAP_flag shall be equal to 1 when the primary coded picture is an intra random access picture (IRAP) picture. gfv_IRAP_flag shall be equal to 0 when the primary coded picture is not an IRAP picture. gfv_max_sub_layers_minus1 plus 1 specifies the maximum number of temporal sublayer that contains GFV SEI. It is noted that if more flexibility is needed to code a key frame, one can use gfv_key_flag instead of gfv_IRAP_flag. gfv_key_flag shall be equal to 1 when the primary coded picture is key picture and used for face generation for driving picture. gfv_key_flag shall be equal to 0 when the primary coded picture is not a key picture. Primary coded picture characteristics and driving picture handling (gfv_drive_pic_idc) [00060] In example embodiments, there are four types of a primary coded picture: 1) key frame: coded as intra picture; 2) dummy picture: this picture cannot be used for display or other post processing. The driving picture is generated using face features carried in SEI and generated by the Generator.3) driving picture, to be fused with a generative driving picture from the Generator. The coded driving picture can be used to improve face details or handle background change and the like.4) driving picture, to be displayed alone. It is not intended to be fused with a generative driving image. A key frame can be derived from a primary coded picture type or use the gfv_IRAP_flag. For a driving picture, a new syntax element is proposed (gfv_drive_pic_idc ) to differentiate among the three types. gfv_drive_pic_idc indicates the driving picture type as specified in Table 5. Table 5. Example description of gfv drive pic idc values
[00061] In general use cases, by default, a key frame is used to generate a driving frame (picture). In some cases, however, compared to a key frame, a driving frame can serve as a better “key” frame, e.g., face expression, poses, and the like, which are closer to the current frame. Alternatively, a driving frame can work with a key frame for a better face generation, e.g., to be smoother, handling background better, and the like. In this case, one needs to indicate if the generated driving frame needs to replace the decoded one in DPB. The syntax in SEI messaging also needs to signal which frame (uni_pred) or frames (bi_pred), or any number of reference pictures in DPB is used. One can use the delta_poc (difference between picture order count (POC)) between the current frame and the reference frame (ref_frame) to indicate which reference picture is being used. An example is given in Table 6. Table 6. Example syntax for handling a driving picture
gfv_drive_pic_adv_flag equal to 1 indicates the syntax related to the advanced handling of drive picture is present. gfv_drive_pic_adv_flag equal to 0 indicates the syntax related to the advanced handling of drive picture is not present. gfv_drive_pic_ref_flag equal to 1 indicates that generated driving picture might serve as reference picture for the face generation for the following picture in the
decoded order. gfv_drive_pic_ref_flag equal to 0 indicates that generated driving picture does not serve as reference picture for the face generation for the following picture in the decoded order. gfv_bipred_flag equal to 0 specifies that face generation uses one reference picture. gfv_bipred_flag equal to 1 specifies that face generation uses two reference pictures._ delta_poc0 and delta_poc1 specifies the POC difference of current picture and 0th and 1st reference picture, respectively. [00062] To handle backward compatibility with an existing decoder which does not understand the GFV SEI message, or if the GFV SEI message is lost, several methods can be applied. Method 1: mark decoded pictures as non-output pictures [00063] If the current decoded drive picture is a dummy picture or is to be used for the purpose of fusion with a generative picture, one can mark the decoded pictures as non-output pictures to prevent an ordinary decoder from outputting them. Therefore, an ordinary decoder will only output the key pictures. Method 2: VUI signalling change One can re-use the vui_non_packed_constraint_flag. vui_non_packed_constraint_flag equal to 1 specifies that there shall not be any frame packing arrangement SEI messages or any GFV SEI present in the bitstream that apply to the CLVS. vui_non_packed_constraint_flag equal to 0 does not impose such a constraint. Background handling (gfv_bg_flag) [00064] In a video conference, the background may occupy a large portion of an image. Sometimes, even if there is small background motion between frames, or if a person’s head pose changes between frames, it will cause unpleasant artifact, e.g., occlusion. Several ways can be used to handle background change. As discussed
earlier, one can use the coded driving picture to handle the background change. In another method, one can use an affine transformation matrix to predict the background change. The two methods can work together, so one can code less bits in coded driving pictures. In general, the background handling is much simpler than the human’s face and head. It is proposed to explicitly signal the affine transform matrix for background. For example, the gfv_bg_flag flag may be used to signal the presence of the background handling. gfv_bg_flag equal to 1 indicate that syntax related to background features corresponding generative face video is present. gfv_bg_flag equal to 0 syntax related to background features corresponding generative face video is not present. Persistence of GFV SEI [00065] For GFV in communications, it is expected that the features will be updated frame by frame, thus it is appropriate to consider that the persistence of a GFV SEI message is for the current frame only. However, there might be cases where one want to save bits by not sending the updated features or the feature differences from the previous frame are too minor to send. In that case, one either can drop the frame completely using temporal sampling or apply SEI persistence. Facial Features [00066] In this section, after discussing the face features not covered in Refs [23- 24], a revised syntax for facial features will be proposed. Table 2 listed features supported in Refs. [23-24]. It is suggested to add two new features in the GFV SEI. - FACS: Earlier, the facial action coding system was discussed. Using FACS one can manually code nearly any anatomically possible facial expression. This feature is not present in Table 2. - Background features: as explained earlier, there are benefits to list features in the background. The usage for background features can be very different from the face. Thus, it is proposed to separate those features from the facial features. [00067] In addition, a new design of syntax elements for all the features is proposed. In Ref.[24], it mixes all the matrices together using a matrix_idx variable,
as shown in Table 7. Some parameters in the table can work with keypoints, such as the affine translation matrix or the covariance matrix, but some parameters cannot, such as the head rotation and translation matrix or the compact feature matrix. For easier understanding or a clearer presentation and combination, a separate syntax is proposed. Table 7. Definition of matrix_idx (Ref.[24])
[00068] In an embodiment, the features are categorized into six classes: 1) Keypoints; 2) Region matrices; 3) 3D head transformation matrices; 4) Compact features; 5) FACS and facial semantics; and 6) Background features. Furthermore, while Refs. [23-24] coded directly all features using ue(v) (unsigned integer using Exp-Golomb coding) a more efficient coding scheme is proposed that reduces the bitrate of GFV SEI messaging. The proposed compression process is depicted in FIG. 4 and is explained in more detail later. [00069] As depicted in FIG.4, at the encoder side, after optional feature preprocessing (405) there are three main steps: 1) quantization of coded feature parameters (410); 2) entropy reduction of quantized parameters (415) (for example, using differential coding to code feature residuals instead of coding features directly); 3) entropy coding (420) of quantized feature parameters to generate a coded bitstream (422) [00070] At the decoder side, given the coded bitstream (422), the process is reversed: 1) Feature entropy-decoding (425);
2) restoration of entropy reduction (430) (if entropy reduction (415) was applied in the encoder); 3) inverse quantization (435), followed by optional feature postprocessing (440) [00071] In an alternative method, the order of steps 1) (410) and 2) (415) at the encoder can be exchanged. The difference is mainly on whether reduced entropy coding occurs before quantization or after quantization. As depicted in FIG.4, feature preprocessing (405) may be added at the encoder before step 1) (410) to preprocess or reshape the feature data for easier compression. An example of preprocessing is luma- mapping or “image reshaping” (e.g., as luma-mapping chroma scaling (LMCS) in VVC) to improve coding efficiency. [00072] The proposed six major categories and corresponding syntax elements will be discussed next: 1) Keypoints in 2D and 3D Space (gfv_kp_flag) 2) Region matrices in 2D and 3D Space (gfv_rm_flag) 3) 3D head transformation matrices (gfv_ht_flag) 4) Compact Features (gfv_cf_flag) 5) FACS and facial Semantics (gfv_st_flag) 6) Background handling features (gfv_bg_flag) 1) Keypoints in 2D and 3D Space (gfv_kp_flag) [00073] Keypoints can be facial landmarks or other keypoints without semantic meaning. Compared to syntax design in Ref.[24], it is proposed: 1) to be more efficient to code keypoints or to simplify the GFV parsing at the decoder, it is beneficial to signal if the keypoints are facial landmarks or not. 2) for each keypoint, one can optionally support an additional local matrix, such as an affine or covariance matrix. 3) improve keypoints compression to save bits, furthermore, allow 2D and 3D keypoints to be either independently (exclusively) specified or jointly specified. [00074] As an example, only 2D keypoints are considered; however, the techniques can easily be extended to 3D keypoints. For 3D keypoints, one may also need to specify the coordinate system, e.g., using an OMAF (ISO/IEC 23000-20 Omnidirectional Media Format) coordinate system and the like.
[00075] At the encoder, 2D keypoints are specified in grid [-11] x [-11]. Suppose a quantization factor is specified using q_step (q_step is an integer value), the quantization at the encoder for each coordinate is specified as: KP_quant = Round( (KP + 1) * q_step ) Dequantization at the decoder is specified as KP_dec = ((float) KP_quant / (float) q_step) – 1, where KP is the floating point value in [-11], q_step is a positive integer quantization factor, KP_quant is quantized integer value in the range of [0, 2*q_step], and KP_dec is a dequantized floating point value in the range of [-1, 1]. [00076] One can code KP_quant directly with ue(v). Alternatively, one can also apply differential coding or residual coding before entropy coding. At the encoder, one can apply spatial prediction, temporal prediction, or template prediction to generate a residue specified as: KP_residue = KP_quant – ref_quant, where ref_quant can be a quantized spatial predictor, temporal predictor, or template predictor. At the decoder: KP_quant = KP_residue + ref_quant . [00077] For spatial prediction, the previous numbered keypoint for the current picture can serve as a reference (ref_quant) for the current keypoint. Sometimes, one can cluster keypoints in groups. So spatial prediction only happens within groups. For example, if 2D facial landmarks are used for keypoints, using the Dlib library as an example, one can cluster the 68 facial landmarks as 6 groups: 1-17, 18-27, 2-36, 37- 42, 43-48, 49-68. With each group, the first keypoint is either absolutely coded or spatially predicted by a designated KP, for example, for 18, it is specified to use 1 as spatial predictor. For other keypoints in the group, the previous numbering keypoint is used as spatial predictor. [00078] For temporal prediction, in literature (Refs.[3-6]), the features from previously reconstructed pictures are used as reference. The previous picture can be
either the closest temporally decoded picture or adaptively selected from previously decoded picture in the DPB. In a coding standard, for error resilience issue, such scheme is not preferred. It is proposed to use features from the Key frame as temporal predictor/reference: 1) Key frame has source image coded. If it is lost, the GFV system is broken anyway.2) if Key frame SEI is lost, at decoder, one can detect keypoints using a coded image as error concealment. [00079] Besides spatial and temporal prediction, a third method, called template predictor, is proposed. One can predefine the positions of all keypoints in the normalized template, which serves as a reference for the prediction. In one embodiment, one can allow each keypoint or a group of keypoints to adaptively select the reference predictors. For the residue coding, especially when the number of keypoints is relatively large or temporal prediction is used, there might be lots of zeros. Then, one only needs to code the keypoint number which has non-zero residue values. In one embodiment, traditional run-length coding can be applied. In another embodiment, one may code “non-zero residue” keypoint number and residue values separately. [00080] For entropy coding, it is preferred to apply Exp-Golomb coding. Typically, a 0th-order Exp-Golomb-code is used, such as ue(v) and se(v). A residue value can be negative, 0, or positive. One can either code it as se(v) or code the absolute value first. If the absolute value is not equal to zero, then one codes its sign as well. An example of syntax to code keypoints is specified in Table 8. Table 8. Example syntax for coding keypoints
[00081] For the syntax example shown in Table 8, it is assumed that the number of keypoints for facial landmarks is generally larger than the number for other cases. Thus, if facial landmarks are used or the number of keypoints is larger than a pre- defined threshold TH_KP (e.g., 10, or one can add a syntax to signal TH_KP), if it is IRAP picture, one may use spatial residue coding, otherwise, one may apply temporal
residue coding. If keypoints are not facial landmarks, if it is IRAP picture, one uses direct coding, otherwise, one uses temporal residue coding for all keypoints. [00082] It is noted that keypoints are allowed to include both facial landmarks and non-facial landmarks. If standard facial landmarks are used for spatial group predictions, one can put this case as non-facial landmark categories. [00083] It is noted that when scanning a matrix from a 2D to a 1D representation, one uses row-wise scanning. [00084] It is noted that for the quantization factor (q_step), one can select values of power of 2 to replace division with binary shifts. In that case, one only needs to signal log2(q_step) to save more bits. [00085] Example semantics in Table 8 are defined as follows. gfv_kp_flag equal to 1 specifies that the syntax elements related to keypoints are present in the SEI. gfv_kp_flag equal to 0 specifies that the syntax elements related to keypoints are not present in the SEI. gfv_log2_kp_quant_factor specifies the value of keypoint quantization factor qKp.
gfv_facial_landmark_flag equal to 1 specifies that keypoints are facial landmarks. gfv_facial_landmark_flag equal to 0 specifies that keypoints are not facial landmarks. gfv_3D_kp_flag equal to 1 specifies that the keypoints are 3D coordinates. gfv_3D_kp_flag equal to 0 specifies that the keypoints are 2D coordinates. gfv_num_kp_minus1 plus 1 specifies the total number of keypoints. Each keypoints is associated with a numbered identifier in the range of 0 to gfv_num_kp_minus1, identified as kpID. gfv_kp_local_matrix_flag equal to 1 specifies that the keypoint has a local matrix associated. gfv_kp_local_matrix_flag equal to 0 specifies that the keypoint does not have a local matrix associated.
delta_x_coordinator, delta_y_coordinator, delta_z_coordinator indicates that the keypoint x, y, z coordinates are coded as residue. x_coordinator, y_coordinator, z_coordinator indicates that the keypoint x, y, z coordinates are directly coded. gfv_num_kp_update_minus1 plus1 specifies the number of keypoints whose coordinate values are different from the reference. gfv_kp_id[ j ] specifies the kpID number for the j-th updated keypoint. gfv_kp_local_matrix[ i ][ j ] specifies the value of i-th keypoint j-th local matrix coefficient. 2) Region Matrix in 2D and 3D Space (gfv_rm_flag) [00086] Features of each region or segmentation can be represented by a matrix. To save bits, one does not need to code region-segmentation details. The decoder can derive the same region/segmentation information from the key (intra) frame. The improvements compared to Ref.[24] are as follows: 1) Ref.[24] only considered the 2D cases, and 3D cases are now supported. 2) Ref.[24] mixed all matrices together, such as keypoint based, region based, head pose together. Matrices are separated to improve clarity. An example of the proposed syntax is shown in Table 9. Table 9. Example syntax of Region Matrices
Table 10. Definition of rmID
[00087] It is noted for the above example that matrix coefficients are coded directly; however, differential coding can be used as in keypoints. The semantics are defined as follows. gfv_rm_flag equal to 1 specifies that the syntax elements related to Region Matrix are present in the SEI. gfv_rm_flag equal to 0 specifies that the syntax elements related to Region Matrix are not present in the SEI. gfv_log2_rm_quant_factor specifies the value of Region Matrix quantization factor qRm = 2( gfv_log2_rm_quant_factor )
gfv_rm_3D_flag equal to 1 specifies a 3D region matrix. gfv_rm_3D_flag equal to 0 specifies a 2D region matrix. gfv_num_region_minus1 plus 1 specifies the total number of regions. Each region is associated with a numbered identifier in the range of 0 to gfv_num_region_minus1, identified as regionID. gfv_num_rm_minus1 plus 1 specifies the total number of region matrices. Each region matrix is associated with a numbered identifier in the range of 0 to gfv_num_rm_minus1. gfv_rm_id[ k ] specifies the rmID for the k-th region matrix. rmID is defined as in Table 10. gfv_rm_coeff[ i ][ k ][ j ] specifies the value of the i-th region, k-th matrix, j-th coefficient. 3) 3D head transformation matrices (gfv_ht_flag) [00088] 3D head transformation plays an import role in face video.3D head transformation matrices generally include a rotation matrix (3x3) and a translation matrix (1x3), so in total there are 12 parameters.3D head transformation matrices can work with other features. An example syntax is shown in Table 11. Table 11. Example syntax of 3D head transformation matrices
[00089] It is noted for the above example, that coefficients for the 3D head transformation matrices are coded directly. Differential coding can be used as in keypoints. The semantics is defined as follows. gfv_ht_flag equal to 1 specifies that the syntax elements related to 3D Head Transformation Matrices are present in the SEI. gfv_ht_flag equal to 0 specifies that the syntax elements related to 3D Head Transformation matrices are not present in the SEI. gfv_log2_ht_quant_factor specifies the value of 3D Head Transformation quantization factor qHt.
gfv_ht_coeff[ i ] specifies the value of i-th coefficient. 4) Compact Features (gfv_cf_flag) [00090] Compact features generally refer to the learning-based features from neural networks. An alternative term could be “compact learned features.” Compared to Ref.[24], one should allow the flexibility to specify the dimensionality of the compact features and the parameter values. An example syntax is shown in Table 12. Table 12. Example syntax of compact features
[00091] It is noted for the above example, that compact feature parameters are coded directly. Differential coding can be used as in keypoints. The semantics are defined as follows. gfv_cf_flag equal to 1 specifies that the syntax elements related to Compact Features are present in the SEI. gfv_cf_flag equal to 0 specifies that the syntax elements related to Compact Features are not present in the SEI. gfv_log2_cf_quant_factor specifies the value of Compact Feature quantization factor qCf.
gfv_cf_length_minus1 plus 1 specifies the length of compact feature. gfv_cf_coeff[ i ] specifies the value of i-th coefficient. 5) FACS and Facial Semantics (gfv_st_flag) [00092] One may code FACS parameters based on Ref.[4]. There are 99 AU numbers, where some might have multiple sub-AUs such as an M prefix. There are 5 intensity scores: A, B, C, D, E (for minimal-maximal intensity). There are 4 Modifiers: L, R, U, A. A straightforward coding with u(v) is to assign 7 bits for the AU number, 1 bit for M prefix, 3 bits for intensity score, and 2 bits for a modifier. Thus, in total one needs 13 bits to code it. [00093] Facial semantics used in Ref.[24] is considered as a simpler case for FACS, as shown in Table 1. It is based on a 3D Morphable Model (3DMM). For head, it used 1x3 dimension (yaw, roll and pitch) to identify head rotation, and 1x1 dimension to locate face from input image. The representation is different from head pose matrix. When facial semantics is used as in Table 1, gfv_hp_flag shall be set to 0. It is noted the matrix dimension is fixed in Table 1. Thus, there is no need to signal the matrix dimension as in Ref.[24]. An example syntax is shown in Table 13. Table 13. Example syntax of facial semantics (including FACS)
[00094] It is noted for the above example that code facial semantics coefficients are coded directly. Differential coding can be used as in keypoints. The semantics are defined as follows. gfv_st_flag equal to 1 specifies that the syntax related to facial semantics is present in the SEI. gfv_st_flag equal to 0 specifies that the syntax related to facial semantics is not present in the SEI. gfv_log2_st_quant_factor specifies the value of facial semantics quantization factor qSt.
gfv_facs_flag equal to 1 specifies that the syntax related to FACS is present. gfv_facs_flag equal to 0 specifies that the syntax related to FACS is not present. gfc_facs_num_minus1 plus 1 specifies the total number of FACS parameters. gfv_facs_coeff[ i ] specifies the i-th FACS parameter. gfv_fs_coeff[ i ] [ j ] specifies the value of i-th semantics and j-th coeffient.
6) Background handling features (gfv_bg_flag) [00095] For background, it is proposed to use affine background transformation. The dimension is 2x3, thus 6 parameters total. An example syntax is shown in Table 14. Table 14. Example syntax of background matrix
[00096] It is noted for the above example, that the background matrix coefficients are coded directly. Differential coding can be used as in keypoints. The semantics are defined as follows. gfv_bg_flag equal to 1 indicates that syntax related to background features corresponding generative face video is present. gfv_bg_flag equal to 0 indicates that syntax related to background features corresponding generative face video is not present. gfv_log2_bg_quant_factor specifies the value of Background Matrix quantization factor qBg
gfv_bg_coeff[ i ] specifies the value of the i-th coefficient. [00097] It is noted that speech/audio features can also be added into the syntax to support generative face video; however, such syntax is beyond the scope of the proposed embodiments.
Summary of Overall Syntax [00098] Bringing all syntax elements together, Table 15 provides an example of an overall syntax scheme for GFV metadata using SEI messaging. Table 15. Example 1 syntax for GFV SEI messaging
[00099] In another embodiment, Table 16 provides an alternative, simplified version of Table 15. Table 16. Example 2 syntax for GFV SEI messaging
[000100] The generative face video (GFV) SEI message indicates face feature information for source pictures prior to encoding and specifies a neural network, denoted as Generator( ), that may be used to generate novel output pictures using the indicated face feature information and previously decoded output pictures. For example, A Generator( ) may be a generative adversarial network (GAN). [000101] Use of this SEI message requires the definition of the following variables: – Input picture width and height in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively. – Luma sample array keyCroppedYPic and chroma sample arrays keyCroppedCbPic and keyCroppedCrPic for a decoded output picture, denoted as KeyPicture, corresponding to a source key picture. – Luma sample array driveCroppedYPic and chroma sample arrays driveCroppedCbPic and driveCroppedCrPic for a decoded output picture, denoted as DrivePicture, corresponding to a source driving picture. – Bit depth BitDepthY for the luma sample array of the input pictures. – Bit depth BitDepthC for the chroma sample arrays, if any, of the input pictures. – A chroma format indicator, denoted herein by ChromaFormatIdc, as described in subclause 7.3 of H.274 (VSEI) (Ref.[27]).
The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified by Table 2 of H.274 (VSEI) (Ref.[27]). gfv_id contains an identifying number that may be used to identify face feature information and specify a neural network that may be used as Generator( ). The value of gfv_id shall be in the range of 0 to 232–− 2, inclusive. Values of gfv_id from 256 to 511, inclusive, and from 231 to 232–− 2, inclusive, are reserved for future use by ITU- T | ISO/IEC. Decoders conforming to this edition of this document encountering a GFV SEI message with gfv_id in the range of 256 to 511, inclusive, or in the range of 231 to 232–− 2, inclusive, shall ignore the SEI message. NOTE – Different values of gfv_id in different GFV SEI messages could be used to identify different faces when more than one face is present in an output picture, for example. gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_payload_byte[ i ] specify a neural network that may be used as a Generator( ). gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_payload_byte[ i ] have the same syntax and semantics nnpfc_base_flag, nnpfc_mode_idc, nnpfc_reserved_zero_bit_a, nnpfc_tag_uri, nnpfc_uri, nnpfc_payload_byte[ i ], respectively. gfv_key_pic_flag equal to 1 indicates the current decoded output picture corresponds to a key picture. gfv_key_pic_flag equal to 0 indicates the current decoded output picture does not correspond to a key picture. gfv_drive_pic_flag, when present, equal to 1 indicates the current decoded output picture, which corresponds to a driving picture, may be input to Generator( ). gfv_drive_pic_flag equal to 0 indicates the current decoded output picture should not be input to Generator( ). NOTE – A gfv_drive_pic_flag value of 1 could be used to indicate the current decoded output picture could be used to improve face details or handle background changes, as examples.
Face feature information is coded as 2D keypoints specified using ( x , y ) coordinate pairs in the grid [ − 1, + 1] x [ − 1, + 1 ]. gfv_log2_kp_quant_factor specifies a quantization factor qKp determined as follows:
The dequantization function DeqKp( ) is specified as follows: DeqKp( x ) = Clip3( − 1, 1, ( x ÷ qKp ) – 1) gfv_num_kp_minus1 plus 1 specifies the number of 2D keypoints specified in the GFV SEI message. gfv_kp_pred_flag equal to 1 indicates syntax elements gfv_kp_grid_pos_dx[ i ] and gfv_kp_grid_pos_dy[ i ] are present in the GFV SEI message and the values of syntax elements gfv_kp_grid_pos_x[ i ] and gfv_kp_grid_pos_y[ i ] are derived. gfv_kp_pred_flag equal to 0 indicates syntax elements gfv_kp_grid_pos_dx[ i ] and gfv_kp_grid_pos_dy[ i ] are not present and syntax elements gfv_kp_grid_pos_x[ i ] and gfv_kp_grid_pos_y[ i ] are present in the GFV SEI message. gfv_kp_grid_pos_x[ i ] specifies the x component of the ( x , y ) coordinate pair for the i-th 2D keypoint. gfv_kp_grid_pos_y[ i ] specifies the y component of the ( x , y ) coordinate pair for the i-th 2D keypoint. gfv_kp_grid_pos_dx[ i ], when present, specifies the delta-x component of the ( x, y ) coordinate pair for the i-th 2D keypoint. gfv_kp_grid_pos_dy[ i ], when present, specifies the delta-y component of the ( x, y ) coordinate pair for the i-th 2D keypoint. When gfv_kp_pred_flag is equal to 1, gfv_kp_grid_pos_x[ i ] and gfv_kp_grid_pos_y[ i ] are derived as follows: if( gfv_key_pic_flag ) { predX = ( i > 0 ) ? gfv_kp_grid_pos_x[ i – 1 ] : 0 predY = ( i > 0 ) ? gfv_kp_grid_pos_y[ i – 1 ] : 0
gfv_kp_grid_pos_x[ i ] = gfv_kp_grid_pos_dx[ i ] + predX gfv_kp_grid_pos_y[ i ] = gfv_kp_grid_pos_dy[ i ] + predY } else{ predX = KeyKpGridPosX[ i ] predY = KeyKpGridPosY[ i ] gfv_kp_grid_pos_x[ i ] = gfv_kp_grid_pos_dx[ i ] + predX gfv_kp_grid_pos_y[ i ] = gfv_kp_grid_pos_dy[ i ] + predY } The x and y components of the (x , y ) coordinate pair for the i-th 2D keypoint for the key picture, denoted as KeyKpGridPosX[ i ] and KeyKpGridPosY[ i ], respectively, where i = 0..gfv_num_kp_minus1, are specified as follows: if( gfv_key_pic_flag ) { KeyKpGridPosX[ i ] = gfv_kp_grid_pos_x[ i ] KeyKpGridPosY[ i ] = gfv_kp_grid_pos_y[ i ] } The dequantized x component and dequantized y component of the ( x , y ) coordinate pair for the i-th 2D keypoint for the current picture, denoted as deqKpGridPosX[ i ] and deqKpGridPosY[ i ], respectively, where i = 0..gfv_num_kp_minus1, are specified as follows: deqKpGridPosX[ i ] = DeqKp( gfv_kp_grid_pos_x[ i ] ) deqKpGridPosY[ i ] = DeqKp( gfv_kp_grid_pos_y[ i ] ) gfv_kp_local_matrix_flag equal to 1 indicates that syntax element gfv_kp_local_matrix[ i ][ j ] is present in the GFV SEI message. gfv_kp_local_matrix_flag equal to 0 indicates syntax element gfv_kp_local_matrix[ i ][ j ] is not present in the GFV SEI message. gfv_kp_local_matrix[ i ][ j ] specifies the value of j-th local matrix coefficient for the i-th 2D keypoint used to derive a feature array. The dequantized j-th local matrix coefficient for the i-th 2D keypoint, denoted as deqLocalMat[ i ][ j ], is derived as follows: deqLocalMat[ i ][ j ] = DeqKp( gfv_kp_local_matrix[ i ][ j ]) where i = 0..gfv_num_kp_minus1, j = 0..3
The feature array for the current picture, denoted as deqFeature[ i ][ j ], is derived as follows: deqFeature[ i ][ 0 ] = deqKpGridPosX[ i ] deqFeature[ i ][ 1 ] = deqKpGridPosY[ i ] if( gfv_kp_local_matrix_flag ) { deqFeature[ i ][ 2 ] = deqLocalMat[ i ][ 0 ] deqFeature[ i ][ 3 ] = deqLocalMat[ i ][ 1 ] deqFeature[ i ][ 4 ] = deqLocalMat[ i ][ 2 ] deqFeature[ i ][ 5 ] = deqLocalMat[ i ][ 3 ] } else { deqFeature[ i ][ 2 ] = 0 deqFeature[ i ][ 3 ] = 0 deqFeature[ i ][ 4 ] = 0 deqFeature[ i ][ 5 ] = 0 } where i = 0..gfv_num_kp_minus1. The feature array for the key picture, denoted as KeyDeqFeature[ i ][ j ], is derived as follows: if( gfv_key_picture_flag ) { KeyDeqFeature[ i ][ 0 ] = deqFeature[ i ][ 0 ] KeyDeqFeature[ i ][ 1 ] = deqFeature[ i ][ 1 ] KeyDeqFeature[ i ][ 2 ] = deqFeature[ i ][ 2 ] KeyDeqFeature[ i ][ 3 ] = deqFeature[ i ][ 3 ] KeyDeqFeature[ i ][ 4 ] = deqFeature[ i ][ 4 ] KeyDeqFeature[ i ][ 5 ] = deqFeature[ i ][ 5 ] } where i = 0..gfv_num_kp_minus1. [000102] In an embodiment, the following syntax elements could also be part of the SEI message. fdi_type_idc indicates the type of feature descriptor information, as determined by the content provider, as specified, for example, in Table 17.
Table 17. Example description of fdi type idc
fdi_purpose_idc indicates the purpose of feature descriptors, as determined by the content provider, as specified, for example, in Table 18. Table 18. Example description of fdi type idc
Description of inputs and outputs of Generator( ) [000103] Input values to Generator( ) are real numbers and the functions InpY( ) and InpC( ) are specified as follows: InpY( x ) = x ÷ ( ( 1 << BitDepthY ) – 1 ) InpC( x ) = x ÷ ( ( 1 << BitDepthC ) – 1 ) Output values from Generator( ) are real numbers and the functions OutY( ) and OutC( ) are specified as follows: OutY( x ) = Clip3( 0, ( 1 << BitDepthY ) – 1 , x * ( ( 1 << BitDepthY ) – 1 ) OutC( x ) = Clip3( 0, ( 1 << BitDepthC ) – 1 , x * ( ( 1 << BitDepthC ) – 1 ) [000104] When gfv_key_pic_flag is equal to 1, the KeyPicture luma sample array KeyY is derived as follows: KeyY[ x ][ y ] = InpY( keyCroppedYPic[ x ][ y ] ). where x = 0..CroppedWidth − 1, y = 0..CroppedHeight – 1 When gfv_key_pic_flag is equal to 1 and chromaIDC is not equal to 0, the KeyPicture chroma sample arrays KeyCb and KeyCr are derived as follows:
KeyCb[ x ][ y ] = InpC( keyCroppedCbPic[ x ][ y ] ). where x = 0..CroppedWidth / SubWidthC − 1, y = 0..CroppedHeight /SubHeightC − 1 KeyCr[ x ][ y ] = InpC( keyCroppedCrPic[ x ][ y ] ). where x = 0..CroppedWidth / SubWidthC − 1, y = 0..CroppedHeight /SubHeightC − 1 [000105] When gfv_drive_pic_flag is equal to 1, the DrivePicture luma sample array driveY is derived as follows: driveY[ x ][ y ] = InpY( driveCroppedYPic[ x ][ y ] ). where x = 0..CroppedWidth − 1, y = 0..CroppedHeight – 1 When gfv_drive_pic_flag is equal to 1 and chromaIDC is not equal to 0, the DrivePicture chroma sample arrays driveCb and driveCr are derived as follows: driveCb[ x ][ y ] = InpC( driveCroppedCbPic[ x ][ y ] ). where x = 0..CroppedWidth / SubWidthC − 1, y = 0..CroppedHeight /SubHeightC − 1 driveCr[ x ][ y ] = InpC( driveCroppedCrPic[ x ][ y ] ). where x = 0..CroppedWidth / SubWidthC − 1, y = 0..CroppedHeight /SubHeightC − 1 [000106] Input to Generator() are: – When gfv_key_pic_flag is equal to 0 and gfv_drive_pic_flag is equal to 0, deqFeature, KeyDeqFeature, KeyY, KeyCb, KeyCr When gfv_key_pic_flag is equal to 0 and gfv_drive_pic_flag is equal to 1, deqFeature, KeyDeqFeature, KeyY, KeyCb, KeyCr, driveY, driveCb, and driveCr Output of Generator( ) are: – When chromaIDC is equal to 0, a luma sample array genY – When chromaIDC is not equal to 0, a luma sample array genY and chroma sample arrays genCb and genCr. Arrays outYPic[ x ][ y ], outYPic[ x ][ y ], and outYPic[ x ][ y ] are derived as follows: – If gfv_key_pic_flag is equal to 0, the following applies: – outYPic[ x ][ y ] = OutY( genY[ x ][ y ] ). where x = 0..CroppedWidth − 1, y = 0..CroppedHeight – 1 – if chromaIDC is not equal to 0, the following applies: – outCbPic[ x ][ y ] = OutC( genCb[ x ][ y ] ). where x = 0..CroppedWidth / SubWidthC − 1, y = 0..CroppedHeight /SubHeightC − 1
– outCrPic[ x ][ y ] = OutC( genCr[ x ][ y ] ). where x = 0..CroppedWidth / SubWidthC − 1, y = 0..CroppedHeight /SubHeightC − 1 – Otherwise (gfv_key_pic_flag is equal to 1), the following applies: – outYPic[ x ][ y ] = keyCroppedYPic[ x ][ y ] – if chromaIDC is not equal to 0, the following applies: – outCbPic[ x ][ y ] = keyCroppedCbPic[ x ][ y ] – outCrPic[ x ][ y ] = keyCroppedCrPic[ x ][ y ] [000107] In an embodiment, example pseudocode for use of the Generator( ) may comprise: parse and process the GFV SEI message { if the current decoded output picture corresponds to a key picture { store the decoded 2D keypoints representing the face feature information for the key picture; store the current decoded output picture corresponding to the key picture; } else /*the current decoded output picture corresponds to a driving picture*/ { store decoded 2D keypoints representing the face feature information for the driving picture; optionally, store the current decoded output picture corresponding to the key picture; prepare input to Generator( ): decoded picture corresponding to key picture; decoded 2D keypoints for key picture; decoded 2D keypoints for driving picture; optionally, decoded picture corresponding to driving picture apply inputs to Generator( ) to generate an output picture that approximates the driving picture;
store generated output picture } } [000108] Table 15 uses multiple flags to detect whether there is support for the various features. An alternative way is to use bit-masking. This option is easier for the extension and combination of the features. gfv_feature indicates the features as specified in Table 15, where ( gfv_feature & bitmask) not equal to 0 indicates that GFV uses the feature associated with the bitMask value in Table 19. Table 19. Example interpretations of gfv_feature
[000109] It is noted that most of the proposed features, such as keypoints, region matrices, compact features, and the like, may not be limited to the face. Additional features could be used for generative image/video compression. The proposed SEI message can be generalized as a generative video SEI message (GV SEI). A new syntax gv_purpose_idc can be defined to indicate the generative video purpose. gv_purpose_idc indicates the purpose of the GV message as specified, for example, in Table 20. Table 20. Example specifications of gv purpose idc
Relation with Neural-Network Post-Filter SEI Messages [000110] In order to use GFV SEI to generate the generative face video, it may be required to access the neural-network Generator model. For inter-operability, one may need to specify how to access the network model according to a standard. NNPFC SEI (Ref.[26]) provides such capability using syntax element nnpfc_mode_idc. In order to use NNPFC SEI for GFV, the following aspects need to be updated. [000111] In the current proposed amendment to the VSEI (H.274) specification (Ref.[26]), neural-network post-filter SEI messages are designed to enable post filter processing using neural networks. It contains two SEIs: NNPFC ( neural-network post-filter characteristics SEI message ) and NNPFA (neural- network post-filter activation). The neural-network post-filter characteristics (NNPFC) SEI message specifies a neural network that may be used as a post- processing filter. The use of specified neural-network post-processing filters (NNPFs) for specific pictures is indicated with neural-network post-filter activation (NNPFA) SEI messages. [000112] Aspect 1: In NNPFC, there is a syntax nnpfc_purpose. As described in Ref.[26]: “nnpfc_purpose indicates the purpose of the NNPF as specified in Table 20 (copied here as Table 21), where ( nnpfc_purpose & bitMask ) not equal to 0 indicates that the NNPF has the purpose associated with the bitMask value in Table 20. When nnpfc_purpose is greater than 0 and ( nnpfc_purpose & bitMask ) is equal to 0, the purpose associated with the bitMask value is not applicable to the NNPF. When nnpfc_pupose is equal to 0, the NNPF may be used as determined by the application. The value of nnpfc_purpose shall be in the range of 0 to 63, inclusive, in bitstreams conforming to this edition of this document. Values of 64 to 65535, inclusive, for nnpfc_purpose are reserved for future use by ITU-T | ISO/IEC and shall not be present in bitstreams conforming to this edition of this document. Decoders conforming to this edition of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535, inclusive.”
Table 21. Table 20 - Definition of nnpfc_purpose from Ref.[26]
[000113] To enable the usage of generative face video, it is proposed to add one purpose of generative face video for bitMark 0x40, shown in Bold Italics in the following Table. To have more broader usage, as an example, the syntax name “Generative AI” may be used. Table 22. Proposed update of Table 20 – Definition of nnpfc purpose in Ref.[26]
[000114] Aspect 2: NNPFC is designed without considering additional input of facial features. It is proposed to use syntax nnpfc_auxiliary_inp_idc to carry facial feature information from GFV_SEI. From Ref.[26]: “nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensor of the NNPF. nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that auxiliary input data is derived as specified in Formula 85. The value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 1, inclusive, in bitstreams conforming to this edition of this document. Values of 2 to 255, inclusive, for nnpfc_auxiliary_inp_idc are reserved for future use by ITU-T | ISO/IEC and shall not be present in bitstreams conforming to this edition of this document. Decoders
conforming to this edition of this document shall ignore NNPFC SEI messages with nnpfc_auxiliary_inp_idc in the range of 2 to 255, inclusive. Values of nnpfc_auxiliary_inp_idc greater than 255 shall not be present in bitstreams conforming to this edition of this document and are not reserved for future use.” Using bold Italics to denote the proposed updated syntax, the syntax of nnpfc_auxiliary_inp_idc can be updated as: nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that auxiliary input data is derived as specified in Formula 85. nnpfc_auxiliary_inp_idc equal to 2 specifies that auxiliary input data is derived from GFV SEI. The value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 2, inclusive, in bitstreams conforming to this edition of this document. Values of 3 to 255, inclusive, for nnpfc_auxiliary_inp_idc are reserved for future use by ITU-T | ISO/IEC and shall not be present in bitstreams conforming to this edition of this document. Decoders conforming to this edition of this document shall ignore NNPFC SEI messages with nnpfc_auxiliary_inp_idc in the range of 3 to 255, inclusive. Values of nnpfc_auxiliary_inp_idc greater than 255 shall not be present in bitstreams conforming to this edition of this document and are not reserved for future use. [000115] NNPFC SEI also specifies NNPF properties with the nnpfc_property_present_flag, which includes input resolution change and frame rate change, thus, one can use the syntax in NNPFC SEI instead to signal them in GFV SEI. Alternatively, one can implement GFV SEI without relying on NNPFC SEI. Then, one can incorporate nnpfc_mode_idc related syntax and semantics into GFV SEI, as shown in Table 23. Table 23. Example GFV SEI syntax related to the NN Generator network
[000116] Aspect 3: Reuse picture rate up-sampling plus GV SEI messaging as part of neural-networks post processing. [000117] Generative face video, as described in this application, can be seen as a form of picture rate up-sampling, but requiring addition features, such as keypoints, region matrices, and the like. In an example, POC0 and POC30 may be key pictures, while POC1, POC2, … , and POC29 may be marked as interpolated pictures using generative video. A GFV SEI message needs to be associated with POC1, POC2, .. , and POC29. Since there is no primary coded picture associated with the interpolated picture, it is proposed to add syntax element gfv_poc_delta in each GFV SEI message to associate timing with those GFV SEIs. The POC for the generative video using one GFV SEI can be computed as the base_POC + gfv_poc_delta, where base_POC is the picture order count (POC) for the current decoded picture. [000118] Note: while GFV video-related parameters are described herein in terms of SEI messaging, such parameter representation could be easily adapted in alternative metadata and file formats, suitable for alternative video codecs, such as AV1, AVS, and the like. Other Applications of GFV SEI messaging [000119] As described so far, the GFV SEI message (e.g., see Table 15 or Table 16) conveys two types (or parts) of information:
1) coded spatial and temporal metadata to identify and represent key parts of the human body 2) information that identifies driving pictures and corresponding properties in the coded video [000120] For communications using generative video, both parts #1 and #2 are used; however, the information conveyed in part #1 could be used alone for other applications, such as machine analysis (e.g., athlete tracking, lip synch monitoring, conference participant engagements, and the like) and/or for the improvement of coded video (e.g., face-aware super resolution, frame rate upsampling, and the like). [000121] The information conveyed in part #1 could be signaled by a separate SEI, say a “human body feature information” (HBFI) SEI message or could be signaled in the GFV SEI message using conditioning flags so that the driving-picture information (part #2) need not to be signaled. Another alternative is to signal part #2 in a standalone SEI message, say, a “generative driving picture information” (GDPI) SEI message that could be used in coordination with the HBFI SEI message to achieve the full functionality of the GFV SEI message. [000122] For example, the syntax and semantics of the HBFI SEI message can be as described earlier, starting with the section “Facial Features,” and including syntax elements described in Tables 7 to Table 13. (i.e., the parts of the document that relate to coding of landmarks, keypoints, etc.). Details for two example applications are described next. Human body feature information (HBFI) SEI message for machine analysis [000123] FIG.5 depicts an example process using GFV-related messaging to facilitate machine analysis according to an example embodiment. As depicted in FIG. 5, in such a scenario, feature extraction and coding is performed on the encoder side to reduce processing demands at the receiver. While the decoded bitstream could be also displayed, it is not required; the extracted HBFI SEI message can be fed directly to a machine analysis block to generate information of relevance, such as viewer attention detection (e.g., during driving), lip-synch monitoring, security analysis, and the like. Human body feature information (HBFI) SEI message for video enhancement
[000124] FIG.6 depicts an example process using GFV-related messaging to facilitate video improvements or enhancements according to another embodiment. The GFV-related messaging can be applied to improve image quality or enhance picture characteristics, such as frame rate, spatial resolution, depth of field, automated pan-and-scan (tracking individual faces of athletes, etc.), and the like. In an embodiment, a decoder may also include neural-network post filtering (NNPF) signaled by NNPC SEI messages and activated by NNPFA SEI messages. In the NNPF-related use cases, the information conveyed by an HBFI SEI message can be considered auxiliary information, as noted in paragraph “Aspect 2” after Table 22. SEI messaging implementation considerations [000125] During the January 2024 JVET meeting, a proposal for an SEI message for generative face video was adopted as technology under consideration (Refs. [28-29]). In that proposal, a gfv_drive_pic_fusion_flag is defined as: gfv_drive_pic_fusion_flag, when present, equal to 1 indicates the current decoded picture, which corresponds to a driving picture that may be used for fusion, may be input to GenerativeNN( ). gfv_drive_pic_fusion_flag equal to 0 indicates the current decoded picture should not be input to GenerativeNN( ). NOTE 3 – A gfv_drive_pic_fusion_flag value of 1 can be used, for example, to indicate that the current decoded picture can be used to improve face details or handle background changes. NOTE 4 – Fusion takes the three inputs: the base picture, features from keypoints and/or matrices carried in the GFV SEI message, and the current decoded picture, and outputs a picture. NOTE 5 – When current decoded picture corresponds to a driving picture, it should be marked as not for output purpose. [000126] The purpose of this flag is similar to setting gfv_drive_pic_idc, as defined earlier in Table 5, to 1. An issue with NOTE 5 is that in the current VSEI specification (Ref.[27]) there is no clear definition on how to conform to this note. In other words,
consider a fused drive picture; next, it is required to decode the primary coded picture so it can be fused together, which means that in the codec (say, VVC) one needs to set or infer for that picture that ph_pic_output_flag = 1, that is, it needs to be outputted from the decoder but not to be displayed on the display, which is currently not supported by the current VSEI specification. [000127] In HEVC (H.265), a “no_display“ SEI message is defined as:
“The no display SEI message indicates that the current picture should not be displayed.” [000128] Introducing the same SEI message to VSEI should resolve the way to indicate the behaviour described in NOTE 5; moreover, it is also suggested that NOTE 5 should be replaced with the following requirement: “When current decoded picture corresponds to a driving picture, it is required to send a “no_display” SEI message for the associated picture.” [000129] It is noted that in an embodiment the “no_display” message could be nested in another SEI message, e.g.:
or, alternatively, one or more SEI messages could be nested in a “no_display” SEI message, as in:
[000130] In another embodiment, given the proliferation of possible post-processing operations even when a decoded frame is not displayed (e.g., generative fusion, auxiliary input to a neural network, pattern matching for copyright, provenance or authentication, and the like), an alternative version of the no_display SEI message may be expressed as in Table 24: Table 24. Example of no display SEI message with a post-processing field
The syntax parameter no_display_idc indicates that the current picture may be available for post processing, for example, as specified in Table 25. Table 25. Example description of the no display_idc parameter
[000131] Alternatively, for simplicity, syntax parameter no_display_idc can be replaced by a no_display_flag flag with possible values 0 and 1 corresponding to the 0 and 1 values of Table 25. [000132] In Ref. [30], the syntax element gfv_drive_pic_fusion_flag, when present, and when equal to 1, indicates that the current decoded picture, denoted as a driving picture, may be input to a face picture generator neural network, denoted as GenerativeNN( ), to improve background texture and facial details. [000133] In some cases, the driving picture is intended only as input to the GenerativeNN( ) and is not intended to be displayed. For example, a driving picture representing only background texture might be input to the GenerativeNN( ) for fusion with a generated face. In such a scenario, it would be beneficial to indicate that the unfused background texture should not be displayed directly.
[000134] In an embodiment, instead of sending a separate no-display SEI message, as discussed earlier, the no-display flag may be part of the syntax elements of the generative_face_video() message. For example, as depicted in Table 26, a non-display message may be indicated by the gfv_drive_pic_no_display_flag, defined as follows: gfv_drive_pic_no_display_flag, when present, equal to 1 indicates that the current decoded picture, which corresponds to a driving picture, should not be displayed. gfv_drive_pic_no_display_flag equal to 0 indicates that the current decoded picture could be displayed. For example, a gfv_drive_pic_no_display_flag value of 1 can be used to indicate that the current decoded picture should be used only as input to the GenerativeNN( ). Table 26. Example messaging for a no-display flag for a driving picture
[000135] From Ref. [30], the rest of syntax parameters in Table 26 are given by: gfv_base_pic_flag equal to 1 indicates the current decoded output picture corresponds to a base picture. gfv_base_pic_flag equal to 0 indicates the current decoded output picture does not correspond to a base picture or this SEI message does not specify syntax elements for a base picture. When gfv_base_pic_flag is not present, it is inferred to be equal to 0. gfv_drive_pic_fusion_flag, when present, equal to 1 indicates the current decoded picture, which corresponds to a driving picture that may be used for fusion, may be input to GenerativeNN( ). gfv_drive_pic_fusion_flag equal to 0 indicates the current decoded picture should not be input to GenerativeNN( ).
NOTE – A gfv_drive_pic_fusion_flag value of 1 can be used, for example, to indicate that the current decoded picture can be used to improve face details or handle background changes. NOTE – Fusion takes the three inputs: the base picture, features from keypoints and/or matrices carried in the GFV SEI message, and the current decoded picture, and outputs a picture. References Each of these references is included by reference in its entirety. JVET refers to the Joint Video Experts Team by ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29. [1] A. Siarohin, et al., “First order motion model for image animation,” Advances in Neural Information Processing Systems, 2019, 32. [2] A. Siarohin, et al., “Motion representations for articulated animation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2021: 13653-13662, arXiv:2104.11280v1, 22 Apr.2021. [3] T-C. Wang, et al., “One-shot free-view neural talking-head synthesis for video conferencing,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.2021: 10039-10049, arXiv:2011.15126v3, 2 April 2021. [4] B. Chen, et al. “Beyond keypoint coding: Temporal evolution inference with compact feature representation for talking face video compression,” 2022 Data Compression Conference (DCC). IEEE, 2022: 13-22. [5] B. Chen, et al. “Interactive face video coding: a generative compression Framework,” arXiv:2302.09919v1, 20 Feb.2023. [6] D. Feng, et al. “A generative compression framework for low bandwidth video conference,” 2021 IEEE International Conference on Multimedia & Expo Workshops (ICMEW). IEEE, 2021: 1-6. [7] M. Oquab, et al., “Low Bandwidth video-chat compression using deep generative models,” CVPRW 2021. [8] A. Tang, et al., “Generative compression for face video: a hybrid scheme,” ICME 2022, arXiv:2204.10055v2, 26 Apr.2022.
[9] G. Konuko, et al., “Ultra-low bitrate video conferencing using deep image animation, IEEE ICASSP 2021, arXiv:2012.00346v1, 1 Dec.2020. [10] G. Konuko, et al., “A hybrid deep animation codec for low bitrate video conferencing,” IEEE ICIP 2022. [11] M. Agarwal, et al., “Compressing video calls using synthetic talking heads,” BMVC 2022. [12] M. Agarwal, et al., “Audio-visual face reenactment,” WACV 2023, arXiv:2210.02755v1, 6 Oct., 2022. [13] K.R. Prajwal et al., “A lip sync expert is all you need for speech to lip generation in the wild,” ACM MM 2020, arXiv:2008.10010v1, 23 Aug.2020. [14] Y. Zhou, et al., “MakeitTalk: Speaker-aware talking head animation,” SIGGRAPH Asia 2020, arXiv:2004.12992v3, 25 Feb.2021. [15] S. B. Hegde, et al., “Extreme-scale talking-face video upsampling with audio- visual priors,” ACM MM 2022, arXiv: 2208.08118v1, 17 Aug.2022. [16] Y. Guo, et al., “AD_NeRF: Audio driven neural radiance fields for talking head synthesis,” ICCV 2021, arXiv:2103.11078v3, 19 Aug.2021. [17] X. Ji, et al., “Audio-driven emotional video portraits,” CVPR 2021, arXiv:2104.07452v2, 20 May, 2021. [18] O. Ronneberger, et al. “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI, 2015, arXiv:1505.04597v1, 18 May 2015. [19] J. Johnson, et al., “Perceptual losses for real-time style transfer and super resolution,” in ECCV, 2016, arXiv:1603.08155v1, 27 Mar.2016. [20] T. Park, et al., “Semantic image synthesis with spatially-adaptive normalization,” in CVPR 2019, arXiv:1903.07291v2, 5 Nov.2019. [21] J. M. D. Barros, et al., "Fusion of keypoint tracking and facial landmark detection for real-time head pose estimation," 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA, 2018, pp. 2028-2037, doi: 10.1109/WACV.2018.00224. [22] J. Balle’, et al., “Density modeling of images using a generalized normalization transformation,” in International Conference on Learning Representations, 2016, arXiv:1511.06281v4, 29 Feb.2016. [23] B. Chen, et al., “AHG9: Generative face video SEI Message,” doc. no. JVET- AC0088, 29th meeting, by teleconference, Jan.11-20, 2023.
[24] B. Chen, et al. “AHG9: Common SEI Message of Generative Face Video,”, doc. no. JVET-AD0051, 30th meeting, Antalya, Turkey, 21-28 April 2023. [25] H.264, “Infrastructure of audiovisual services – coding of moving video,” ITU-T (08/2021). [26] S. McCarthy, et al., “Additional SEI messages for VSEI (Draft 4),” doc. No. JVET-AD2006, 30th meeting, Antalya, Turkey, May 2023. [27] H.274,”Versatile supplemental enhancement information messages for coded video bitstreams,” ITU-T (05/2022). [28] J. Chen, et al., “AHG9/AHG16: Common text for proposed generative face video SEI message,” doc. No. JVET-AG0203, 33d JVET meeting, by teleconference, 17-26 Jan 2024. [29] S. McCarthy, et al., “Technologies under consideration for future extensions of VSEI (version 3),” output document No. JVET-AG2032, 33d JVET meeting, by teleconference, uploaded March 29, 2024. [30] S. McCarthy, et al., “Technologies under consideration for future extensions of VSEI (version 4),” output document No. JVET-AH2032, 34d JVET meeting, Rennes, FR, uploaded June 5, 2024 EXAMPLE COMPUTER SYSTEM IMPLEMENTATION [000136] Embodiments of the present invention may be implemented with a computer system, systems configured in electronic circuitry and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and/or apparatus that includes one or more of such systems, devices or components. The computer and/or IC may perform, control, or execute instructions related to video coding and metadata using GFV, such as those described herein. The computer and/or IC may compute any of a variety of parameters or values that relate video
coding and metadata using GFV processes described herein. The image and video embodiments may be implemented in hardware, software, firmware and various combinations thereof. [000137] Certain implementations of the invention comprise computer processors which execute software instructions which cause the processors to perform a method of the invention. For example, one or more processors in a display, an encoder, a set top box, a transcoder or the like may implement methods related to video coding and metadata using GFV processes as described above by executing software instructions in a program memory accessible to the processors. The invention may also be provided in the form of a program product. The program product may comprise any tangible and non-transitory medium which carries a set of computer-readable signals comprising instructions which, when executed by a data processor, cause the data processor to execute a method of the invention. Program products according to the invention may be in any of a wide variety of tangible forms. The program product may comprise, for example, physical media such as magnetic data storage media including floppy diskettes, hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAM, or the like. The computer-readable signals on the program product may optionally be compressed or encrypted. [000138] Where a component (e.g. a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including a reference to a "means") should be interpreted as including as equivalents of that component any component which performs the function of the described component (e.g., that is functionally equivalent), including components which are not structurally equivalent to the disclosed structure which performs the function in the illustrated example embodiments of the invention. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS [000139] Example embodiments that relate to video coding and metadata using GFV processes are thus described. In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is the invention, and what is intended by the applicants to be the invention, is the set of claims that issue from this application, in the specific form in
which such claims issue, including any subsequent correction. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
Claims
CLAIMS What is claimed is: 1. A method to generate a video bitstream using generative face video (GFV), the method comprising: receiving a sequence of video pictures; generating a face video picture based on the sequence of video pictures; generating (210) facial features for the face video picture; generating metadata (222) representing the facial features; encoding the face video picture using a video encoder to generate a coded face video picture; and multiplexing the coded face video picture and the metadata to generate a GFV-coded picture.
2. The method of claim 1, further comprising applying spatial and/or temporal downsampling to the sequence of video pictures before generating the face video picture and the facial features.
3. The method of claim 1, wherein the face video picture comprises one of a key frame picture, a driving frame picture, or a dummy frame picture.
4. The method of claim 1, further comprising: applying temporal downsampling (206) on the sequence of video pictures before generating the face video picture and the facial features for the face video picture; and applying spatial downsampling (207) before encoding the face video picture to generate the coded face picture.
5. A method to generate a decoded picture from a generative face video (GFV)-coded bitstream, the method comprising: extracting a GFV-coded picture from the GFV-coded bitstream; generating from the GFV-coded picture a decoded face video picture and metadata representing facial features in the decoded face video picture; and generating an output face picture based on the decoded face video picture and the metadata.
6. The method of claim 5, further comprising: generating an intermediate face picture based on the decoded face video picture and the metadata; and applying spatial and/or temporal upsampling (245) on the intermediate face picture to generate the output face picture.
7. The method of claim 5, further comprising: applying spatial upsampling to the decoded face video picture to generate an upsampled face video picture; generating an intermediate face picture based on the upsampled face video picture and the metadata; and applying temporal upsampling (246) on the intermediate face picture to generate the output face picture.
8. The method of any one of claims 1-7, wherein the metadata comprise supplemental enhancement information (SEI) messaging via a GFV SEI message.
9. The method of claim 8, wherein the GFV SEI message comprises syntax elements describing face features and at least one or more of: presence of a single or multiple faces, spatial sampling, temporal sampling, primary code picture characteristics and driving-picture handling, background handling, persistence of SEI, and compression parameters for the face features.
10. The method of claim 8, wherein the GFV SEI message further comprises parameters representing a neural network post-filtering message.
11. The method of claim 9, wherein the presence of a single or multiple faces is described by a syntax parameter specifying how many faces are described in the GFV SEI message and each described face is associated with a face ID.
12. The method of claim 9, wherein the spatial sampling is described by syntax parameters specifying width and height of input pictures to a GFV network and/or width and heigh of output pictures of the GFV network.
13. The method of claim 9, wherein the temporal sampling is described by parameters comprising one or more of: a syntax parameter indicating whether temporal sampling is enabled or not; a syntax parameter indicating a maximum number of temporal sublayers containing GFV SEI messages; a GFV IRAP flag indicating when a primary coded picture is an Intra-coded picture; or a GFV key flag indicating when a primary coded picture is a key picture.
14. The method of claim 9, wherein the primary code picture characteristics and driving-picture handling is described by parameters comprising one or more of: a syntax parameter indicating whether a driving picture comprises a dummy picture, or a picture intended to be fused or not with a generative picture; a flag indicating whether advanced handling of a driving picture is needed; a flag indicating that a generated driving picture might serve as reference picture for face generation for the following picture in decoded order; a flag indicating that face generation is using one or two reference pictures; or a delta value indicating a picture order count difference between a current picture and a reference picture.
15. The method of claim 9, wherein the background handling is described by parameters comprising a flag indicating that syntax related to background features corresponding generative face video is present or not.
16. The method of claim 9, wherein the persistence of SEI is described by parameters comprising a flag indicating that parameters related to GFV SEI persist for more than one picture.
17. The method of claim 9, wherein the face features are described by parameters describing one or more of: keypoints, region matrices, 3D head transformation matrices, compact features, facial semantics, or background features.
18. The method of claim 17, wherein the keypoints are described by parameters comprising one or more of:
a keypoint quantization factor to improve coding efficiency of the keypoints; a flag indicating whether a keypoint is a facial landmark or not; a flag indicating whether the keypoints are represented in 2D or 3D space; a flag indicating whether a local matrix associated with it or not; residue coordinates for coding the keypoints using differential coding; or indices related to coding the keypoints using differential coding.
19. The method of claim 17, wherein the region matrices are described by parameters comprising one or more of: a region matrix ID; a flag indicating whether a region matrix is represented in 2D or 3D space; a variable indicating a total number of region matrices; an array rmIDs for two or more region matrices; or a coefficient array for each of the region matrices.
20. The method of claim 17, wherein the 3D head transformation matrices are described by parameters comprising one or more of: a quantization factor for the 3D head transformation matrices; or coefficient parameters for the 3D head transformation matrices.
21. The method of claim 17, wherein the compact features are described by parameters comprising: a variable indicating dimensionality of a compact feature; and an array of coefficients for the compact feature.
22. The method of claim 17, wherein the facial semantics are described by parameters comprising one or more of: a flag indicating whether the facial semantics are related to a facial action coding system (FACS) or not; a number indicating a total of FACS semantics. an array of FACS parameters; or an array of FACS coefficients for the FACS parameters.
23. The method of claim 17, wherein the background features are described by parameters comprising an array of background coefficients representing a 2 x 3 transformation matrix for the background.
24. The method of claim 17, wherein support for specific facial features is indicated via bitmasking.
25. The method of claim 5, wherein generating the output face picture further comprises: applying the metadata and the decoded face video picture to a neural-network to generate the output face picture, wherein the neural network is determined from neural-network post filtering metadata received in the GFV-coded bitstream.
26. The method of claim 5, further comprising performing machine analysis of the facial features based on the metadata received in the GFV-coded bitstream.
27. The method of any one of claims 1-26, wherein a neural network model for generative video is communicated to a decoder via a neural-network post-filter characteristics (NNPFC) SEI message and/or a neural-network post-filter activation (NNPFA) message.
28. The method of claim 27, wherein the NNPFC message includes a NNPFC- purpose syntax element to indicate use of generative video.
29. The method of claim 6, wherein temporal upsampling is performed using neural- network-based post processing.
30. The method of any one of claims 1-26, wherein an encoded or decoded picture comprises a human face, a human body, an animal, or animation video.
31. The method of claim 27, wherein processing facial features is replaced by processing human body features or vehicle features.
32. The method of claim 14, wherein if a syntax parameter indicates that a driving picture should not be displayed, then further transmitting to a decoder via metadata a no-display message for the driving picture.
33. The method of claim 32, wherein the no-display message comprises a syntax parameter indicating whether the driving picture is a) not intended for use by post processing, or b) available for use by post processing.
34. The method of claim 32, wherein the no-display message is a separate SEI message from the GFV SEI message.
35. The method of claim 32, wherein the syntax parameter indicating that a driving picture should not be displayed is part of the GFV SEI message.
36. An apparatus comprising a processor and configured to perform the method recited in any one of claims 1-27.
37. A non-transitory computer-readable storage medium having stored thereon computer-executable instruction for executing a method with one or more processors in accordance with any one of claims 1-27.
Applications Claiming Priority (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363509119P | 2023-06-20 | 2023-06-20 | |
| US202363511827P | 2023-07-03 | 2023-07-03 | |
| US202363587703P | 2023-10-03 | 2023-10-03 | |
| US202463572783P | 2024-04-01 | 2024-04-01 | |
| PCT/US2024/034624 WO2024263644A2 (en) | 2023-06-20 | 2024-06-19 | Coding techniques and metadata for video communications using generative face video |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4732539A2 true EP4732539A2 (en) | 2026-04-29 |
Family
ID=91950168
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24743096.0A Pending EP4732539A2 (en) | 2023-06-20 | 2024-06-19 | Coding techniques and metadata for video communications using generative face video |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4732539A2 (en) |
| KR (1) | KR20260027253A (en) |
| CN (1) | CN121587021A (en) |
| WO (1) | WO2024263644A2 (en) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2025149229A1 (en) * | 2024-01-10 | 2025-07-17 | Nokia Technologies Oy | Signaling information for temporal extrapolation |
| WO2025245132A1 (en) * | 2024-05-22 | 2025-11-27 | Bytedance Inc. | Persistence of a generative face video sei message |
| WO2025245309A1 (en) * | 2024-05-23 | 2025-11-27 | Bytedance Inc. | Presence of and relationship between pictures related to generative face video (gfv) supplemental enhancement information (sei) messages |
| WO2025250580A1 (en) * | 2024-05-29 | 2025-12-04 | Bytedance Inc. | Signalling of facial parameter dimensions in a generative face video (gfv) supplemental enhancement information (sei) message |
-
2024
- 2024-06-19 EP EP24743096.0A patent/EP4732539A2/en active Pending
- 2024-06-19 CN CN202480047455.0A patent/CN121587021A/en active Pending
- 2024-06-19 KR KR1020267001650A patent/KR20260027253A/en active Pending
- 2024-06-19 WO PCT/US2024/034624 patent/WO2024263644A2/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024263644A2 (en) | 2024-12-26 |
| WO2024263644A3 (en) | 2025-03-06 |
| KR20260027253A (en) | 2026-02-27 |
| CN121587021A (en) | 2026-02-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4732539A2 (en) | Coding techniques and metadata for video communications using generative face video | |
| Ma et al. | Overview of intelligent video coding: from model-based to learning-based approaches | |
| Chen et al. | Generative face video coding techniques and standardization efforts: A review | |
| JP7806305B2 (en) | Parallel Processing of Image Domains Using Neural Networks: Decoding, Post-Filtering, and RDOQ | |
| CN120019657A (en) | Scalable 3D scene representation using neural field modeling | |
| WO2025072500A1 (en) | Method, apparatus, and medium for visual data processing | |
| WO2024193709A9 (en) | Method, apparatus, and medium for visual data processing | |
| WO2024226920A1 (en) | Syntax for image/video compression with generic codebook-based representation | |
| CN120419185A (en) | Method, apparatus and medium for visual data processing | |
| WO2025149889A1 (en) | Using generative artifical intelligence for decoding media data | |
| WO2025149229A1 (en) | Signaling information for temporal extrapolation | |
| KR20250172613A (en) | SEI messages for generative face videos | |
| TW202416712A (en) | Parallel processing of image regions with neural networks – decoding, post filtering, and rdoq | |
| Ostermann et al. | Natural and synthetic video in MPEG-4 | |
| Chen et al. | Pleno-Generation: A Scalable Generative Face Video Compression Framework with Bandwidth Intelligence | |
| Wang et al. | DSCVC: Deep Screen Content Video Compression | |
| EP4511805A1 (en) | Wavelet coding and decoding of dynamic meshes based on video components and metadata | |
| WO2025131051A1 (en) | Method, apparatus, and medium for visual data processing | |
| US20260012646A1 (en) | Supplemental enhancement information (sei) message for generative face video | |
| US12621494B2 (en) | SEI message for generative face video | |
| WO2025077746A1 (en) | Method, apparatus, and medium for visual data processing | |
| WO2025077744A1 (en) | Method, apparatus, and medium for visual data processing | |
| Lim et al. | Adaptive Patch-Wise Depth Range Linear Scaling Method for MPEG Immersive Video Coding | |
| WO2025217233A1 (en) | Generative face video coding with disentangled background | |
| WO2025137147A1 (en) | Method, apparatus, and medium for visual data processing |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20260115 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |