WO2021248432A1 - Systems and methods for performing motion transfer using a learning model - Google Patents
Systems and methods for performing motion transfer using a learning model Download PDFInfo
- Publication number
- WO2021248432A1 WO2021248432A1 PCT/CN2020/095755 CN2020095755W WO2021248432A1 WO 2021248432 A1 WO2021248432 A1 WO 2021248432A1 CN 2020095755 W CN2020095755 W CN 2020095755W WO 2021248432 A1 WO2021248432 A1 WO 2021248432A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- motion
- learning model
- loss
- movable object
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
- G06T7/246—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
- G06T7/251—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments involving models
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
- G06T7/246—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/0895—Weakly supervised learning, e.g. semi-supervised or self-supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/094—Adversarial learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10016—Video; Image sequence
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20076—Probabilistic image processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30196—Human being; Person
Definitions
- the present disclosure relates to systems and methods for performing motion transfer using a learning model, and more particularly to, systems and methods for synthesizing a motion information of a first image with a static information of a second image using a learning model.
- the motion information may vary if the limb ratio between the target person and source person is different, e.g., an adult has longer arms and legs than a child does. Besides that, the distance between the person and camera would also alter the ratio of the person displayed in the image.
- Embodiments of the disclosure address the above problems by providing methods and systems for synthesizing a motion information of a first image with a static information of a second image using a learning model.
- Embodiments of the disclosure provide a system for performing motion transfer using a learning model.
- An exemplary system may include a communication interface configured to receive a first image including a first movable object and a second image including a second movable object.
- the system may also include at least one processor coupled to the interface.
- the at least one processor may be configured to extract a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extract a first set of static features of the second movable object from the second image using a second encoder of the learning model.
- the at least one processor may also be configured to generate a third image by synthesizing the first set of motion features and the first set of static features.
- Embodiments of the disclosure also provide a method for motion transfer using a learning model.
- An exemplary method may include receiving, by a communication interface, a first image including a first movable object and a second image including a second movable object.
- the method may also include extracting, by at least one processor, a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extracting, by the at least one processor, a first set of static features of the second movable object from the second image using a second encoder of the learning model.
- the method may further include generating, by the at least one processor, a third image by synthesizing the first set of motion features and the first set of static features.
- Embodiments of the disclosure further provide a non-transitory computer-readable medium storing instruction that, when executed by one or more processors, cause the one or more processors to perform a method for motion transfer using a learning model.
- the method may include receiving a first image including a first movable object and a second image including a second movable object.
- the method may also include extracting a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extracting a first set of static features of the second movable object from the second image using a second encoder of the learning model.
- the method may further include generating a third image by synthesizing the first set of motion features and the first set of static features.
- FIG. 1 illustrates a schematic diagram of an exemplary motion transfer system, according to embodiments of the disclosure.
- FIG. 2 illustrates a block diagram of an exemplary motion transfer device, according to embodiments of the disclosure.
- FIG. 3 illustrates a flowchart of an exemplary method for motion transfer, according to embodiments of the disclosure.
- FIG. 4 illustrates a schematic diagram of an exemplary learning model for motion transfer, according to embodiments of the disclosure.
- FIG. 5 illustrates a flowchart of an exemplary method for training the exemplary learning model, according to embodiments of the disclosure.
- FIG. 6 illustrates a schematic diagram illustrating training of an exemplary learning model, according to embodiments of the disclosure.
- FIG. 1 illustrates a schematic diagram of an exemplary motion transfer system 100, according to embodiments of the disclosure.
- motion transfer system 100 is configured to transfer motions from one object to another (e.g., transfer the motion of an object in an image 101 to an object in an image 102 received from user device 160) based on a learning model (e.g., a learning model 105) trained by model training device 120 using training data (e.g., a training image 103, a training image 103’ and a training image 104) .
- the objects may be movable objects, such as persons, animals, robots, and animated characters, etc.
- training image 103 may include/depict the same object as that in training image 103’ (e.g., having similar static information but different motion information)
- training image 104 may include/depict an object different from the one in training image 103 and training image 103’.
- motion transfer system 100 may include components shown in FIG. 1, including a motion transfer device 110, a model training device 120, a training database 140, a database/repository 150, a user device 160, and a network 170 to facilitate communications among the various components.
- motion transfer system 100 may optionally include a display device 130 to display the motion transfer result (e.g., a synthesized image 107) . It is to be contemplated that motion transfer system 100 may include more or less components compared to those shown in FIG. 1.
- motion transfer system 100 may transfer the motion of a first object (e.g., a first human being) included/depicted in image 101 to a second object (e.g., a second human being, being the same or different from the first human being) included/depicted in image 102 using motion transfer device 110.
- a first object e.g., a first human being
- a second object e.g., a second human being, being the same or different from the first human being
- motion transfer device 110 may use a motion feature encoder to extract pose information of the first object (e.g., generate keypoint map (s) representing a probability that a keypoint exists at each pixel and part affinity field of a body part of the first object) in image 101.
- motion transfer device 110 may further use a static feature encoder to extract static information of the second object (e.g., the appearance and environment background) included/depicted in image 102.
- motion transfer device 110 may additionally use an image generator for generating synthesized image 107 using the pose information of the first object and the static information of the second object.
- the motion transfer operations may be performed based on learning model 105, trained by model training devise 120.
- motion transfer system 100 may display the motion transfer result (e.g., synthesized image 107) on display device 130.
- motion transfer system 100 may include only motion transfer device 110, database/repository 150, and optionally display device 130 to perform motion transfer related functions.
- Motion transfer system 100 may optionally include network 170 to facilitate the communication among the various components of motion transfer system 100, such as databases 140 and 150, devices 110, 120 and 160.
- network 170 may be a local area network (LAN) , a wireless network, a personal area network (PAN) , metropolitan area network (MAN) , a wide area network (WAN) , etc.
- LAN local area network
- PAN personal area network
- MAN metropolitan area network
- WAN wide area network
- network 170 may be replaced by wired data communication systems or devices.
- the various components of motion transfer system 100 may be remote from each other or in different locations and be connected through network 170 as shown in FIG. 1.
- certain components of motion transfer system 100 may be located on the same site or inside one device.
- training database 140 may be located on-site with or be part of model training device 120.
- model training device 120 and motion transfer device 110 may be inside the same computer or processing device.
- motion transfer system 100 may store images including a movable object (e.g., a human being, an animal, a machine with different moving parts, or an animated character, etc. ) .
- images 103 and 103’ including/depicting the same object with different motion information.
- an image 104 including/depicting an object different from the object in images 103 and 103’ may also be stored in training database 140.
- target and source images for transferring of motions e.g., images 101 and 102
- database/repository 150 may be stored in database/repository 150.
- the various images may be images captured by user device 160 such as a camera, a smartphone, or any other electronic device with photo capturing functions, etc.
- the images may be created/generated by user device 160 using image processing programs or software, e.g., when the object is an animated character.
- the images can be a frame extracted from an image sequence in a video clip.
- the object included/depicted in each image can be any suitable object capable of moving (i.e., capable of transferring a motion to/or from) such as a robot, a machine, a human being, an animal, etc.
- training database 140 may store training images 103 and 103’ including/depicting the same object, and training image 104 including/depicting a different object.
- training images 103 and 103’ may have similar/the same static information (e.g., objects with same appearance but depict from different angles, and/or different background) , but different motion information (e.g., different pose and/or location information) .
- Training image 104 may have different static information and motion information than either training image103 or 103’.
- training images 103 or 103’ may be used for training learning model 105 based on minimizing a joint loss.
- training image 104 may be used as a support image for further improving the generalization ability of learning model 105 by further adding a support loss to the joint loss.
- learning model 105 may have an architecture that includes multiple sub-networks (e.g., a motion feature encoder, a static feature encoder and an image generator) .
- Each sub-network may include multiple convolutional blocks, residual blocks and/or transposed convolution blocks for performing functions such as extracting feature vectors (e.g., representing the motion features and/or the static features) and generating images (e.g., synthesizing the motion features and the static features extracted from different images) .
- the motion feature encoder may include a pose estimator (e.g., a pre-trained VGG-19 network) , a keypoint amplifier, and a motion refiner network (e.g., a network having residual blocks) for extracting the motion features.
- the static feature encoder may include convolutional blocks with down-sampling modules, and some residual blocks (e.g., 3 convolutional blocks with down-sampling modules and 5 residual blocks) for extracting the static features.
- the image generator may include residual blocks and transposed convolution blocks (e.g., 4 residual blocks and 3 transposed convolution blocks) for generating the output image (e.g., synthesized image 107) in the same size as the input images (e.g., images 101 and 102) .
- the model training process is performed by model training device 120. It is contemplated that some of the sub-networks of learning model may be pretrained, e.g., ahead of time before the rest parts of the learning model are trained.
- pose estimator 106 may be pretrained either by model training device 120 or by another device and provided to model training device 120.
- model training device 120 may receive pretrained pose estimator 106 through network 107, instead of training it jointly with the rest of learning model 105.
- pose estimator 106 may be trained for extracting human pose information by estimating keypoints of a human body (e.g., the PoseNet vision model) .
- pose estimator 106 may also be trained with specifically designed training set for exacting pose information of living creatures other than a human being (e.g., an animal) , a machine capable of moving (e.g., a robot, a vehicle, etc. ) , or an animated character.
- a human being e.g., an animal
- a machine capable of moving e.g., a robot, a vehicle, etc.
- training a learning model refers to determining one or more parameters of at least one layer of a block in the learning model.
- a convolutional layer of the static feature encoder may include at least one filter or kernel.
- One or more parameters, such as kernel weights, size, shape, and structure, of the at least one filter may be determined by e.g., an adversarial-based training process.
- learning model 105 may be trained based on supervised, semi-supervised, or non-supervised methods.
- motion transfer device 110 may receive learning model 105 from model training device 120.
- Motion transfer device 110 may include a processor and a non-transitory computer-readable medium (not shown) .
- the processor may perform instructions of a motion transfer process stored in the medium.
- Motion transfer device 110 may additionally include input and output interfaces to communicate with database/repository 150, user device 160, network 170 and/or a user interface of display device 130.
- the input interface may be used for selecting an image (e.g., image 101 and/or 102) for motion transfer.
- the output interface may be used for providing the motion transfer result (e.g., synthesized image 107) to display device 130.
- Model training device 120 may communicate with training data base 140 to receive one or more set of training data (e.g., training images 103, 103’ and 104) , and may receive pretrained pose estimator 106 through network 107. Each set of the training data may include training images 103 and 103’ including/depicting the same object with different motion information, and training image 104 including/depicting a different object. Model training device 120 may use each training data set received from training database 140 to train learning model 105 (the training process is described in greater detail in connection with FIGs. 5 and 6 below) . Model training device 120 may be implemented with hardware specially programmed by software that performs the training process. For example, model training device 120 may include a processor and a non-transitory computer-readable medium (not shown) .
- Model training device 120 may additionally include input and output interfaces to communicate with training database 140, network 170, and/or a user interface (not shown) .
- the user interface may be used for selecting sets of training data, adjusting one or more parameters of the training process, selecting or modifying a framework of learning model 105, and/or manually or semi-automatically providing training images.
- motion transfer system 100 may optionally include display 130 for displaying the motion transfer result, e.g., synthesized image 107.
- Display 130 may include a display such as a Liquid Crystal Display (LCD) , a Light Emitting Diode Display (LED) , a plasma display, or any other type of display, and provide a Graphical User Interface (GUI) presented on the display for user input and data depiction.
- the display may include a number of different types of materials, such as plastic or glass, and may be touch-sensitive to receive inputs from the user.
- the display may include a touch-sensitive material that is substantially rigid, such as Gorilla Glass TM , or substantially pliable, such as Willow Glass TM .
- display 130 may be part of motion transfer device 110.
- FIG. 2 illustrates a block diagram of an exemplary motion transfer device 110, according to embodiments of the disclosure.
- motion transfer device 110 may include a communication interface 202, a processor 204, a memory 206, and a storage 208.
- motion transfer device 110 may have different modules in a single device, such as an integrated circuit (IC) chip (e.g., implemented as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA) ) , or separate devices with dedicated functions.
- IC integrated circuit
- ASIC application-specific integrated circuit
- FPGA field-programmable gate array
- one or more components of motion transfer device 110 may be located in a cloud or may be alternatively in a single location (such as inside a mobile device) or distributed locations.
- motion transfer device 110 may be in an integrated device or distributed at different locations but communicate with each other through a network (not shown) . Consistent with the present disclosure, motion transfer device 110 may be configured to synthesize motion information (e.g., motion features of the object) extracted from image 101 with static information (e.g., static features of the object) extracted from image 102, and generate synthesized image 107 as an output.
- motion information e.g., motion features of the object
- static information e.g., static features of the object
- Communication interface 202 may send data to and receive data from components such as database/repository 150, user device 160, model training device 120 and display device 130 via communication cables, a Wireless Local Area Network (WLAN) , a Wide Area Network (WAN) , wireless networks such as radio waves, a cellular network, and/or a local or short-range wireless network (e.g., Bluetooth TM ) , or other communication methods.
- communication interface 202 may include an integrated service digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection.
- ISDN integrated service digital network
- communication interface 202 may include a local area network (LAN) card to provide a data communication connection to a compatible LAN.
- Wireless links can also be implemented by communication interface 202.
- communication interface 202 can send and receive electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.
- communication interface 202 may receive learning network 105 from model training device 120, and images 101 and 102 to be processed from database/repository 150. Communication interface 202 may further provide images 101 and 102 and learning model 105 to memory 206 and/or storage 208 for storage or to processor 204 for processing.
- Processor 204 may include any appropriate type of general-purpose or special-purpose microprocessor, digital signal processor, or microcontroller. Processor 204 may be configured as a separate processor module dedicated to motion transfer, e.g., synthesizing motion information of a first object extracted from one image with static information of a second object extracted from another image using a learning model. Alternatively, processor 204 may be configured as a shared processor module for performing other functions in addition to motion transfer.
- Memory 206 and storage 208 may include any appropriate type of mass storage provided to store any type of information that processor 204 may need to operate.
- Memory 206 and storage 208 may be a volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, or other type of storage device or tangible (i.e., non-transitory) computer-readable medium including, but not limited to, a ROM, a flash memory, a dynamic RAM, and a static RAM.
- Memory 206 and/or storage 208 may be configured to store one or more computer programs that may be executed by processor 204 to perform functions disclosed herein.
- memory 206 and/or storage 208 may be configured to store program (s) that may be executed by processor 204 to transfer motion based on images 101 and 102.
- memory 206 and/or storage 208 may also store intermediate data such as keypoint heatmaps, part affinity fields of body parts of the object, extracted motion features, and extracted static features, etc.
- Memory 206 and/or storage 208 may additionally store various sub-learning models (e.g., sub-networks included in learning model 105) including their model parameters and model configurations, such as pre-trained pose estimator 106 (e.g., a pre-trained VGG-19 network) , the motion feature extracting blocks (e.g., motion feature encoder) including the keypoint amplifier, the motion refiner, the static feature extracting blocks (e.g., and static feature encoder) , and the image generator blocks, etc.
- pre-trained pose estimator 106 e.g., a pre-trained VGG-19 network
- the motion feature extracting blocks e.g., motion feature encoder
- static feature extracting blocks e.g., and static feature encoder
- image generator blocks etc.
- processor 204 may include multiple modules, such as a motion feature extraction unit 240, a static feature extraction unit 242, and an image generation unit 244, and the like. These modules (and any corresponding sub-modules or sub-units) can be hardware units (e.g., portions of an integrated circuit) of processor 204 designed for use with other components or software units implemented by processor 204 through executing at least part of a program.
- the program may be stored on a computer-readable medium, and when executed by processor 204, it may perform one or more functions.
- FIG. 2 shows units 240-244 all within one processor 204, it is contemplated that these units may be distributed among different processors located closely or remotely with each other.
- FIG. 3 illustrates a flowchart of an exemplary method 300 for motion transfer, according to embodiments of the disclosure.
- Method 300 may be implemented by motion transfer device 110 and particularly processor 204 or a separate processor not shown in FIG. 2 using learning model 105.
- Method 300 may include steps S302-S310 as described below. It is to be appreciated that some of the steps may be performed simultaneously, or in a different order than shown in FIG. 3.
- communication interface 202 may receive images 101 and 102 acquired/generated by user device 160 from database/repository 150.
- user device 160 may acquire/generate an image including/depicting an object by using a camera.
- user device 160 may be a smart phone with a camera configured to take pictures or sequence of pictures (e.g., a video clip) .
- the object may be a living creature (e.g., an animal, a human, etc. ) or a machine capable of moving (e.g., a robot, a vehicle, etc. ) .
- User device 160 may also generate the image of the object (e.g., an animated character) using image/photo processing software.
- image 101 and/or 102 may be an image being part of a drawn figure or a sequence of drawn figures (e.g., an animation clip) .
- Database/repository 150 may store the images and transmit the images to communication interface 202 for motion transfer.
- motion feature extraction unit 240 may extract motion features (i.e., pose information and location) of a first object included/depicted in image 101 (also referred to as source image x s ) using a motion feature encoder.
- FIG. 4 illustrates a schematic diagram of exemplary learning model 105 for motion transfer, according to embodiments of the disclosure.
- learning model 105 may include a motion feature encoder 410, a static feature encoder 420, and an image generator 430.
- motion feature encoder 410 may include a pose estimator 412 (e.g., pre-trained pose estimator 106) , a keypoint amplifier 414 and a motion refiner 416.
- pose estimator 412 may extract keypoint heatmaps of the object representing a probability that a keypoint exists at each pixel, and part affinity fields of different body parts of the object showing the orientation of the body part.
- p may have 38 (19 ⁇ 2) channels, and can be a set of 2D vectors indicating the location and orientation by x-y coordinates for each channel of the keypoint heatmap h.
- the generated keypoint heatmap h may only keep the first 18 of total 19 channels and discard the last channel (e.g., the heatmap of the background) .
- Both h and p may be downsampled multiple times in order to reduce the size of the image. For example, both h and p may be downsampled in half for 3 times, resulting in an image 1/2 3 the size of the original input images (e.g., images 101 and 102) .
- the keypoints may correspond to joints of the object, such as elbows, wrists, etc. of a human being.
- keypoint amplifier 414 may denoise the extracted keypoint heatmap values and obtain the amplified the keypoint heatmap h’. For example, keypoint amplifier 414 may apply a softmax function with a relatively small temperature T as the keypoint amplifier to the keypoint heatmap h according to equation (1) :
- T can be set as 0.01 such that the gap between large values and small values in the keypoint heatmap h can be enlarged. This can reduce the effect caused by the noise.
- motion refiner 416 may generate the encoded motion feature vector M (x s ) representing the motion feature of the object based on refining both the part affinity fields p and the amplified keypoint heatmaps h’.
- motion refiner 416 may include 5 residual blocks. Accordingly, the motion features extracted from pose estimator 412 may be refined such that the influence caused by different body part ratios (e.g., limb ratios) and/or camera angles and/or distances can be reduced.
- static feature extraction unit 420 may extract static features S (x t ) (e.g., appearance and environment background) of a second object included in image 102 (also referred to as target image x t ) .
- static feature extraction unit 420 may apply a static encoder 420 for extracting the background, the appearance, etc., of the second object.
- static encoder 420 may include 3 convolutional blocks with down-sampling modules and 5 residual blocks.
- image generation unit 244 may generate synthesized image 107 by synthesizing the motion features M (x s ) extracted from image 101 and the static features S (x t ) extracted from image 102.
- image generation unit 244 may apply an image generator 430 to M (x s ) and S (x t ) according to equation (2) :
- image generator 430 may include 4 residual blocks and 3 transposed convolution blocks such that the output of image generator 430 (e.g., synthesized image 107) may have the same size as that of the input images (e.g., image 101 and/or 102) .
- step S310 the output of image generator 430 (e.g., synthesized image 107) may be transmitted to display device 130 for display.
- image generator 430 e.g., synthesized image 107
- learning model 105 may be trained by model training device 120 before being used by motion transfer device 110 for motion transfer.
- FIG. 5 illustrates a flowchart of an exemplary method 500 for training learning model 105, according to embodiments of the disclosure.
- Method 500 may be implemented by model training device 120 for training learning model 105.
- Method 500 may include steps S502-S518 as described below. It is to be appreciated that some of the steps may be performed simultaneously, or in a different order than shown in FIG. 5.
- learning model 105 may be trained using training images 103 (e.g., including a target object x t ) and 103’ (e.g., including a source object x s ) that include a same object (e.g., target object x t and source object x s being a same movable object in the training images) in the same environment (e.g., same place, same lighting condition, etc. ) with different motion information.
- images 103 and 103’ may be extracted from the same video clip.
- learning model 105 may be trained on premises that the motion features extracted from the synthesized image may be a reconstruction of (e.g., approximately equals to) the motion features extracted from image 103’, i.e.,
- model training device 120 may further adopt a support group during the training, to further improve the generalization ability/performance and stability of learning model 105.
- the support group may include an image (e.g., image 104) depicting an object different from that in images 103 and 103’ to train learning model 105.
- model training device 120 may receive training images 103 and 103’ that include/depict a same object in the same environment with different motion information.
- FIG. 6. illustrates a schematic diagram illustrating training an exemplary learning model 105, according to embodiments of the disclosure.
- image 103 may include a target object x t (e.g., target person x t ) and image 103’ may include a source object x s (e.g., source person x t ) .
- the background of each image is not shown in FIG. 6 for simplification and illustrative purposes.
- images 103 and 103’ are selected such that target person x t and source person x s are the same person in the same environment. Accordingly, images 103 and 103’ may have similar/same static information (e.g., having the same appearance, but different in camera angles for taking the appearance and/or having different backgrounds in images 103 and 103’) , i.e., However, the same person may have different gestures or movements so that images 103 and 103’ may contain different motion information.
- step S504 the motion features M (x s ) of source person x s , and the static features S (x t ) of the target person x t , are extracted from images 103’ and 103 respectively, using motion feature encoder 410 and static feature encoder 420 of learning model 105, similar to steps S304 and S306 in method 300.
- a synthesized image (e.g., including/depicting a synthesized object x syn , synthesized based on the motion features of x s and the static features of x t ) may be generated using image generator 430 of learning model 105, similar to step S308 in method 300.
- step S508 motion features and static features of synthesized object x syn , M (x syn ) and S (x syn ) may be extracted from the synthesized image using motion feature encoder 410 and static feature encoder 420 of learning model 105 respectively, similar to steps S304 and S306 in method 300.
- model training device 120 may implement an adversarial-based training approach.
- model training device 120 may calculate an adversarial loss L adv to discern image 103’ (e.g., including/depicting the source object x s ) and the synthesized image (e.g., including/depicting the synthesized object x syn ) .
- model training device 120 may apply an image discriminator D to discern between the real sample source object x s and the synthesized object x syn , conditioned on the motion features M (x s ) extracted from the source image (image 103’) .
- the adversarial loss can be calculated according to equations (3) , (4) and (5) :
- a discriminator feature matching loss L fm may be calculated.
- the discriminator feature matching loss L fm may be calculated based on a weighted sum of multiple feature losses from each of the different layers of image discriminator D.
- image discriminator D may include 5 different layers and discriminator feature matching loss L fm may be the weighted sum of a L 1 distance between the corresponding features of x s and x syn at each layer of image discriminator D.
- model training device 120 may calculate feature-level consistency losses indicative of a difference between features extracted from the synthesized image (e.g., the motion features and the static features) and the corresponding features extracted from images 103 and 103’. This may insure that the synthesized object (e.g., x syn ) has the same static features of the target object (e.g., x t from image 103) and the same motion features as the source object (e.g., x s from image 103’) .
- model training device 120 may calculate a motion consistency loss L mc indicating a difference (e.g., a L 1 distance) between the motion features extracted from the synthesized image and the motion features extracted from image 103’.
- model training device 120 may calculate a static consistency loss L sc indicating a difference (e.g., a L 1 distance) between the static features extracted from the synthesized image and the static features extracted from image 103.
- L sc a difference (e.g., a L 1 distance) between the static features extracted from the synthesized image and the static features extracted from image 103.
- the motion consistency loss and the static consistency loss can be calculated according to equations (6) and (7) :
- model training device 120 may calculate a perpetual loss L per based on image 103’ and the synthesized image.
- the perpetual loss may be calculated using a pre-trained deep convolutional network for object recognition (e.g., a VGG network) .
- the perpetual loss may be added to the full object to improve the stability and quality of the training.
- model training device 120 may further calculate a support loss based on a support set.
- the support set may include images of different objects as the source object for training, e.g., image 104 including an object different from that of images 103 and 103’. Images in the support set provide many kinds of unseen motions and various static information.
- a support loss L sup may be calculated using the support set (e.g., image 104) as a target image (e.g., including a target object) .
- the synthesized image x syn obtained based on the support set may not be a reconstruction of the source image x s . Accordingly, when calculating the support loss L sup , the ground truth image of the target object performing the motion of the source object is not available. Thus, L + adv , L fm and L per , for calculating the support loss L sup are not available.
- the support loss L sup may include a feature-level consistency loss L mc indicative of a difference between the motion features extracted from the synthesized image and the motion features extracted from source image 103’. In some embodiments, the support loss may further include a feature-level consistency loss L sc indicative of a difference between the static features extracted from the synthesized image and the static features extracted from target image 103. In some embodiments, the support loss may also include a negative adversarial loss L - adv determined based on the image 103’ and the synthesized image. In some embodiments, the support loss L sup may be calculated as a weighted sum of L sc , L mc and L - adv .
- model training device 120 may train learning model 105 by jointly training the sub-networks of learning model 105 (e.g., jointly training keypoint amplifier 414, motion refiner network 416, static feature encoder 420 and image generator 430) based on minimizing the joint loss.
- pre-trained pose estimator 106 may remain the same throughout the optimization process.
- model training device 120 may minimize a joint loss L full that includes some or all of the losses calculated above.
- the joint loss L full may be a weighted sum of L adv , L fm , L per , L mc and L sc .
- the joint loss L full may be calculated according to equation (8) :
- ⁇ adv , ⁇ fm , ⁇ per , ⁇ mc and ⁇ sc are the weights assigned for the respective losses, as calculated in previous steps.
- the weights may be selected to reflect the relative importance of the respective losses. For example, ⁇ adv , ⁇ fm , ⁇ per , ⁇ mc and ⁇ sc may be set to 1, 10, 10, 0.1, 0.01 respectively.
- the support loss L sup calculated in step S518 may be added to the joint loss in order to improve the generalization ability of learning model 105.
- the support loss L sup may be calculated as a weighted sum of L sc , L mc and L - adv according to equation (9) and be added to the joint loss L full of equation (8) :
- ⁇ sc , ⁇ mc and ⁇ adv are the weights for L sc , L mc and L - adv respectively and ⁇ sup represents the weight assigned to support loss L sup when calculating the joint loss L full .
- the weight ⁇ sup can be set to 0.001 while other weights may remain the same as for calculating the overall objective joint loss L full .
- the computer-readable medium may include volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, or other types of computer-readable medium or computer-readable storage devices.
- the computer-readable medium may be the storage device or the memory module having the computer instructions stored thereon, as disclosed.
- the computer-readable medium may be a disc or a flash drive having the computer instructions stored thereon.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Software Systems (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- Data Mining & Analysis (AREA)
- Mathematical Physics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computational Linguistics (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Multimedia (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Image Analysis (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Operations Research (AREA)
- Probability & Statistics with Applications (AREA)
- Algebra (AREA)
Abstract
Embodiments of the disclosure provide systems and methods for performing motion transfer using a learning model. An exemplary system may include a communication interface configured to receive a first image including a first movable object and a second image including a second movable object. The system may also include at least one processor coupled to the communication interface. The at least one processor may be configured to extract a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extract a first set of static features of the second movable object from the second image using a second encoder of the learning model. The at least one processor may also be configured to generate a third image by synthesizing the first set of motion features and the first set of static features.
Description
The present disclosure relates to systems and methods for performing motion transfer using a learning model, and more particularly to, systems and methods for synthesizing a motion information of a first image with a static information of a second image using a learning model.
Recent deep generative models have made great progress in synthesizing images with arbitrary object (e.g., human beings) motions and transferring motions of one object to the others. However, existing approaches require generating skeleton images using pose estimators and image processing operations as an intermediary to form a paired data set with the original images when making the motion transfer. The pose estimator first finds the locations of person keypoints and the image processing operation then connects person keypoints to form a skeleton image. Since the image processing operations, which involve drawing a line between two points, are usually not differentiable, the learning networks used by existing methods cannot be trained in an end-to-end manner. This reduces the availability and compatibility of the model and makes the model impractical in many applications.
Moreover, existing approaches fail to leverage the feature level motion and static information of the real image (s) and synthesized image (s) . This causes the model to generate inaccurate motion information, making the model difficult to generate suitable motions for the target. For example, the motion information may vary if the limb ratio between the target person and source person is different, e.g., an adult has longer arms and legs than a child does. Besides that, the distance between the person and camera would also alter the ratio of the person displayed in the image.
Embodiments of the disclosure address the above problems by providing methods and systems for synthesizing a motion information of a first image with a static information of a second image using a learning model.
SUMMARY
Embodiments of the disclosure provide a system for performing motion transfer using a learning model. An exemplary system may include a communication interface configured to receive a first image including a first movable object and a second image including a second movable object. The system may also include at least one processor coupled to the interface. The at least one processor may be configured to extract a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extract a first set of static features of the second movable object from the second image using a second encoder of the learning model. The at least one processor may also be configured to generate a third image by synthesizing the first set of motion features and the first set of static features.
Embodiments of the disclosure also provide a method for motion transfer using a learning model. An exemplary method may include receiving, by a communication interface, a first image including a first movable object and a second image including a second movable object. The method may also include extracting, by at least one processor, a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extracting, by the at least one processor, a first set of static features of the second movable object from the second image using a second encoder of the learning model. The method may further include generating, by the at least one processor, a third image by synthesizing the first set of motion features and the first set of static features.
Embodiments of the disclosure further provide a non-transitory computer-readable medium storing instruction that, when executed by one or more processors, cause the one or more processors to perform a method for motion transfer using a learning model. The method may include receiving a first image including a first movable object and a second image including a second movable object. The method may also include extracting a first set of motion features of the first movable object from the first image using a first encoder of the learning model and extracting a first set of static features of the second movable object from the second image using a second encoder of the learning model. The method may further include generating a third image by synthesizing the first set of motion features and the first set of static features.
It is to be understood that both the foregoing general descriptions and the following detailed descriptions are exemplary and explanatory only and are not restrictive of the invention, as claimed.
FIG. 1 illustrates a schematic diagram of an exemplary motion transfer system, according to embodiments of the disclosure.
FIG. 2 illustrates a block diagram of an exemplary motion transfer device, according to embodiments of the disclosure.
FIG. 3 illustrates a flowchart of an exemplary method for motion transfer, according to embodiments of the disclosure.
FIG. 4 illustrates a schematic diagram of an exemplary learning model for motion transfer, according to embodiments of the disclosure.
FIG. 5 illustrates a flowchart of an exemplary method for training the exemplary learning model, according to embodiments of the disclosure.
FIG. 6 illustrates a schematic diagram illustrating training of an exemplary learning model, according to embodiments of the disclosure.
Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts.
FIG. 1 illustrates a schematic diagram of an exemplary motion transfer system 100, according to embodiments of the disclosure. Consistent with the present disclosure, motion transfer system 100 is configured to transfer motions from one object to another (e.g., transfer the motion of an object in an image 101 to an object in an image 102 received from user device 160) based on a learning model (e.g., a learning model 105) trained by model training device 120 using training data (e.g., a training image 103, a training image 103’ and a training image 104) . The objects may be movable objects, such as persons, animals, robots, and animated characters, etc. In some embodiments, training image 103 may include/depict the same object as that in training image 103’ (e.g., having similar static information but different motion information) , and training image 104 may include/depict an object different from the one in training image 103 and training image 103’.
In some embodiments, motion transfer system 100 may include components shown in FIG. 1, including a motion transfer device 110, a model training device 120, a training database 140, a database/repository 150, a user device 160, and a network 170 to facilitate communications among the various components. In some embodiments, motion transfer system 100 may optionally include a display device 130 to display the motion transfer result (e.g., a synthesized image 107) . It is to be contemplated that motion transfer system 100 may include more or less components compared to those shown in FIG. 1.
As shown in FIG. 1, motion transfer system 100 may transfer the motion of a first object (e.g., a first human being) included/depicted in image 101 to a second object (e.g., a second human being, being the same or different from the first human being) included/depicted in image 102 using motion transfer device 110.
In some embodiments, motion transfer device 110 may use a motion feature encoder to extract pose information of the first object (e.g., generate keypoint map (s) representing a probability that a keypoint exists at each pixel and part affinity field of a body part of the first object) in image 101. In some embodiments, motion transfer device 110 may further use a static feature encoder to extract static information of the second object (e.g., the appearance and environment background) included/depicted in image 102. In some embodiments, motion transfer device 110 may additionally use an image generator for generating synthesized image 107 using the pose information of the first object and the static information of the second object. In some embodiments, the motion transfer operations may be performed based on learning model 105, trained by model training devise 120. In some embodiments, motion transfer system 100 may display the motion transfer result (e.g., synthesized image 107) on display device 130. In some embodiments, when a learning model (e.g., learning model 105) is pre-trained for motion transfer, motion transfer system 100 may include only motion transfer device 110, database/repository 150, and optionally display device 130 to perform motion transfer related functions.
In some embodiments, the various components of motion transfer system 100 may be remote from each other or in different locations and be connected through network 170 as shown in FIG. 1. In some alternative embodiments, certain components of motion transfer system 100 may be located on the same site or inside one device. For example, training database 140 may be located on-site with or be part of model training device 120. As another example, model training device 120 and motion transfer device 110 may be inside the same computer or processing device.
Consistent with the present disclosure, motion transfer system 100 may store images including a movable object (e.g., a human being, an animal, a machine with different moving parts, or an animated character, etc. ) . For example, images 103 and 103’ including/depicting the same object with different motion information. In some embodiments, an image 104 including/depicting an object different from the object in images 103 and 103’ may also be stored in training database 140. On the other hand, target and source images for transferring of motions (e.g., images 101 and 102) may be stored in database/repository 150.
The various images (e.g., images 101, 102, 103, 103’ and 104) may be images captured by user device 160 such as a camera, a smartphone, or any other electronic device with photo capturing functions, etc. The images may be created/generated by user device 160 using image processing programs or software, e.g., when the object is an animated character. In some embodiments, the images can be a frame extracted from an image sequence in a video clip. The object included/depicted in each image can be any suitable object capable of moving (i.e., capable of transferring a motion to/or from) such as a robot, a machine, a human being, an animal, etc.
In some embodiments, training database 140 may store training images 103 and 103’ including/depicting the same object, and training image 104 including/depicting a different object. In some embodiments, training images 103 and 103’ may have similar/the same static information (e.g., objects with same appearance but depict from different angles, and/or different background) , but different motion information (e.g., different pose and/or location information) . Training image 104 may have different static information and motion information than either training image103 or 103’. In some embodiments, training images 103 or 103’ may be used for training learning model 105 based on minimizing a joint loss. In some embodiments, training image 104 may be used as a support image for further improving the generalization ability of learning model 105 by further adding a support loss to the joint loss.
In some embodiments, learning model 105 may have an architecture that includes multiple sub-networks (e.g., a motion feature encoder, a static feature encoder and an image generator) . Each sub-network may include multiple convolutional blocks, residual blocks and/or transposed convolution blocks for performing functions such as extracting feature vectors (e.g., representing the motion features and/or the static features) and generating images (e.g., synthesizing the motion features and the static features extracted from different images) . For example, the motion feature encoder may include a pose estimator (e.g., a pre-trained VGG-19 network) , a keypoint amplifier, and a motion refiner network (e.g., a network having residual blocks) for extracting the motion features. In an example, the static feature encoder may include convolutional blocks with down-sampling modules, and some residual blocks (e.g., 3 convolutional blocks with down-sampling modules and 5 residual blocks) for extracting the static features. In another example, the image generator may include residual blocks and transposed convolution blocks (e.g., 4 residual blocks and 3 transposed convolution blocks) for generating the output image (e.g., synthesized image 107) in the same size as the input images (e.g., images 101 and 102) .
In some embodiments, the model training process is performed by model training device 120. It is contemplated that some of the sub-networks of learning model may be pretrained, e.g., ahead of time before the rest parts of the learning model are trained. For example, pose estimator 106 may be pretrained either by model training device 120 or by another device and provided to model training device 120. For example, model training device 120 may receive pretrained pose estimator 106 through network 107, instead of training it jointly with the rest of learning model 105. In some embodiments, pose estimator 106 may be trained for extracting human pose information by estimating keypoints of a human body (e.g., the PoseNet vision model) . In some other embodiments, pose estimator 106 may also be trained with specifically designed training set for exacting pose information of living creatures other than a human being (e.g., an animal) , a machine capable of moving (e.g., a robot, a vehicle, etc. ) , or an animated character.
As used herein, “training” a learning model refers to determining one or more parameters of at least one layer of a block in the learning model. For example, a convolutional layer of the static feature encoder may include at least one filter or kernel. One or more parameters, such as kernel weights, size, shape, and structure, of the at least one filter may be determined by e.g., an adversarial-based training process. Consistent with some embodiments, learning model 105 may be trained based on supervised, semi-supervised, or non-supervised methods.
As show in FIG. 1, motion transfer device 110 may receive learning model 105 from model training device 120. Motion transfer device 110 may include a processor and a non-transitory computer-readable medium (not shown) . The processor may perform instructions of a motion transfer process stored in the medium. Motion transfer device 110 may additionally include input and output interfaces to communicate with database/repository 150, user device 160, network 170 and/or a user interface of display device 130. The input interface may be used for selecting an image (e.g., image 101 and/or 102) for motion transfer. The output interface may be used for providing the motion transfer result (e.g., synthesized image 107) to display device 130.
In some embodiments, motion transfer system 100 may optionally include display 130 for displaying the motion transfer result, e.g., synthesized image 107. Display 130 may include a display such as a Liquid Crystal Display (LCD) , a Light Emitting Diode Display (LED) , a plasma display, or any other type of display, and provide a Graphical User Interface (GUI) presented on the display for user input and data depiction. The display may include a number of different types of materials, such as plastic or glass, and may be touch-sensitive to receive inputs from the user. For example, the display may include a touch-sensitive material that is substantially rigid, such as Gorilla Glass
TM, or substantially pliable, such as Willow Glass
TM. In some embodiments, display 130 may be part of motion transfer device 110.
FIG. 2 illustrates a block diagram of an exemplary motion transfer device 110, according to embodiments of the disclosure. In some embodiments, as shown in FIG. 2, motion transfer device 110 may include a communication interface 202, a processor 204, a memory 206, and a storage 208. In some embodiments, motion transfer device 110 may have different modules in a single device, such as an integrated circuit (IC) chip (e.g., implemented as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA) ) , or separate devices with dedicated functions. In some embodiments, one or more components of motion transfer device 110 may be located in a cloud or may be alternatively in a single location (such as inside a mobile device) or distributed locations. Components of motion transfer device 110 may be in an integrated device or distributed at different locations but communicate with each other through a network (not shown) . Consistent with the present disclosure, motion transfer device 110 may be configured to synthesize motion information (e.g., motion features of the object) extracted from image 101 with static information (e.g., static features of the object) extracted from image 102, and generate synthesized image 107 as an output.
Consistent with some embodiments, communication interface 202 may receive learning network 105 from model training device 120, and images 101 and 102 to be processed from database/repository 150. Communication interface 202 may further provide images 101 and 102 and learning model 105 to memory 206 and/or storage 208 for storage or to processor 204 for processing.
In some embodiments, memory 206 and/or storage 208 may also store intermediate data such as keypoint heatmaps, part affinity fields of body parts of the object, extracted motion features, and extracted static features, etc. Memory 206 and/or storage 208 may additionally store various sub-learning models (e.g., sub-networks included in learning model 105) including their model parameters and model configurations, such as pre-trained pose estimator 106 (e.g., a pre-trained VGG-19 network) , the motion feature extracting blocks (e.g., motion feature encoder) including the keypoint amplifier, the motion refiner, the static feature extracting blocks (e.g., and static feature encoder) , and the image generator blocks, etc.
As shown in FIG. 2, processor 204 may include multiple modules, such as a motion feature extraction unit 240, a static feature extraction unit 242, and an image generation unit 244, and the like. These modules (and any corresponding sub-modules or sub-units) can be hardware units (e.g., portions of an integrated circuit) of processor 204 designed for use with other components or software units implemented by processor 204 through executing at least part of a program. The program may be stored on a computer-readable medium, and when executed by processor 204, it may perform one or more functions. Although FIG. 2 shows units 240-244 all within one processor 204, it is contemplated that these units may be distributed among different processors located closely or remotely with each other.
In some embodiments, units 240-244 of FIG. 2 may execute computer instructions to perform the motion transfer. For example, FIG. 3 illustrates a flowchart of an exemplary method 300 for motion transfer, according to embodiments of the disclosure. Method 300 may be implemented by motion transfer device 110 and particularly processor 204 or a separate processor not shown in FIG. 2 using learning model 105. Method 300 may include steps S302-S310 as described below. It is to be appreciated that some of the steps may be performed simultaneously, or in a different order than shown in FIG. 3.
In step S302, communication interface 202 may receive images 101 and 102 acquired/generated by user device 160 from database/repository 150. In some embodiments, user device 160 may acquire/generate an image including/depicting an object by using a camera. For example, user device 160 may be a smart phone with a camera configured to take pictures or sequence of pictures (e.g., a video clip) . The object may be a living creature (e.g., an animal, a human, etc. ) or a machine capable of moving (e.g., a robot, a vehicle, etc. ) . User device 160 may also generate the image of the object (e.g., an animated character) using image/photo processing software. For example, image 101 and/or 102 may be an image being part of a drawn figure or a sequence of drawn figures (e.g., an animation clip) . Database/repository 150 may store the images and transmit the images to communication interface 202 for motion transfer.
In step S304, motion feature extraction unit 240 may extract motion features (i.e., pose information and location) of a first object included/depicted in image 101 (also referred to as source image x
s) using a motion feature encoder. For example, FIG. 4 illustrates a schematic diagram of exemplary learning model 105 for motion transfer, according to embodiments of the disclosure. In the shown embodiments, learning model 105 may include a motion feature encoder 410, a static feature encoder 420, and an image generator 430.
In some embodiments, as shown in FIG. 4, motion feature encoder 410 may include a pose estimator 412 (e.g., pre-trained pose estimator 106) , a keypoint amplifier 414 and a motion refiner 416. For example, pose estimator 412 may extract keypoint heatmaps
of the object representing a probability that a keypoint exists at each pixel, and part affinity fields
of different body parts of the object showing the orientation of the body part. In some embodiments, p may have 38 (19×2) channels, and can be a set of 2D vectors indicating the location and orientation by x-y coordinates for each channel of the keypoint heatmap h. In some embodiments, the generated keypoint heatmap h may only keep the first 18 of total 19 channels and discard the last channel (e.g., the heatmap of the background) . Both h and p may be downsampled multiple times in order to reduce the size of the image. For example, both h and p may be downsampled in half for 3 times, resulting in an image 1/2
3 the size of the original input images (e.g., images 101 and 102) . In some embodiments, the keypoints may correspond to joints of the object, such as elbows, wrists, etc. of a human being.
In some embodiments, keypoint amplifier 414 may denoise the extracted keypoint heatmap values and obtain the amplified the keypoint heatmap h’. For example, keypoint amplifier 414 may apply a softmax function with a relatively small temperature T as the keypoint amplifier to the keypoint heatmap h according to equation (1) :
For example, T can be set as 0.01 such that the gap between large values and small values in the keypoint heatmap h can be enlarged. This can reduce the effect caused by the noise.
In some embodiments, motion refiner 416 may generate the encoded motion feature vector M (x
s) representing the motion feature of the object based on refining both the part affinity fields p and the amplified keypoint heatmaps h’. For example, motion refiner 416 may include 5 residual blocks. Accordingly, the motion features extracted from pose estimator 412 may be refined such that the influence caused by different body part ratios (e.g., limb ratios) and/or camera angles and/or distances can be reduced.
In step S306, static feature extraction unit 420 may extract static features S (x
t) (e.g., appearance and environment background) of a second object included in image 102 (also referred to as target image x
t) . In some embodiments, static feature extraction unit 420 may apply a static encoder 420 for extracting the background, the appearance, etc., of the second object. For example, static encoder 420 may include 3 convolutional blocks with down-sampling modules and 5 residual blocks.
In step S308, image generation unit 244 may generate synthesized image 107 by synthesizing the motion features M (x
s) extracted from image 101 and the static features S (x
t) extracted from image 102. For example, image generation unit 244 may apply an image generator 430 to M (x
s) and S (x
t) according to equation (2) :
x
syn=G (S (x
t) , M (x
s) ) (2)
where G (·) represents a function performed by image generator 430 and x
syn represents synthesized image 107. In some embodiments, image generator 430 may include 4 residual blocks and 3 transposed convolution blocks such that the output of image generator 430 (e.g., synthesized image 107) may have the same size as that of the input images (e.g., image 101 and/or 102) .
In step S310, the output of image generator 430 (e.g., synthesized image 107) may be transmitted to display device 130 for display.
In some embodiments, learning model 105 may be trained by model training device 120 before being used by motion transfer device 110 for motion transfer. For example, FIG. 5 illustrates a flowchart of an exemplary method 500 for training learning model 105, according to embodiments of the disclosure. Method 500 may be implemented by model training device 120 for training learning model 105. Method 500 may include steps S502-S518 as described below. It is to be appreciated that some of the steps may be performed simultaneously, or in a different order than shown in FIG. 5.
In some embodiments, learning model 105 may be trained using training images 103 (e.g., including a target object x
t) and 103’ (e.g., including a source object x
s) that include a same object (e.g., target object x
t and source object x
s being a same movable object in the training images) in the same environment (e.g., same place, same lighting condition, etc. ) with different motion information. In some embodiments, images 103 and 103’ may be extracted from the same video clip. As the same object in images 103 and 103’ may have similar static information (e.g., having the same appearance, but different in camera angles for taking the appearance and/or having different backgrounds in images 103 and 103’) , i.e.,
learning model 105 may be trained on premises that the motion features extracted from the synthesized image may be a reconstruction of (e.g., approximately equals to) the motion features extracted from image 103’, i.e.,
In some embodiments, model training device 120 may further adopt a support group during the training, to further improve the generalization ability/performance and stability of learning model 105. The support group may include an image (e.g., image 104) depicting an object different from that in images 103 and 103’ to train learning model 105.
Specifically, as illustrated in FIG. 5, in step S502, model training device 120 may receive training images 103 and 103’ that include/depict a same object in the same environment with different motion information. For example, FIG. 6. illustrates a schematic diagram illustrating training an exemplary learning model 105, according to embodiments of the disclosure. As illustrated in FIG. 6, image 103 may include a target object x
t (e.g., target person x
t) and image 103’ may include a source object x
s (e.g., source person x
t) . The background of each image is not shown in FIG. 6 for simplification and illustrative purposes. In some embodiments, to train learning model 105, images 103 and 103’ are selected such that target person x
t and source person x
s are the same person in the same environment. Accordingly, images 103 and 103’ may have similar/same static information (e.g., having the same appearance, but different in camera angles for taking the appearance and/or having different backgrounds in images 103 and 103’) , i.e.,
However, the same person may have different gestures or movements so that images 103 and 103’ may contain different motion information.
In step S504, the motion features M (x
s) of source person x
s, and the static features S (x
t) of the target person x
t, are extracted from images 103’ and 103 respectively, using motion feature encoder 410 and static feature encoder 420 of learning model 105, similar to steps S304 and S306 in method 300.
In step S506, a synthesized image (e.g., including/depicting a synthesized object x
syn, synthesized based on the motion features of x
s and the static features of x
t) may be generated using image generator 430 of learning model 105, similar to step S308 in method 300.
In step S508, motion features and static features of synthesized object x
syn, M (x
syn) and S (x
syn) may be extracted from the synthesized image using motion feature encoder 410 and static feature encoder 420 of learning model 105 respectively, similar to steps S304 and S306 in method 300.
In step S510, model training device 120 may implement an adversarial-based training approach. In some embodiments, model training device 120 may calculate an adversarial loss L
adv to discern image 103’ (e.g., including/depicting the source object x
s) and the synthesized image (e.g., including/depicting the synthesized object x
syn) . For example, model training device 120 may apply an image discriminator D to discern between the real sample source object x
s and the synthesized object x
syn, conditioned on the motion features M (x
s) extracted from the source image (image 103’) . In some embodiments, image discriminator D may take image 103’ as a real sample labeled with 1 and the synthesized image as a fake sample labeled with 0, where D (x
s , M (x
s) ) = 1 and D (x
syn , M (x
s) ) = 0. For example, the adversarial loss can be calculated according to equations (3) , (4) and (5) :
where
In some embodiments, image discriminator D may be a multi-scale discriminator D= (D
1, D
2) . In some embodiments, a discriminator feature matching loss L
fm may be calculated. In some embodiments, the discriminator feature matching loss L
fm may be calculated based on a weighted sum of multiple feature losses from each of the different layers of image discriminator D. For example, image discriminator D may include 5 different layers and discriminator feature matching loss L
fm may be the weighted sum of a L
1 distance between the corresponding features of x
s and x
syn at each layer of image discriminator D.
In step S512, model training device 120 may calculate feature-level consistency losses indicative of a difference between features extracted from the synthesized image (e.g., the motion features and the static features) and the corresponding features extracted from images 103 and 103’. This may insure that the synthesized object (e.g., x
syn) has the same static features of the target object (e.g., x
t from image 103) and the same motion features as the source object (e.g., x
s from image 103’) . For example, model training device 120 may calculate a motion consistency loss L
mc indicating a difference (e.g., a L
1 distance) between the motion features extracted from the synthesized image and the motion features extracted from image 103’. Similarly, model training device 120 may calculate a static consistency loss L
sc indicating a difference (e.g., a L
1 distance) between the static features extracted from the synthesized image and the static features extracted from image 103. For example, the motion consistency loss and the static consistency loss can be calculated according to equations (6) and (7) :
In step S514, model training device 120 may calculate a perpetual loss L
per based on image 103’ and the synthesized image. In some embodiments, the perpetual loss may be calculated using a pre-trained deep convolutional network for object recognition (e.g., a VGG network) . The perpetual loss may be added to the full object to improve the stability and quality of the training.
In step S516, model training device 120 may further calculate a support loss based on a support set. In some embodiments, the support set may include images of different objects as the source object for training, e.g., image 104 including an object different from that of images 103 and 103’. Images in the support set provide many kinds of unseen motions and various static information. In some embodiments, a support loss L
sup may be calculated using the support set (e.g., image 104) as a target image (e.g., including a target object) .
When training with the support set, because the objects included in the target image x
t and the source image x
s are different, they do not share the same static features, i.e., S (x
t) ≠ S (x
s) . Meanwhile, the synthesized image x
syn obtained based on the support set may not be a reconstruction of the source image x
s. Accordingly, when calculating the support loss L
sup, the ground truth image of the target object performing the motion of the source object is not available. Thus, L
+
adv, L
fm and L
per, for calculating the support loss L
sup are not available. In some embodiments, the support loss L
sup may include a feature-level consistency loss L
mc indicative of a difference between the motion features extracted from the synthesized image and the motion features extracted from source image 103’. In some embodiments, the support loss may further include a feature-level consistency loss L
sc indicative of a difference between the static features extracted from the synthesized image and the static features extracted from target image 103. In some embodiments, the support loss may also include a negative adversarial loss L
-
adv determined based on the image 103’ and the synthesized image. In some embodiments, the support loss L
sup may be calculated as a weighted sum of L
sc, L
mc and L
-
adv.
In step S518, model training device 120 may train learning model 105 by jointly training the sub-networks of learning model 105 (e.g., jointly training keypoint amplifier 414, motion refiner network 416, static feature encoder 420 and image generator 430) based on minimizing the joint loss. In some embodiments, pre-trained pose estimator 106 may remain the same throughout the optimization process. For example, model training device 120 may minimize a joint loss L
full that includes some or all of the losses calculated above. In some embodiments, the joint loss L
full may be a weighted sum of L
adv, L
fm, L
per, L
mc and L
sc. For example, the joint loss L
full may be calculated according to equation (8) :
where λ
adv, λ
fm, λ
per, λ
mc and λ
sc are the weights assigned for the respective losses, as calculated in previous steps. In some embodiments, the weights may be selected to reflect the relative importance of the respective losses. For example, λ
adv, λ
fm, λ
per, λ
mc and λ
sc may be set to 1, 10, 10, 0.1, 0.01 respectively.
In some embodiments, the support loss L
sup calculated in step S518 may be added to the joint loss in order to improve the generalization ability of learning model 105. For example, when training learning model 105, the support loss L
sup may be calculated as a weighted sum of L
sc, L
mc and L
-
adv according to equation (9) and be added to the joint loss L
full of equation (8) :
where λ
sc, λ
mc and λ
adv are the weights for L
sc, L
mc and L
-
adv respectively and λ
sup represents the weight assigned to support loss L
sup when calculating the joint loss L
full. For example, the weight λ
sup can be set to 0.001 while other weights may remain the same as for calculating the overall objective joint loss L
full.
Another aspect of the disclosure is directed to a non-transitory computer-readable medium storing instruction which, when executed, cause one or more processors to perform the methods, as discussed above. The computer-readable medium may include volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, or other types of computer-readable medium or computer-readable storage devices. For example, the computer-readable medium may be the storage device or the memory module having the computer instructions stored thereon, as disclosed. In some embodiments, the computer-readable medium may be a disc or a flash drive having the computer instructions stored thereon.
It will be apparent to those skilled in the art that various modifications and variations can be made to the disclosed system and related methods. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the disclosed system and related methods.
It is intended that the specification and examples be considered as exemplary only, with a true scope being indicated by the following claims and their equivalents.
Claims (20)
- A system for performing motion transfer using a learning model, comprising:a communication interface configured to receive a first image including a first movable object and a second image including a second movable object; andat least one processor coupled to the communication interface and configured to:extract a first set of motion features of the first movable object from the first image using a first encoder of the learning model;extract a first set of static features of the second movable object from the second image using a second encoder of the learning model; andgenerate a third image by synthesizing the first set of motion features and the first set of static features.
- The system of claim 1, wherein the first encoder of the learning model includes a pretrained pose estimator configured to extract pose information of the first movable object and a motion refiner configured to generate a motion feature vector representing the first set of motion features.
- The system of claim 2, wherein to extract the first set of motion features from the first image, the pretrained pose estimator is further configured to:determine a keypoint heatmap representing a probability that a keypoint exists at each pixel; anddetermine a part affinity field of a body part of the first movable object.
- The system of claim 3, wherein to extract the first set of motion features from the first image, the first encoder further includes a keypoint amplifier configured to amplify the keypoint heatmap.
- The system of claim 4, wherein to generate the motion feature vector, the motion refiner is further configured to:refine the amplified keypoint heatmap and the part affinity field.
- The system of claim 1, wherein the learning model is trained using a joint loss comprising an adversarial loss and at least one feature-level consistency loss.
- The system of claim 6, wherein the adversarial loss is determined by applying an image discriminator to discern between the first image and the third image, conditioned on the first set of motion features extracted from the first image.
- The system of claim 7, wherein the image discriminator comprises multiple layers, and wherein the joint loss further comprises:a discriminator feature matching loss indicative of a weighted sum of differences between corresponding features of the first image and the third image at each layer of the image discriminator.
- The system of claim 6, wherein the at least one feature-level consistency loss further comprises:a first feature-level consistency loss indicative of a difference between a second set of motion features extracted from the third image and the first set of motion features extracted from the first image; anda second feature-level consistency loss indicative of a difference between a second set of static features extracted from the third image and the first set of static features extracted from the second image.
- The system of claim 6, wherein the joint loss further comprises:a perceptual loss determined based on applying a pretrained deep convolutional network for object recognition to the first and the third images.
- The system of claim 1, wherein the learning model is trained using a support set including a fourth image including a third movable object different from the first object or the second object.
- The system of claim 11, wherein the learning model is trained using a support loss determined based on the fourth image, wherein the support loss comprises:a third feature-level consistency loss indicative of a difference between a third set of motion features extracted from the fourth image and the first set of motion features extracted from the first image;a fourth feature-level consistency loss indicative of a difference between a third set of static features extracted from the fourth image and the first set of static features extracted from the second image; anda negative adversarial loss determined based on the first image and the fourth image.
- The system of claim 12, wherein the support loss is a weighted sum of the third feature-level consistency loss, the fourth feature-level consistency loss, and the negative adversarial loss.
- A method for motion transfer using a learning model, comprising:receiving, by a communication interface, a first image including a first movable object and a second image including a second movable object;extracting, by at least one processor, a first set of motion features of the first movable object from the first image using a first encoder of the learning model;extracting, by the at least one processor, a first set of static features of the second movable object from the second image using a second encoder of the learning model; andgenerating, by the at least one processor, a third image by synthesizing the first set of motion features and the first set of static features.
- The method of claim 14, further comprising:determining a keypoint heatmap representing a probability that a keypoint exists at each pixel; anddetermining a part affinity field of a body part of the first movable object.
- The method of claim 15, further comprising:amplifying the keypoint heatmap using an amplifier; andgenerating a motion vector representing the first set of motion features based on refine the amplified keypoint heatmap and the part affinity field.
- The method of claim 14, wherein the learning model is trained using a joint loss comprising an adversarial loss and at least one feature-level consistency loss.
- The method of claim 17, wherein the at least one feature-level consistency loss further comprises:a first feature-level consistency loss indicative of a difference between a second set of motion features extracted from the third image and the first set of motion features extracted from the first image; anda second feature-level consistency loss indicative of a difference between a second set of static features extracted from the third image and the first set of static features extracted from the second image.
- The method of claim 14, wherein the learning model is trained using a support set including a fourth image including a third movable object different from the first object or the second object.
- A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for motion transfer using a learning model, comprising:receiving a first image including a first movable object and a second image including a second movable object;extracting a first set of motion features of the first movable object from the first image using a first encoder of the learning model;extracting a first set of static features of the second movable object from the second image using a second encoder of the learning model; andgenerating a third image by synthesizing the first set of motion features and the first set of static features.
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202080019200.5A CN114144778B (en) | 2020-06-12 | 2020-06-12 | Systems and methods for motion transfer using learned models |
| PCT/CN2020/095755 WO2021248432A1 (en) | 2020-06-12 | 2020-06-12 | Systems and methods for performing motion transfer using a learning model |
| US17/020,668 US11830204B2 (en) | 2020-06-12 | 2020-09-14 | Systems and methods for performing motion transfer using a learning model |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2020/095755 WO2021248432A1 (en) | 2020-06-12 | 2020-06-12 | Systems and methods for performing motion transfer using a learning model |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/020,668 Continuation US11830204B2 (en) | 2020-06-12 | 2020-09-14 | Systems and methods for performing motion transfer using a learning model |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021248432A1 true WO2021248432A1 (en) | 2021-12-16 |
Family
ID=78825769
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/095755 Ceased WO2021248432A1 (en) | 2020-06-12 | 2020-06-12 | Systems and methods for performing motion transfer using a learning model |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11830204B2 (en) |
| CN (1) | CN114144778B (en) |
| WO (1) | WO2021248432A1 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210343201A1 (en) * | 2020-05-01 | 2021-11-04 | AWL, Inc. | Signage control system and non-transitory computer-readable recording medium for recording signage control program |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US12067659B2 (en) * | 2021-10-15 | 2024-08-20 | Adobe Inc. | Generating animated digital videos utilizing a character animation neural network informed by pose and motion embeddings |
| KR20240168358A (en) * | 2022-12-15 | 2024-11-29 | 엘지전자 주식회사 | Artificial intelligence device and method for creating its three-dimensional agency |
| KR20240102204A (en) * | 2022-12-26 | 2024-07-03 | 광주과학기술원 | Apparatus and method of data augmentation for action recognition via self supervised learning based on objects |
| US20240290025A1 (en) * | 2023-02-27 | 2024-08-29 | Google Llc | Avatar based on monocular images |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102682302A (en) * | 2012-03-12 | 2012-09-19 | 浙江工业大学 | Human body posture identification method based on multi-characteristic fusion of key frame |
| CN104253994A (en) * | 2014-09-03 | 2014-12-31 | 电子科技大学 | Night monitored video real-time enhancement method based on sparse code fusion |
| CN104581437A (en) * | 2014-12-26 | 2015-04-29 | 中通服公众信息产业股份有限公司 | Video abstract generation and video backtracking method and system |
| CN105187801A (en) * | 2015-09-17 | 2015-12-23 | 桂林远望智能通信科技有限公司 | Condensed video generation system and method |
| CN106599907A (en) * | 2016-11-29 | 2017-04-26 | 北京航空航天大学 | Multi-feature fusion-based dynamic scene classification method and apparatus |
| CN106791380A (en) * | 2016-12-06 | 2017-05-31 | 周民 | The image pickup method and device of a kind of vivid photograph |
| US10600158B2 (en) * | 2017-12-04 | 2020-03-24 | Canon Kabushiki Kaisha | Method of video stabilization using background subtraction |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20160093809A (en) * | 2015-01-29 | 2016-08-09 | 한국전자통신연구원 | Method and apparatus for detecting object based on frame image and motion vector |
| WO2017137948A1 (en) * | 2016-02-10 | 2017-08-17 | Vats Nitin | Producing realistic body movement using body images |
| CN109598671A (en) * | 2018-11-29 | 2019-04-09 | 北京市商汤科技开发有限公司 | Image generating method, device, equipment and medium |
| CN110047119B (en) * | 2019-03-20 | 2021-04-13 | 北京字节跳动网络技术有限公司 | Animation generation method and device comprising dynamic background and electronic equipment |
| CN109977847B (en) * | 2019-03-22 | 2021-07-16 | 北京市商汤科技开发有限公司 | Image generation method and device, electronic equipment and storage medium |
-
2020
- 2020-06-12 WO PCT/CN2020/095755 patent/WO2021248432A1/en not_active Ceased
- 2020-06-12 CN CN202080019200.5A patent/CN114144778B/en active Active
- 2020-09-14 US US17/020,668 patent/US11830204B2/en active Active
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102682302A (en) * | 2012-03-12 | 2012-09-19 | 浙江工业大学 | Human body posture identification method based on multi-characteristic fusion of key frame |
| CN104253994A (en) * | 2014-09-03 | 2014-12-31 | 电子科技大学 | Night monitored video real-time enhancement method based on sparse code fusion |
| CN104581437A (en) * | 2014-12-26 | 2015-04-29 | 中通服公众信息产业股份有限公司 | Video abstract generation and video backtracking method and system |
| CN105187801A (en) * | 2015-09-17 | 2015-12-23 | 桂林远望智能通信科技有限公司 | Condensed video generation system and method |
| CN106599907A (en) * | 2016-11-29 | 2017-04-26 | 北京航空航天大学 | Multi-feature fusion-based dynamic scene classification method and apparatus |
| CN106791380A (en) * | 2016-12-06 | 2017-05-31 | 周民 | The image pickup method and device of a kind of vivid photograph |
| US10600158B2 (en) * | 2017-12-04 | 2020-03-24 | Canon Kabushiki Kaisha | Method of video stabilization using background subtraction |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210343201A1 (en) * | 2020-05-01 | 2021-11-04 | AWL, Inc. | Signage control system and non-transitory computer-readable recording medium for recording signage control program |
| US11682037B2 (en) * | 2020-05-01 | 2023-06-20 | AWL, Inc. | Signage control system and non-transitory computer-readable recording medium for recording signage control program |
Also Published As
| Publication number | Publication date |
|---|---|
| CN114144778B (en) | 2024-07-19 |
| CN114144778A (en) | 2022-03-04 |
| US20210390713A1 (en) | 2021-12-16 |
| US11830204B2 (en) | 2023-11-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11830204B2 (en) | Systems and methods for performing motion transfer using a learning model | |
| CN113822977B (en) | Image rendering methods, apparatus, devices and storage media | |
| US11417014B2 (en) | Method and apparatus for constructing map | |
| US10475207B2 (en) | Forecasting multiple poses based on a graphical image | |
| CN113487608B (en) | Endoscope image detection method, endoscope image detection device, storage medium, and electronic apparatus | |
| CN110503703B (en) | Methods and apparatus for generating images | |
| US11024060B1 (en) | Generating neutral-pose transformations of self-portrait images | |
| CN115115713A (en) | Unified space-time fusion all-around aerial view perception method | |
| CN115761565B (en) | Video generation method, device, equipment and computer readable storage medium | |
| CN111402122A (en) | Image mapping processing method and device, readable medium and electronic equipment | |
| CN110490959B (en) | Three-dimensional image processing method and device, virtual image generating method and electronic equipment | |
| KR20210032678A (en) | Method and system for estimating position and direction of image | |
| CN110275968A (en) | Image processing method and device | |
| CN117252914A (en) | Training methods, devices, electronic equipment and storage media for depth estimation networks | |
| CN114445676B (en) | A gesture image processing method, storage medium and device | |
| KR20210040702A (en) | Mosaic generation apparatus and method thereof | |
| CN116310408B (en) | Method and device for establishing data association between event camera and frame camera | |
| CN115018979B (en) | Image reconstruction method, device, electronic device, storage medium and program product | |
| CN112270242A (en) | Track display method, device, readable medium and electronic device | |
| CN115205325B (en) | Target tracking method and device | |
| CN110084306B (en) | Method and apparatus for generating dynamic image | |
| CN113920023A (en) | Image processing method and device, computer readable medium and electronic device | |
| CN117916773A (en) | Method and system for simultaneous pose reconstruction and parameterization of 3D mannequins in mobile devices | |
| CN118172476A (en) | Lighting estimation method and device | |
| CN115512038A (en) | Real-time rendering method for free viewpoint synthesis, electronic device and readable storage medium |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20939550 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20939550 Country of ref document: EP Kind code of ref document: A1 |
