EP4681170A1 - Methods and apparatuses for immersive videoconference - Google Patents

Methods and apparatuses for immersive videoconference

Info

Publication number
EP4681170A1
EP4681170A1 EP24709747.0A EP24709747A EP4681170A1 EP 4681170 A1 EP4681170 A1 EP 4681170A1 EP 24709747 A EP24709747 A EP 24709747A EP 4681170 A1 EP4681170 A1 EP 4681170A1
Authority
EP
European Patent Office
Prior art keywords
user
head
image
hair
model
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24709747.0A
Other languages
German (de)
French (fr)
Inventor
Francois Le Clerc
Quentin AVRIL
Philippe Henri GOSSELIN
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
InterDigital CE Patent Holdings SAS
Original Assignee
InterDigital CE Patent Holdings SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by InterDigital CE Patent Holdings SAS filed Critical InterDigital CE Patent Holdings SAS
Publication of EP4681170A1 publication Critical patent/EP4681170A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T17/00Three-dimensional [3D] modelling for computer graphics
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2200/00Indexing scheme for image data processing or generation, in general
    • G06T2200/08Indexing scheme for image data processing or generation, in general involving all processing steps from image acquisition to 3D model generation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/26Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation

Definitions

  • the present embodiments generally relate to a method and an apparatus for encoding/decoding semantic description data representative of a 3D hair model for immersive telepresence.
  • the present embodiments also generally relate to methods and apparatuses for encoding or decoding based on a neural network.
  • Telepresence refers to the use of virtual reality technology, for instance for apparent participation in distant events.
  • a popular application is found in a telepresence videoconferencing system that immerses the participants in a single common environment.
  • such systems are meant to ensure that a user sitting at a table in a boardroom gets the impression that the other participants are sitting at the same table in the same boardroom, and directly looking at him or her when s/he talks.
  • An immersive telepresence system requires some computer vision processing on the captures of distant participants to achieve its goals.
  • these captures are 2D videos obtained by commodity cameras.
  • a first issue is the head pose, that is the position and orientation of the head of the distant participant in the received images need to be changed at the receiver end to establish eye contact with the user in his/her viewing device.
  • a second issue is the rendering of the hair region, for instance to fit with the head pose or to be faithful to the real appearance of the distant participant.
  • An efficient immersive telepresence system is therefore desirable that addresses two well-known problems for telepresence, (a) achieving proper pose and eye position of the rendered face to support proper eye contact, and (b) faithful rendering of the hair region.
  • the head model at least comprises a 3D model of a face of a user and 3D model of the hair of the user.
  • a method comprising receiving semantic description data representative of a 3D model of a face region of a user’s head in an image; receiving semantic description data representative of a 3D model of hair region of the user’s head in the image; synthesizing an image of a face region of the user’s head; synthesizing an image of a hair region of the user’s head; and generating an image of a head of the user from the synthesized image of the face region of the user’s head and the synthesized image of the hair region of the user’s head.
  • a method comprising receiving an input image comprising a head of a user; determining, by applying neural networks to the image, semantic description data representative of a 3D model of a face region of the user’s head in the input image; determining, by applying neural networks to the input image, semantic description data representative of a 3D model of hair region of the user’s head in the input image; and providing the semantic description data representative of a 3D model of the face region of the user head and the semantic description data representative of a 3D model of the hair region of the user head to synthesize an image of the user’s head.
  • up to two networks are used to determine semantic description data representative of a 3D model of a face region.
  • up to three networks are used to determine semantic description data representative of a 3D model of a hair region.
  • One or more embodiments also provide an apparatus comprising one or more processors configured for performing any one of the embodiments of the methods cited above.
  • One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform any one of the methods according to any of the embodiments described above.
  • One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for editing a video shot, encoding at least one image or a video or decoding at least one image or a video according to the any of the embodiments described above.
  • One or more embodiments also provide a bitstream comprising image or video data encoded according to any one of the embodiments of the encoding method cited above.
  • One or more of the present embodiments also provide a computer readable storage medium having stored thereon a bitstream described above.
  • One or more embodiments also provide a method for transmitting a bitstream comprising image or video data encoded according to any one of the embodiments of the encoding method described herein.
  • One or more embodiments also provide an apparatus for receiving a bitstream comprising image or video data encoded according to any one of the embodiments of the encoding method described herein.
  • FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to an embodiment.
  • FIG. 2 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to another embodiment.
  • FIG. 3 illustrates schematically a telepresence system within which aspects of the present embodiments may be implemented, according to an embodiment.
  • FIG. 4 illustrates schematically a telepresence system with its encoder and its decoder, within which aspects of the present embodiments may be implemented, according to an embodiment.
  • FIG. 5 illustrates schematically a telepresence system with its encoder and its decoder, implementing the method according to the present invention.
  • FIG. 6 illustrates examples of the components of a 3D environment model according to an embodiment.
  • FIG. 7 illustrates a method for decoding at least one image according to an embodiment.
  • FIG. 8 illustrates a method for encoding at least one image according to an embodiment.
  • FIG. 9 shows two remote devices communicating over a communication network in accordance with an example of present principles.
  • FIG. 10 shows the syntax of a signal in accordance with an example of present principles.
  • the present principles will be now described in the particular case of an immersive videoconference.
  • the present principles are not limited to videoconferencing, but could be directly and non-ambiguously derived to any telepresence system where the representation of the user is driven by a distant capture of his or her head by a camera.
  • Such systems include, but are not limited to, gaming frameworks where the user is represented by an avatar whose motion and expressions are driven by the distant live video capture, or more generally frameworks falling in the scope of the Metaverse where participants interact in a virtual environment through their embodiments as avatars and the avatar appearance, motion and expression is driven by a distant live video capture of the participants' heads.
  • Such frameworks can host commercial applications such as e-learning, e-tourism and e-commerce, to name a few.
  • FIG. 3 illustrates schematically a telepresence system within which aspects of the present embodiments may be implemented, according to an embodiment.
  • the telepresence system of FIG. 3 comprises three communication apparatus D1 , D2, D3 connected through a communication network.
  • the communication apparatus D1 comprises a camera for capturing a scene as a succession of images forming video data.
  • the captured scene is constituted here at least by the head of a first user P1 for instance in front of a table T.
  • the communication apparatus D1 further comprises a display for rendering a video in which the remote users P2 and P3 are displayed in an immersive environment, for instance in front of a representation of a same table T and apparently directly looking at the user P1 .
  • the user P1 is also represented in the immersive environment by the display of the communication apparatus D2 or D3, for instance in front of a representation of the same table T and also apparently looking at the user P2 or P3.
  • the communication apparatus D1 comprises a transmitter/encoder used to process the captured video data as described below and to provide description data for synthesizing the head of the user P1 in the remote communication device D2 or D3.
  • the communication apparatus D1 also comprises a receiver/decoder for receiving and processing description data provided by the remote communication devices D2, D3 and rendering the head of the users P2 and P3 in the immersive video displayed by the communication device D1.
  • the communication apparatus D2 is capable of reproducing a video with immersive effect thanks to which user P2 viewing immersive video reproduced on apparatus D2 will get the impression that the user P1 is sitting at the same table T, and directly looking at him or her when s/he talks.
  • receiver/decoder of the apparatus D2 which for instance implements a semantic compression scheme, as illustrated on FIG.4, which extracts, encodes, transmits and decodes a 3D model of the participants’ faces.
  • FIG.4 illustrates a semantic compression scheme
  • the transmitter/encoder of the apparatus D1 processes the 2D video data to obtain, by applying an encoder to the video data, semantic description data representative of a 3D model, for instance a 3D geometric and photometric model, of the head of the first user in the first video data; and provides the semantic description data for rendering the head of the first user in an immersive video.
  • the head may comprise the face of the user, or the head may comprise the face of the user and the hair of the user.
  • the embodiments of FIG. 4 OR FIG. 5 allow to significantly reduce the amount of data to transmit on the communication network.
  • FIG. 4 illustrates schematically a telepresence system with its encoder and its decoder, within which aspects of the present embodiments may be implemented, according to an embodiment.
  • One possible semantic compression scheme for videoconferencing is described in the EP patent application number 22306339.7, filed on Sept. 12, 2022 by the same applicant and represented in FIG. 4.
  • the face region is detected 420.
  • a pre-trained autoencoder neural network made up of a face model encoder 430 and a face model decoder 450 is leveraged to map the detected face sub-images to a semantic 3D model.
  • This autoencoder is person-generic as it is trained on a large collection of photographs with a large variety of identities, facial expressions, lighting and head poses.
  • the face model encoder also referred to as transmitter, comprises an encoding module 430 which performs a task corresponding to the 3D model extraction module. From the captured 2D video of the sender’s face, it extracts semantic description data representative of a parametric 3D face model 440 of the sender’s face.
  • the model addresses the interior part of the face, including the eyes, nose and mouth but, at this stage, not the hair.
  • facial hair such as beard, eyebrow, mustache, is not managed by the hair model, but by the face region model (as a specific texture).
  • the sender semantic description data at least comprises:
  • the face model is transmitted over the network then decoded at the receiver end to reconstruct the face sub-image captured at the transmitter end.
  • the semantic face model only addresses the interior region of the face, i.e. , the region surrounding the eyes, the nose and the mouth, but does not encode the hair and the bust regions.
  • a face model decoder 450 retrieve semantic description data of 3D model of the sender's face and generates an image of the sender’s face interior 460.
  • the head pose of the 3D model of the sender's face and the appearance representative of the texture of the sender’s face are obtained to further drive the computation of the sender's face image, so that this image can be overlaid on the rendering of the environment.
  • the generator module 470 at the back-end of the semantic compression pipeline of FIG. 4 reconstructs the image 480 to be displayed to the receiving participant.
  • the generator 470 is person-specific.
  • the generator 470 is a neural network that has been trained offline on images or videos of the face of the considered person.
  • the missing elements that it generates are determined by the contents of the training images and videos and not by the contents of the live capture in videoconferencing sessions.
  • a first issue in the semantic compression scheme of FIG. 4 is that, in general, only the face interior region in the reconstructed image at the receiver end is true to the appearance of the person at the transmitting end. Indeed, the other elements of the head, in particular the hair region, are hallucinated by the generator based on the contents of the images of videos of the person it has been trained on offline. A person’s haircut as recorded in the offline training videos and images will be hallucinated by the generator and displayed to the participants of videoconference sessions, even if the person has changed his/her haircut in the meantime.
  • This is a major downside of the compression scheme, as it potentially introduces a discrepancy between the ground truth appearance of the person as captured by his/her transmitting device and the appearance of the person that is displayed to the other participants of videoconference sessions.
  • a second issue of the semantic compression scheme of FIG.4 is that, unlike the face interior region, the hair region may not be edited, for instance to adjust the head pose to achieve eye contact. Indeed, the hair region in the reconstructed image is hallucinated by the generator module and not based on a 3D model of the person’s hair.
  • the generator may, to some extent, learn to adapt the rendering of the hair region to the head pose in its input face interior image, but it is unlikely to be as faithful to the ground truth as if a 3D model of the hair was transmitted and its geometrical component edited to match a target head pose.
  • At least one embodiment circumvents these two issues by further estimating, encoding and transmitting a semantic 3D model for the person’s hair, in addition to and separately from the semantic 3D model of the face.
  • a reconstruction of the hair region in the head image can be composited with the reconstruction of the face interior region to provide the generator at its input with an almost complete synthetic picture of the head.
  • This ensures that the appearance of hair in the synthesized image matches the hair features in the transmitted image.
  • it relieves the generator of the task of hallucinating the hair region in addition to the image background and other face regions, thereby potentially improving the quality of the reconstructed image at the receiver end.
  • the transmitted 3D semantic model of hair can be edited at the receiver end, in particular to adjust its position, scale and orientation to ensure the hair regions in the reconstructed images are consistent with the desired rendering viewpoints in the receiving devices of the videoconference participants.
  • FIG. 5 illustrates schematically a telepresence system with its encoder and its decoder, implementing the method according to at least one embodiment.
  • the telepresence scheme will focus on a scenario with just two participants, a sender and a receiver.
  • the upper path in the diagram describes the encoding and decoding of the face interior region in the detected head region in the input image as presented in FIG. 4 and described in the EP patent application number 22306339.7, filed on Sept. 12, 2022.
  • receiving an input image 510 comprising a head of a user is fed into the system.
  • a neural network face autoencoder consisting of a neural network face encoder followed by a neural network face decoder, is trained on a large collection of faces with various physiognomies, head poses, expressions and illuminants.
  • the face encoder 530 outputs a hand-crafted semantic 3D face model 540 that consists of several vectors of numbers representing the following components: • The translation, rotation and scale defining the viewpoint of the face with reference to a fronto-parallel viewpoint where the face is seen head-on;
  • This autoencoder network is trained on a dataset of guide curves such as the guide curve set output by step 524. It consists of a guide curve set encoder 535 and a guide curve set decoder 552.
  • the encoder network 535 encodes the guide curve set at the output of step 524 into a latent code of a predetermined dimension that is substantially smaller than the dimension of the input guide curve set, thereby providing a compact representation of said set.
  • This latent code is sent over the network, then decoded back at the receiver end by the guide curve set decoder 552 into a close approximation of the encoded guide curve set.
  • the compressed representation of the set of guide curves is decoded and the 3D model of the user’s hair comprising a set of guide curves representative of the shapes of hair wisps and a dominant hair color is reconstructed.
  • a neural network performing the inverse functionality of step 524 is used to render a hair image from the - possibly densified - guide curve set at the output of step 552.
  • This network 556 is trained on the same dataset as in step 524, where the inputs and outputs are reversed. It takes as input the decoded guide curve set providing the 3D model for the hair of the subject represented in the input image 510, as well as information representative of the color of the hair that was transmitted through the network, and synthesizes an image of the hair region.
  • a segmentation mask of the hair region image inside its bounding box may be sent in parallel to the guide curve set geometry and used as an extra input to the neural network, to guarantee that the hair pixels are synthesized in the correct region of the output image.
  • a Generative Adversarial Network 570 to the synthesized image of the user’s head to generate an image of a head of the user.
  • the output of the composition is fed to a generator network, typically complemented by a discriminator network to form a Generative Adversarial Network.
  • the generator and discriminator are trained jointly.
  • the generator is fed with the composited face interior and hair image, which acts as a conditioning input and drives the estimation of the complete image at its output.
  • the generator hallucinates the missing regions to form a plausible and photorealistic head image where the appearance of the face interior and hair regions at its input are roughly preserved.
  • the semantic description data representative of a 3D model of the face region of a user head and the semantic description data representative of a 3D model of the hair region of the user head are transmitted from an encoder to a decoder to synthesize an image of the user’s head.
  • the generated image 580 of a head of the user is faithful to the input image of the head of the user used to obtain the semantic description data representative of a 3D model of the face of the user and the semantic description data representative of a 3D model of the hair of the user.
  • the image is part of a video, and the steps of the method are repeated for each input image.
  • FIG. 6 schematically illustrates examples of the components of a 3D environment model according to an embodiment.
  • the immersive environment in which participants in the videoconference are represented is obtained from a predetermined 3D model of a scene.
  • this scene could represent a room with a floor, walls and windows, further comprising a table 610 and chairs 620 around this table 610.
  • each user would be assigned a predetermined chair 620 on which s/he would be represented sitting in the immersive video displayed on the receiver devices of the other participants.
  • the image of the virtual environment is computed by rendering the projection of the aforementioned 3D environment model on the image plane of a virtual camera 630 whose position, attitude and optical parameters, comprising in particular its focal length, are predetermined.
  • FIG. 7 illustrates a generic method 700 for decoding semantic description data representative of a head of a user and generating an image of a head of the user in an immersive environment according to an embodiment.
  • the semantic description data representative of a head of a user such as P1 in FIG. 3 comprises 2 parts: semantic description data representative of a 3D model of the face of a user, and semantic description data representative of a 3D model of the hair of the user. For instance, both semantic description data form part of a bitstream.
  • semantic description data representative of a 3D model of the face (i.e. face interior) of a user is received.
  • the 3D model of a face of a user comprises an indication of an identity representative of a physiognomy of the sender P1 of FIG. 3 with a neutral expression; an indication of an expression representative of an emotional expression with respect to the neutral expression of the sender P1 ; an indication of an appearance representative of the color of the face of the sender P1 , for instance a texture on the surface of the 3D mesh.
  • an image of the sender’s face is synthesized from received semantic description data representative of the 3D geometric and photometric model of the face of a remote user P1 and from the rigid head pose of the face of the remote user.
  • the rigid head pose 650 is computed so that the image resulting from the overlay represents the sender sitting on the predetermined seat that was assigned to him or her, with his or her face looking at the virtual camera.
  • the 3D head pose model consists of scale, translation and rotation components, defined in the 3D coordinate system 640 of the predetermined 3D scene model as shown on FIG. 6.
  • semantic description data representative of a 3D model of the hair of the user is received.
  • semantic description data representative of a 3D model of the user’s hair comprises a set of guide curves representative of the shapes of hair wisps and a dominant hair color.
  • a compressed representation of the set of guide curves is decoded in a subsequent step (not shown on FIG. 7).
  • the set of guide curves may be interpolated to obtain a denser set of guide curves (not shown on FIG. 7).
  • an image of the user’s hair is synthesized from received semantic description data representative of the 3D hair model.
  • synthesizing the image of the user’s hair may further take as input the rigid head pose of the face of the remote user for rendering the user’s head with a desired head pose.
  • a photorealistic image of the user head is generated from the synthesized face and hair regions.
  • the generated image of a head of the user is faithful to the appearance of the face and hair of the user in the input image used to obtain the semantic description data representative of a 3D model of the face of the user and the semantic description data representative of a 3D model of the hair of the user.
  • the generating step 750 is performed using a Generative Adversarial Network GAN.
  • the GAN fulfills two functions: firstly, hallucinating the parts of the input image that are not encoded and transmitted through the network, such as the background, and secondly, making the renderings of the face and head image models synthesized in steps 730 and 740 more photorealistic.
  • This GAN is a person-specific neural network that has been trained offline on images or videos of the face of the considered person.
  • the missing elements that it generates are determined by the contents of the training images and videos and not by the contents of the live capture in the videoconferencing session.
  • the foreground regions representing the user and excluding the background in the images output by the GAN are segmented out from the images output by the GAN and overlaid on a predetermined rendering of a virtual environment to generate the final rendered image displayed to the user at the receiver end.
  • the training images and videos for the GAN are captured against a uniform background.
  • the foreground region representing the sender in the rendered images can be effectively and efficiently extracted using color keying techniques known from the state of art.
  • the present principles are not limited to a GAN for the generating step, the skilled in the art will appreciate that the GAN is currently the most effective implementation to generate such photorealistic images from synthetic views.
  • the method 700 consequently reduces the amount of data to transmit in a videoconference system by decoding the semantic description data of a 3D head model instead a 2D video data.
  • the method advantageously achieves proper pose of the rendered face to support proper eye contact and provides a faithful reconstruction of both user’s face and user’s hair captured in the input image.
  • the generated image is part of a video and the decoding is repeated for each frame/image of the video.
  • FIG. 8 illustrates a generic method 800 for encoding semantic description data representative of a head of a user according to an embodiment.
  • a first step 810 an input image is received, the input image comprising the face of a user.
  • the input image is a frame of a video.
  • a task corresponding to the 3D face model NN encoder of FIG. 4 or of FIG. 5 is applied to the input image to obtain semantic description data representative of the 3D geometric and photometric model of the face of the user.
  • the model addresses only the interior part of the head, i.e the region including the eyes, nose and mouth but not the hair.
  • a preliminary step of detection and cropping of the face in the input image is performed before the encoding step.
  • the semantic description data at the output of the encoding step 820 comprises:
  • a task corresponding to the NN based encoding of semantic 3D hair model of FIG. 5 is applied to the input image to obtain semantic description data representative of the 3D geometric and photometric model of the hair of the user.
  • a NN guide curve estimator is applied to the input image to generate semantic description data representative of a 3D model of the hair region of the user’s head in the input image.
  • the encoding step 830 may comprise segmenting the hair region in the input image, and providing the segmented hair region to a NN guide curve estimator to generate a set of guide curves representative of the shapes of hair wisps in the input image and a dominant hair color.
  • the encoding step 830 may further comprise applying a guide curves NN encoder to the set of guide curves to reduce the volume of the semantic description data to be transmitted.
  • the semantic description data representative of a 3D model of the face region of the user head and the semantic description data representative of a 3D model of the hair region of the user head are provided to a remote decoder for synthesizing and rendering of the user’s head in the input image into a virtual environment.
  • FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to an embodiment.
  • FIG. 1 shows schematically a communication apparatus, for instance the videoconference device of FIG. 3 according to an embodiment.
  • the methods described above are implemented as instructions causing one or more processors to perform the methods steps.
  • FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments described above can be implemented.
  • System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers.
  • Elements of system 100 singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and/or discrete components.
  • the processing and encoder/decoder elements of system 100 are distributed across multiple ICs and/or discrete components.
  • system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports.
  • system 100 is configured to implement one or more of the aspects described in this application.
  • the system 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application.
  • Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art.
  • the system 100 includes at least one memory 120 (e.g., a volatile memory device, and/or a non-volatile memory device).
  • System 100 includes a storage device 140, which may include non-volatile memory and/or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and/or optical disk drive.
  • the storage device 140 may include an internal storage device, an attached storage device, and/or a network accessible storage device, as non-limiting examples.
  • system 100 includes an encoder/decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder/decoder module 130 may include its own processor and memory.
  • the encoder/decoder module 130 represents module(s) that may be included in a device to perform encoding and/or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder/decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
  • Program code to be loaded onto processor 110 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110.
  • one or more of processor 110, memory 120, storage device 140, and encoder/decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, one of more input video shots, mosaic images, warpings, 3D models, color transform information, visibility maps, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
  • memory inside of the processor 110 and/or the encoder/decoder module 130 is used to store instructions and to provide working memory for processing that is needed during pre-processing steps of the method described herein and/or video editing.
  • a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder/decoder module 130) is used for one or more of these functions.
  • the external memory may be the memory 120 and/or the storage device 140, for example, a dynamic volatile memory and/or a non-volatile flash memory.
  • the input to the elements of system 100 may be provided through various input devices as indicated in block 105.
  • Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal, (iii) a USB input terminal, and/or (iv) an HDMI input terminal.
  • the input devices of block 105 have associated respective input processing elements as known in the art.
  • the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets.
  • the RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, bandlimiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers.
  • the RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband.
  • the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band.
  • Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog- to-digital converter.
  • the RF portion includes an antenna.
  • USB and/or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and/or HDMI connections.
  • various aspects of input processing for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary.
  • aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary.
  • the demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder/decoder 130 operating in combination with the memory and storage elements to process the data-stream as necessary for presentation on an output device.
  • connection arrangement 115 for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
  • the system 100 includes communication interface 150 that enables communication with other devices via communication channel 190.
  • the communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190.
  • the communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and/or a wireless medium.
  • Data is streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.11.
  • the Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications.
  • the communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications.
  • Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105.
  • Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.
  • the system 100 may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185.
  • the other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100.
  • control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention.
  • the output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180.
  • the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150.
  • the display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television.
  • the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
  • the display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box.
  • the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
  • FIG. 2 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to another embodiment.
  • FIG. 2 shows schematically a communication apparatus, for instance the communication device of FIG. 3 according to an embodiment.
  • FIG. 2 shows one embodiment of an apparatus using the aforementioned methods.
  • the apparatus comprises Processor 210 and can be interconnected to a memory 220 through at least one port. Both Processor 210 and memory 220 can also have one or more additional interconnections to external connections.
  • Processor 210 is also configured to either receive an image or output a generated image and encode at least one image or decode at least one image, using the aforementioned methods.
  • the device A comprises a processor in relation with memory RAM and ROM which are configured to implement any one of the embodiments of the method for encoding at least one image as described in relation with the FIGs. 4, 5 or 8 and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement any one of the embodiments of the method for decoding at least one image as described in relation with FIGs 4, 5, or 7.
  • the network is a broadcast network, adapted to broadcast/transmit encoded images from device A to decoding devices including the device B.
  • a signal, intended to be transmitted by the device A carries at least one bitstream comprising coded data representative of at least one image.
  • FIG. 10 shows an example of the syntax of such a signal when the at least one coded image is transmitted over a packet-based transmission protocol.
  • Each transmitted packet P comprises a header H and a payload PAYLOAD.
  • each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
  • the implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program).
  • An apparatus may be implemented in, for example, appropriate hardware, software, and firmware.
  • the methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable/personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
  • PDAs portable/personal digital assistants
  • references to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment.
  • the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
  • Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
  • Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
  • this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
  • any of the following “and/or”, and “at least one of”, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B).
  • such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C).
  • This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
  • the word “signal” refers to, among other things, indicating something to a corresponding decoder.
  • the same parameter is used at both the encoder side and the decoder side.
  • an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter.
  • signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways.
  • one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
  • implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted.
  • the information may include, for example, instructions for performing a method, or data produced by one of the described implementations.
  • a signal may be formatted to carry the bitstream of a described embodiment.
  • Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal.
  • the formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream.
  • the information that the signal carries may be, for example, analog or digital information.
  • the signal may be transmitted over a variety of different wired or wireless links, as is known.
  • the signal may be stored on a processor-readable medium.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Computer Graphics (AREA)
  • Geometry (AREA)
  • Software Systems (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Processing Or Creating Images (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

Methods and apparatuses for encoding/decoding semantic description data representative of a 3D head model including face and hair modeling are provided. Such methods and apparatuses implement neural networks. In an embodiment, an image (510) comprising a head of a user is encoded (840) by extracting semantic description data representative of a 3D geometric and photometric model of the face of the user's head (520, 530, 540, 820) and extracting semantic description data representative of a 3D model of a hair region of the user's head (522, 524, 535, 830). For instance, semantic description data representative of the 3D model of the user's hair comprises a set of guide curves (535) representative of shapes of hair wisps in the input image and a dominant hair colour. In another embodiment, an image (580) of a head of the user in a virtual environment is generated (570, 750) from a synthesized image of the user's face (550, 730) and a synthesized image of the user's hair (552, 554, 556, 740) thanks to the received semantic description data (710, 720) representative of a 3D head model.

Description

Methods and apparatuses for immersive videoconference
CROSS REFERENCE TO RELATED APPLICATIONS
This application claims the benefit of European Patent Application No. 23305338.8, filed on March 13, 2023, which is incorporated herein by reference in its entirety.
TECHNICAL FIELD
The present embodiments generally relate to a method and an apparatus for encoding/decoding semantic description data representative of a 3D hair model for immersive telepresence. The present embodiments also generally relate to methods and apparatuses for encoding or decoding based on a neural network.
BACKGROUND
Telepresence refers to the use of virtual reality technology, for instance for apparent participation in distant events. A popular application is found in a telepresence videoconferencing system that immerses the participants in a single common environment. Specifically, in a typical use case, such systems are meant to ensure that a user sitting at a table in a boardroom gets the impression that the other participants are sitting at the same table in the same boardroom, and directly looking at him or her when s/he talks.
An immersive telepresence system requires some computer vision processing on the captures of distant participants to achieve its goals. Typically, these captures are 2D videos obtained by commodity cameras. A first issue is the head pose, that is the position and orientation of the head of the distant participant in the received images need to be changed at the receiver end to establish eye contact with the user in his/her viewing device. A second issue is the rendering of the hair region, for instance to fit with the head pose or to be faithful to the real appearance of the distant participant.
An efficient immersive telepresence system is therefore desirable that addresses two well-known problems for telepresence, (a) achieving proper pose and eye position of the rendered face to support proper eye contact, and (b) faithful rendering of the hair region.
SUMMARY
According to various embodiments methods and apparatuses for encoding/decoding semantic description data representative of a 3D geometric and photometric head model for immersive telepresence are provided. The head model at least comprises a 3D model of a face of a user and 3D model of the hair of the user.
According to an embodiment, a method is provided wherein the method comprises receiving semantic description data representative of a 3D model of a face region of a user’s head in an image; receiving semantic description data representative of a 3D model of hair region of the user’s head in the image; synthesizing an image of a face region of the user’s head; synthesizing an image of a hair region of the user’s head; and generating an image of a head of the user from the synthesized image of the face region of the user’s head and the synthesized image of the hair region of the user’s head.
According to another embodiment, a method is provided wherein the method comprises receiving an input image comprising a head of a user; determining, by applying neural networks to the image, semantic description data representative of a 3D model of a face region of the user’s head in the input image; determining, by applying neural networks to the input image, semantic description data representative of a 3D model of hair region of the user’s head in the input image; and providing the semantic description data representative of a 3D model of the face region of the user head and the semantic description data representative of a 3D model of the hair region of the user head to synthesize an image of the user’s head. According to an embodiment, up to two networks, respectively for face detection and face encoding, are used to determine semantic description data representative of a 3D model of a face region. According to an embodiment, up to three networks, respectively for hair segmentation, guide curve estimation and guide curve encoding, are used to determine semantic description data representative of a 3D model of a hair region.
One or more embodiments also provide an apparatus comprising one or more processors configured for performing any one of the embodiments of the methods cited above.
One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform any one of the methods according to any of the embodiments described above. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for editing a video shot, encoding at least one image or a video or decoding at least one image or a video according to the any of the embodiments described above.
One or more embodiments also provide a bitstream comprising image or video data encoded according to any one of the embodiments of the encoding method cited above. One or more of the present embodiments also provide a computer readable storage medium having stored thereon a bitstream described above. One or more embodiments also provide a method for transmitting a bitstream comprising image or video data encoded according to any one of the embodiments of the encoding method described herein. One or more embodiments also provide an apparatus for receiving a bitstream comprising image or video data encoded according to any one of the embodiments of the encoding method described herein.
BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to an embodiment.
FIG. 2 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to another embodiment.
FIG. 3 illustrates schematically a telepresence system within which aspects of the present embodiments may be implemented, according to an embodiment.
FIG. 4 illustrates schematically a telepresence system with its encoder and its decoder, within which aspects of the present embodiments may be implemented, according to an embodiment.
FIG. 5 illustrates schematically a telepresence system with its encoder and its decoder, implementing the method according to the present invention.
FIG. 6 illustrates examples of the components of a 3D environment model according to an embodiment.
FIG. 7 illustrates a method for decoding at least one image according to an embodiment.
FIG. 8 illustrates a method for encoding at least one image according to an embodiment.
FIG. 9 shows two remote devices communicating over a communication network in accordance with an example of present principles.
FIG. 10 shows the syntax of a signal in accordance with an example of present principles.
DETAILED DESCRIPTION
The present principles will be now described in the particular case of an immersive videoconference. However, the present principles are not limited to videoconferencing, but could be directly and non-ambiguously derived to any telepresence system where the representation of the user is driven by a distant capture of his or her head by a camera. Such systems include, but are not limited to, gaming frameworks where the user is represented by an avatar whose motion and expressions are driven by the distant live video capture, or more generally frameworks falling in the scope of the Metaverse where participants interact in a virtual environment through their embodiments as avatars and the avatar appearance, motion and expression is driven by a distant live video capture of the participants' heads. Such frameworks can host commercial applications such as e-learning, e-tourism and e-commerce, to name a few.
FIG. 3 illustrates schematically a telepresence system within which aspects of the present embodiments may be implemented, according to an embodiment.
The telepresence system of FIG. 3 comprises three communication apparatus D1 , D2, D3 connected through a communication network. The communication apparatus D1 comprises a camera for capturing a scene as a succession of images forming video data. The captured scene is constituted here at least by the head of a first user P1 for instance in front of a table T. The communication apparatus D1 further comprises a display for rendering a video in which the remote users P2 and P3 are displayed in an immersive environment, for instance in front of a representation of a same table T and apparently directly looking at the user P1 . As represented on the right part of FIG.3, the user P1 is also represented in the immersive environment by the display of the communication apparatus D2 or D3, for instance in front of a representation of the same table T and also apparently looking at the user P2 or P3.
To achieve the rendering of a common immersive environment in the telepresence system, the communication apparatus D1 comprises a transmitter/encoder used to process the captured video data as described below and to provide description data for synthesizing the head of the user P1 in the remote communication device D2 or D3. The communication apparatus D1 also comprises a receiver/decoder for receiving and processing description data provided by the remote communication devices D2, D3 and rendering the head of the users P2 and P3 in the immersive video displayed by the communication device D1.
Similarly, the communication apparatus D2 is capable of reproducing a video with immersive effect thanks to which user P2 viewing immersive video reproduced on apparatus D2 will get the impression that the user P1 is sitting at the same table T, and directly looking at him or her when s/he talks. This is made possible thanks to receiver/decoder of the apparatus D2 which for instance implements a semantic compression scheme, as illustrated on FIG.4, which extracts, encodes, transmits and decodes a 3D model of the participants’ faces. This is in contrast to traditional compression systems where images are encoded as arrays of pixels without regard for their semantic content. Thus, according to the embodiment of FIG.4 or FIG.5, instead of transmitting/encoding 2D video data representing the head of a first user, the transmitter/encoder of the apparatus D1 processes the 2D video data to obtain, by applying an encoder to the video data, semantic description data representative of a 3D model, for instance a 3D geometric and photometric model, of the head of the first user in the first video data; and provides the semantic description data for rendering the head of the first user in an immersive video. According to different variant embodiments illustrated in FIG. 4 or FIG. 5, the head may comprise the face of the user, or the head may comprise the face of the user and the hair of the user. Advantageously, the embodiments of FIG. 4 OR FIG. 5 allow to significantly reduce the amount of data to transmit on the communication network. Beside low bitrate, the skilled in the art will appreciate that the semantic data are independent of the resolution of the displayed video data, therefore the compression efficiency is all the more important that the resolution of the displayed video is high. Advantageously, such methods could be either embarked on a user smartphone, a user laptop or deployed on the cloud of social networks.
FIG. 4 illustrates schematically a telepresence system with its encoder and its decoder, within which aspects of the present embodiments may be implemented, according to an embodiment. One possible semantic compression scheme for videoconferencing is described in the EP patent application number 22306339.7, filed on Sept. 12, 2022 by the same applicant and represented in FIG. 4. In each captured image 410 at the transmitter end, the face region is detected 420. A pre-trained autoencoder neural network made up of a face model encoder 430 and a face model decoder 450 is leveraged to map the detected face sub-images to a semantic 3D model. This autoencoder is person-generic as it is trained on a large collection of photographs with a large variety of identities, facial expressions, lighting and head poses.
The face model encoder, also referred to as transmitter, comprises an encoding module 430 which performs a task corresponding to the 3D model extraction module. From the captured 2D video of the sender’s face, it extracts semantic description data representative of a parametric 3D face model 440 of the sender’s face. The model addresses the interior part of the face, including the eyes, nose and mouth but, at this stage, not the hair. However, according to a variant embodiment, facial hair, such as beard, eyebrow, mustache, is not managed by the hair model, but by the face region model (as a specific texture). The sender semantic description data at least comprises:
• an indication of the head pose of the sender, i.e. , the 3D rotation and translation of the face in the image, with respect to the fronto-parallel viewpoint
• an indication of an identity representative of a physiognomy of the sender with a neutral expression, for instance a 3D mesh representing the 3D geometry of the sender’s face with a neutral expression;
• an indication of an expression representative of an emotional expression with respect to the neutral expression of the sender, typically as a result of showing an emotion and/or uttering speech; for instance a plurality of displacements of vertices of the 3D mesh incurred by the facial expression; and
• an indication of an appearance representative of the texture of the sender face.
The face model is transmitted over the network then decoded at the receiver end to reconstruct the face sub-image captured at the transmitter end. Importantly, the semantic face model only addresses the interior region of the face, i.e. , the region surrounding the eyes, the nose and the mouth, but does not encode the hair and the bust regions.
At each receiving device, a face model decoder 450 retrieve semantic description data of 3D model of the sender's face and generates an image of the sender’s face interior 460. According to a particular feature of the present embodiment, the head pose of the 3D model of the sender's face and the appearance representative of the texture of the sender’s face are obtained to further drive the computation of the sender's face image, so that this image can be overlaid on the rendering of the environment. The generator module 470 at the back-end of the semantic compression pipeline of FIG. 4 reconstructs the image 480 to be displayed to the receiving participant. Its purpose is twofold: firstly, it adds realism to the synthetic reconstruction of the face interior region 460, driven by the semantic 3D face model; secondly, it “hallucinates” the missing elements to reconstruct a full image, essentially the hair and bust regions and the background. The generator 470 is person-specific. The generator 470 is a neural network that has been trained offline on images or videos of the face of the considered person. Thus, the missing elements that it generates are determined by the contents of the training images and videos and not by the contents of the live capture in videoconferencing sessions.
Besides allowing an edition of the components of the 3D model, e.g., head pose, at the receiver end, another advantage of semantic face compression lies in the compacity of the transmitted models, which allows transmission at very low bitrates. Importantly, the model contents, hence the bitrate, is independent of the input image resolution.
However, this semantic compression scheme still presents some issues. A first issue in the semantic compression scheme of FIG. 4 is that, in general, only the face interior region in the reconstructed image at the receiver end is true to the appearance of the person at the transmitting end. Indeed, the other elements of the head, in particular the hair region, are hallucinated by the generator based on the contents of the images of videos of the person it has been trained on offline. A person’s haircut as recorded in the offline training videos and images will be hallucinated by the generator and displayed to the participants of videoconference sessions, even if the person has changed his/her haircut in the meantime. This is a major downside of the compression scheme, as it potentially introduces a discrepancy between the ground truth appearance of the person as captured by his/her transmitting device and the appearance of the person that is displayed to the other participants of videoconference sessions.
A second issue of the semantic compression scheme of FIG.4 is that, unlike the face interior region, the hair region may not be edited, for instance to adjust the head pose to achieve eye contact. Indeed, the hair region in the reconstructed image is hallucinated by the generator module and not based on a 3D model of the person’s hair. The generator may, to some extent, learn to adapt the rendering of the hair region to the head pose in its input face interior image, but it is unlikely to be as faithful to the ground truth as if a 3D model of the hair was transmitted and its geometrical component edited to match a target head pose.
At least one embodiment circumvents these two issues by further estimating, encoding and transmitting a semantic 3D model for the person’s hair, in addition to and separately from the semantic 3D model of the face. As a result, a reconstruction of the hair region in the head image can be composited with the reconstruction of the face interior region to provide the generator at its input with an almost complete synthetic picture of the head. This ensures that the appearance of hair in the synthesized image matches the hair features in the transmitted image. Besides, it relieves the generator of the task of hallucinating the hair region in addition to the image background and other face regions, thereby potentially improving the quality of the reconstructed image at the receiver end. Besides, the transmitted 3D semantic model of hair can be edited at the receiver end, in particular to adjust its position, scale and orientation to ensure the hair regions in the reconstructed images are consistent with the desired rendering viewpoints in the receiving devices of the videoconference participants.
FIG. 5 illustrates schematically a telepresence system with its encoder and its decoder, implementing the method according to at least one embodiment. For clarity, but without loss of generality, the telepresence scheme will focus on a scenario with just two participants, a sender and a receiver.
In the arrangement of FIG. 5, the upper path in the diagram describes the encoding and decoding of the face interior region in the detected head region in the input image as presented in FIG. 4 and described in the EP patent application number 22306339.7, filed on Sept. 12, 2022. In a preliminary step, receiving an input image 510 comprising a head of a user is fed into the system. A neural network face autoencoder, consisting of a neural network face encoder followed by a neural network face decoder, is trained on a large collection of faces with various physiognomies, head poses, expressions and illuminants. The face encoder 530 outputs a hand-crafted semantic 3D face model 540 that consists of several vectors of numbers representing the following components: • The translation, rotation and scale defining the viewpoint of the face with reference to a fronto-parallel viewpoint where the face is seen head-on;
• an indication of an identity representative of a physiognomy of the sender with a neutral expression, for instance a 3D mesh representing the 3D geometry of the sender’s face with a neutral expression;
• an indication of an expression representative of an emotional expression with respect to the neutral expression of the sender, typically as a result of showing an emotion and/or uttering speech; for instance, a plurality of displacements of vertices of the 3D mesh incurred by the facial expression;
• an indication of an appearance representative of the texture of the user face on the surface of the 3D mesh.
According to the present embodiment, the head pose of the 3D model of the sender's face is not transmitted through the network but is determined at the decoding to achieve eye contact in the immersive environment. Thus, as previously described with FIG. 4, in steps 520, 530 semantic description data representative of a 3D model of a face region of the user’s head in the input image are determined by applying neural networks to the image.
The face decoder implements a differentiable image formation model that reconstructs the interior face region of the head image.
According to a salient characteristic of the present embodiment, the lower path is the counterpart of the upper path for the hair region. Hair consists of a large number, of the order of 100,000, of thin fibers (strands) rooted on the scalp. A simplified model of this complex 3D geometry may be obtained by considering that the strands are grouped into wisps, and that each wisp can be thought of as a cylinder inside which strands have similar 3D shapes. The central strand in each wisp is called a guide curve. The set of guide curves for all the wisps of a haircut forms a guide curve set and provides a simplified 3D hair model from which the hair image may be reconstructed. The skilled in the art will note that wisps may be made as small as needed to model isolated strands or small groups of strands.
According to at least one embodiment, the hair region in the input image of a videoconferencing system is encoded into a guide curve set, which is compressed and transmitted as a compact hair model over the network. At the receiver end, the guide curve set is decompressed and mapped to a rendering of the hair region by a neural network. Thus, with reference to the lower path in the block diagram of FIG. 5, semantic description data representative of a 3D model of hair region of the user’s head in the input image are determined by applying neural networks to the input image. According to at least one embodiment, semantic description data representative of a 3D model of the user’s hair comprises a set of guide curves representative of the shapes of hair wisps in the input image. Into more detail, the processing of the hair part in the input image comprises a first segmentation step 522 of the hair region in the input image. For instance, the hair region in the input image is segmented using a dedicated neural network, trained on pairs, wherein a pair comprises a face image as a first element of the pair, a corresponding hair segmentation mask as a second element of a pair. For instance, such network may be implemented as described in "Two-stage human hair segmentation in the wild using deep shape prior" by Y. Yan, S. Duffner, X. Naturel, A. Berthelier, C. Garcia, C. Blanc and T. Chateau, in Pattern Recognition Letters, vol. 136, pp. 293-300, 2020. The segmented hair region is fed to a guide curve estimator NN to generate the set of guide curves in step 524. Information representative of the color of the hair in the segmented hair region, such as the dominant hair color, is extracted from the segmented hair region and transmitted through the network to the decoder for rendering the hair region. In yet others variant embodiments, a segmentation mask of the hair region is transmitted through the network to the decoder as side information to assist the rendering step. A neural network trained with a dataset is used for this purpose where the dataset comprises pairs of image hair region with corresponding guide curve set. For instance, the NN Guide curve estimator 524 may be implemented as described in "HairNet: Single-View Hair Reconstruction using Convolutional Neural Networks" by Y. Zhou, L. Hu, W. Chen, H. Kung, X. Tong and H. Li, in European Conference on Computer Vision, 2018. In this dataset, guide curve sets are typically handcrafted by artists after photographs of haircuts taken from different viewpoints. Some of these photographs provide the image hair region items in the pairs. Advantageously, the guide curve sets for all the haircuts in the dataset are resampled to the same number of guide curves at fixed, pre-determined root positions on the scalp surface. This normalization of the guide curve estimation network output data makes its convergence easier. According to at least one variant embodiment, a guide curve set autoencoder network is used to reduce the volume of data to be sent over the network. Indeed, A guide curve set typically consists of several hundred curves, each represented by a fairly large number of parameters. However, in practice neighboring guide curves have similar shapes. This redundancy can be exploited to reduce the dimensionality of the information in a guide curve set. This autoencoder network is trained on a dataset of guide curves such as the guide curve set output by step 524. It consists of a guide curve set encoder 535 and a guide curve set decoder 552. The encoder network 535 encodes the guide curve set at the output of step 524 into a latent code of a predetermined dimension that is substantially smaller than the dimension of the input guide curve set, thereby providing a compact representation of said set. This latent code is sent over the network, then decoded back at the receiver end by the guide curve set decoder 552 into a close approximation of the encoded guide curve set. Thus, in the step 552, the compressed representation of the set of guide curves is decoded and the 3D model of the user’s hair comprising a set of guide curves representative of the shapes of hair wisps and a dominant hair color is reconstructed.
On the decoder side, a neural network performing the inverse functionality of step 524 is used to render a hair image from the - possibly densified - guide curve set at the output of step 552.
This network 556 is trained on the same dataset as in step 524, where the inputs and outputs are reversed. It takes as input the decoded guide curve set providing the 3D model for the hair of the subject represented in the input image 510, as well as information representative of the color of the hair that was transmitted through the network, and synthesizes an image of the hair region. Optionally, to assist in the reconstruction process, a segmentation mask of the hair region image inside its bounding box may be sent in parallel to the guide curve set geometry and used as an extra input to the neural network, to guarantee that the hair pixels are synthesized in the correct region of the output image. In an optional variant, if the density of the guide curves at the output of step 552 is too low to represent the 3D hair model, a guide curve interpolation block 554 is used to produce a denser hair model at the receiver. The synthesized image of the user’s face and the synthesized image of the user’s hair are composited 560 into a synthesized image of the user’s head. That is, the face interior image region output by the face decoder 550 of the upper path and the hair region image output by the hair model Tenderer 556 of the lower path are composited to form a more complete, but still incomplete, reconstruction of the full head image. Still missing after this composition are face regions in-between the face interior and the hair, as well as the chin and neck region and the image background. This issue is solved by applying a Generative Adversarial Network 570 to the synthesized image of the user’s head to generate an image of a head of the user. In a variant embodiment, the output of the composition is fed to a generator network, typically complemented by a discriminator network to form a Generative Adversarial Network. The generator and discriminator are trained jointly. The generator is fed with the composited face interior and hair image, which acts as a conditioning input and drives the estimation of the complete image at its output. The generator hallucinates the missing regions to form a plausible and photorealistic head image where the appearance of the face interior and hair regions at its input are roughly preserved.
According to the present principles, the semantic description data representative of a 3D model of the face region of a user head and the semantic description data representative of a 3D model of the hair region of the user head are transmitted from an encoder to a decoder to synthesize an image of the user’s head. According to a particular embodiment, the generated image 580 of a head of the user is faithful to the input image of the head of the user used to obtain the semantic description data representative of a 3D model of the face of the user and the semantic description data representative of a 3D model of the hair of the user. According to a particular embodiment, the image is part of a video, and the steps of the method are repeated for each input image.
FIG. 6 schematically illustrates examples of the components of a 3D environment model according to an embodiment. The immersive environment in which participants in the videoconference are represented is obtained from a predetermined 3D model of a scene. For example, this scene could represent a room with a floor, walls and windows, further comprising a table 610 and chairs 620 around this table 610. In this example, in case of multiple participants in the videoconference system, each user would be assigned a predetermined chair 620 on which s/he would be represented sitting in the immersive video displayed on the receiver devices of the other participants. At each receiver device, the image of the virtual environment is computed by rendering the projection of the aforementioned 3D environment model on the image plane of a virtual camera 630 whose position, attitude and optical parameters, comprising in particular its focal length, are predetermined.
FIG. 7 illustrates a generic method 700 for decoding semantic description data representative of a head of a user and generating an image of a head of the user in an immersive environment according to an embodiment. The semantic description data representative of a head of a user such as P1 in FIG. 3 comprises 2 parts: semantic description data representative of a 3D model of the face of a user, and semantic description data representative of a 3D model of the hair of the user. For instance, both semantic description data form part of a bitstream. In a step 710, semantic description data representative of a 3D model of the face (i.e. face interior) of a user is received. For instance, the 3D model of a face of a user comprises an indication of an identity representative of a physiognomy of the sender P1 of FIG. 3 with a neutral expression; an indication of an expression representative of an emotional expression with respect to the neutral expression of the sender P1 ; an indication of an appearance representative of the color of the face of the sender P1 , for instance a texture on the surface of the 3D mesh. In a step 730, an image of the sender’s face is synthesized from received semantic description data representative of the 3D geometric and photometric model of the face of a remote user P1 and from the rigid head pose of the face of the remote user. In the aforementioned example scene of FIG. 6, the rigid head pose 650 is computed so that the image resulting from the overlay represents the sender sitting on the predetermined seat that was assigned to him or her, with his or her face looking at the virtual camera. For instance, the 3D head pose model consists of scale, translation and rotation components, defined in the 3D coordinate system 640 of the predetermined 3D scene model as shown on FIG. 6. In a step 720, semantic description data representative of a 3D model of the hair of the user is received. For instance, semantic description data representative of a 3D model of the user’s hair comprises a set of guide curves representative of the shapes of hair wisps and a dominant hair color. Advantageously, a compressed representation of the set of guide curves is decoded in a subsequent step (not shown on FIG. 7). In a variant, the set of guide curves may be interpolated to obtain a denser set of guide curves (not shown on FIG. 7). Then, in a step 740, an image of the user’s hair is synthesized from received semantic description data representative of the 3D hair model. In a variant, synthesizing the image of the user’s hair may further take as input the rigid head pose of the face of the remote user for rendering the user’s head with a desired head pose. Finally, in a step 750, a photorealistic image of the user head is generated from the synthesized face and hair regions. According to the present principles, the generated image of a head of the user is faithful to the appearance of the face and hair of the user in the input image used to obtain the semantic description data representative of a 3D model of the face of the user and the semantic description data representative of a 3D model of the hair of the user. According to a variant, the generating step 750 is performed using a Generative Adversarial Network GAN. The GAN fulfills two functions: firstly, hallucinating the parts of the input image that are not encoded and transmitted through the network, such as the background, and secondly, making the renderings of the face and head image models synthesized in steps 730 and 740 more photorealistic. This GAN is a person-specific neural network that has been trained offline on images or videos of the face of the considered person. Thus, the missing elements that it generates are determined by the contents of the training images and videos and not by the contents of the live capture in the videoconferencing session. According to another variant, the foreground regions representing the user and excluding the background in the images output by the GAN are segmented out from the images output by the GAN and overlaid on a predetermined rendering of a virtual environment to generate the final rendered image displayed to the user at the receiver end. Advantageously, the training images and videos for the GAN are captured against a uniform background. Since the frames output by the GAN replicate this uniform background, the foreground region representing the sender in the rendered images can be effectively and efficiently extracted using color keying techniques known from the state of art. Although the present principles are not limited to a GAN for the generating step, the skilled in the art will appreciate that the GAN is currently the most effective implementation to generate such photorealistic images from synthetic views.
Advantageously, the method 700 consequently reduces the amount of data to transmit in a videoconference system by decoding the semantic description data of a 3D head model instead a 2D video data. Besides, the method advantageously achieves proper pose of the rendered face to support proper eye contact and provides a faithful reconstruction of both user’s face and user’s hair captured in the input image. In yet another variant, the generated image is part of a video and the decoding is repeated for each frame/image of the video.
FIG. 8 illustrates a generic method 800 for encoding semantic description data representative of a head of a user according to an embodiment. In a first step 810, an input image is received, the input image comprising the face of a user. For instance, the input image is a frame of a video. In an encoding step 820, a task corresponding to the 3D face model NN encoder of FIG. 4 or of FIG. 5 is applied to the input image to obtain semantic description data representative of the 3D geometric and photometric model of the face of the user. The model addresses only the interior part of the head, i.e the region including the eyes, nose and mouth but not the hair. According to a variant, a preliminary step of detection and cropping of the face in the input image is performed before the encoding step. According to a variant, the semantic description data at the output of the encoding step 820 comprises:
• an indication of an identity representative of a physiognomy of the user with a neutral expression, for instance a 3D mesh representing the 3D geometry of the sender’s face with a neutral expression;
• an indication of an expression representative of an emotional expression with respect to the neutral expression of the user, for instance a plurality of displacements of vertices of the 3D mesh incurred by the facial expression;
• an indication of an appearance representative of the texture of the user face on the surface of the 3D mesh; and
• an indication of a rigid head pose representative of the 3D rotation and translation of the face in the input video image with respect to the fronto-parallel viewpoint.
In principle, these components are extracted because the complete face model is needed for the autoencoder according to the variant of FIG. 4 to work, but they do not all need be transmitted as some of them are replaced by components of the immersive environment. Advantageously, even more bitrate is saved.
In a step 830, a task corresponding to the NN based encoding of semantic 3D hair model of FIG. 5 is applied to the input image to obtain semantic description data representative of the 3D geometric and photometric model of the hair of the user. A NN guide curve estimator is applied to the input image to generate semantic description data representative of a 3D model of the hair region of the user’s head in the input image. According to a variant embodiment, the encoding step 830 may comprise segmenting the hair region in the input image, and providing the segmented hair region to a NN guide curve estimator to generate a set of guide curves representative of the shapes of hair wisps in the input image and a dominant hair color. According to yet another variant, the encoding step 830 may further comprise applying a guide curves NN encoder to the set of guide curves to reduce the volume of the semantic description data to be transmitted. In a step 840, the semantic description data representative of a 3D model of the face region of the user head and the semantic description data representative of a 3D model of the hair region of the user head are provided to a remote decoder for synthesizing and rendering of the user’s head in the input image into a virtual environment.
FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to an embodiment. FIG. 1 shows schematically a communication apparatus, for instance the videoconference device of FIG. 3 according to an embodiment.
According to an embodiment, the methods described above are implemented as instructions causing one or more processors to perform the methods steps.
According to an embodiment, FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments described above can be implemented. System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and/or discrete components. For example, in at least one embodiment, the processing and encoder/decoder elements of system 100 are distributed across multiple ICs and/or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.
The system 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device, and/or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and/or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and/or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and/or a network accessible storage device, as non-limiting examples.
According to an embodiment, system 100 includes an encoder/decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder/decoder module 130 may include its own processor and memory. The encoder/decoder module 130 represents module(s) that may be included in a device to perform encoding and/or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder/decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
Program code to be loaded onto processor 110 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder/decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, one of more input video shots, mosaic images, warpings, 3D models, color transform information, visibility maps, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
In several embodiments, memory inside of the processor 110 and/or the encoder/decoder module 130 is used to store instructions and to provide working memory for processing that is needed during pre-processing steps of the method described herein and/or video editing. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder/decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and/or the storage device 140, for example, a dynamic volatile memory and/or a non-volatile flash memory.
The input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal, (iii) a USB input terminal, and/or (iv) an HDMI input terminal.
In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, bandlimiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and/or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog- to-digital converter. In various embodiments, the RF portion includes an antenna.
Additionally, the USB and/or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and/or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder/decoder 130 operating in combination with the memory and storage elements to process the data-stream as necessary for presentation on an output device.
Various elements of system 100 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and/or a wireless medium.
Data is streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.
The system 100 may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
FIG. 2 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to another embodiment. FIG. 2 shows schematically a communication apparatus, for instance the communication device of FIG. 3 according to an embodiment. FIG. 2 shows one embodiment of an apparatus using the aforementioned methods. The apparatus comprises Processor 210 and can be interconnected to a memory 220 through at least one port. Both Processor 210 and memory 220 can also have one or more additional interconnections to external connections. Processor 210 is also configured to either receive an image or output a generated image and encode at least one image or decode at least one image, using the aforementioned methods.
According to an example of the present principles, illustrated in FIG. 9, in a transmission context between two remote devices A and B over a communication network NET, the device A comprises a processor in relation with memory RAM and ROM which are configured to implement any one of the embodiments of the method for encoding at least one image as described in relation with the FIGs. 4, 5 or 8 and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement any one of the embodiments of the method for decoding at least one image as described in relation with FIGs 4, 5, or 7. In accordance with an example, the network is a broadcast network, adapted to broadcast/transmit encoded images from device A to decoding devices including the device B. A signal, intended to be transmitted by the device A, carries at least one bitstream comprising coded data representative of at least one image.
FIG. 10 shows an example of the syntax of such a signal when the at least one coded image is transmitted over a packet-based transmission protocol. Each transmitted packet P comprises a header H and a payload PAYLOAD.
Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination. Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.
The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable/personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
It is to be appreciated that the use of any of the following “and/or”, and “at least one of”, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and/or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

Claims

1. A method comprising: receiving semantic description data representative of a 3D model of a face region of a user’s head in an image; receiving semantic description data representative of a 3D model of a hair region of the user’s head in the image; synthesizing an image of a face region of the user’s head; synthesizing an image of a hair region of the user’s head; and generating an image of a head of the user from the synthesized image of the face region of the user’s head and the synthesized image of the hair region of the user’s head.
2. The method of claim 1 wherein semantic description data representative of a 3D model of the hair region comprises a set of guide curves representative of shapes of hair wisps and a dominant hair color.
3. The method of claim 2 further comprising decoding a compressed representation of the set of guide curves.
4. The method of claim 2 further comprising interpolating the set of guide curves to obtain a denser set of guide curves.
5. The method of claim 1 wherein the generating of an image further comprises: compositing the synthesized image of the face region and the synthesized image of the hair region into a synthesized image of the user’s head; and applying a Generative Adversarial Network to the synthesized image of the user’s head to generate an image of a head of the user.
6. The method of claim 5 further comprising: receiving a segmentation mask of the hair region of the user’s head in the image; and wherein the synthesizing an image of a hair region of the user’s head uses the segmentation mask.
7. The method of any of claims 1-6 wherein the generated image is part of a video.
8. The method of claim 1 wherein the generated image of a head of the user is faithful to an input image of a head of the user used to obtain the semantic description data representative of a 3D model of the face of the user and the semantic description data representative of a 3D model of the hair of the user.
9. A method comprising: receiving an input image comprising a head of a user; determining, by applying an NN encoder to the image, semantic description data representative of a 3D model of a face region of the user’s head in the input image; determining, by applying an NN guide curve estimator to the input image, semantic description data representative of a 3D model of hair region of the user’s head in the input image; and providing the semantic description data representative of a 3D model of the face region of the user head and the semantic description data representative of a 3D model of the hair region of the user head to synthesize an image of the user’s head.
10. The method of claim 9 wherein semantic description data representative of a 3D model of the user’s hair comprises a set of guide curves representative of shapes of hair wisps in the input image and a dominant hair color.
11. The method of claim 10 wherein determining semantic description data representative of a 3D model of the user’s hair in the input image further comprises: segmenting the hair region in the input image; and feeding the segmented hair region to the NN guide curve estimator to generate the set of guide curves.
12. The method of claim 11 , wherein providing semantic description data to synthesize an image of the head of the user further comprises: applying a guide curves NN encoder to the set of guide curves to reduce a volume of semantic description data to be transmitted.
13. The method of claim 11 further comprising providing a segmentation mask of the hair region.
14. The method of any of claims 9-13 wherein the input image is part of a video.
15. The method of claim 9 wherein the input image is part of a video and wherein the semantic description data representative of a 3D model of the user’s hair are determined for each input image in the video.
16. An apparatus, comprising one or more processors configured to: receive semantic description data representative of a 3D model of a face region of a user’s head in an image; receive semantic description data representative of a 3D model of a hair region of the user’s head in the image; synthesize an image of a face region of the user’s head; synthesize an image of a hair region of the user’s head; and generate an image of a head of the user from the synthesized image of the face region of the user’s head and the synthesized image of the hair region of the user’s head.
17. An apparatus, comprising one or more processors configured to: receive an input image comprising a head of a user; determine, by applying an NN encoder to the image, semantic description data representative of a 3D model of a face region of the user’s head in the input image; determine, by applying a NN guide curve estimator to the input image, semantic description data representative of a 3D model of a hair region of the user’s head in the input image; and provide the semantic description data representative of a 3D model of the face region of the user head and the semantic description data representative of a 3D model of the hair region of the user head to synthesize an image of the user’s head.
18. A device in a videoconference system comprising: an apparatus according to claim 17; and at least one apparatus according to claim 16.
19. A device in a videoconference system comprising: an apparatus according to claim 17; at least one apparatus according to claim 16; an antenna configured to receive a bitstream, the bitstream including semantic description data representative of a 3D model of a head of a remote user; and a display configured to display the generated image comprising the synthesized head of the remote user in an immersive video.
20. A bitstream comprising semantic description data for rendering a head of a user from an input image, the bitstream comprising semantic description data representative of a 3D model of a hair region of the user’s head in the input image wherein semantic description data representative of a 3D model of a hair region comprises a set of guide curves representative of shapes of hair wisps in the input image and a dominant hair color.
21. The bitstream of claim 20 wherein the bitstream further comprises semantic description data representative of a 3D model of a face region of the user’s head in the input image, wherein semantic description data representative of a 3D model of face region comprises: an indication of an identity representative of a physiognomy of the user with a neutral expression; an indication of an expression representative of an emotional expression or a deformation of the face incurred by uttering speech with respect to the neutral expression of the user; and an indication of an appearance representative of the texture of the user face.
22. A computer readable medium having stored thereon a bitstream according to any of claims 20 or 21.
23. A computer readable storage medium having stored thereon instructions for causing one or more processors to perform the method of any one of claims 1 to 15.
EP24709747.0A 2023-03-13 2024-03-08 Methods and apparatuses for immersive videoconference Pending EP4681170A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23305338 2023-03-13
PCT/EP2024/056126 WO2024188838A1 (en) 2023-03-13 2024-03-08 Methods and apparatuses for immersive videoconference

Publications (1)

Publication Number Publication Date
EP4681170A1 true EP4681170A1 (en) 2026-01-21

Family

ID=85778950

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24709747.0A Pending EP4681170A1 (en) 2023-03-13 2024-03-08 Methods and apparatuses for immersive videoconference

Country Status (4)

Country Link
EP (1) EP4681170A1 (en)
JP (1) JP2026509906A (en)
CN (1) CN120826712A (en)
WO (1) WO2024188838A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019226494A1 (en) * 2018-05-21 2019-11-28 Magic Leap, Inc. Generating textured polygon strip hair from strand-based hair for a virtual character
US11580395B2 (en) * 2018-11-14 2023-02-14 Nvidia Corporation Generative adversarial neural network assisted video reconstruction

Also Published As

Publication number Publication date
JP2026509906A (en) 2026-03-25
WO2024188838A1 (en) 2024-09-19
CN120826712A (en) 2025-10-21

Similar Documents

Publication Publication Date Title
WO2024078243A1 (en) Training method and apparatus for video generation model, and storage medium and computer device
EP3526966B1 (en) Decoder-centric uv codec for free-viewpoint video streaming
CN108449569B (en) Virtual meeting method, system, device, computer device and storage medium
JP6283108B2 (en) Image processing method and apparatus
CN111402399B (en) Face driving and live broadcasting method and device, electronic equipment and storage medium
RU2421933C2 (en) System and method to generate and reproduce 3d video image
CN101651841B (en) Method, system and equipment for realizing stereo video communication
CN101742349B (en) Method for expressing three-dimensional scenes and television system thereof
CN110663257B (en) Method and system for providing virtual reality content using 2D captured images of a scene
CN113507627A (en) Video generation method and device, electronic equipment and storage medium
CN110401810B (en) Virtual picture processing method, device and system, electronic equipment and storage medium
US20250086842A1 (en) Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
US20250014256A1 (en) Decoder, encoder, decoding method, and encoding method
JP7202087B2 (en) Video processing device
CN109769143A (en) Video image processing method, video image processing device, video system, video equipment and storage medium
WO2025007761A1 (en) Method for providing digital human, system, and computing device cluster
US10937462B2 (en) Using sharding to generate virtual reality content
JP2020005201A (en) Transmitting device and receiving device
WO2024188838A1 (en) Methods and apparatuses for immersive videoconference
US20260065583A1 (en) Methods and apparatuses for immersive videoconference
CN118555421A (en) Video processing method and device, electronic equipment and storage medium
CN114170379B (en) A three-dimensional model reconstruction method, device and equipment
CN115190289A (en) 3D holographic video screen communication method, cloud server, storage medium and electronic device
CN117836815A (en) Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device and point cloud data receiving method
US20250336097A1 (en) Encoder, decoder, encoding method, and decoding method

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250911

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR