EP4690121A1 - Avatar metadata representation - Google Patents

Avatar metadata representation

Info

Publication number
EP4690121A1
EP4690121A1 EP24712806.9A EP24712806A EP4690121A1 EP 4690121 A1 EP4690121 A1 EP 4690121A1 EP 24712806 A EP24712806 A EP 24712806A EP 4690121 A1 EP4690121 A1 EP 4690121A1
Authority
EP
European Patent Office
Prior art keywords
avatar
scene
objects
parameter
interaction
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24712806.9A
Other languages
German (de)
French (fr)
Inventor
João Pedro COVA REGATEIRO
Quentin AVRIL
Philippe Guillotel
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
InterDigital CE Patent Holdings SAS
Original Assignee
InterDigital CE Patent Holdings SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by InterDigital CE Patent Holdings SAS filed Critical InterDigital CE Patent Holdings SAS
Publication of EP4690121A1 publication Critical patent/EP4690121A1/en
Pending legal-status Critical Current

Links

Classifications

    • AHUMAN NECESSITIES
    • A63SPORTS; GAMES; AMUSEMENTS
    • A63FCARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F13/00Video games, i.e. games using an electronically generated display having two or more dimensions
    • A63F13/55Controlling game characters or game objects based on the game progress
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/20Three-dimensional [3D] animation
    • G06T13/40Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T9/00Image coding
    • G06T9/001Model-based coding, e.g. wire frame
    • AHUMAN NECESSITIES
    • A63SPORTS; GAMES; AMUSEMENTS
    • A63FCARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F2300/00Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game
    • A63F2300/50Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers
    • A63F2300/55Details of game data or player data management
    • A63F2300/5546Details of game data or player data management using player registration data, e.g. identification, account, preferences, game history
    • A63F2300/5553Details of game data or player data management using player registration data, e.g. identification, account, preferences, game history user representation in the game field, e.g. avatar

Definitions

  • the present embodiments generally relate to digital human representation and interaction within 3D-engineered virtual scenes.
  • Extended reality is a technology enabling interactive experiences where the real- world environment and/or a video content is enhanced by virtual content, which can be defined across multiple sensory modalities, including visual, auditory, haptic, etc.
  • virtual content 3D content or audio/video file for example
  • the virtual content is rendered in real-time in a way that is consistent with the user context (environment, point of view, device, etc.).
  • Scene graphs (such as the one proposed by Khronos / glTF (Graphics Language Transmission Format) and its extensions defined in MPEG Scene Description format or Apple / USDZ for instance) are a possible way to represent the content to be rendered. They combine a declarative description of the scene structure linking real-environment objects and virtual objects on one hand, and binary representations of the virtual content on the other hand. Scene description frameworks ensure that the timed media and the corresponding relevant virtual content are available at any time during the rendering of the application. Scene descriptions can also carry data at scene level describing how a user can interact with the scene objects at runtime for immersive XR experiences.
  • a method comprising: obtaining at least a parameter, from a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and obtaining 3D geometry data and texture associated with said avatar.
  • a method comprising: encoding at least a parameter, in a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that i represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and encoding 3D geometry data and texture associated with said avatar.
  • an apparatus comprising: one or more processors; and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain at least a parameter, from a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and obtain 3D geometry data and texture associated with said avatar.
  • an apparatus comprising: one or more processors; and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: encode at least a parameter, in a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and encode 3D geometry data and texture associated with said avatar.
  • One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the method according to any of the embodiments described herein.
  • One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for processing scene description according to the methods described herein.
  • One or more embodiments also provide a computer readable storage medium having stored thereon a scene description generated according to the methods described above.
  • One or more embodiments also provide a method and apparatus for transmitting or receiving the scene description generated according to the methods described herein.
  • FIG. 1 shows an example architecture of an XR processing engine.
  • FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description.
  • FIG. 3 shows an example graph of an extended reality scene description.
  • FIG. 4 shows an example of an extended reality scene description.
  • FIG. 5 illustrates an architecture of current solutions to exploit synthetic representation.
  • FIG. 6 illustrates a pipeline to introduce additional markers, according to an embodiment.
  • FIG. 7 illustrates MPEG node avatar contribution in MPEG-I SD.
  • FIG. 8 illustrates the processing model to handle such metadata, according to an embodiment.
  • Various XR applications may apply to different context and real or virtual environments.
  • a virtual 3D content item e.g., a piece A of an engine
  • a reference object piece B of an engine
  • the 3D content item is positioned in the real-world with a position and a scale defined relatively to the detected reference object.
  • a 3D model of a furniture is displayed when a given image from the catalog is detected in the input camera view.
  • the 3D content is positioned in the real-world with a position and scale defined relatively to the detected reference image.
  • some audio file might start playing when the user enters an area close to a church (being real or virtually rendered in the extended real environment).
  • an ad jingle file may be played when the user sees a can of a given soda in the real environment.
  • various virtual characters may appear, depending on the semantics of the scenery which is observed by the user.
  • bird characters are suitable for trees, so if the sensors of the XR device detect real objects described by a semantic label ‘tree’, birds can be added flying around the trees.
  • a car noise may be launched in the user’s headset when a car is detected within the field of view of the user camera, in order to warn him of the potential danger.
  • the sound may be spatialized in order to make it arrive from the direction where the car was detected.
  • An XR application may also augment a video content rather than a real environment. The video is displayed on a rendering device and virtual objects described in the node tree are overlaid when timed events are detected in the video. In such a context, the node tree comprises only virtual objects descriptions.
  • FIG. 1 shows an example architecture of an XR processing engine 130 which may be configured to implement the methods described herein.
  • a device according to the architecture of FIG. 1 is linked with other devices via their bus 131 and/or via I/O interface 136.
  • Device 130 comprises following elements that are linked together by a data and address bus 131:
  • microprocessor 132 which is, for example, a DSP (or Digital Signal Processor);
  • ROM Read Only Memory
  • RAM or Random Access Memory
  • a power supply (not represented in FIG. 1), e.g., a battery.
  • the power supply is external to the device.
  • the word “register” used in the specification may correspond to area of small capacity (some bits) or to very large area (e.g., a whole program or large amount of received or decoded data).
  • the ROM 133 comprises at least a program and parameters.
  • the ROM 133 may store algorithms and instructions to perform techniques in accordance with present principles. When switched on, the CPU 132 uploads the program in the RAM and executes the corresponding instructions.
  • the RAM 134 comprises, in a register, the program executed by the CPU 132 and uploaded after switch-on of the device 130, input data in a register, intermediate data in different states of the method in a register, and other variables used for the execution of the method in a register.
  • Device 130 is linked, for example via bus 131 to a set of sensors 137 and to a set of rendering devices 138.
  • Sensors 137 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, IR or UV light sensors or wind sensors.
  • Rendering devices 138 may be, for example, displays, speakers, vibrators, heat, fan, etc.
  • the device 130 is configured to implement a method according to the present principles, and belongs to a set comprising:
  • FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description.
  • FIG. 2 shows an example structure 210 of an XR scene description.
  • the structure consists in a container which organizes the stream in independent elements of syntax.
  • the structure may comprise a header part 220 which is a set of data common to every syntax element of the stream.
  • the header part comprises some of metadata about syntax elements, describing the nature and the role of each of them.
  • the structure also comprises a pay load comprising an element of syntax 230 and an element of syntax 240.
  • Syntax element 230 comprises data representative of the media content items described in the nodes of the scene graph related to virtual elements. Images, meshes and other raw data may have been compressed according to a compression method.
  • Element of syntax 240 is a part of the payload of the data stream and comprises data encoding the scene description as described according to the present principles.
  • FIG. 3 shows an example graph 310 of an extended reality scene description.
  • the scene graph may comprise a description of real objects, for example ‘plane horizontal surface’ (that can be a table or a road) and a description of virtual objects 312, for example an animation of a car.
  • Scene description is organized as an array of nodes.
  • a node can be linked to child nodes to form a scene structure 311.
  • a node can carry a description of a real object (e.g., a semantic description) or a description of a virtual object.
  • node 301 describes a virtual camera located in the 3D volume of the XR application.
  • Node 302 describes a virtual car and comprises an index of a representation of the car, for example an index in an array of 3D meshes.
  • Node 303 is a child of node 302 and comprises a description of one wheel of the car. The same way, it comprises an index to the 3D mesh of the wheel. The same 3D mesh may be used for several objects in the 3D scene as the scale, location and orientation of objects are described in the scene nodes.
  • Scene graph 310 also comprises nodes that are a description of the spatial relation between the real objects and the virtual objects.
  • the scene description itself can be time-evolving to provide the relevant virtual content for each sequence of a media stream.
  • a virtual bottle can be displayed on a table during a video sequence where people are seated around the table. This kind of behavior can be achieved by relying on the framework defined in the Scene Description for MPEG media document.
  • the MPEG-I Scene Description framework uses “behavior” data to augment the time-evolving scene description and provides description of how a user can interact with the scene objects at runtime for immersive XR experiences. These behaviors are related to predefined virtual objects on which runtime interactivity is allowed for user specific XR experiences. These behaviors are also time-evolving and are updated through the existing scene description update mechanism.
  • FIG. 4 shows an example of an extended reality scene description comprising behavior data, stored at scene level, describing how a user can interact with the scene objects, described at node level, at runtime for immersive XR experiences.
  • media content items e.g., meshes of virtual objects visible from the camera
  • the application displays the buffered media content item as described in related scene nodes.
  • the timing is managed by the application according to features detected in the real environment and to the timing of the animation.
  • a node of a scene graph may also comprise no description and only play a role of a parent for child nodes.
  • Behaviors 410 are related to pre-defined virtual objects on which runtime interactivity is allowed for user specific XR experiences. Behavior 410 is also time-evolving and is updated through the scene description update mechanism.
  • a behavior comprises:
  • - triggers 420 defining the conditions to be met for its activation; a trigger control parameter defining logical operations between the defined triggers; actions 430 to be proceeded processed when the triggers are activated; an action control parameter defining the order of execution of the related actions; a priority number enabling the selection of the behavior of highest priority in the case of competition between several behaviors on the same virtual object at the same time; an optional interrupt action that specifies how to terminate this behavior when it is no longer defined in a newly received scene update; for instance, a behavior is no longer defined if a related object does not belong to the new scene or if the behavior is no longer relevant for this current media (e.g., audio or video) sequence.
  • a trigger control parameter defining logical operations between the defined triggers
  • an action control parameter defining the order of execution of the related actions
  • a priority number enabling the selection of the behavior of highest priority in the case of competition between several behaviors on the same virtual object at the same
  • Behavior 410 takes place at scene level.
  • a trigger is linked to nodes and to the nodes’ child nodes.
  • Trigger 1 is linked to nodes 1, 2 and 8.
  • Trigger 1 is linked to node 31.
  • Trigger 1 is also linked to node 14 as a child of node 8.
  • Trigger 2 is linked to node 1. Indeed, a same node may be linked to several triggers.
  • Trigger n is linked to nodes 5, 6 and 7.
  • a behavior may comprise several triggers. For instance, a first behavior may be activated by trigger 1 AND trigger 2, AND being the trigger control parameter of the first behavior.
  • a behavior may have several actions. For instance, the first behavior may perform Action m first and, then action 1, “first and then” being the action control parameter of the first behavior.
  • a second behavior may be activated by trigger n and perform action 1 first and, then action 2, for example.
  • Different formats can be used to represent the node tree.
  • the MPEG-I Scene Description framework using the Khronos glTF extension mechanism may be used for the node tree.
  • an interactivity extension may apply at the glTF scene level and is called MPEG scene interactivity.
  • the corresponding semantic is provided in Table 1, where ‘M’ in ‘Usage’ column indicates that the field is mandatory in a XR scene description format and ‘O’ indicates the field is optional.
  • Digital humans can take the form through model-based reconstruction pipelines, which make assumptions of the captured subset in the form of a 3D template model.
  • This assumption can be a generic human body model with a skeletal structure attached, a subject-specific model or a statistical shape model.
  • These approaches will accurately provide a 3D model as an initialization stage capable of statistically representing different human body shapes in distinct poses.
  • This representation is mostly known as synthetic representation and is easy to manipulate and craft to match a specific body anatomy.
  • the synthetic models also facilitates appearance generalization and stylization, which can be performed by professionals and used for animation and streaming.
  • FIG. 5 A diagram of the architecture of current solutions to exploit synthetic representation is shown in FIG. 5.
  • the entry point is the sensor data, which corresponds to all input data incoming from the user’s devices. These data are then split and processed by different encoding modules: the head encoder (510), the body encoder (520), the hand encoder (530) and the head pose estimator (540). All of this encoded data is then streamed out on the network to an “Avatar Reconstruction and Animation” module (550), which reconstructs and animates the user’s avatar based on an “offline 3D model”. This avatar is then passed to any shared space to be displayed and visualized by other users.
  • This solution may be limited as it does not take into account key information such as gaze, body position and skeleton animation. Furthermore, it does not convey non-morphological and non-bio-mechanical information such as mental state, emotion or personality.
  • the sensor data is limited, not providing information on faces, eyes, skeleton structures, clothing and accessories, feet and semantical queues.
  • this approach is limited when using synthetic models for the objective of streaming and interacting with digital humans.
  • this model does not take into consideration the social behaviors, time-based animations and privacy issues that are common in the real-world and world of social technologies.
  • the encoded data and the 3D model of the receiver need to be defined.
  • all those data are known, but for an open system the format of those data needs to be specified.
  • the proposed method describes the format for humanoids. However, it can be easily extended for any type of character (e.g., animals, plants).
  • the proposed format follows the glTF format and is compatible with the current MPEG-I Scene Description (SD) effort to extend glTF with MPEG extensions.
  • SD Motion Description
  • the meaning and use is generic and can be coded with any other formats (e g., XML, USD).
  • a static representation of an avatar is a description of the attributes and characteristics that are commutable between a user and a 3D digital human (avatar). Tables 2-4 describe such attributes and separate them into four major areas. In particular, “Metadata”, “Geometry”, “Visual” and “Add-ons” are the designated overall areas to characterize an avatar, and each area is described in more detail.
  • Identity Contains all identity information of the avatar such as name, gender, age, weight, health, etc.
  • the identity information here is not necessarily the real name or age of the user, but the one set to the user avatar.
  • an appropriate mechanism could be added for protection (e.g., encryption, etc.).
  • the identity information will be less sensitive.
  • Boxes An avatar bounding region that represents the area of interaction with other avatars or objects in the scene. This can change according to the settings defined at the scene or node level.
  • the Social box corresponds to the attributes that an avatar defines in its social behavior status. This merely indicates whether an avatar wants to interact or not with other avatars.
  • a generic term will be to set a bounding region surrounding the avatar, for example, with a distance of 1 meter, which is named “Social box”. In this region, a flag “interaction” will be set and attached to the avatar node. Consequently, any other avatar/3D object that overlaps the “Social box” will have permission to interact with this avatar. And the opposite can be applied to avoid social interactions with other avatars.
  • the Contact box has some similarities with the “Social box” in terms of distance behavior. But instead of social attributes, it allows physical manipulation/interaction with the 3D objects in proximity/contact.
  • the Contact box is going to be a region around a body part to signal collisions or trigger haptic feedbacks between the avatar body part and the 3D object.
  • the Restriction box defines a region that is related to a level of permission that the avatar has in relation to accessing 3D interactive content.
  • This box can be seen as a permission access feature, such as passwords on email addresses. It can restrict or grant privileged access to the scene elements. Within a 3D virtual environment, this can be seen as content that requires a specific identifier to access or interact with, which facilitates content creators to create private room meetings or restrict user interaction to a predefined space.
  • the Experience box corresponds to the space the user is allowed to move without constraints.
  • a use case is a museum virtual experience, where the avatar users are only allowed to move freely within a certain perimeter.
  • the perimeter where the “art” is located is out of limits for avatar users to move or interact, but not restricted to view.
  • the Parental box limits the interaction with allowed content to protect children and young adults by parental control.
  • Disability Description of disabilities of the avatar senses that can represent the human user. It also can include some missing parts of the user body (e.g., arm, legs etc.). It might also include robotic/prosthesis body parts replacing the user’s body parts.
  • the introduction of a user’s disability is useful to permit the engine to handle/render the appropriated content for the user, e.g., a real user with speech impairment, won’t be able to use microphone input if the application is a video conference tool; if the user has hearing disabilities the application should render visual cues and synthesize text from speech.
  • Capability An attribute that describes the abilities of the avatar, such as, ablility to run, walk, jump, talk, fly, etc. In the case of Disabilities, it could also represent the impact on the capabilities.
  • the type of capability can be restricted to triggered events that only allow certain actions to be performed, e.g., in a meeting room the spectating avatars are only allowed to use speech, and upper body motion, such as, gestures and head motions.
  • This attribute facilitates client applications to assign pre-determined functionalities to specific time-based events.
  • This attribute is defined at the avatar level and can be complementary with other capabilities defined at the scene level, to either restrict or augment the default avatar capabilities. Nevertheless, this represents what the avatar by default is able to perform, e.g., pre-defined set of animations.
  • Personality influences the social behavior of a person, hence impacting the animation of an avatar, for example, calm, introverted, extroverted, outgoing, friendly, etc.
  • personality can influence the box region and impact the area of social interaction and permission for interaction. Therefore, personality attributes have an important impact and role in social interactions and privacy depending on individual characteristics.
  • Emotion describes the current avatar state, such as, happy, sad, angry, frustrated, etc.
  • the type of emotion strongly impacts the motor behavior of an individual, hence client applications can infer predetermined motion behavior depending on this attribute. For example, erratic emotion displays an unregular pattern of movement.
  • Geometry refers to the full body shape and semantics.
  • an article by Q. Avril et al. entitled “Draft Annex to ISO/IEC 23090-14:2021 - MPEG Reference Humanoid Avatar” (ISO/IEC JTC 1/SC 29/WG 3 m61232, hereinafter “Avril”) describes the skeletal anatomy and mesh graph connectivity of the avatar.
  • the full body is separated into two sections, i.e., the upper and lower body parts, and each section is split into further dependent sections.
  • the body model is represented by a mesh, which contains vertices, faces and normals as detailed in Avril.
  • the semantic attribute is a dictionary that correlates each body section with the vertices/faces of the avatar. This facilitates the lower-level information (vertices) to directly access high level data (body regions) and generate avatar to scene and avatar to avatar interactions, such as contact triggers.
  • Table 3 illustrates semantical description of the “Geometry” column in Table 2.
  • Table 4 illustrates semantical description of the “Visual” column in Table 2.
  • Avatars are described as digital representations of humans. There are two purposes for creating a human avatar: user representation or virtual scene interactions.
  • the representation of a human in the digital context can be achieved in two different manners, either volumetric or synthetic.
  • Volumetric representations are 3D objects that naturally encapsulate the human shape and appearance.
  • volumetric content accurately represents the user body anatomy, e.g., shape morphology, limbs dimension, height or a particular wardrobe, and the user appearance, such as, eyes/hair/skin color, subtleties to the natural skin (freckles) or patterns found in different clothing.
  • Synthetic representations are 3D crafted or statistically learned 3D objects of human bodies. Synthetic representations are easier to manipulate and craft to match a specific body anatomy. Synthetic models also facilitate appearance generalization and stylization, which can be performed by professionals and used for animation streaming.
  • the 3D model is extended with facial landmarks, eye tracking information, skeletal structures, and body semantics. In this manner, the model is personalized to the user and ready for scene interactions.
  • FIG. 6 illustrates a proposed pipeline extension to process captured human shape and pose with the objective of animating and representing 3D representation of a user, according to an embodiment. Except blocks 610, 620, 621, 630, 635, 645 and 660, the blocks in FIG. 6 signal the proposed extensions.
  • the sensors (610) are general input devices (cameras, controllers, or IMUs), and the head, body, clothes, hands, and feet encoders (620, 621, 622, 623, 624, 625) are processing models that extract specific information from the sensors and feed it to the avatar reconstruction, retargeting and animation module (650). This process generates a personalized avatar that is fed into the shared space (660) for rendering proposes.
  • Each processing model is divided into sub-processing models (631, 632, 633, 634, 636, 637) to specialize the data from the sensors.
  • the models (630, 631, 632, 633) are related to the head information, and they have the objective of extracting, facial landmarks and position (640), eye gaze and shape (641), hair shape and pose (642), and jaws shape (643).
  • the model (634, 635, 636) are related to the body, and it extracts skeletal hierarchy, shape, pose, spatial position (644), and the relation between body part semantics and respective geometry (644).
  • the model (623) extracts skeletal and pose information (646), for the hands in the sensor data.
  • the model (624) computes the contact between feet and ground (647) to synchronize animations.
  • the model (625) contains additional metadata information that facilitates the personalization of digital avatars.
  • an avatar can take many shapes and forms allowing unique creations of avatars that closely resemble a real human’s physical and mental characteristic. Therefore, it is necessary to reconstruct and define the properties prior to instantiate a dynamic 3D mesh of an Avatar.
  • an avatar should respect the following requirements:
  • the avatar may be reconstructed or animated with a wide range of specifications presented above.
  • the avatar is represented with semantical labels which allow for individual body part interactivity triggers.
  • the avatar’s shape contains a surrounding region with dynamic dimensions. This facilitates interaction triggers with scene, other avatars and implements social barriers between the avatar and the dynamic elements of a scene.
  • the proposed avatar is reconstructed given the user inputs and mapped onto a reference 3D avatar, for example, the 3D avatar “MPEG reference avatar” presented in Avril, which contains a full body-based mesh representation, skeleton and skinning weights, facial blend shapes and landmarks, and additional geometry including eyes, jaws, teeth and tongue.
  • a reference 3D avatar for example, the 3D avatar “MPEG reference avatar” presented in Avril, which contains a full body-based mesh representation, skeleton and skinning weights, facial blend shapes and landmarks, and additional geometry including eyes, jaws, teeth and tongue.
  • Avril contains a full body-based mesh representation, skeleton and skinning weights, facial blend shapes and landmarks, and additional geometry including eyes, jaws, teeth and tongue.
  • This common representation eases animation and mapping of any animation to different morphologies or context.
  • FIG. 7 illustrates MPEG node avatar contribution in MPEG-I SD (see an article by E. Thomas et al., entitled “[SD] Signalling for Avatars in Scene Description”, ISO/IEC JTC 1/SC 29/WG 3 m62027, hereinafter “m62027”).
  • FIG. 7 represents a glTF file structure.
  • the entry point is the “scene” node (730) which contains “node” node(s) (735) that can be either a “camera” (710), a “light” (765) or a “mesh” (740).
  • the MPEG-I SD has defined a new Boolean attribute “MPEG node avatar” (735) indicating whether the node corresponds to an avatar node or not, as illustrated in Table 6. The purpose of this Boolean attribute is to define which node defines the user representation.
  • the metadata contains all elements described before that define/identify the avatar (identity, age, name, gender etc.), as illustrated in Table 7.
  • Table 7
  • Table 8 illustrates format description of the “Geometry” proprieties presented in Table 2.
  • Table 9 illustrates a semantical list of the available names from the avatar model. These names are used on the mapping propriety presented in Table 8.
  • This attribute signals the client about the capabilities of the avatar.
  • the avatar by default is capable of jumping, flying, and running.
  • This signal will provide semantics to the animation engine, so it generates and handles the following capabilities.
  • the engine will have the objective of guaranteeing that this animation is available to the final user. From a format point of view, this signal can be used to evaluate if the conditions of the 3D environment are capable and respect the default capabilities of the avatar.
  • the user/application side will have several manners to interpret the metadata.
  • the common one and supported technique is interaction between the scene object and the avatar.
  • the scene is going to be composed of nodes and each node represents 3D meshes.
  • the avatar node will trigger events when interacting with nodes in the scene, hence there will be some node interaction that can use the metadata information store at the avatar node level.
  • the Parental attribute will be compared against the interactive node, and if the avatar parental attribute is higher than a parental attribute present in such node then the interaction is allowed otherwise stopped. This behavior is chosen by the processing/engine model, and several behaviors can be instantiated from this standard format.
  • Avatars are represented with 3D geometries that make most of the visuals and manipulators to generate realistic content in 3D applications.
  • the issue arrives when such user virtual representation becomes social and gains access to interaction with other virtual users or objects, and needs additional information that facilitates, improves, and restricts such interactions.
  • this document provides additional metadata for avatar representation in virtual 3D environments.
  • FIG. 8 illustrates the processing model to handle such metadata, according to an embodiment.
  • the engine will process all dynamic (sensors input data) and static data (user metadata input and scene description) and generate a unique user avatar reconstruction and animation.
  • the Human Features Encoder (810) relates to the encoders proposed and illustrated in FIG. 6.
  • the Engine (835) refers to an application, such as hot not exclusive Unreal, Unity or blender, that uses the decoded (830) content provided by the Humans Features Encoder and Scene Description + Avatar Metadata (820) to create a personalized avatar.
  • FIG. 8 illustrates processing of dynamic sensor data and user static metadata to generate a personalized user avatar experience.
  • the sensors represent real-world devices information that are fed to a human encoder model to extract relevant information (e.g., human pose, motion, appearance) to generate a personalized avatar.
  • relevant information e.g., human pose, motion, appearance
  • the scene description file and avatar metadata are combined with the encoded features to inform engines how to render (840), reconstruct (850), animate (860) and interact (870) with other virtual objects, hence the definition of personalized avatar.
  • each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
  • the implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program).
  • An apparatus may be implemented in, for example, appropriate hardware, software, and firmware.
  • the methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable/personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
  • PDAs portable/personal digital assistants
  • references to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment.
  • the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
  • this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
  • Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
  • this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
  • such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C).
  • This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
  • implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted.
  • the information may include, for example, instructions for performing a method, or data produced by one of the described implementations.
  • a signal may be formatted to carry the bitstream of a described embodiment.
  • Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal.
  • the formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream.
  • the information that the signal carries may be, for example, analog or digital information.
  • the signal may be transmitted over a variety of different wired or wireless links, as is known.
  • the signal may be stored on a processor-readable medium.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Processing Or Creating Images (AREA)
  • Information Transfer Between Computers (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

In one implementation, additional metadata is proposed to normalize social interactivity between computer generated 3D models in a virtual environment, such as dynamic objects or static objects. For example, "Metadata", "Geometry", "Visual" and "Add-ons" are used to characterize the avatar. In particular, "Metadata" can indicate the identity, vital space, disability, capability, personality and emotion of the avatar; "Geometry" refers to the full body shape and semantics; "Visual" provides descriptions of the texture map, materials used for rendering the avatar, overlay and the depth of detail of the appearance of the avatar; and "Add- ons" can specify the clothes and accessories of the avatar. To generate realistic content in 3D application, an engine will process all dynamic (sensors input data) and static data (user metadata input and scene description) and generate a unique user avatar reconstruction and animation.

Description

AVATAR METADATA REPRESENTATION
TECHNICAL FIELD
[1] The present embodiments generally relate to digital human representation and interaction within 3D-engineered virtual scenes.
BACKGROUND
[2] Extended reality (XR) is a technology enabling interactive experiences where the real- world environment and/or a video content is enhanced by virtual content, which can be defined across multiple sensory modalities, including visual, auditory, haptic, etc. During runtime of the application, the virtual content (3D content or audio/video file for example) is rendered in real-time in a way that is consistent with the user context (environment, point of view, device, etc.).
[3] Scene graphs (such as the one proposed by Khronos / glTF (Graphics Language Transmission Format) and its extensions defined in MPEG Scene Description format or Apple / USDZ for instance) are a possible way to represent the content to be rendered. They combine a declarative description of the scene structure linking real-environment objects and virtual objects on one hand, and binary representations of the virtual content on the other hand. Scene description frameworks ensure that the timed media and the corresponding relevant virtual content are available at any time during the rendering of the application. Scene descriptions can also carry data at scene level describing how a user can interact with the scene objects at runtime for immersive XR experiences.
SUMMARY
[4] According to an embodiment, a method is presented, comprising: obtaining at least a parameter, from a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and obtaining 3D geometry data and texture associated with said avatar.
[5] According to another embodiment, a method is presented, comprising: encoding at least a parameter, in a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that i represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and encoding 3D geometry data and texture associated with said avatar.
[6] According to another embodiment, an apparatus is presented, comprising: one or more processors; and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain at least a parameter, from a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and obtain 3D geometry data and texture associated with said avatar.
[7] According to an embodiment, an apparatus is presented, comprising: one or more processors; and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: encode at least a parameter, in a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following: identity, a bounding region that represents an area of interaction of said avatar with other objects in said scene, disabilities of said avatar, capabilities of said avatar, personality of said avatar, and emotion of said avatar; and encode 3D geometry data and texture associated with said avatar.
[8] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for processing scene description according to the methods described herein.
[9] One or more embodiments also provide a computer readable storage medium having stored thereon a scene description generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the scene description generated according to the methods described herein.
BRIEF DESCRIPTION OF THE DRAWINGS
[10] FIG. 1 shows an example architecture of an XR processing engine.
[11] FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description.
[12] FIG. 3 shows an example graph of an extended reality scene description.
[13] FIG. 4 shows an example of an extended reality scene description.
[14] FIG. 5 illustrates an architecture of current solutions to exploit synthetic representation.
[15] FIG. 6 illustrates a pipeline to introduce additional markers, according to an embodiment.
[16] FIG. 7 illustrates MPEG node avatar contribution in MPEG-I SD.
[17] FIG. 8 illustrates the processing model to handle such metadata, according to an embodiment.
DETAILED DESCRIPTION
[18] Various XR applications may apply to different context and real or virtual environments. For example, in an industrial XR application, a virtual 3D content item (e.g., a piece A of an engine) is displayed when a reference object (piece B of an engine) is detected in the real environment by a camera rigged on a head mounted display device. The 3D content item is positioned in the real-world with a position and a scale defined relatively to the detected reference object.
[19] For example, in an XR application for interior design, a 3D model of a furniture is displayed when a given image from the catalog is detected in the input camera view. The 3D content is positioned in the real-world with a position and scale defined relatively to the detected reference image. In another application, some audio file might start playing when the user enters an area close to a church (being real or virtually rendered in the extended real environment). In another example, an ad jingle file may be played when the user sees a can of a given soda in the real environment. In an outdoor gaming application, various virtual characters may appear, depending on the semantics of the scenery which is observed by the user. For example, bird characters are suitable for trees, so if the sensors of the XR device detect real objects described by a semantic label ‘tree’, birds can be added flying around the trees. In a companion application implemented by smart glasses, a car noise may be launched in the user’s headset when a car is detected within the field of view of the user camera, in order to warn him of the potential danger. Furthermore, the sound may be spatialized in order to make it arrive from the direction where the car was detected. [20] An XR application may also augment a video content rather than a real environment. The video is displayed on a rendering device and virtual objects described in the node tree are overlaid when timed events are detected in the video. In such a context, the node tree comprises only virtual objects descriptions.
[21] FIG. 1 shows an example architecture of an XR processing engine 130 which may be configured to implement the methods described herein. A device according to the architecture of FIG. 1 is linked with other devices via their bus 131 and/or via I/O interface 136.
[22] Device 130 comprises following elements that are linked together by a data and address bus 131:
- a microprocessor 132 (or CPU), which is, for example, a DSP (or Digital Signal Processor);
- a ROM (or Read Only Memory) 133;
- a RAM (or Random Access Memory) 134;
- a storage interface 135;
- an I/O interface 136 for reception of data to transmit, from an application; and
- a power supply (not represented in FIG. 1), e.g., a battery.
[23] In accordance with an example, the power supply is external to the device. In each of mentioned memory, the word “register” used in the specification may correspond to area of small capacity (some bits) or to very large area (e.g., a whole program or large amount of received or decoded data). The ROM 133 comprises at least a program and parameters. The ROM 133 may store algorithms and instructions to perform techniques in accordance with present principles. When switched on, the CPU 132 uploads the program in the RAM and executes the corresponding instructions.
[24] The RAM 134 comprises, in a register, the program executed by the CPU 132 and uploaded after switch-on of the device 130, input data in a register, intermediate data in different states of the method in a register, and other variables used for the execution of the method in a register.
[25] Device 130 is linked, for example via bus 131 to a set of sensors 137 and to a set of rendering devices 138. Sensors 137 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, IR or UV light sensors or wind sensors. Rendering devices 138 may be, for example, displays, speakers, vibrators, heat, fan, etc. [26] In accordance with examples, the device 130 is configured to implement a method according to the present principles, and belongs to a set comprising:
- a mobile device;
- a communication device;
- a game device;
- a tablet (or tablet computer);
- a laptop;
- a still picture camera;
- a video camera.
[27] In XR applications, scene description is used to combine explicit and easy-to-parse description of a scene structure and some binary representations of media content. FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description. FIG. 2 shows an example structure 210 of an XR scene description. The structure consists in a container which organizes the stream in independent elements of syntax. The structure may comprise a header part 220 which is a set of data common to every syntax element of the stream. For example, the header part comprises some of metadata about syntax elements, describing the nature and the role of each of them. The structure also comprises a pay load comprising an element of syntax 230 and an element of syntax 240. Syntax element 230 comprises data representative of the media content items described in the nodes of the scene graph related to virtual elements. Images, meshes and other raw data may have been compressed according to a compression method. Element of syntax 240 is a part of the payload of the data stream and comprises data encoding the scene description as described according to the present principles.
[28] FIG. 3 shows an example graph 310 of an extended reality scene description. In this example, the scene graph may comprise a description of real objects, for example ‘plane horizontal surface’ (that can be a table or a road) and a description of virtual objects 312, for example an animation of a car. Scene description is organized as an array of nodes. A node can be linked to child nodes to form a scene structure 311. A node can carry a description of a real object (e.g., a semantic description) or a description of a virtual object. In the example of FIG. 3, node 301 describes a virtual camera located in the 3D volume of the XR application. Node 302 describes a virtual car and comprises an index of a representation of the car, for example an index in an array of 3D meshes. Node 303 is a child of node 302 and comprises a description of one wheel of the car. The same way, it comprises an index to the 3D mesh of the wheel. The same 3D mesh may be used for several objects in the 3D scene as the scale, location and orientation of objects are described in the scene nodes. Scene graph 310 also comprises nodes that are a description of the spatial relation between the real objects and the virtual objects.
[29] In time-based media streaming, the scene description itself can be time-evolving to provide the relevant virtual content for each sequence of a media stream. For instance, for advertising purpose, a virtual bottle can be displayed on a table during a video sequence where people are seated around the table. This kind of behavior can be achieved by relying on the framework defined in the Scene Description for MPEG media document.
[30] Currently, the MPEG-I Scene Description framework uses “behavior” data to augment the time-evolving scene description and provides description of how a user can interact with the scene objects at runtime for immersive XR experiences. These behaviors are related to predefined virtual objects on which runtime interactivity is allowed for user specific XR experiences. These behaviors are also time-evolving and are updated through the existing scene description update mechanism.
[31] FIG. 4 shows an example of an extended reality scene description comprising behavior data, stored at scene level, describing how a user can interact with the scene objects, described at node level, at runtime for immersive XR experiences. When the XR application is started, media content items (e.g., meshes of virtual objects visible from the camera) are loaded, rendered and buffered to be displayed when triggered. For example, when a plane surface is detected in the real environment by sensors, the application displays the buffered media content item as described in related scene nodes. The timing is managed by the application according to features detected in the real environment and to the timing of the animation. A node of a scene graph may also comprise no description and only play a role of a parent for child nodes. FIG. 4 shows relationships between behaviors that are comprised in the scene description at the scene level and nodes that are components of the scene graph. Behaviors 410 are related to pre-defined virtual objects on which runtime interactivity is allowed for user specific XR experiences. Behavior 410 is also time-evolving and is updated through the scene description update mechanism.
[32] A behavior comprises:
- triggers 420 defining the conditions to be met for its activation; a trigger control parameter defining logical operations between the defined triggers; actions 430 to be proceeded processed when the triggers are activated; an action control parameter defining the order of execution of the related actions; a priority number enabling the selection of the behavior of highest priority in the case of competition between several behaviors on the same virtual object at the same time; an optional interrupt action that specifies how to terminate this behavior when it is no longer defined in a newly received scene update; for instance, a behavior is no longer defined if a related object does not belong to the new scene or if the behavior is no longer relevant for this current media (e.g., audio or video) sequence.
[33] Behavior 410 takes place at scene level. A trigger is linked to nodes and to the nodes’ child nodes. In the example of FIG. 4, Trigger 1 is linked to nodes 1, 2 and 8. As Node 31 is a child of node 1, Trigger 1 is linked to node 31. Trigger 1 is also linked to node 14 as a child of node 8. Trigger 2 is linked to node 1. Indeed, a same node may be linked to several triggers. Trigger n is linked to nodes 5, 6 and 7. A behavior may comprise several triggers. For instance, a first behavior may be activated by trigger 1 AND trigger 2, AND being the trigger control parameter of the first behavior. A behavior may have several actions. For instance, the first behavior may perform Action m first and, then action 1, “first and then” being the action control parameter of the first behavior. A second behavior may be activated by trigger n and perform action 1 first and, then action 2, for example.
[34] Different formats can be used to represent the node tree. For example, the MPEG-I Scene Description framework using the Khronos glTF extension mechanism may be used for the node tree. In this example, an interactivity extension may apply at the glTF scene level and is called MPEG scene interactivity. The corresponding semantic is provided in Table 1, where ‘M’ in ‘Usage’ column indicates that the field is mandatory in a XR scene description format and ‘O’ indicates the field is optional.
TABLE 1 [35] Current solutions for avatar representation
[36] Digital humans can take the form through model-based reconstruction pipelines, which make assumptions of the captured subset in the form of a 3D template model. This assumption can be a generic human body model with a skeletal structure attached, a subject-specific model or a statistical shape model. These approaches will accurately provide a 3D model as an initialization stage capable of statistically representing different human body shapes in distinct poses. This representation is mostly known as synthetic representation and is easy to manipulate and craft to match a specific body anatomy. The synthetic models also facilitates appearance generalization and stylization, which can be performed by professionals and used for animation and streaming.
[37] A diagram of the architecture of current solutions to exploit synthetic representation is shown in FIG. 5. The entry point is the sensor data, which corresponds to all input data incoming from the user’s devices. These data are then split and processed by different encoding modules: the head encoder (510), the body encoder (520), the hand encoder (530) and the head pose estimator (540). All of this encoded data is then streamed out on the network to an “Avatar Reconstruction and Animation” module (550), which reconstructs and animates the user’s avatar based on an “offline 3D model”. This avatar is then passed to any shared space to be displayed and visualized by other users. This solution may be limited as it does not take into account key information such as gaze, body position and skeleton animation. Furthermore, it does not convey non-morphological and non-bio-mechanical information such as mental state, emotion or personality.
[38] As can be seen from FIG. 5, the sensor data is limited, not providing information on faces, eyes, skeleton structures, clothing and accessories, feet and semantical queues. Thus, this approach is limited when using synthetic models for the objective of streaming and interacting with digital humans. In addition, this model does not take into consideration the social behaviors, time-based animations and privacy issues that are common in the real-world and world of social technologies.
[39] In addition, in an actual distribution or communication use case the encoded data and the 3D model of the receiver (as well as the encoder side) need to be defined. In a proprietary system all those data are known, but for an open system the format of those data needs to be specified. In this document, we provide mechanisms to provide the information and associated coding format used for the representation and coding of a user representation (aka avatar). [40] In the following, the proposed method describes the format for humanoids. However, it can be easily extended for any type of character (e.g., animals, plants).
[41] Proposed Avatar Representation
[42] The proposed representation of an avatar, which is intended to be compatible with a scene description (SD) content, can be divided into two main areas:
- the static representation - which contains metadata (ID etc.) and a static representation
- the dynamic representation - which animates and updates the static representation
[43] The following details those elements, with the associated meaning, JSON coding schemes and how it can be used.
[44] In the following description, the proposed format follows the glTF format and is compatible with the current MPEG-I Scene Description (SD) effort to extend glTF with MPEG extensions. However, the meaning and use is generic and can be coded with any other formats (e g., XML, USD).
[45] Static Representation
[46] A static representation of an avatar is a description of the attributes and characteristics that are commutable between a user and a 3D digital human (avatar). Tables 2-4 describe such attributes and separate them into four major areas. In particular, “Metadata”, “Geometry”, “Visual” and “Add-ons” are the designated overall areas to characterize an avatar, and each area is described in more detail.
Table 2 - Static information for a generic representation of an avatar. [47] Metadata
[48] Identity Contains all identity information of the avatar such as name, gender, age, weight, health, etc. The identity information here is not necessarily the real name or age of the user, but the one set to the user avatar. In the case several avatars are used to represent a user, there are several versions of the user. Considering the sensibility of those data, an appropriate mechanism could be added for protection (e.g., encryption, etc.). In the case of pseudo and non-real values, the identity information will be less sensitive.
[49] Boxes An avatar bounding region that represents the area of interaction with other avatars or objects in the scene. This can change according to the settings defined at the scene or node level.
[50] An illustrative use case of such regions can be for instance:
- The Social box corresponds to the attributes that an avatar defines in its social behavior status. This merely indicates whether an avatar wants to interact or not with other avatars. A generic term will be to set a bounding region surrounding the avatar, for example, with a distance of 1 meter, which is named “Social box”. In this region, a flag “interaction” will be set and attached to the avatar node. Consequently, any other avatar/3D object that overlaps the “Social box” will have permission to interact with this avatar. And the opposite can be applied to avoid social interactions with other avatars.
- The Contact box has some similarities with the “Social box” in terms of distance behavior. But instead of social attributes, it allows physical manipulation/interaction with the 3D objects in proximity/contact. The Contact box is going to be a region around a body part to signal collisions or trigger haptic feedbacks between the avatar body part and the 3D object.
- The Restriction box defines a region that is related to a level of permission that the avatar has in relation to accessing 3D interactive content. This box can be seen as a permission access feature, such as passwords on email addresses. It can restrict or grant privileged access to the scene elements. Within a 3D virtual environment, this can be seen as content that requires a specific identifier to access or interact with, which facilitates content creators to create private room meetings or restrict user interaction to a predefined space.
- The Experience box corresponds to the space the user is allowed to move without constraints. A use case is a museum virtual experience, where the avatar users are only allowed to move freely within a certain perimeter. The perimeter where the “art” is located is out of limits for avatar users to move or interact, but not restricted to view. - The Parental box limits the interaction with allowed content to protect children and young adults by parental control.
[51] Disability Description of disabilities of the avatar, such as, impairing on mobility, speech or any other humans’ senses that can represent the human user. It also can include some missing parts of the user body (e.g., arm, legs etc.). It might also include robotic/prosthesis body parts replacing the user’s body parts.
[52] The introduction of a user’s disability is useful to permit the engine to handle/render the appropriated content for the user, e.g., a real user with speech impairment, won’t be able to use microphone input if the application is a video conference tool; if the user has hearing disabilities the application should render visual cues and synthesize text from speech.
[53] Capability An attribute that describes the abilities of the avatar, such as, ablility to run, walk, jump, talk, fly, etc. In the case of Disabilities, it could also represent the impact on the capabilities. The type of capability can be restricted to triggered events that only allow certain actions to be performed, e.g., in a meeting room the spectating avatars are only allowed to use speech, and upper body motion, such as, gestures and head motions. This attribute facilitates client applications to assign pre-determined functionalities to specific time-based events. This attribute is defined at the avatar level and can be complementary with other capabilities defined at the scene level, to either restrict or augment the default avatar capabilities. Nevertheless, this represents what the avatar by default is able to perform, e.g., pre-defined set of animations.
[54] Personality Social attribute influences the social behavior of a person, hence impacting the animation of an avatar, for example, calm, introverted, extroverted, outgoing, friendly, etc. Personality can influence the box region and impact the area of social interaction and permission for interaction. Therefore, personality attributes have an important impact and role in social interactions and privacy depending on individual characteristics.
[55] Emotion The emotion attribute describes the current avatar state, such as, happy, sad, angry, frustrated, etc. The type of emotion strongly impacts the motor behavior of an individual, hence client applications can infer predetermined motion behavior depending on this attribute. For example, erratic emotion displays an unregular pattern of movement.
[56] Geometry
[57] Geometry refers to the full body shape and semantics. For example, an article by Q. Avril et al. entitled “Draft Annex to ISO/IEC 23090-14:2021 - MPEG Reference Humanoid Avatar” (ISO/IEC JTC 1/SC 29/WG 3 m61232, hereinafter “Avril") describes the skeletal anatomy and mesh graph connectivity of the avatar. The full body is separated into two sections, i.e., the upper and lower body parts, and each section is split into further dependent sections. The body model is represented by a mesh, which contains vertices, faces and normals as detailed in Avril.
[58] The semantic attribute is a dictionary that correlates each body section with the vertices/faces of the avatar. This facilitates the lower-level information (vertices) to directly access high level data (body regions) and generate avatar to scene and avatar to avatar interactions, such as contact triggers. [59] Table 3 illustrates semantical description of the “Geometry” column in Table 2.
Table 3
[60] Visuals
[61] Table 4 illustrates semantical description of the “Visual” column in Table 2.
Table 4
[62] Add-ons [63] Table 5 illustrates semantical description of the “Add-ons” column in Table 2.
Table 5
[64] Dynamic Representation [65] Avatars are described as digital representations of humans. There are two purposes for creating a human avatar: user representation or virtual scene interactions. The representation of a human in the digital context can be achieved in two different manners, either volumetric or synthetic. Volumetric representations are 3D objects that naturally encapsulate the human shape and appearance. Hence, volumetric content accurately represents the user body anatomy, e.g., shape morphology, limbs dimension, height or a particular wardrobe, and the user appearance, such as, eyes/hair/skin color, subtleties to the natural skin (freckles) or patterns found in different clothing. Synthetic representations are 3D crafted or statistically learned 3D objects of human bodies. Synthetic representations are easier to manipulate and craft to match a specific body anatomy. Synthetic models also facilitate appearance generalization and stylization, which can be performed by professionals and used for animation streaming.
[66] In the previous example as illustrated in FIG. 5, the exploited data from the sensor data is limited when using synthetic models. Hence, according to an embodiment, we propose to extend the pipeline as illustrated in FIG. 6, to introduce additional markers to allow the generation and creation of a robust avatar, which closely mimics the human anatomy, mechanics and social interaction.
[67] In the proposed example, the 3D model is extended with facial landmarks, eye tracking information, skeletal structures, and body semantics. In this manner, the model is personalized to the user and ready for scene interactions.
[68] In particular, FIG. 6 illustrates a proposed pipeline extension to process captured human shape and pose with the objective of animating and representing 3D representation of a user, according to an embodiment. Except blocks 610, 620, 621, 630, 635, 645 and 660, the blocks in FIG. 6 signal the proposed extensions. The sensors (610) are general input devices (cameras, controllers, or IMUs), and the head, body, clothes, hands, and feet encoders (620, 621, 622, 623, 624, 625) are processing models that extract specific information from the sensors and feed it to the avatar reconstruction, retargeting and animation module (650). This process generates a personalized avatar that is fed into the shared space (660) for rendering proposes. Each processing model is divided into sub-processing models (631, 632, 633, 634, 636, 637) to specialize the data from the sensors. The models (630, 631, 632, 633) are related to the head information, and they have the objective of extracting, facial landmarks and position (640), eye gaze and shape (641), hair shape and pose (642), and jaws shape (643). The model (634, 635, 636) are related to the body, and it extracts skeletal hierarchy, shape, pose, spatial position (644), and the relation between body part semantics and respective geometry (644). Similarly, the model (623) extracts skeletal and pose information (646), for the hands in the sensor data. The model (624) computes the contact between feet and ground (647) to synchronize animations. Lastly, the model (625) contains additional metadata information that facilitates the personalization of digital avatars.
[69] Avatar Reconstruction Standards
[70] As described above, an avatar can take many shapes and forms allowing unique creations of avatars that closely resemble a real human’s physical and mental characteristic. Therefore, it is necessary to reconstruct and define the properties prior to instantiate a dynamic 3D mesh of an Avatar.
[71] In one embodiment, an avatar should respect the following requirements:
1. The avatar may be reconstructed or animated with a wide range of specifications presented above.
2. The reconstruction and animation respect the supported primitives in the scene description.
3. The avatar is represented with semantical labels which allow for individual body part interactivity triggers.
4. The avatar’s shape contains a surrounding region with dynamic dimensions. This facilitates interaction triggers with scene, other avatars and implements social barriers between the avatar and the dynamic elements of a scene.
[72] The proposed avatar is reconstructed given the user inputs and mapped onto a reference 3D avatar, for example, the 3D avatar “MPEG reference avatar” presented in Avril, which contains a full body-based mesh representation, skeleton and skinning weights, facial blend shapes and landmarks, and additional geometry including eyes, jaws, teeth and tongue. This allows to have a shared geometry and features representation for any humanoid character (of course appropriate deformation allows to adapt the shape). This common representation eases animation and mapping of any animation to different morphologies or context.
[73] Avatar in Scene Description
[74] In the scene description we extend the existing glTF node “MPEG node avatar” elements by adding the attributes presented above.
[75] Since glTF standards allow to define 3D meshes, skeleton hierarchies and skinning weights, there is no extension needed to enable the visual appearance of avatars. The proposed extension indicates what type of avatar the node “MPEG node avatar” is referred (i.e. , if it uses the “MPEG reference avatar” as template or another), adds additional information about the avatar representation, and adds interactivity constraints on the avatar and respective elements.
[76] FIG. 7 illustrates MPEG node avatar contribution in MPEG-I SD (see an article by E. Thomas et al., entitled “[SD] Signalling for Avatars in Scene Description”, ISO/IEC JTC 1/SC 29/WG 3 m62027, hereinafter “m62027”).
[77] In particular, FIG. 7 represents a glTF file structure. The entry point is the “scene” node (730) which contains “node” node(s) (735) that can be either a “camera” (710), a “light” (765) or a “mesh” (740). The MPEG-I SD has defined a new Boolean attribute “MPEG node avatar” (735) indicating whether the node corresponds to an avatar node or not, as illustrated in Table 6. The purpose of this Boolean attribute is to define which node defines the user representation.
[78] This approach is limited as it does not provide any visual or animation information on the avatar asset. Such information has to be explicitly defined on the nodes below. For appearance, it is defined as “material” (770) with “texture” (790) that is referenced either by a direct “source” (795) or an “image” (798), “technique” (775), “program” (780) and “shader” (785). For the animation it is defined in “accessor" (745) with “animation” (720) and “skin” (725). They are all displayed using “bufferView” (750) and “buffer” nodes information (755). The “MPEG_media” (760) can redefine some existing attributes, such as geometry, appearance, sound, haptics and/or animation by referring to external data.
Table 6
[79] We propose to complete the contribution presented in m62027 with new extensions of the glTF node element to represent the wide range of possible avatar representations with the node “MPEG node avatar”. This is to normalize the avatar representation in the scene description while maintaining the newly introduced features presented previously.
[80] The metadata contains all elements described before that define/identify the avatar (identity, age, name, gender etc.), as illustrated in Table 7. Table 7
[81] Table 8 illustrates format description of the “Geometry” proprieties presented in Table 2.
Table 8
[82] Table 9 illustrates a semantical list of the available names from the avatar model. These names are used on the mapping propriety presented in Table 8.
Table 9
[83] glTF Schemas Example [84] The following glTF is an example of an instantiation of new avatar attributes in clients that support “MPEG_node_avatar”. Note that a large number of instantiations are possible, depending on the application. Here we give an example for illustration purposes.
[85] The following example illustrates how to signal the level of detail of the model and texture along with a semantical definition of the head relative to the vertices index.
[86] Here we give an example for illustration purposes of a capability extension attribute and how it is defined in the glTF format using MPEG extensions. This attribute signals the client about the capabilities of the avatar. In this scenario, the avatar by default is capable of jumping, flying, and running. This signal will provide semantics to the animation engine, so it generates and handles the following capabilities. The engine will have the objective of guaranteeing that this animation is available to the final user. From a format point of view, this signal can be used to evaluate if the conditions of the 3D environment are capable and respect the default capabilities of the avatar.
[87] Following the same principal described above, the following examples are only syntax examples, and the interpretation should be handled on the processing/engine model side. Here we give an example for illustration purposes of a disability extension attribute.
[88] Here we give an example for illustration purposes of an emotion extension attribute,
[89] Here we give an example for illustration purposes of a personality extension attribute,
[90] Here we give an example for illustration purposes of a complete metadata extension attribute.
[91] For each of the metadata examples shown above, the user/application side will have several manners to interpret the metadata. The common one and supported technique is interaction between the scene object and the avatar. The scene is going to be composed of nodes and each node represents 3D meshes. The avatar node will trigger events when interacting with nodes in the scene, hence there will be some node interaction that can use the metadata information store at the avatar node level. A practical example is, the Parental attribute will be compared against the interactive node, and if the avatar parental attribute is higher than a parental attribute present in such node then the interaction is allowed otherwise stopped. This behavior is chosen by the processing/engine model, and several behaviors can be instantiated from this standard format.
[92] Application
[93] Avatars are represented with 3D geometries that make most of the visuals and manipulators to generate realistic content in 3D applications. The issue arrives when such user virtual representation becomes social and gains access to interaction with other virtual users or objects, and needs additional information that facilitates, improves, and restricts such interactions. Hence this document provides additional metadata for avatar representation in virtual 3D environments.
[94] FIG. 8 illustrates the processing model to handle such metadata, according to an embodiment. The engine will process all dynamic (sensors input data) and static data (user metadata input and scene description) and generate a unique user avatar reconstruction and animation. In particular, the Human Features Encoder (810) relates to the encoders proposed and illustrated in FIG. 6. The Engine (835) refers to an application, such as hot not exclusive Unreal, Unity or blender, that uses the decoded (830) content provided by the Humans Features Encoder and Scene Description + Avatar Metadata (820) to create a personalized avatar.
[95] More specifically, FIG. 8 illustrates processing of dynamic sensor data and user static metadata to generate a personalized user avatar experience. The sensors represent real-world devices information that are fed to a human encoder model to extract relevant information (e.g., human pose, motion, appearance) to generate a personalized avatar. The scene description file and avatar metadata are combined with the encoded features to inform engines how to render (840), reconstruct (850), animate (860) and interact (870) with other virtual objects, hence the definition of personalized avatar.
[96] Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.
[97] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[98] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable/personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
[99] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
[100] Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[101] Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[102] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[103] It is to be appreciated that the use of any of the following “and/or”, and “at least one of’, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and/or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[104] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

Claims

1. A method, comprising: obtaining at least a parameter, from a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following:
- identity,
- a bounding region that represents an area of interaction of said avatar with other objects in said scene,
- disabilities of said avatar,
- capabilities of said avatar,
- personality of said avatar, and
- emotion of said avatar; and obtaining 3D geometry data and texture associated with said avatar.
2. A method, comprising: encoding at least a parameter, in a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following:
- identity,
- a bounding region that represents an area of interaction of said avatar with other objects in said scene,
- disabilities of said avatar,
- capabilities of said avatar,
- personality of said avatar, and
- emotion of said avatar; and encoding 3D geometry data and texture associated with said avatar.
3. The method of claim 1 or 2, wherein said bounding region corresponds to one of (1) an attribute that said avatar defines in a social behavior status; (2) a region that allows physical manipulation or interaction of said avatar with 3D objects in said scene; (3) a space that restricts user interaction with said avatar; (4) a space said avatar is allowed to move in; and (5) a region set by parental control.
4. The method of any one of claims 1-3, wherein said 3D geometry data includes information about a level of detail.
5. The method of any one of claims 1-4, wherein said 3D geometry data includes at least one of an eye model and a hair model.
6. The method of any one of claims 1-5, wherein said texture includes information about a level of detail.
7. The method of any one of claims 1-6, wherein representation of said avatar is based on a reference template model.
8. The method of any one of claims 1-7, wherein said at least a parameter further includes attributes that defines texture maps overlaying a character’s skin.
9. The method of any one of claims 1-8, further comprising: obtaining input data from sensors; and generating a reconstruction and animation of said avatar based on said input from sensors.
10. An apparatus, comprising: one or more processors; and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain at least a parameter, from a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following:
- identity,
- a bounding region that represents an area of interaction of said avatar with other objects in said scene,
- disabilities of said avatar,
- capabilities of said avatar,
- personality of said avatar, and
- emotion of said avatar; and obtain 3D geometry data and texture associated with said avatar.
11. An apparatus, comprising: one or more processors; and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: encode at least a parameter, in a description for an extended reality scene, used to represent an avatar, wherein said at least a parameter includes one or more of the following:
- identity,
- a bounding region that represents an area of interaction of said avatar with other objects in said scene,
- disabilities of said avatar,
- capabilities of said avatar,
- personality of said avatar, and
- emotion of said avatar; and encode 3D geometry data and texture associated with said avatar.
12. The apparatus of claim 10 or 11, wherein said bounding region corresponds to one of (1) an attribute that said avatar defines in a social behavior status; (2) a region that allows physical manipulation or interaction of said avatar with 3D objects in said scene; (3) a space that restricts user interaction with said avatar; (4) a space said avatar is allowed to move in; and (5) a region set by parental control.
13. The apparatus of any one of claims 10-12, wherein said 3D geometry data includes information about a level of detail.
14. The apparatus of any one of claims 10-13, wherein said 3D geometry data includes at least one of an eye model and a hair model.
15. The apparatus of any one of claims 10-14, wherein said texture includes information about a level of detail.
16. The apparatus of any one of claims 10-15, wherein representation of said avatar is based on a reference template model.
17. The apparatus of any one of claims 10-16, wherein said at least a parameter further includes attributes that defines texture maps overlaying a character’s skin.
18. The apparatus of any one of claims 10-17, wherein said one or more processors are further configured: obtain input data from sensors; and generate a reconstruction and animation of said avatar based on said input from sensors.
19. A non-transitory computer readable medium comprising instructions which, when the instructions are executed by a computer, cause the computer to perform the method of any of claims 1-9.
EP24712806.9A 2023-03-24 2024-03-18 Avatar metadata representation Pending EP4690121A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23305406 2023-03-24
PCT/EP2024/057130 WO2024200063A1 (en) 2023-03-24 2024-03-18 Avatar metadata representation

Publications (1)

Publication Number Publication Date
EP4690121A1 true EP4690121A1 (en) 2026-02-11

Family

ID=86052004

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24712806.9A Pending EP4690121A1 (en) 2023-03-24 2024-03-18 Avatar metadata representation

Country Status (8)

Country Link
EP (1) EP4690121A1 (en)
JP (1) JP2026511175A (en)
KR (1) KR20250164736A (en)
CN (1) CN120958488A (en)
AU (1) AU2024254295A1 (en)
MX (1) MX2025011269A (en)
TW (1) TW202440204A (en)
WO (1) WO2024200063A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107765852A (en) * 2017-10-11 2018-03-06 北京光年无限科技有限公司 Multi-modal interaction processing method and system based on visual human
JP7218979B1 (en) * 2022-07-07 2023-02-07 Jp Games株式会社 Information processing device, information processing method and program

Also Published As

Publication number Publication date
TW202440204A (en) 2024-10-16
AU2024254295A1 (en) 2025-10-09
JP2026511175A (en) 2026-04-10
CN120958488A (en) 2025-11-14
MX2025011269A (en) 2025-10-01
KR20250164736A (en) 2025-11-25
WO2024200063A1 (en) 2024-10-03

Similar Documents

Publication Publication Date Title
CN102458595B (en) The system of control object, method and recording medium in virtual world
US20230130535A1 (en) User Representations in Artificial Reality
CN107656615B (en) Massively simultaneous remote digital presentation of the world
JP6888096B2 (en) Robot, server and human-machine interaction methods
CN106484115B (en) Systems and methods for augmented and virtual reality
US9094576B1 (en) Rendered audiovisual communication
CN103916621A (en) Method and device for video communication
CN108668168B (en) Android VR video player based on Unity3D and design method thereof
CN114779948B (en) Method, device and equipment for controlling instant interaction of animation characters based on facial recognition
CN112673400A (en) Avatar animation
US12614354B2 (en) Animatable garment extraction through volumetric reconstruction
EP4688192A1 (en) Avatar actions and behaviors in virtual environments
WO2025165783A1 (en) N-flows estimation for garment transfer
US20250131669A1 (en) Efficient avatar creation with mesh penetration avoidance
Preda et al. Avatar interoperability and control in virtual worlds
CN117097919B (en) Virtual character rendering method, device, equipment, storage medium and program product
WO2025149308A1 (en) Avatar joint semantics representation
EP4690121A1 (en) Avatar metadata representation
US20240179291A1 (en) Generating 3d video using 2d images and audio with background keyed to 2d image-derived metadata
EP4689839A1 (en) Avatar signaling in scene description
WO2025190653A1 (en) Avatars with parametric textures in scene descriptions
EP4722861A1 (en) Semantic multilinear geometrical eye model
US20250329118A1 (en) Generative ai experience with movement
KR20240168331A (en) Proximity trigger for scene description
CN120543711A (en) Three-dimensional digital human streaming media quality assessment method, device, equipment and storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250926

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR