EP4736461A1 - Avatar mesh masking - Google Patents

Avatar mesh masking

Info

Publication number
EP4736461A1
EP4736461A1 EP24734914.5A EP24734914A EP4736461A1 EP 4736461 A1 EP4736461 A1 EP 4736461A1 EP 24734914 A EP24734914 A EP 24734914A EP 4736461 A1 EP4736461 A1 EP 4736461A1
Authority
EP
European Patent Office
Prior art keywords
mesh
avatar
scene
portions
node
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24734914.5A
Other languages
German (de)
French (fr)
Inventor
João Pedro COVA REGATEIRO
Philippe Henri GOSSELIN
Francois Le Clerc
Quentin AVRIL
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
InterDigital CE Patent Holdings SAS
Original Assignee
InterDigital CE Patent Holdings SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by InterDigital CE Patent Holdings SAS filed Critical InterDigital CE Patent Holdings SAS
Publication of EP4736461A1 publication Critical patent/EP4736461A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T17/00Three-dimensional [3D] modelling for computer graphics
    • G06T17/005Tree description, e.g. octree, quadtree
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/20Three-dimensional [3D] animation
    • G06T13/40Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • H04N21/23412Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs for generating or manipulating the scene composition of objects, e.g. MPEG-4 objects
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/81Monomedia components thereof
    • H04N21/816Monomedia components thereof involving special video data, e.g 3D video
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring
    • H04N21/85406Content authoring involving a specific file format, e.g. MP4 format
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring
    • H04N21/8543Content authoring using a description language, e.g. Multimedia and Hypermedia information coding Expert Group [MHEG], eXtensible Markup Language [XML]
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring
    • H04N21/8545Content authoring for generating interactive applications
    • AHUMAN NECESSITIES
    • A63SPORTS; GAMES; AMUSEMENTS
    • A63FCARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F2300/00Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game
    • A63F2300/50Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers
    • A63F2300/55Details of game data or player data management
    • A63F2300/5546Details of game data or player data management using player registration data, e.g. identification, account, preferences, game history
    • A63F2300/5553Details of game data or player data management using player registration data, e.g. identification, account, preferences, game history user representation in the game field, e.g. avatar
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2210/00Indexing scheme for image generation or computer graphics
    • G06T2210/61Scene description

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Security & Cryptography (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Computer Graphics (AREA)
  • Geometry (AREA)
  • Software Systems (AREA)
  • Processing Or Creating Images (AREA)

Abstract

Some embodiments of a method may include: obtaining scene description data for a 3D scene, wherein the scene description data comprises: scene element information describing each of a plurality of scene elements in the scene, a node associated with a mesh object, wherein one of the scene elements is an avatar associated with the node, and wherein at least two portions of the avatar are associated with the mesh; a mesh mapping associated with the mesh object, wherein the mesh mapping indicates how to partition the mesh object for the at least two portions of the avatar; and partitioning the mesh object using the mesh mapping for the at least two portions of the avatar. For some embodiments, the method is compatible with an extension of Moving Pictures Expert Group-1 Scene Description (MPEG-1 SD) and/or Graphics Language Transmission Format (glTF) scene description format.

Description

AVATAR MESH MASKING
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims benefit of European Patent Application No. EP23306091 , entitled “AVATAR MESH MASKING” and filed June 30, 2023, which is hereby incorporated by reference in its entirety.
BACKGROUND
[0002] This application applies to 3D scene and user representation of interactions within immersive environments. Representations of users such as avatars are in common usage in 3D immersive environments. MPEG-I Scene Description (SD) (23090-14) provides an interactivity framework in support of use of user representations (e.g., avatars) in these environments, and in extended reality (XR) such as virtual reality (VR), augmented reality (AR), and/or mixed reality (MR).
SUMMARY
[0003] Embodiments described herein include methods that are used in video encoding and decoding (collectively “coding”).
[0004] An example method in accordance with some embodiments may include: obtaining scene description data for a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, a node associated with a mesh object, wherein one of the scene elements is an avatar associated with the node, and wherein at least two portions of the avatar are associated with the mesh object; a mesh mapping associated with the mesh object, wherein the mesh mapping indicates how to partition the mesh object for the at least two portions of the avatar; and partitioning the mesh object using the mesh mapping for the at least two portions of the avatar.
[0005] Some embodiments of the example method may further include processing the scene description data associated with the avatar.
[0006] For some embodiments of the example method, processing the scene description data associated with the avatar may include rendering the avatar as part of the scene using the scene description data.
[0007] For some embodiments of the example method, partitioning the mesh object may include: repeating a process for each of the at least two portions of the avatar associated with the mesh object, wherein the process may include: obtaining a current portion selected from the at least two portions of the avatar; obtaining a current mesh map associated with the current portion; and masking the current mesh map with the mesh object to obtain current position information for the current portion.
[0008] For some embodiments of the example method, obtaining the current mesh map may be performed using the mesh mapping.
[0009] For some embodiments of the example method, the method may be compatible with an extension of Moving Pictures Expert Group-1 Standard Definition (MPEG-1 SD) scene description format.
[0010] For some embodiments of the example method, the method may be compatible with an extension of Graphics Language Transmission Format (gITF) scene description format.
[0011] For some embodiments of the example method, rendering the avatar may retain the mesh object as a unitary mesh object for each of the at least two portions of the avatar.
[0012] For some embodiments of the example method, the mesh mapping may include an integer index referencing a data structure.
[0013] For some embodiments of the example method, the mesh mapping may include: information indicating the node; and information indicating a path to information related to the node.
[0014] For some embodiments of the example method, the mesh mapping may include a pointer-indexed data structure.
[0015] For some embodiments of the example method, the at least two portions of the avatar may include non-overlapping portions of the avatar.
[0016] For some embodiments of the example method, the scene description data may further include: a second mesh object associated with the node, wherein the at least two portions of the avatar are associated with the second mesh object; and a second mesh mapping associated with the second mesh object, wherein the second mesh mapping indicates how to partition the second mesh object for the at least two portions of the avatar, and wherein the mesh object is different from the second mesh object.
[0017] For some embodiments of the example method, the scene description data may further include: a third mesh object associated with the node, wherein at least two further portions of the avatar are associated with the third mesh object, and wherein the at least two further portions of the avatar are different than the at least two portions of the avatar; and a third mesh mapping associated with the third mesh object, wherein the third mesh mapping indicates how to partition the third mesh object for the at least two further portions of the avatar.
[0018] For some embodiments of the example method, wherein the at least two portions of the avatar may include a first portion and second portion of the avatar, wherein the method may further include: obtaining a first mesh map associated with the first portion of the avatar; and obtaining a second mesh map associated with the second portion of the avatar, wherein the first mesh map shares a common element with the second mesh map.
[0019] For some embodiments of the example method, wherein the at least two portions of the avatar comprise a first portion and second portion of the avatar, wherein the method may further include: obtaining a first mesh map associated with the first portion of the avatar; and obtaining a second mesh map associated with the second portion of the avatar, wherein the first mesh map is unique compared to the second mesh map.
[0020] An example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any of the methods listed above.
[0021] Another example method in accordance with some embodiments may include performing a mesh masking process to facilitate mesh segmentation of a single, three- dimensional mesh object corresponding to a node hierarchy delineated in scene description data.
[0022] For some embodiments of another example method, the mesh object is associated with a non-avatar environment.
[0023] For some embodiments of another example method, the mesh object is associated with an avatar environment.
[0024] Another example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any of the methods listed above.
[0025] An example apparatus in accordance with some embodiments may include at least one processor configured to perform any one of the methods listed above.
[0026] An example apparatus in accordance with some embodiments may include a computer- readable medium storing instructions for causing one or more processors to perform any one of the methods listed above. [0027] An example apparatus in accordance with some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any one of the methods listed above.
[0028] An example signal in accordance with some embodiments may include a bitstream comprising scene description data generated according to any one of the methods listed above.
[0029] In additional embodiments, encoder and decoder apparatus are provided to perform the methods described herein. An encoder or decoder apparatus may include a processor configured to perform the methods described herein. The apparatus may include a computer-readable medium (e.g. a non-transitory medium) storing instructions for performing the methods described herein. In some embodiments, a computer-readable medium (e.g. a non-transitory medium) stores a video encoded using any of the methods described herein.
[0030] One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for performing bi-directional optical flow, encoding or decoding video data according to any of the methods described above. The present embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described above. The present embodiments also provide a method and apparatus for transmitting the bitstream generated according to the methods described above. The present embodiments also provide a computer program product including instructions for performing any of the methods described.
BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIG. 1 A is a schematic side view illustrating an example waveguide display that may be used with extended reality (XR) applications according to some embodiments.
[0032] FIG. 1 B is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments.
[0033] FIG. 1 C is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments.
[0034] FIG. 1 D is a system diagram illustrating an example set of interfaces for a system according to some embodiments.
[0035] FIG. 1 E is a system diagram illustrating an example set of interfaces for a scene description (stored as an item in gITF.json), three video tracks, an audio track, and a JSON patch update track in an ISOBMFF file according to some embodiments. [0036] FIG. 2 is a system diagram illustrating an example set of interfaces for an MPEG-I node hierarchy supporting elements of scene interactivity according to some embodiments.
[0037] FIG. 3 is a block diagram showing an example of logical relationships between trigger information (describing triggers 1 through n), action information (describing actions 1 through m), and behavior information (describing relationships between the triggers and actions) in which triggers and actions may refer to one or more nodes in a scene description, such as a hierarchical scene graph, according to some embodiments.
[0038] FIG. 4 is a system diagram illustrating an example set of interfaces for a scene graph supporting an avatar according to some embodiments.
[0039] FIG. 5 is a system diagram illustrating an example set of interfaces for a scene graph according to some embodiments.
[0040] FIG. 6 is a system diagram illustrating an example set of interfaces for a scene graph referencing a single mesh with mesh object segmentation according to some embodiments.
[0041] FIG. 7 is a flowchart illustrating an example process for an avatar using a mask attribute according to some embodiments.
[0042] FIG. 8A is a schematic perspective view illustrating an example cube avatar according to some embodiments.
[0043] FIG. 8B is a schematic plan view illustrating an example unfolded cube avatar according to some embodiments.
[0044] FIG. 9A is a table showing example avatar position information and mesh masks according to some embodiments.
[0045] FIG. 9B is a table showing example avatar position information and mesh masks according to some embodiments.
[0046] FIG. 10A is a table showing example avatar mesh masks according to some embodiments.
[0047] FIG. 10B is a table showing example avatar position information and mesh masks according to some embodiments.
[0048] FIG. 1 1A is a table showing example avatar position information and mesh masks according to some embodiments.
[0049] FIG. 1 1 B is a table showing example avatar position information and mesh masks according to some embodiments. [0050] FIG. 1 10 is a table showing example avatar mesh masks according to some embodiments.
[0051] FIG. 1 1 D is a table showing example avatar position information and mesh masks according to some embodiments.
[0052] FIG. 12 is a schematic illustration showing an example avatar according to some embodiments.
[0053] FIG. 13 is a flowchart illustrating an example process for sharing a unitary mesh object among multiple avatar elements according to some embodiments.
[0054] The entities, connections, arrangements, and the like that are depicted in — and described in connection with — the various figures are presented by way of example and not by way of limitation. As such, any and all statements or other indications as to what a particular figure “depicts,” what a particular element or entity in a particular figure “is” or “has,” and any and all similar statements — that may in isolation and out of context be read as absolute and therefore limiting — may only properly be read as being constructively preceded by a clause such as “In at least one embodiment, ... " For brevity and clarity of presentation, this implied leading clause is not repeated ad nauseum in the detailed description. FIG. 1 A is a system diagram illustrating an example set of interfaces for a system according to some embodiments.
DETAILED DESCRIPTION
[0055] FIG. 1 A is a schematic side view illustrating an example waveguide display that may be used with extended reality (XR) applications according to some embodiments. An image is projected by an image generator 102. The image generator 102 may use one or more of various techniques for projecting an image. For example, the image generator 102 may be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light-emitting diode (LED) display (including an organic LED (OLED) or micro LED (pLED) display), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other type of image generator or light engine.
[0056] Light representing an image 1 12 generated by the image generator 102 is coupled into a waveguide 104 by a diffractive in-coupler 106. The in-coupler 106 diffracts the light representing the image 112 into one or more diffractive orders. For example, light ray 108, which is one of the light rays representing a portion of the bottom of the image, is diffracted by the in-coupler 106, and one of the diffracted orders 110 (e.g. the second order) is at an angle that is capable of being propagated through the waveguide 104 by total internal reflection. The image generator 102 displays images as directed by a control module 124, which operates to render image data, video data, point cloud data, or other displayable data. [0057] At least a portion of the light 110 that has been coupled into the waveguide 104 by the diffractive in-coupler 106 is coupled out of the waveguide by a diffractive out-coupler 114. At least some of the light coupled out of the waveguide 104 replicates the incident angle of light coupled into the waveguide. For example, in the illustration, out-coupled light rays 1 16a, 1 16b, and 116c replicate the angle of the in-coupled light ray 108. Because light exiting the out-coupler replicates the directions of light that entered the in-coupler, the waveguide substantially replicates the original image 112. A user’s eye 118 can focus on the replicated image.
[0058] In the example of FIG. 1 A, the out-coupler 1 14 out-couples only a portion of the light with each reflection allowing a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). In this way, at least some of the light originating from each portion of the image is likely to reach the user’s eye even if the eye is not perfectly aligned with the center of the out-coupler. For example, if the eye 1 18 were to move downward, beam 116c may enter the eye even if beams 1 16a and 1 16b do not, so the user can still perceive the bottom of the image 1 12 despite the shift in position. The out-coupler 1 14 thus operates in part as an exit pupil expander in the vertical direction. The waveguide may also include one or more additional exit pupil expanders (not shown in FIG. 1A) to expand the exit pupil in the horizontal direction.
[0059] In some embodiments, the waveguide 104 is at least partly transparent with respect to light originating outside the waveguide display. For example, at least some of the light 120 from real-world objects (such as object 122) traverses the waveguide 104, allowing the user to see the real-world objects while using the waveguide display. As light 120 from real-world objects also goes through the diffraction grating 1 14, there will be multiple diffraction orders and hence multiple images. To minimize the visibility of multiple images, it is desirable for the diffraction order zero (no deviation by 1 14) to have a great diffraction efficiency for light 120 and order zero, while higher diffraction orders are lower in energy. Thus, in addition to expanding and out-coupling the virtual image, the out-coupler 114 is preferably configured to let through the zero order of the real image. In such embodiments, images displayed by the waveguide display may appear to be superimposed on the real world.
[0060] FIG. 1 B is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments. In an XR headmounted display device 130, a control module 132 controls a display 134, which may be an LCD, to display an image. The head-mounted display includes a partly-reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) the image displayed on the LCD to make the image visible to the user. The partly-reflective surface 136 also allows the passage of at least some exterior light, permitting the user to see their surroundings. [0061] FIG. 1 C is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments. In an XR headmounted display device 140, a control module 142 controls a display 144, which may be an LCD, to display an image. The image is focused by one or more lenses of display optics 146 to make the image visible to the user. In the example of FIG. 1 C, exterior light does not reach the user’s eyes directly. However, in some such embodiments, an exterior camera 148 may be used to capture images of the exterior environment and display such images on the display 144 together with any virtual content that may also be displayed.
[0062] The embodiments described herein are not limited to any particular type or structure of XR display device.
[0063] FIG. 1 D is a system diagram illustrating an example set of interfaces for a system according to some embodiments. An extended reality display device, together with its control electronics, may be implemented using a system such as the system of FIG. 1 D. System 150 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 150, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and/or discrete components. For example, in at least one embodiment, the processing and encoder/decoder elements of system 150 are distributed across multiple ICs and/or discrete components. In various embodiments, the system 150 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports. In various embodiments, the system 1000 is configured to implement one or more of the aspects described in this document.
[0064] The system 150 includes at least one processor 152 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 152 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 150 includes at least one memory 154 (e.g., a volatile memory device, and/or a non-volatile memory device). System 150 may include a storage device 158, which can include non-volatile memory and/or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and/or optical disk drive. The storage device 158 can include an internal storage device, an attached storage device (including detachable and non-detachable storage devices), and/or a network accessible storage device, as non-limiting examples.
[0065] System 150 includes an encoder/decoder module 156 configured, for example, to process data to provide an encoded video or decoded video, and the encoder/decoder module 156 can include its own processor and memory. The encoder/decoder module 156 represents module(s) that can be included in a device to perform the encoding and/or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, encoder/decoder module 156 can be implemented as a separate element of system 150 or can be incorporated within processor 152 as a combination of hardware and software as known to those skilled in the art.
[0066] Program code to be loaded onto processor 152 or encoder/decoder 156 to perform the various aspects described in this document can be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. In accordance with various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder/decoder module 156 can store one or more of various items during the performance of the processes described in this document. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0067] In some embodiments, memory inside of the processor 152 and/or the encoder/decoder module 156 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device can be either the processor 152 or the encoder/decoder module 152) is used for one or more of these functions. The external memory can be the memory 154 and/or the storage device 158, for example, a dynamic volatile memory and/or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of, for example, a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO/IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team). [0068] The input to the elements of system 150 can be provided through various input devices as indicated in block 172. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and/or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. 1 C, include composite video.
[0069] In various embodiments, the input devices of block 172 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, downconverting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and/or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0070] Additionally, the USB and/or HDMI terminals can include respective interface processors for connecting system 150 to other electronic devices across USB and/or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within processor 152 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within processor 152 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 152, and encoder/decoder 156 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0071] Various elements of system 150 can be provided within an integrated housing, Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangement 174, for example, an internal bus as known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards.
[0072] The system 150 includes communication interface 160 that enables communication with other devices via communication channel 162. The communication interface 160 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 162. The communication interface 160 can include, but is not limited to, a modem or network card and the communication channel 162 can be implemented, for example, within a wired and/or a wireless medium.
[0073] Data is streamed, or otherwise provided, to the system 150, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.1 1 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 162 and the communications interface 160 which are adapted for Wi-Fi communications. The communications channel 162 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 150 using a set-top box that delivers the data over the HDMI connection of the input block 172. Still other embodiments provide streamed data to the system 150 using the RF connection of the input block 172. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
[0074] The system 150 can provide an output signal to various output devices, including a display 176, speakers 178, and other peripheral devices 180. The display 176 of various embodiments includes one or more of, for example, a touchscreen display, an organic lightemitting diode (OLED) display, a curved display, and/or a foldable display. The display 176 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 176 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 180 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and/or a lighting system. Various embodiments use one or more peripheral devices 180 that provide a function based on the output of the system 150. For example, a disk player performs the function of playing the output of the system 150.
[0075] In various embodiments, control signals are communicated between the system 150 and the display 176, speakers 178, or other peripheral devices 180 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to- device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 164, 166, and 168. Alternatively, the output devices can be connected to system 150 using the communications channel 162 via the communications interface 160. The display 176 and speakers 178 can be integrated in a single unit with the other components of system 150 in an electronic device such as, for example, a television. In various embodiments, the display interface 164 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0076] The display 176 and speaker 178 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 172 is part of a separate set-top box. In various embodiments in which the display 176 and speakers 178 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0077] The system 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscopic sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and/or magnetometers. Such sensors may be used to determine information such as user’s position and orientation. Where the system 150 is used as the control module for an extended reality display (such as control modules 124, 132), the user’s position and orientation may be used in determining how to render image data such that the user perceives the correct portion of a virtual object or virtual scene from the correct point of view. In the case of head-mounted display devices, the position and orientation of the device itself may be used to determine the position and orientation of the user for the purpose of rendering virtual content. In the case of other display devices, such as a phone, a tablet, a computer monitor, or a television, other inputs may be used to determine the position and orientation of the user for the purpose of rendering content. For example, a user may select and/or adjust a desired viewpoint and/or viewing direction with the use of a touch screen, keypad or keyboard, trackball, joystick, or other input. Where the display device has sensors such as accelerometers and/or gyroscopes, the viewpoint and orientation used for the purpose of rendering content may be selected and/or adjusted based on motion of the display device. [0078] The embodiments can be carried out by computer software implemented by the processor 152 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 154 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as nonlimiting examples. The processor 152 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
Scene Description Framework for XR
[0079] The present principles generally relate to the domain of rendering of extended reality scene description and extended reality rendering. The present document is also understood in the context of the formatting and the playing of extended reality applications when rendered on end-user devices such as mobile devices or Head-Mounted Displays (HMD).
[0080] In XR applications, a scene description is used to combine explicit and easy-to-parse description of a scene structure and some binary representations of media content.
[0081] In time-based media streaming, the scene description itself can be time-evolving to provide the relevant virtual content for each sequence of a media stream. For instance, for advertising purposes, a virtual bottle can be displayed during a video sequence where people are drinking.
[0082] This kind of behavior can be achieved by relying on the framework defined in the Scene Description for MPEG media document, Information technology - Coded representation of immersive media - Parti 4: Scene Description for MPEG media, ISO/IEC DIS 23090-14 :2021 (E). A scene update mechanism based on the JSON Patch protocol as defined in IETF RFC 6902 may be used to synchronize virtual content to MPEG media streams.
[0083] FIG. 1 E is a system diagram illustrating an example set of interfaces for a scene description (stored as an item in gITF.json), three video tracks, an audio track, and a JSON patch update track in an ISOBMFF file 182 according to some embodiments.
[0084] Although the MPEG-I Scene Description framework ensures that the timed media and the corresponding relevant virtual content are available at any time, it does not provide a description of how a user can interact with the scene objects at runtime for immersive XR experiences. Hence, there is no support of user specific XR experiences for consuming the immersive media. [0085] Example embodiments as described herein may be used to provide a scene description that includes a virtual object or light source but that does not necessarily display or render the virtual object or light source even if available. In some embodiments, one or more of the following aspects may be considered in determining whether to display a virtual object or light source.
[0086] A spatial aspect may be considered in determining whether to display a virtual object or light source. For example, if the user environment is not suited (e.g. the user is too far from the rendered timed media’s location) or if the user is not looking toward the right direction, or if the virtual object should be displayed on a user-specific area (e.g. above his left hand which is not yet detected), then the virtual object or light source may not be displayed.
[0087] A temporal aspect may be considered in determining whether to display a virtual object or light source. For example, if the user is not yet ready or wants to trigger himself the display of the object (e.g. using a specific gesture), the virtual object or light source may not be displayed until the appropriate trigger is detected.
[0088] In some embodiments, it is specified in the scene description which objects or light sources the user is allowed to manipulate or to interact with through potential haptic feedbacks.
Runtime Interactivity
[0089] FIG. 2 is a system diagram illustrating an example set of interfaces for an MPEG-I node hierarchy 200 supporting elements of scene interactivity according to some embodiments. According to the present principles, in addition to a node tree as described in relation to FIG. 4, behavior metadata items (herein called ‘behaviors’) are added to the scene description. In example embodiments, the time-evolving scene description is augmented by adding information identifying behaviors. These behaviors may be related to pre-defined virtual objects on which runtime interactivity is allowed for user specific XR experiences.
[0090] In some embodiments, these behaviors are time-evolving. In such embodiments, the behaviors may be updated through the already-existing scene description update mechanism.
[0091] In example embodiments, a behavior is characterized by one or more of the following properties:
• One or more triggers defining the conditions to be met for activation.
• A trigger control parameter defining the logical operations between the defined triggers.
• Actions to be implemented in response to the activation of the triggers.
• An action control parameter defining the order of execution of the defined actions.
• A priority number enabling the selection of the behavior of highest priority in the case of concurrence of several behaviors on the same virtual object at the same time. • An optional interrupt action to specify how to terminate this behavior when the behavior is no longer defined in a newly received scene update. For instance, a behavior is no longer defined if the related object has been removed or if the behavior is no longer relevant for this current media (e.g. audio or video) sequence.
[0092] With the addition of these behaviors, time-dependent user interactivity in immersive content for XR experiences may be defined.
[0093] When a second scene description is received, some of the behaviors of the first scene description may be “on-going”, that is they are triggered, and their actions are running. The second scene description may be provided as update metadata, that is metadata describing the differences between the first scene description and the second description. The second scene description includes a node tree describing objects that may be common or different than objects of the first scene descriptions. Objects of the node tree of the first scene description may be no longer present in the second description. If the objects related to the running actions of the ongoing behaviors are missing in the second scene description, then, these on-going behaviors are no longer appliable. The same way, if an on-going behavior is not defined in the second description, the on-going behavior is no longer appliable. The interrupt action field describes how to correctly interrupt the running actions on the on-going behavior.
[0094] FIG. 3 is a block diagram showing an example of logical relationships between trigger information (describing triggers 1 through n), action information (describing actions 1 through m), and behavior information (describing relationships between the triggers and actions) in which triggers and actions may refer to one or more nodes in a scene description, such as a hierarchical scene graph, according to some embodiments.
[0095] In XR applications, a scene description is used to combine explicit and easy-to-parse description of a scene structure and some binary representations of media content. The above sections describe action mechanisms for scene descriptions. These behaviors are related to predefined virtual objects on which runtime interactivity is allowed for user specific XR experiences. FIG. 3 illustrates the structure of an example behavior mechanism. The example structure 300 shows example trigger information 302 and action information 304. Within the trigger information 302 are example triggers 1 (306), 2 (308), ... , n (310). Within the action information 304 are example triggers 1 (312), 2 (314), ... , n (316). The triggers 306, 308, 310 and actions 312, 314, 316 are shown with example relationships to various nodes 318.
[0096] This application applies to 3D scene and user representation of interactions within immersive environments in accordance with some embodiments. A new extension under a user representation (e.g., avatar) node is described herein in accordance with some embodiments to provide a mesh masking mechanism to facilitate mesh partition/segmentation, and node hierarchy definition in the presence of a single or multiple 3D mesh object.
[0097] In accordance with some embodiments, this application introduces a new extension for an “MPEG_node_avatar”. See the current extension International Organization for Standardization / International Electrotechnical Commission (ISO/IEC) Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 3 m63505 Extension (ISO/IEC JTC 1/SC 29/WG 3 m63505) (“ MPEG_node_avatar Extension”) of the MPEG-I Scene Description (SD) (ISO/IEC 23090-14) standard regarding avatar mesh partition/segmentation in interactive 3D environments. The current MPEG_node_avatar Extension, under the node level in gITF format, does not permit avatar mesh segmentation in the presence of a single gITF object (accessor) for an avatar representation. The segmentation of a 3D avatar mesh and hierarchical definition currently present in the MPEG_node_avatar Extension is not compatible with a single gITF object (accessor) definition of avatar/user representations.
[0098] In order to render an avatar, a device such as a client device generally may need to know what parts of the avatar mesh correspond to different body parts. MPEG Scene Description (23090-14) provides - in the MPEG_node_avatar extension - a way to map between body parts (as identified using 'path' labels) to sub-parts of the avatar mesh. However, the existing mapping function points into a single node in the avatar mesh. As such, the existing mapping function may only be useful in the case where the avatar mesh is already segmented into well-defined submeshes corresponding to the body parts. For the case where the avatar mesh is more monolithic, e.g., in the case of a unitary avatar mesh, it would be advantageous to provide a new solution. As described in more detail herein, in accordance with some embodiments, is a syntax addition that allows the mapping to be done using a new 'mask' element. In accordance with some embodiments, this mask element includes, e.g., an indicator per vertex (or in other embodiments, per surface triangle) of the avatar mesh. For example, in the extended mapping structure 'path' might point to the label 'Left Hand' and the corresponding 'mask' element would have binary (or equivalent) indicators which specify which vertices of the avatar mesh are part of the left hand (e.g., value=1 ), and which are not (e.g., value=0).
[0099] Table 1 describes the Avatar properties of the current MPEG_node_avatar Extension. Currently, the properties available are the Avatar scheme type and the semantic to node mapping.
Table 1
[0100] Table 1 lists proprieties of the "Mapping" object from MPEG_node_avatar Extension. This representation may be used to hierarchically represent an avatar structure, e.g., skeletal framework or semantical description of the shape, through the “Mapping” object. The “Mapping” object may refer to a “path” object, which is a string path to a node that refers to an element of the avatar morphology. The “node” object points to the 3D mesh that represents this element.
Table 2
[0101] FIG. 4 is a system diagram illustrating an example set of interfaces for a scene graph supporting an avatar according to some embodiments. FIG. 4 illustrates the scenario currently available in the MPEG_node_avatar Extension. The example shown in FIG. 4 shows a hierarchical body representation 400 in which several meshes 402, 404, 406, 408 are available for each body part. FIG. 4 shows avatar scene graph and node object relationships in which the mesh may be split or several meshes may be used to represent body parts corresponding to the mapping hierarchy. FIG. 4 shows multiple “mapping” objects 410, 412. The Mappings[mapping] box 414 represents an array of “mapping {path, node}” objects 410, 412. An index indicates which “mapping {path, node}” object to use. A set of "mapping(path, node)" structures are included in "Mappingsfl" array. Each "mapping(path, node)" element has a different index, but a single index (index 0) is used for the "Mappingsfl” array 414. "mapping" is an object, and "Mappingsfl" is an array of objects of type "mapping". For some embodiments, the nomenclature may be expressed as:
Mappings[mapping(path, node), mapping(path, node), ... mapping(path, node)]
[0102] For the first set of items, the mesh 402 for the left hand node 416 is index 0, and the mesh 404 for the right hand node 418 is index 1 . Continuing for the first set of meshes, the mesh 406 for the left foot node 420 is index 2, and the mesh 408 for the right foot node 422 is index 3. These meshes are distinct from each other and not combined into a single mesh. For the current MPEG_node_avatar Extension, if an avatar/user representation is entirely represented by a single “node” and a single 3D mesh, the “Mapping” object becomes unusable.
[0103] FIG. 5 is a system diagram illustrating an example set of interfaces for a scene graph referencing a single mesh according to some embodiments. FIG. 5 shows a node graph 500 referencing a single body mesh object 502 with an index of 0 in which the mesh object 502 is not split into multiple body parts or several mesh objects. As a result, the hierarchical representation of the shape is lost, and how to manipulate such a single object without manual segmentation of the asset is unclear.
[0104] In FIG. 5, four example nodes 504, 506, 508, 510 point to the same mesh, which is index 0 of a mesh array 502. The single mesh representation includes the 3D geometry of the avatar/user representation. Because of this architecture, representation or identification of individual body parts becomes difficult to perform. To allow such behavior, the avatar object needs to be separated into individual, 3D body mesh objects as shown in FIG. 4. Discussed herein is a mesh that may be partitioned without duplicating the 3D geometry information stored in buffers in the context of the MPEG_node_avatar Extension.
[0105] FIG. 6 is a system diagram illustrating an example set of interfaces for a scene graph referencing a single mesh with mesh object segmentation according to some embodiments. FIG. 6 shows a node/scene graph 600 in which a new “mask” attribute 602, 604 is added to provide a context of what part of the geometry of the mesh each node refers to. The “mask” attribute enables segmentation of a single mesh object. This new mesh masking mechanism is added to the MPEG_node_avatar Extension.
MPEG_node_avatar Mesh Mask [0106] The sections below describe the new semantic and mask attributes, provide JSON coding schemes, and describe how the coding schemes may be used within the MPEG-I SD standard. The new semantic and mask attributes are described to conform to the gITF format and are compatible with the current MPEG effort to extend gITF with MPEG extensions. For some embodiments, the semantic and mask attributes may be encoded using other formats such as
Extensible Markup Language (XML) or Universal Scene Description (USD).
[0107] Table 3 introduces the new node attribute “mask” under the MPEG_node_avatar Extension. Added are the new semantic and mask objects. For this example, a single mask attribute is shown. Table 3
[0108] The new attribute “mask” is an integer index to an accessor created at the scene level. The accessor points to a chunk of UNSIGNED BYTE integers with values in {0,1} (or any other type of data where zero relates to “FALSE” and any other value relates to “TRUE”). The accessor count must be the same as the number of elements in the attributes of the avatar mesh primitives. The accessor count represents the number of vertices in the avatar mesh. A non-zero value for a vertex in the buffer pointed to by the mask indicates that the vertex belongs to the mesh part pointed to by the considered node. Correspondingly, a value of 0 for a vertex in the buffer pointed to by the mask indicates that the vertex does not belong to the mesh part pointed to by the considered node. The new attribute “semantic” is a string that indicates to which primitive from the mesh (if present) the mask applies. These new attributes added to Table 2 (and thereby shown in Table 3) enrich the ability to use a more generic avatar hierarchical structure.
Avatar Mask Processing Model
[0109] FIG. 7 is a flowchart illustrating an example process for an MPEG node avatar using a mask attribute according to some embodiments. FIG. 9 shows how the added attributes may be used in immersive and interactive systems. The model process flow 700 processes 702 a node and verifies 704 if the avatar extension is present and if the node contains the attribute that references to a mesh object (“mesh”). If both of these conditions are TRUE, then the mesh uses its own accessor in combination with the new attribute “mask” that references another accessor to render the body part specified. Such rendering is done using the “path” attribute referenced in Table 3. In this manner, the engine accurately renders different body parts described in “mapping” while using a single mesh object. The masking is handled by the engine, which is described below. For some embodiments, if a check for the MPEG_node_avatar extension is TRUE and a mesh exists and is associated with the node, then a mesh mask is applied to the mesh object.
[0110] For some embodiments, if the check 704 for the MPEG_node_avatar extensions is false, then the application exits 706 this process and continues with other processes. If the MPEG_node_avatar extensions check 704 is true, then the process loops 708 over the Mappings[] array. The loop checks 710 to see if the mappings array length is greater than a count. If no, then the application exits 706 this process and continues with other processes. Otherwise, the current process continues to access 712 nodej for mapping i. A check 714 is performed to see if a «mesh» exists in node i. If yes, then the new «mask_i» is applied to the nodej. mesh buffer with the key «semanticj», and the count is incremented. The loop then returns to the top to check 710 to see if the mappings array length is greater than the count. gITF Schemas Example
[0111] FIG. 8A is a schematic perspective view illustrating an example cube avatar according to some embodiments. The following gITF is an non-exhaustive example of the discussed extension. Numerous other examples could be presented, depending on the application. These examples are for illustration purposes. [0112] To simplify the examples, an avatar will be represented as a cube in FIGs. 8A and 8B. To further simplify things, these examples will focus on the left and right sides of the cube. In the color version of FIGs. 8A and 8B, the left side 802 of the cube 800 is blue, and the right side 804 of the cube 800 is green. The remaining parts of the cube remain neutral for the purpose of this example, but the same concepts may be applied to those sides. The numbers shown in FIGs. 8A and 8B are the vertex indices. In FIG. 8A, the vertices for the left side 802 of the cube 800 are vertices 0 to 3. The vertices for the right side 804 of the cube 800 are vertices 4 to 7.
[0113] FIG. 8B is a schematic plan view illustrating an example unfolded cube avatar 850 according to some embodiments. The same vertices of the left and right sides of the cube of FIG. 8A are shown in the flattened version of the cube 850 of FIG. 8B. FIG. 8B also shows each of the six faces of the cube being divided into two triangles.
[0114] FIG. 9A is a table showing example avatar position information and mesh masks according to some embodiments. For the example cube of FIG. 8A, FIG. 9A shows three buffers. The first (top) buffer 900 indicates the left side of the cube as vertices 0 to 3. The first (top) buffer 900 also indicates the right side of the cube as vertices 4 to 7. The second (middle) buffer 902 shows the mask used to select the left side of the cube. In this example, the second buffer 902 selects vertices 0 to 3. The third (bottom) buffer 904 shows the mask used to select the right side of the cube. In this example, the third buffer 904 selects vertices 4 to 7.
[0115] FIG. 9B is a table showing example avatar position information and mesh masks according to some embodiments. FIG. 9B shows three buffers. The first buffer 950 shows the coordinates of the cube vertex positions. For this example, the first buffer 950 is shown as three octets listing the positions for the mesh attribute. The second buffer 952 (Mask 1 Blue) is the mask for the left side of the cube of FIG. 8A. The third buffer 954 (Mask 2 Green) is the mask for the right side of the cube of FIG. 8A.
[0116] The first mask (second buffer 952) keeps only the left (blue) vertices, and the second mask (third buffer 954) keeps only the right (green) vertices of the cube. For rendering, these masks are applied to the first buffer 950 to select the corresponding coordinates listed in the first buffer 950.
[0117] In FIG. 9B, the vertices 0 to 7 are applied as a “column” across each of the three buffers 950, 952, 954. For example, the first buffer 950 shows vertex 2 with the xyz-coordinates of x=1 .0, y=1 .0, and z=-1 .0. For this example, the origin is assigned to the center of the cube 800 of FIG. 8A. As depicted in FIG. 8A, the x-axis points into the paper (or towards the upper right corner of the paper) for positive values. The y-axis points towards the bottom of the paper for positive values. The z-axis is shown going to the left for positive values. Hence, vertex 2 is shown as (x, y, z) = (1 , 1 , -1 ). Furthermore, mask 1 selects vertices 0 to 3 to select the left side of the cube, and mask 2 selects vertices 4 to 7 to select the right side of the cube.
[0118] FIG. 10A is a table showing example avatar mesh masks 1000, 1002 according to some embodiments. In this example, the left side of the cube includes triangles 0 and 1 , while the right side of the cube includes triangles 4 and 5.
[0119] FIG . 10B is a table showing example avatar position information 1050 and mesh masks 1052, 1054 according to some embodiments. FIG. 10B illustrates the same example as FIG. 9B except, instead of masking the vertices, the faces/triangles of the cube shown in FIGs. 8A and 8B are masked. In this example, the left side of the cube includes triangles 0 and 1 , while the right side of the cube includes triangles 4 and 5. For the left side, triangle 0 is defined by vertices 0, 1 and 3, and triangle 1 is defined by vertices 1 , 2, and 3. For the right side, triangle 4 is defined by vertices 4, 5 and 6, and triangle 5 is defined by vertices 4, 6, and 7.
[0120] Code listing 1 below shows a code listing for a cube according to some embodiments. The example code listing shown below is a gITF format example of the scenario depicted in FIGs. 9A and 9B, in which the "positions" mesh attribute is selected by a mask. The avatar is one node with a name of “Cube”. In this node, the mapping points to the left and right sides of the cube using the “mask” attribute, while sharing the same node ID (node: 1 ). These items are shown in this example near the end of the code listing. Using a mask, a renderer or engine may be able to determine what parts of the same mesh correspond to the left and right sides of the cube avatar node. Masking based on the “positions” mesh attribute may be encoded in a gITF file using the MPEG_node_avatar Extension as shown in code listing 1 . The code listing 1 follows the gITF 2.0 standard specification (registry<dot>khronos<dot>org/glTF/specs/2.0/glTF-2.0<dot>html) (“gITF 2.0 Standard’), except for the MPEG_node_avatar extension in the "nodes/0/extensions/MPEG_node_avatar" section, which extends the specification. Considering standard properties, the data is contained in "accessors", which references items in "bufferViews", which references items in "buffers". Following the specification, sets of typed values (like a list of integers or a list of 3D vectors) may be deduced from each accessor. See the gITF 2.0 Standard for more details on how to run this decoding.
[0121] A code parser decodes the accessors/buffer views/buffers properties, which leads to the following data:
• Accessor 0 (lines 3-10 in code listing 1 ), which refers to buffer view 0 (lines 49-53 in code listing 1 ), which refers buffer 0 (lines 80-82 in code listing 1 ) with the name "positions": eight 3D vectors with values (same as vertex positions in FIG. 9B): [-1.00000, -1.00000, -1.00000], [1.00000, -1.00000, -1.00000], [1.00000, 1.00000, -1.00000],
[-1.00000, 1.00000, -1.00000], [-1.00000, -1.00000, 1.00000], [1.00000, -1.00000, 1.00000], [1.00000, 1.00000, 1.00000], [-1.00000, 1.00000, 1.00000]
• Accessor 1 (lines 12-19 in code listing 1), which refers to buffer view 1 (lines 55-59 in code listing 1 ), which refers to buffer 1 (lines 84-86 in code listing 1) with the name "indices": 36 values (same as Triangles in FIG. 10B):
0,3,1 ,
3.2.1 ,
1.2.5,
2.6.5,
5.6.4,
6.7.4, 4,7,0, 7,3,0,
3.7.2,
7.6.2, 4,0,5, 0,1 ,5
• Accessor 2 (lines 21 -28 in code listing 1 ), which refers to buffer view 2 (lines 61 -65 in code listing 1 ), which refers to buffer 2 (lines 88-90 in code listing 1) with the name "colors": eight 3D vectors with values:
[0.00000, 1.00000, 0.00000], [0.00000, 1.00000, 0.00000], [0.00000, 1.00000, 0.00000], [0.00000, 1.00000, 0.00000], [1.00000, 0.00000, 0.00000], [1.00000, 0.00000, 0.00000], [1.00000, 0.00000, 0.00000], [1.00000, 0.00000, 0.00000]
• Accessor 3 (lines 30-37 in code listing 1), which refers to buffer view 3 (lines 67-71 in code listing 1 ), which refers to buffer 3 (lines 92-94 in code listing 1)with the name "positionsMasksl" (same as “Mask 1 Blue” in FIG. 9B): eight values:
1 , 1 , 1 , 1 , 0, 0, 0, 0
• Accessor 4 (lines 39-46 in code listing 1), which refers to buffer view 4 (lines 73-77 in code listing 1 ), which refers to buffer 4 (lines 96-98 in code listing 1) with the name "positionsMasks2" (same as “Mask 2 Green” in FIG. 9B): eight values:
0, 0, 0, 0, 1 , 1 , 1 , 1
[0122] The code parser decodes the "scenes" properties (lines 132-135 in code listing 1 ), which leads to parsing of the first node (node index 0) of the "nodes" property (lines 111 -128 in code listing 1 ). Node 0 contains an "MPEG_node_avatar" extension, with two items in the "mappings" structure (lines 114-127 in code listing 1 ). The first item defines the left part of the cube ("cube/left", lines 1 17-120 in code listing 1 ) and the second item defines the right part of the cube ("cube/right", lines 122-125 in code listing 1 ). Both items refer to the second node (node index 1 , line 130 in code listing 1 ), which points to the first mesh (lines 101 -107 in code listing 1 ) of the "meshes" main property (lines 100-109 in code listing 1 ). However, since each item has a "mask" property (lines 120 and 125 in code listing 1 ), the items do not use the whole mesh object but only a part of the mesh object. Attribute "semantic" is set to "attributes" (lines 119 and 124 in code listing 1 ). The segments are based on attribute masking, and the resulting meshes are:
• "cube/left": o "attributes/positions": four 3D vectors with these values (same as columns 0-3 of “Vertex positions” of FIG. 9B): [-1.00000, -1.00000, -1.00000], [1.00000, -1.00000, -1.00000], [1.00000, 1.00000, -1.00000], [-1.00000, 1.00000, -1.00000] o "attributes/colors": four 3D vectors with these values:
[0.00000, 1.00000, 0.00000],
[0.00000, 1.00000, 0.00000],
[0.00000, 1.00000, 0.00000],
[0.00000, 1.00000, 0.00000] o "indices": six values (same as columns 0 and 1 of “Triangles” of FIG. 10B):
0, 3, 1 ,
3, 2, 1
• "cube/right": o "attributes/positions": four 3D vectors with these values (same as columns 4-7 of “Vertex positions” of FIG. 9B): [-1.00000, -1.00000, 1.00000], [1.00000, -1.00000, 1.00000], [1.00000, 1.00000, 1.00000], [-1.00000, 1.00000, 1.00000] o “attributes/colors”: four 3D vectors with these values:
[1.00000, 0.00000, 0.00000],
[1.00000, 0.00000, 0.00000],
[1.00000, 0.00000, 0.00000],
[1.00000, 0.00000, 0.00000] o “indices”: six values (same as columns 4 and 5 of “Triangles” of FIG. 10B):
5, 6, 4,
6, 7, 4 [0123] The same result may be obtained using the masking based on "indices" mesh property.
In this case, the gITF is the same as the previous one, except that:
• Accessor 3 (-> buffer view 3 -> buffer 3) is replaced with a mask for indices (same as “Mask 1 Blue” in FIG. 10B): twelve values: 1 , 1 , 0, 0, 0, 0, 0, 0, 0, 0, 0, 0
• Accessor 4 (-> buffer view 4 -> buffer 4) is replaced with a mask for indices (same as “Mask 2 Green” in FIG. 10B): twelve values: 0, 0, 0, 0, 1 , 1 , 0, 0, 0, 0, 0, 0
[0124] The "semantic" property of the "mappings" items in "MPEG_node_avatar" is "indices".
[0125] FIGS. 11 A-11 D and Code Listing 2 provide similar examples to those of FIGS. 9A, 9B, 10A, 10B and Code Listing 1 , with a difference being that the vertex indices and faces/triangles are shuffled.
[0126] FIG. 11 A is a table showing example avatar mesh masks 1102, 1104 and position information 1100 according to some embodiments. The example of FIG.11 shows vertices that are not shown in numerical order. The left side of the cube includes vertices 0, 1 , 2, and 3, while the right side of the cube includes vertices 4, 5, 6, and 7. This example is the same as FIGs. 9A and 9B except that the vertices are shuffled in non-numerical order.
[0127] FIG. 11 B is a table showing example avatar position information 1120 and mesh masks 1122, 1124 according to some embodiments. FIG. 11 B shows the same scenario as in FIGs. 9A and 9B but instead with shuffled vertex indices.
[0128] FIG. 11 C is a table showing example avatar mesh masks 1140, 1142 according to some embodiments. FIG. 11C also shows example masks for a left “face” and a “right” face. The left face includes triangles 1 and 10, while the right face includes triangles 2 and 5.
[0129] FIG. 11 D is a table showing example avatar position information 1160 and mesh masks 1162, 1164 according to some embodiments. In this example, the left side of the cube includes triangles 1 and 10, while the right side of the cube includes triangles 2 and 5. For the left side, triangle 1 is defined by vertices 1 , 3 and 6, and triangle 10 is defined by vertices 3, 4, and 6. For the right side, triangle 2 is defined by vertices 0, 2 and 7, and triangle 5 is defined by vertices 0, 5, and 7. FIG.11 D shows the same scenario as in FIGs. 10A and 10B but instead with shuffled faces/triangles. For some embodiments, a different numbering of the cube than the numbering shown in FIGs. 8A and 8B may be used.
[0130] Code listing 2 below shows another example code listing for a cube according to some embodiments. For code listing 2, the gITF content remains the same as code listing 1 , except for the content of the positions, indices, and mask buffers, and the "MPEG_node_avatar" mappings items reference "indices". Compare code listing 2 with the end of code listing 1 . The value of the "semantic" property is changed from “position” to “ indices” for the left and right sides of the cube.
Code Listing 2 [0131] FIG. 12 is a schematic illustration showing an example avatar according to some embodiments. FIG. 12 shows an example use case for a mesh mask using an avatar model. This example is based on the concept of semantically partitioning an avatar mesh object with the objective of providing information to an engine for rendering, animating, and/or interacting. A mesh may be partitioned as many times as needed for a correct representation of an avatar. For example, a first “mapping” may be used with a “mask” to segment an entire hand model, and a second “mapping” may be used to segment individual fingers of the hand. This scenario permits an application/engine to have fine and coarse control of an avatar mesh given an appropriate level of segmentation.
[0132] FIG. 12 shows a use case for an avatar body model 1200, in which each body part (colored different colors in the color version of FIG. 12) (e.g., the lower left leg 1202) is associated with a single “mapping” component (e.g., lower left leg mapping 1204) and a “mask” (e.g., lower left leg mask 1206) to allow body partitioning. The mask permits identification of a geometric region, e.g., vertices, faces, or other mesh primitives, that belong to a body part pointed to by the “path” attribute.
[0133] This strategy facilitates engines to isolate body parts for rendering, animating and interaction. Given that each body part pointed to by the “path” attribute has an underlying skeletal structure, the application may animate through traditional skinning techniques. Additionally, the application may displace the vertices associated with the “mapping” that are given by the “mask” attribute. In other words, for some embodiments, the application may animate or edit the associated vertices with the “mapping” that are given by the “mask” attribute. This process may be done through many techniques. For some embodiments, the process causes the vertices to be positioned in 3D space. For some embodiments, the same scenario applies to animation and interaction use cases. Because partitioning/segmenting of a mesh is attached to a mapping, several mappings may be generated to segment several parts of a mesh/avatar body. For some embodiments, interactive scene applications may include rendering body parts 1250, animation of individual body components 1252, and interaction of individual body components 1254.
[0134] FIG. 13 is a flowchart illustrating an example process for sharing a unitary mesh object among multiple avatar elements according to some embodiments. For some embodiments, an example process 1300 may include obtaining 1302 scene description data for a 3D scene. For some embodiments of the example process 1300, the scene description data 1304 may include scene element information describing each of a plurality of scene elements in the scene. For some embodiments of the example process 1300, the scene description data may further include a node associated with a mesh object. For some embodiments of the example process 1300, one of the scene elements may be an avatar associated with the node. For some embodiments of the example process 1300, at least two portions of the avatar may be associated with the mesh object. For some embodiments of the example process 1300, the scene description data may further include a mesh mapping associated with the mesh object. For some embodiments of the example process 1300, the mesh mapping may indicate how to partition the mesh object for the at least two portions of the avatar. For some embodiments, the example process 1300 may further include partitioning 1306 the mesh object using the mesh mapping for the at least two portions of the avatar. [0135] While the methods and systems in accordance with some embodiments are generally discussed in context of extended reality (XR), some embodiments may be applied to any XR contexts such as, e.g., virtual reality (VR) / mixed reality (MR) / augmented reality (AR) contexts. Also, although the term “head mounted display (HMD)” is used herein in accordance with some embodiments, some embodiments may be applied to a wearable device (which may or may not be attached to the head) capable of, e.g., XR, VR, AR, and/or MR for some embodiments.
[0136] An example method in accordance with some embodiments may include: obtaining scene description data for a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, a node associated with a mesh object, wherein one of the scene elements is an avatar associated with the node, and wherein at least two portions of the avatar are associated with the mesh object; a mesh mapping associated with the mesh object, wherein the mesh mapping indicates how to partition the mesh object for the at least two portions of the avatar; and partitioning the mesh object using the mesh mapping for the at least two portions of the avatar.
[0137] Some embodiments of the example method may further include processing the scene description data associated with the avatar.
[0138] For some embodiments of the example method, processing the scene description data associated with the avatar may include rendering the avatar as part of the scene using the scene description data.
[0139] For some embodiments of the example method, partitioning the mesh object may include: repeating a process for each of the at least two portions of the avatar associated with the mesh object, wherein the process may include: obtaining a current portion selected from the at least two portions of the avatar; obtaining a current mesh map associated with the current portion; and masking the current mesh map with the mesh object to obtain current position information for the current portion.
[0140] For some embodiments of the example method, obtaining the current mesh map may be performed using the mesh mapping.
[0141] For some embodiments of the example method, the method may be compatible with an extension of Moving Pictures Expert Group-1 Scene Description (MPEG-1 SD) format.
[0142] For some embodiments of the example method, the method may be compatible with an extension of Graphics Language Transmission Format (gITF) scene description format.
[0143] For some embodiments of the example method, rendering the avatar may retain the mesh object as a unitary mesh object for each of the at least two portions of the avatar. [0144] For some embodiments of the example method, the mesh mapping may include an integer index referencing a data structure.
[0145] For some embodiments of the example method, the mesh mapping may include: information indicating the node; and information indicating a path to information related to the node.
[0146] For some embodiments of the example method, the mesh mapping may include a pointer-indexed data structure.
[0147] For some embodiments of the example method, the at least two portions of the avatar may include non-overlapping portions of the avatar.
[0148] For some embodiments of the example method, the scene description data may further include: a second mesh object associated with the node, wherein the at least two portions of the avatar are associated with the second mesh object; and a second mesh mapping associated with the second mesh object, wherein the second mesh mapping indicates how to partition the second mesh object for the at least two portions of the avatar, and wherein the mesh object is different from the second mesh object.
[0149] For some embodiments of the example method, the scene description data may further include: a third mesh object associated with the node, wherein at least two further portions of the avatar are associated with the third mesh object, and wherein the at least two further portions of the avatar are different than the at least two portions of the avatar; and a third mesh mapping associated with the third mesh object, wherein the third mesh mapping indicates how to partition the third mesh object for the at least two further portions of the avatar.
[0150] For some embodiments of the example method, wherein the at least two portions of the avatar may include a first portion and second portion of the avatar, wherein the method may further include: obtaining a first mesh map associated with the first portion of the avatar; and obtaining a second mesh map associated with the second portion of the avatar, wherein the first mesh map shares a common element with the second mesh map.
[0151] For some embodiments of the example method, wherein the at least two portions of the avatar comprise a first portion and second portion of the avatar, wherein the method may further include: obtaining a first mesh map associated with the first portion of the avatar; and obtaining a second mesh map associated with the second portion of the avatar, wherein the first mesh map is unique compared to the second mesh map. [0152] An example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any of the methods listed above.
[0153] Another example method in accordance with some embodiments may include performing a mesh masking process to facilitate mesh segmentation of a single, three- dimensional mesh object corresponding to a node hierarchy delineated in scene description data.
[0154] For some embodiments of another example method, the mesh object is associated with a non-avatar environment.
[0155] For some embodiments of another example method, the mesh object is associated with an avatar environment.
[0156] Although certain examples have been presented in accordance some embodiments that relate to avatar, avatar meshes, and avatar environments, the example techniques (relating to, e.g., meshes) described herein are not limited to avatar, avatar meshes, avatar environments and may for example be applied to any non-avatar mesh.
[0157] Another example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any of the methods listed above.
[0158] An example apparatus in accordance with some embodiments may include at least one processor configured to perform any one of the methods listed above.
[0159] An example apparatus in accordance with some embodiments may include a computer- readable medium storing instructions for causing one or more processors to perform any one of the methods listed above.
[0160] An example apparatus in accordance with some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any one of the methods listed above.
[0161] An example signal in accordance with some embodiments may include a bitstream comprising scene description data generated according to any one of the methods listed above.
[0162] This disclosure describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the disclosure or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
[0163] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are illustrated specifically, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described, and/or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
[0164] In the present disclosure, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture” and “frame” may be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side.
[0165] The terms HDR (high dynamic range) and SDR (standard dynamic range) often convey specific values of dynamic range to those of ordinary skill in the art. However, additional embodiments are also intended in which a reference to HDR is understood to mean “higher dynamic range” and a reference to SDR is understood to mean “lower dynamic range.” Such additional embodiments are not constrained by any specific values of dynamic range that might often be associated with the terms “high dynamic range” and “standard dynamic range.”
[0166] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0167] Various numeric values may be used in the present disclosure, for example. The specific values are for example purposes and the aspects described are not limited to these specific values. [0168] Embodiments described herein may be carried out by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a nonlimiting example, the embodiments can be implemented by one or more integrated circuits. The processor can be of any type appropriate to the technical environment and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0169] Various implementations involve decoding. “Decoding”, as used in this disclosure, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this disclosure, for example, extracting a picture from a tiled (packed) picture, determining an upsampling filter to use and then upsampling a picture, and flipping a picture back to its intended orientation.
[0170] As further examples, in one embodiment “decoding” refers only to entropy decoding, in another embodiment “decoding” refers only to differential decoding, and in another embodiment “decoding” refers to a combination of entropy decoding and differential decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions.
[0171] Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this disclosure can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this disclosure.
[0172] As further examples, in one embodiment “encoding” refers only to entropy encoding, in another embodiment “encoding” refers only to differential encoding, and in another embodiment “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions. [0173] Various embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. The rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. A mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion.
[0174] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method/process.
[0175] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants (“PDAs”), and other devices that facilitate communication of information between endusers.
[0176] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment.
[0177] Additionally, this disclosure may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0178] Further, this disclosure may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0179] Additionally, this disclosure may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0180] It is to be appreciated that the use of any of the following 7”, “and/or”, and “at least one of”, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and/or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended for as many items as are listed.
[0181] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a particular one of a plurality of parameters for region-based filter parameter selection for deartifact filtering. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0182] Implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0183] We describe a number of embodiments. Features of these embodiments can be provided alone or in any combination, across various claim categories and types. Further, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types:
• A bitstream or signal that includes one or more of the described syntax elements, or variations thereof.
• A bitstream or signal that includes syntax conveying information generated according to any of the embodiments described.
• Creating and/or transmitting and/or receiving and/or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof.
• Creating and/or transmitting and/or receiving and/or decoding according to any of the embodiments described.
• A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the embodiments described.
[0184] Note that various hardware elements of one or more of the described embodiments are referred to as “modules” that carry out (i.e. , perform, execute, and the like) various functions that are described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by those of skill in the relevant art for a given implementation. Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions could take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and/or the like, and may be stored in any suitable non-transitory computer- readable medium or media, such as commonly referred to as RAM, ROM, etc.
[0185] Although features and elements are described above in particular combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer- readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, a read only memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1 . A method comprising: obtaining scene description data for a 3D scene, wherein the scene description data comprises: scene element information describing each of a plurality of scene elements in the scene, a node associated with a mesh object, wherein one of the scene elements is an avatar associated with the node, and wherein at least two portions of the avatar are associated with the mesh object; a mesh mapping associated with the mesh object, wherein the mesh mapping indicates how to partition the mesh object for the at least two portions of the avatar; and partitioning the mesh object using the mesh mapping for the at least two portions of the avatar.
2. The method of claim 1 , further comprising processing the scene description data associated with the avatar.
3. The method of claim 2, wherein processing the scene description data associated with the avatar comprises rendering the avatar as part of the scene using the scene description data.
4. The method of claim 1 , wherein partitioning the mesh object comprises: repeating a process for each of the at least two portions of the avatar associated with the mesh object, wherein the process comprises: obtaining a current portion selected from the at least two portions of the avatar; obtaining a current mesh map associated with the current portion; and masking the current mesh map with the mesh object to obtain current position information for the current portion.
5. The method of claim 4, wherein obtaining the current mesh map is performed using the mesh mapping.
6. The method of claim 1 , wherein the method is compatible with an extension of Moving Pictures Expert Group-1 Scene Description(MPEG-1 SD) format.
7. The method of claim 1 , wherein the method is compatible with an extension of Graphics
Language Transmission Format (gITF) scene description format.
8. The method of claim 1 , wherein rendering the avatar retains the mesh object as a unitary mesh object for each of the at least two portions of the avatar.
9. The method of claim 1 , wherein the mesh mapping comprises an integer index referencing a data structure.
10. The method of claim 1 , wherein the mesh mapping comprises: information indicating the node; and information indicating a path to information related to the node.
11. The method of claim 1 , wherein the mesh mapping comprises a pointer-indexed data structure.
12. The method of claim 1 , wherein the at least two portions of the avatar comprise nonoverlapping portions of the avatar.
13. The method of claim 1 , wherein the scene description data further comprises: a second mesh object associated with the node, wherein the at least two portions of the avatar are associated with the second mesh object; and a second mesh mapping associated with the second mesh object, wherein the second mesh mapping indicates how to partition the second mesh object for the at least two portions of the avatar, and wherein the mesh object is different from the second mesh object.
14 The method of claim 1 , wherein the scene description data further comprises: a third mesh object associated with the node, wherein at least two further portions of the avatar are associated with the third mesh object, and wherein the at least two further portions of the avatar are different than the at least two portions of the avatar; and a third mesh mapping associated with the third mesh object, wherein the third mesh mapping indicates how to partition the third mesh object for the at least two further portions of the avatar.
15. The method of claim 1 , wherein the at least two portions of the avatar comprise a first portion and second portion of the avatar, wherein the method further comprises: obtaining a first mesh map associated with the first portion of the avatar; and obtaining a second mesh map associated with the second portion of the avatar, wherein the first mesh map shares a common element with the second mesh map.
16. The method of claim 1 , wherein the at least two portions of the avatar comprise a first portion and second portion of the avatar, wherein the method further comprises: obtaining a first mesh map associated with the first portion of the avatar; and obtaining a second mesh map associated with the second portion of the avatar, wherein the first mesh map is unique compared to the second mesh map.
17. An apparatus comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform the method of any one of claims 1 through 16.
18. A method comprising: performing a mesh masking process to facilitate mesh segmentation of a single, three- dimensional mesh object corresponding to a node hierarchy delineated in scene description data.
19. The method of claim 18, wherein the mesh object is associated with a non-avatar environment.
20. The method of claim 18, wherein the mesh object is associated with an avatar environment.
21 . An apparatus comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform the method of any one of claims 18-20.
22. An apparatus comprising at least one processor configured to perform the method of any one of claims 1 -20.
23. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method of any one of claims 1 -20.
24. An apparatus comprising at least one processor and at least one non-transitory computer- readable medium storing instructions for causing the at least one processor to perform the method of any one of claims 1 -20.
25. A signal including a bitstream comprising scene description data generated according to any one of claims 1 -20.
EP24734914.5A 2023-06-30 2024-06-26 Avatar mesh masking Pending EP4736461A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23306091 2023-06-30
PCT/EP2024/067914 WO2025003196A1 (en) 2023-06-30 2024-06-26 Avatar mesh masking

Publications (1)

Publication Number Publication Date
EP4736461A1 true EP4736461A1 (en) 2026-05-06

Family

ID=87429544

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24734914.5A Pending EP4736461A1 (en) 2023-06-30 2024-06-26 Avatar mesh masking

Country Status (4)

Country Link
EP (1) EP4736461A1 (en)
KR (1) KR20260030888A (en)
CN (1) CN121794989A (en)
WO (1) WO2025003196A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20230079256A (en) * 2020-10-02 2023-06-05 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. Data Streams, Devices and Methods for Volumetric Video Data
US20230128878A1 (en) * 2021-10-27 2023-04-27 Illusio, Inc. Three-dimensional content optimization for uniform object to object comparison

Also Published As

Publication number Publication date
WO2025003196A1 (en) 2025-01-02
KR20260030888A (en) 2026-03-06
CN121794989A (en) 2026-04-03

Similar Documents

Publication Publication Date Title
WO2025210015A1 (en) 3d gaussians splatting in scene description
US20260099191A1 (en) Trigger activation mechanism in time-evolving scene description
EP4736461A1 (en) Avatar mesh masking
EP4629174A1 (en) Polygons encoding in scene descriptions
EP4668765A1 (en) Efficient representation of curve sets
EP4686211A1 (en) Signaling avatar landmarks in scene description
WO2025012259A1 (en) Generic conditional trigger in virtual environments
EP4694140A1 (en) Avatar json interchange file format for capture parameters encoding
WO2026002425A1 (en) 3d gaussians splatting in scene description
WO2025068063A1 (en) Real objects anchoring in a scene description
WO2025008261A1 (en) Shared event-based update in scene description
WO2025149465A1 (en) Avatar format validation in scene descriptions
EP4718852A1 (en) Stitched gaussian mixture in scene description
WO2025012249A1 (en) Generic avatar trigger in virtual environments
US20260112135A1 (en) Dynamic gaussian splatting learned from hierarchical motion model
CA3268477A1 (en) Trigger activation mechanism in time-evolving scene description
WO2025073615A1 (en) System and method of mesh primative semantical coding
WO2026033001A1 (en) General signaling of model properties encoding in scene and avatar descriptions
WO2024126025A1 (en) Event based update in scene description
EP4515498A1 (en) Systems and methods for providing interactivity with light sources in a scene description
WO2025011964A1 (en) Controllers for 3d scene descriptions
WO2025215061A1 (en) Avatar metadata format
WO2024184076A1 (en) Scene description update system for game engines

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE