WO2026006752A1 - Conditional, probabilistic generation of 3d virtual environment based on incomplete information available to an extended reality device - Google Patents
Conditional, probabilistic generation of 3d virtual environment based on incomplete information available to an extended reality deviceInfo
- Publication number
- WO2026006752A1 WO2026006752A1 PCT/US2025/035725 US2025035725W WO2026006752A1 WO 2026006752 A1 WO2026006752 A1 WO 2026006752A1 US 2025035725 W US2025035725 W US 2025035725W WO 2026006752 A1 WO2026006752 A1 WO 2026006752A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- keyframe
- processor
- keyframes
- dimensional
- acts
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T19/00—Manipulating three-dimensional [3D] models or images for computer graphics
- G06T19/006—Mixed reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N7/00—Computing arrangements based on specific mathematical models
- G06N7/01—Probabilistic graphical models, e.g. probabilistic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T17/00—Three-dimensional [3D] modelling for computer graphics
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T19/00—Manipulating three-dimensional [3D] models or images for computer graphics
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/10—Segmentation; Edge detection
- G06T7/11—Region-based segmentation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/50—Depth or shape recovery
- G06T7/55—Depth or shape recovery from multiple images
- G06T7/579—Depth or shape recovery from multiple images from motion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/26—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/762—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using clustering, e.g. of similar faces in social networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/20—Scenes; Scene-specific elements in augmented reality scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/60—Type of objects
- G06V20/64—Three-dimensional [3D] objects
- G06V20/647—Three-dimensional [3D] objects by matching two-dimensional images to three-dimensional objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10024—Color image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
Definitions
- VR virtual-reality
- AR augmented reality
- MR mixed-reality
- XR extended-reality
- a VR scenario typically involves presentation of digital or virtual image information without transparency to other actual real-world visual input
- an AR or MR scenario typically involves presentation of digital or virtual image information as an augmentation to visualization of the real world around the user such that the digital or virtual image (e.g., virtual content) may appear to be a part of the real world.
- MR may integrate the virtual content in a contextually meaningful way, whereas AR may not.
- An extended-reality system has the capabilities to create virtual objects that appear to be, or are perceived as, real. Such capabilities, when applied to the Internet technologies, may further expand and enhance the capability of the Internet as well as the user experiences so that using the web resources is no longer limited by the planar, two-dimensional representation of web pages.
- Neural radiance fields are techniques that generate 3D representations of an object or scene from sparse two-dimensional (2D) images by using machine learning. These techniques involve encoding an object or scene into an artificial neural network, which predicts the light intensity -- or radiance -- at any point in the 2D image to generate novel 3D views from different angles.
- the NeRF model enables learning of novel view synthesis, scene geometry, and the reflectance properties of the scene.
- Gaussian splatting is a method for representing 3D scenes and rendering novel views.
- Gaussian splatting a 3D world is represented with a set of 3D points where each point is a 3D Gaussian with its own unique parameters that are fitted per scene such that renders of this scene match closely to the known dataset images.
- 3D Gaussian splatting is, therefore, analogous to triangle rasterization in computer graphics, which is used to draw many triangles on the screen. However, instead of drawing triangles, they are Gaussian. Therefore, a Gaussian is described by parameters such as position, covariance measuring how a Gaussian stretches and/or scales, color such as Red, Green, and Blue, and alpha measuring the transparency of a Gaussian, etc.
- 3D Gaussian splatted radiance field methods directly embed language-based semantics into the Gaussians in their attempt to address the integration of segmentation and open vocabulary semantics into 3D Gaussian Radiance Fields.
- the 3D representations may not be semantically disjoint to allow manipulation (e.g., user interactions in an extended-reality application).
- Some legacy approaches uses simultaneous localization and mapping (SLAM) for constructing or updating a map of an environment while keeping track of the user’s location in the environment.
- SLAM simultaneous localization and mapping
- Some of these legacy SLAM approaches have attempted to combine Gaussian spatting. Nonetheless, these legacy approaches either require RDG-D (red, blue, and green for color plus depth) data or are incompatible with real-time or nearly real-time applications.
- the extended-reality system includes a wearable eyepiece that presents virtual contents to a user, a belt pack operatively coupled to the wearable eyepiece, a processor; and a non- transitory computer readable medium storing thereupon a sequence of instructions which, when executed by a model with the processor, causes the processor to estimate a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device, to generate a semantically annotated, manipulatable three-dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece, and to modify the semantically annotated, manipulatable 3D representation in real-time or nearly real-time in response to a user interaction.
- 3D three-dimensional
- the set of acts executed by the model further includes receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; determining, from the video or the sequence of RGB images, the plurality of keyframes; and estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images.
- RGB Red, Green, Blue
- the set of acts executed by the model further includes performing a global bundle adjustment that corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
- the model further executes the set of acts to identify a non-keyframe from the video or the sequence of RGB images.
- a nonkeyframe pose may be determined for the non-keyframe, and a training data set may be augmented at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
- the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to generate a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
- the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to establish a semantic association among the plurality of object segmentation masks at least by using a zeroshot, one-shot, or few-shot neural network or a multi-view geometry.
- a three-dimensional (3D) point set may be generated based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
- the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to improve 3D Gaussian splatted neural radiance field at least by attaching a lower-dimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three- dimensional (3D) representation.
- Some embodiments are directed to a method for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems.
- a plurality of keyframes, a plurality of keyframe poses, and depth data may be estimated or otherwise determined from a plurality of captures that is captured by at least the extended reality device.
- a semantically annotated, manipulatable three-dimensional (3D) representation may be generated in a physical environment for perception by the user wearing the wearable eyepiece. The semantically annotated, manipulatable 3D representation may then be modified in real-time or nearly real-time in response to a user interaction.
- a video or a sequence of RGB (Red, Green, Blue) images may be received for estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, wherein the video or the sequence of RGB images are not required to include depth information.
- the plurality of keyframes may be determined from the video or the sequence of RGB images. At least one keyframe of the plurality of keyframes may be estimated across the video or the sequence of RGB images.
- a global bundle adjustment may be performed to correct a camera pose corresponding to at least one of the plurality of keyframe poses or to perform a loop closure for estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data.
- a non-keyframe may be identified from the video or the sequence of RGB images for estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data.
- a non-keyframe pose may be determined for the non-keyframe.
- a training data set may be augmented at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
- generating the semantically annotated, manipulatable three-dimensional (3D) representation may include establishing a semantic association among the plurality of object segmentation masks at least by using a zero-shot, one-shot, or few-shot neural network or a multi-view geometry.
- a three- dimensional (3D) point set may be generated based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
- 3D Gaussian splatted neural radiance field may be improved at least by attaching a lower-dimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three-dimensional (3D) representation.
- Some embodiments are directed at a hardware system that may be invoked to perform any of the methods, processes, or sub-processes disclosed herein.
- the hardware system may include or involve an extended-reality system having at least one processor or at least one processor core, which executes one or more threads of execution to perform any of the methods, processes, or sub-processes disclosed herein in some embodiments.
- the hardware system may further include one or more forms of non-transitory machine-readable storage media or devices to temporarily or persistently store various types of data or information.
- Some embodiments are directed at an article of manufacture that includes a non-transitory computer readable medium having stored thereupon a sequence of instructions which, when executed by at least one processor or at least one processor core, causes the at least one processor or the at least one processor core to perform any of the methods, processes, or sub-processes disclosed herein.
- Some exemplary forms of the non-transitory machine-readable storage media may also be found in the System Architecture Overview section below.
- the non-transitory computer readable medium stores thereupon a sequence of instructions which, when executed by a model with a processor, causes the processor to estimate a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device, to generate a semantically annotated, manipulatable three- dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece, and to modify the semantically annotated, manipulatable 3D representation in real-time or nearly real-time in response to a user interaction.
- 3D three- dimensional
- the set of acts executed by the model further includes receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; determining, from the video or the sequence of RGB images, the plurality of keyframes; and estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images.
- RGB Red, Green, Blue
- the set of acts executed by the model further includes performing a global bundle adjustment that corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
- the model further executes the set of acts to identify a non-keyframe from the video or the sequence of RGB images.
- a nonkeyframe pose may be determined for the non-keyframe, and a training data set may be augmented at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
- the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to generate a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
- the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to establish a semantic association among the plurality of object segmentation masks at least by using a zeroshot, one-shot, or few-shot neural network or a multi-view geometry.
- a three-dimensional (3D) point set may be generated based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
- (3D) representation may further execute the set of acts to improve 3D Gaussian splatted neural radiance field at least by attaching a lower-dimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three- dimensional (3D) representation.
- FIG. 1 illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- FIG. 1A illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1 in some embodiments.
- FIG. 1 B illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1 in some embodiments.
- FIG. 10 illustrates a more detailed block diagram conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- FIG. 2A illustrates a high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- FIG. 2B illustrates another high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments.
- FIG. 3A illustrates a simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments.
- FIG. 3B illustrates another simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- FIG. 3C illustrates more details about a portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments.
- FIG. 3D illustrates more details about another portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments.
- FIG. 3E illustrates a simplified block diagram of a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- FIG. 3F illustrates a block diagram that continues from FIG. 3E and generates segmentation masks, according to some embodiments.
- FIG. 3G illustrates an example block diagram for learning semantic labels in 3D, according to some embodiments.
- FIG. 4 illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- FIG. 5A illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments.
- FIGS. 5B-5C illustrate a more detailed block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments.
- FIG. 6A illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields without pre-training, according to some embodiments.
- FIG. 6B illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields with pre-training, according to some embodiments.
- FIG. 7A illustrates s simplified schematic diagram for a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments.
- FIG. 7B illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments.
- FIG. 7D illustrates some example modifications of the simplified schematic diagram illustrated in FIG. 7C, according to some embodiments.
- FIG. 7F illustrates a simplified block diagram for an example text encoder that may be employed in FIG. 7A or 7B, according to some embodiments.
- FIGS. 7G-7H illustrate an example attention mechanism that may be employed for a sequence transduction module in a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments.
- FIG. 8 illustrates a simplified example of a wearable extended-reality (XR) device with a belt pack external to the XR glasses in some embodiments.
- XR extended-reality
- FIG. 9 illustrates a computerized system on which some of the methods described herein may be implemented.
- FIG. 1 illustrates a block diagram for learning semantic labels for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More specifically, FIG. 1 as well as other embodiments described herein illustrates an online reconstruction of instance-aware Gaussian splatted neural radiance field technique that may be used in conditional, probabilistic generation of 3D virtual environment in response to a prompt by a user wearing an extended-reality device in real-time or nearly real-time based at least in part upon incomplete information made available to the extended-reality device.
- the system When a system renders the virtual object using rasterization techniques to “draw” the cluster, grouping, or instance of one or more Gaussians, the system is made aware of the cluster, grouping, or instance (as well as the semantics) and may thus render the cluster, grouping, or instance of one or more Gaussians as an undivided whole to make the rendered cluster, instance, or grouping manipulatable or interactable (e.g., in response to prompts).
- a plurality of keyframes, respective camera poses (may also be referred to as keyframe poses) for the plurality of keyframes, and/or depth information may be determined or estimated at 102 from a video or a sequence of images (e.g., photographs) that is acquired by or provided by (e.g., obtained from another XR device or server) an XR device worn by a user.
- keyframe poses respective camera poses
- depth information may be determined or estimated at 102 from a video or a sequence of images (e.g., photographs) that is acquired by or provided by (e.g., obtained from another XR device or server) an XR device worn by a user.
- a semantically annotated, editable 3D representation may be generated at 104 and placed in a physical environment as perceived by a user wearing the XR device.
- a global bundle adjustment may be performed globally (e.g., to the entire video, the entire sequence of RGB images, or a portion thereof) at 106A to correct one or more camera poses and/or to perform loop closure in some embodiments.
- one or more non-keyframes in the video or the sequence may be embedded at 108A, and the one or more corresponding camera poses for the one or more embedded non-keyframes may be estimated at 110A.
- a training data set may be augmented at 112A at least by embedding the one or more non-keyframes and the one or more corresponding camera poses for the one or more non-keyframes into the training data set, wherein the training data set is used to train a model that performs the acts described above with reference to FIG. 1 .
- FIG. 1 B illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1A in some embodiments. More particularly, FIG. 1 B illustrates more details about generating a semantically annotated, editable 3D representation at 104 in FIG. 1A.
- one or more object segmentation masks may be generated at 102B at least by analyzing the plurality of keyframes. In some embodiments where one or more non-keyframes have been embedded, one or more object segmentation masks may be generated at 102B at least by analyzing the plurality of keyframes as well as the one or more embedded non-keyframes. More details about object segmentation masks will be described below with reference to multiple drawing figures.
- Semantic association may be established among the one or more object segmentation masks at least by using a zero-shot, one-shot, or few-shot tracker neural network or multi-view geometry at 104B.
- a 3D initial potin set may be generated at 106B based at least in part upon an inverse projection of the estimated or determined camera poses, depth data, and/or the associated depth covariance (e.g., uncertainty).
- the 3D initial point set may be generated by selecting points with low uncertainty (e.g., uncertainty below a certain threshold) and by reverse-projecting the selected points in 3D using the corresponding pose and depth information.
- a lower dimensional semantic vector may be attached to a lower dimensional grouping of Gaussians at 108B to improve Gaussian splatted radiance field with semantics.
- a semantic vector may include a compact representation of a true or predicted class label (hence in a lower dimensional space) for an object.
- a low dimensional grouping of Gaussians may be generated before attaching semantic vectors or labels so as to form semantically disjointed objects.
- various embodiments generate a lower-dimensional grouping of Gaussians and then attach semantic meaning to each such low-dimensional grouping of Gaussians.
- 3D object representations generated by various techniques described herein may be semantically disjoint from each other so that a user wearing an XR device may interact with or manipulate such 3D object representations independently and individually.
- a 3D space may be defined as a set of Gaussians where each Gaussian is described by a set of parameters that are calculated or determined by machine learning. These parameters may include position, covariance measuring how a Gaussian stretches and/or scales, color such as Red, Green, and Blue, and alpha measuring the transparency of a Gaussian, etc.
- Gaussian rasterization may be deemed analogous to triangle rasterization in legacy computer graphics that draw triangles on a display device whereas Gaussian rasterization draws Gaussians, instead of triangles.
- Gaussian splatting is a rendering technique and includes a rasterization technique for real-time (or nearly real-time) 3D reconstruction and rendering of images taken from multiple points of view.
- 3D Gaussian splatting represents a 3D scene as a large number of particles (Gaussians), and each 3D Gaussian is associated with its position, orientation, scale, opacity (or transparency), and color.
- the Gaussian may be first converted into the 2D space and then organized for rendering.
- One method to render Gaussians may include creating a point cloud from one or more images; converting each point in the point cloud to a Gaussian by inferring or determining one or more of the aforementioned Gaussian parameters from the data or metadata associated with the point in the point cloud to enable rasterization; training a network to produce high-quality results by using, for example yet without limitation, stochastic gradient descent and by adjusting the Gaussian parameters according to the loss; and performing differentiable Gaussian rasterization by projecting the 2D Gaussian from the camera’s perspective (e.g., pose) and by repeated backward and forward combination for each Gaussian.
- perspective e.g., pose
- 1 C illustrates a block diagram for learning semantic labels in 3D.
- a rendering process may be invoked at 102C, and a number (N) of labels from one or more segmentation masks in 3D may be projected onto a 2D space at 104C to generate a number (N) of projected labels.
- the number (N) of projected labels may be decompressed at 106C into a number (N) of decompressed labels.
- decompressing a projected semantic label into a decompressed semantic label may be performed using a linear layer, a dense layer, or a fully connected (FC) layer where every input neuron or node is connected to every output neuron or node, and the linear layer maps an input to an output with a weight (W or a weight matrix) and a bias (b or a bias vector).
- the number (N) of decompressed labels may be expanded at 108C into the number plus one (N+1 ) labels, and these N+1 decompressed labels may be transformed into a probabilistic distribution of N+1 outcomes at 110C.
- a hierarchical part-whole structure of objects in an input image may be learned at 112C using a decoder based at least in part upon a frequency schedule. These embodiments learn a hierarchical part-whole structure of objects, rather than learning a semantic class per pixel as most of legacy approaches otherwise learn.
- a frequency schedule embeds one or more objects into different latent spaces (also referred to as embedding space or latent feature space) based at least in part upon the one or more respective frequencies where a latent space comprises an embedding of a set of items within a manifold where similar items are positioned closer to each other.
- a plurality of Gaussians may be clustered at 114C into one or more clusters based at least in part upon respective object identifiers. For example, a group of Gaussians corresponding to the same object may be clustered into the same cluster.
- a semantic vector may be embedded at 116C to each of the one or more clusters. In some embodiments, embedding a semantic vector to a cluster may be achieved by running a segmenter (e.g., a contrastive language-image pre-training or CLIP model) on a plurality of masked image crops.
- a segmenter e.g., a contrastive language-image pre-training or CLIP model
- An object may be selected at 118C at least by querying the embedded semantic vectors obtained from 116C using a user input with a textual model and/or speech-to-text model.
- the object may be provisioned by an extended reality device worn by a user in an extended reality session to facilitate user interactions with the rendered object at 120C, and the 3D representation of the object may be updated in real-time or nearly real-time to reflect the effects of the user interactions upon the object.
- the object is semantically disjoint from other objects or 3D representations generated by the extended reality device to allow such user interactions and manipulations.
- FIG. 2A illustrates a high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More particularly, FIG. 2A illustrates provisioning reconstruction of 3D representations of a scene using neural radiance field (NeRF). NeRF is one of several techniques that may be used in various embodiments described herein for reconstructing 3D scenes without contours to address the shortcomings of photometry that is unable to represent scenes that do not have contours.
- a prompt or input may be received at 202A for reconstructing a 3D representation for interaction by a user using an extended reality device. For example, a user wearing an XR device may provide a prompt as a condition for generating a 3D representation for a scene or a portion thereof.
- a prompt may include a textual input, a voice input, and/or an input of an image although different inputs may be processed with different encoders to generate respective embeddings.
- a textual input e.g., by entering textual instructions in natural language
- an image input may be processed by an image encoder
- a voice input may be processed by a speech encoder or by a combination of a speech-to-text encoder and a textual encoder.
- a 3D image or 3D model may be generated at 204A using one or more captures in response to the prompt.
- at least one of the one or more captures may be captured by using one or more sensors.
- an XR device may invoke its image sensor(s) to capture one or more 2D captures (e.g., photographs or other types of images) of a scene for the generation of the 3D image or 3D model at 204A.
- the XR device may invoke an out-painting module to expand the 3D image or 3D model beyond the wall or room by predicting what may lie beyond the wall or outside the room.
- FIG. 2B illustrates another high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments. More particularly, FIG. 2B illustrates more details about neural radiance field that represents the light field as integrated radiance rays through the 3D space.
- a plurality of 2D captures may be received at 202B where the plurality of 2D captures comprises color information such as the RGB data but not depth information.
- FIG. 2D are in sharp contrast with conventional SLAM-based Gaussian splatting techniques that require RGB-D (color information in RGB plus the depth information) for reconstruction of 3D representations.
- a 2D capture may be deemed a five-dimensional (5D) input including its spatial location (x, y, z) and a viewing direction (in terms of theta for azimuth and phi for elevation).
- legacy approaches require six-dimensional (6D) input - the spatial location including depth (x, y, z, d) and viewing direction (theta and phi).
- the volume density at a location includes the differential probability of a ray that terminates at an infinitesimal particle at that location in some embodiments.
- An output including view-dependent emitted radiance may be generated at 204B in response to the aforementioned input at the spatial location (emitted color in terms of R, G, B) and the volume density (alpha measuring the transparency or opacity at the location).
- the output view-dependent emitted radiance may be generated at 204B by compositing the color information and the volume density into the view-dependent emitted radiance.
- Depth information may be determined at206B from the 2D captures.
- a view may be synthesized at 208B at least by querying or sampling the aforementioned 5D coordinates in the input along a camera ray and further by using volume rendering techniques to project the queried output colors and the volume density information (alpha) to a specific pixel represented in the 3D representation of the view where the process illustrated in FIG. 2B represents the light field as integrated radiance along light rays through the 3D space.
- synthesizing and rendering a view includes estimating the expected color of a camera ray with near and far bounds, and the expected color may be expressed as an integral to represent the accumulated transmittance along the ray from the near bound to the far bound without hitting any other particles.
- the expected color for the output may be represented as an integral for a camera ray traced through each pixel of a desired virtual camera for the 3D representation.
- the synthesized view may be expanded into a 3D expanded view at 21 OB at least by using the depth information determined at 206B from the 2D captures.
- the volume rendering techniques may be optionally improved at 212B at least by reducing a residual or loss between the 3D expanded view and a corresponding ground truth image during training of the model for generating the aforementioned 3D expanded views for 3D representation.
- FIG. 3A illustrates a simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments. More particularly, these embodiments illustrated in FIG. 3A comprise a model that may be invoked by various techniques described herein to fill a “hole” or “gap” (e.g., an occluded area in a scene). In these embodiments, the example network receives an input image 302A (e.g., a photograph captured by an XR device when a user wearing the XR device enters a scene) which may be turned into a masked input 304A.
- an input image 302A e.g., a photograph captured by an XR device when a user wearing the XR device enters a scene
- the masked input 304A may be sent to a coarse network 306A for processing.
- the coarse network 306A may comprise a plurality of dilated convolution layers 326A that may be used to provide sufficiently large receptive fields.
- the plurality of dilated convolution layers thus enables the coarse network to capture a wider context, without significant increase in computational costs.
- the plurality of dilated convolution layers may include a first dilated convolution layer with a dilation rate of 2, a second dilated convolution layer with a dilation rate of 4, a third dilated convolution layer with a dilation rate of 8, and a fourth dilated convolution layer with a dilation rate of 16, each with 1x1 stride for 256 outputs each, in some embodiments.
- the refinement network 31 OA may include a plurality of dilated convolution layers 326A to expand the receptive field as well as a contextual attention mechanism.
- the coarse network 306A may include, from the input, a first layer having a 5x5 kernel with one dilation and 1x1 stride, a second layer having a 3x3 kernel with one dilation and 2x2 stride, a third layer having a 3x3 kernel with one dilation and 1x1 stride, a fourth layer having a 3x3 kernel with one dilation and 2x2 stride, a fifth layer having a 3x3 kernel with one dilation and 1x1 stride, and a sixth layer having a 3x3 kernel with one dilation and 1x1 stride.
- the sixth layer feeds into the aforementioned four dilated convolution layers.
- the sixth layers receiving outputs from the four dilated convolution layers may be the aforementioned six layers arranged in the reverse order.
- the coarse network 306A processes the masked input 304A to generate the coarse output 308A which may be forwarded to a refinement network 31 OA which the generates the inpainting results 312A to fill the “hole” or “gap” that is masked in 304A by a mask or patch (e.g., for the purposes of training the network).
- the inpainting result 312A may include the global image 318A (e.g., the masked input image 304A plus the inpainting result for the “hole” or “gap”) and a local image 314A (e.g., the local image for the inpainting result for the “hole” or “gap”).
- the global image 318A may be fed to a global discriminator 320A that determines the loss (e.g.,. cross-entropy loss) between the global image 318A and the corresponding ground truth, and the local inpainting result 314A may be processed by a local discriminator 316A to determine the loss (e.g., cross-entropy loss).
- the losses from the global discriminator 320A and the local discriminator 316A may be forwarded to the loss module 322A that determines the reconstruction loss and the GAN (generative adversarial network) loss both of which may then be propagated to tune the coarse network output 308A and the refinement network output (the inpainting result) 312A.
- FIG. 3B illustrates another simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More particularly, FIG. 3B illustrates more details about the network illustrated in FIG. 3A.
- an input 302B e.g., a 3D input image or a 2.5D input image having 2D image with depth information
- a mask 304B e.g., a binary mask
- a textual or voice input 324B into a masked output 306B, a global spatial loss 326B, and a local spatial loss 327B.
- the masked output 306B may be provided to a dilated convolution network 308B, which may also receive the mask 304B, to increase the receptive field.
- the output of the dilated convolution network 308B may be provided to the coarse prediction module 31 OB that also receives the global spatial loss 326B, if available, to generate a coarse inpainting result 360B.
- the output including the masked output 306B and the coarse inpainting result 360B may be provided to a refinement network 312B which receives contextual attention 326B and generates a refined prediction 314B.
- the refined prediction 314B including the masked output 306B with a refined inpainting result 362B may be provided to a global critic 316B (e.g., the global discriminator 320A in FIG. 3A) to generate a global output 320B.
- the refined inpainting result 362B may be provided to a local critic 318B (e.g., the local discriminator 316A in FIG. 3A) that generates a local output 322B.
- the global output 320B and the local output 322B may be provided to a loss module 332B that determines the reconstruction loss 321 B (e.g., a weighted Mean Squared Error (MSE) loss) and the GAN loss 323B.
- MSE weighted Mean Squared Error
- the reconstruction loss 321 B is also to enhance training stability.
- the GAN loss 323B and the spatial loss 327B may be provided for training 328B that propagates the losses (e.g., by backward propagation) to fine turn the refinement network 312B in order to produce high-fidelity inpainting results.
- the reconstruction loss 321 B and the spatial loss 326B may be provided to training 328B to fine tune the dilated convolution network 308B and the coarse prediction network 310B.
- the global discriminator takes the entire image as input, while the local discriminator takes only a small region around the completed area as input.
- both the global and local discriminators are trained to determine whether an image (e.g., the output image with inpainting result) is real or completed by the completion network (not shown), while the completion network is trained to fool both discriminator networks.
- FIG. 3C illustrates more details about a portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments. More particularly, FIG. 3C illustrates more details about the coarse prediction (e.g., 310B in FIG. 3B or 306A in FIG. 3A).
- the input 302B may be masked by a mask or patch 304B (e.g., a binary mask) to generate the masked input 306B which may be fed to a dilated convolution network 308B, and the input 302B may be further optionally supplemented with textual input 324B as similarly shown in FIG. 3B.
- a mask or patch 304B e.g., a binary mask
- the dilated convolution network 308B may process the masked input 306B to generate an intermediate output that may be further processed by a coarse prediction network 31 OB to generate a coarse inpainting result or a patch 302C.
- the coarse prediction output including the masked input 306B and the patch 302C may be forwarded to a global critic 316B that generates the global output 320; and the coarse inpainting prediction or the patch 302C may be processed by a local critic 318B to generate a local output 322B.
- Both the global output 320B and the local output 322B may be forwarded to a loss module 322B that generates the reconstruction loss 321 B and the GAN loss 323B.
- the reconstruction loss 321 B may be backward propagated 304C to the masked output 306B in order to train the dilated convolution network 308B and/or the coarse prediction network 310B.
- the global critic 318B and the local critic 316B are included for training purposes but are not included for testing purposes.
- FIG. 3D illustrates more details about another portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments. More particularly, FIG. 3D illustrates more details about a refinement network.
- a first input 302D may be provided to an extraction module 304D that produces foreground input features 306D and background input features 308D.
- the background input features 308D may be further processed by a filter generation module 310D to generate one or more filters 312D for subsequent convolution.
- Both the foreground input features 306D and the one or more filters 312D may be provided to a convolution network 314D that generates a matching score 316D.
- attention score may be determined at 316D for each pixel in the foreground input features 306D.
- a vector-to-probabilistic distribution may be generated at 318D based at least in part upon the matching score determined at 316D.
- a channel-wise softmax may be performed to compare and determine the attention score 320D for each pixel and the attention map 322D to transform the values into the probabilistic distribution of possible outcomes at 318D.
- Both the attention scores 320D and the attention map 322D may be provided to an attention fusion mechanism 324D that generates the attention score 326D.
- the attention fusion mechanism 324D performs a left-right propagation to compute an intermediate attention score followed by top-down propagation with a kernel (e.g., an identity matrix) of a certain size to compute the attention score 326D in some of these embodiments.
- the attention score 326D may be forwarded to a deconvolution layer 328D to generate the second output 330D which comprises the output (e.g., the refined prediction 314B) of the refinement network (e.g., 312B).
- FIG. 3E illustrates a simplified block diagram of a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More particularly, FIG. 3E illustrates a block diagram of a 3D Gaussian splatted radiance field module.
- a video or a sequence of RGB images 304E may be provided to two simultaneous localization and mapping (SLAM) module 306E.
- SLAM simultaneous localization and mapping
- the devices e.g., an XR device is also not required to include any depth sensors.
- a SLAM module 306E processes the RGB images 304E to determine and track keyframes, to generate the depth data estimate 308E, to perform a global bundle adjustment operation 310E (e.g., to globally correct camera poses and/or to perform loop closure), to embed one or more non-keyframes in the video or the sequence of RGB images 312E, and/or to estimate camera poses for the respective keyframes 314E.
- the SLAM 306E thus generates a plurality of keyframes 316E, a plurality of camera poses 318E for the plurality of keyframes 316E, and depth information 320E.
- the respective outputs of these two SLAM modules 306E are provided to a 3D Gaussian splatted radiance field module 322E to generate one or more new scenes 324E (e.g., reconstruction of 3D representations) including one or more individually, independently manipulatable objects 326E by generating these one or more semantically disjointed objects.
- these embodiments instead of directly embedding language-based semantics onto the Gaussians, these embodiments first generate a low dimensional grouping of Gaussians (e.g., based on the object identifier to which the Gaussians correspond) into semantically disjoint objects before attaching semantic meaning to each low dimensional grouping.
- the output of one of the two SLAM modules 306E may be provided to initialize the 3D scene 314E.
- a plurality of 3D scenes 302E may be captured by one or more devices such as cameras, video cameras, XR devices, etc., and the plurality of 3D scenes may also be provided to the 3D Gaussian splatted radiance field module 322E for the generation or reconstruction of 3D representations 324E.
- FIG. 3F illustrates a block diagram that continues from FIG. 3E and generates segmentation masks, according to some embodiments.
- the process illustrated in FIG. 3F occurs after acquiring the keyframes, the camera poses, and the depth information described above with reference to FIG. 3E.
- an image, a sequence of images, or a video may be received at 302F, and the image, the sequence, or the video may be analyzed at 304F using a 2D grid.
- the image, the sequence of images, or the video may be acquired by a device such as an XR device worn by a user.
- a grid may include, for example without limitation, a 32 x 32 grid by using a segmentation process.
- the segmentation process returns a part-whole hierarchy of masks, rather than learning a semantic class per pixel as most of legacy approaches otherwise learn.
- a plurality of masks may be determined at 306F for each point in the 2D grid, and semantic association may be established among the plurality of masks at 308F.
- semantic association may be established at 308F by using a zeroshot, one-shot, or few-shot neural network, or multi-view geometry.
- One or more points may be selected at 31 OF from a plurality of points in the 2D grid based at least in part upon an inverse projection of the estimated or determined camera poses, depth data, and/or the associated depth covariance (e.g., uncertainty) some or all of which may be determined according to the block diagram illustrated in FIG. 3E.
- points with low uncertainty below a threshold may be selected from the 2D grid at 391 F.
- a sparse 3D point set may be generated at 312F for the selected points.
- the sparse 3D point set may be determined at 312F based at least in part upon an inverse projection of the estimated or determined camera poses, depth data, and/or the associated depth covariance (e.g., uncertainty) some or all of which may be determined according to the block diagram illustrated in FIG. 3E.
- a selected point may be unprojected into 3D by using, for example, the corresponding camera pose, the depth information, and/or the depth covariance which measures how the point is stretched and/or scaled.
- the 3D Gaussian splatted neural radiance field may be improved or optimized at 314F at least by attaching a low-dimensional semantic vector representing a true class label (e.g., a 3D semantic label) to each cluster of Gaussians or to each Gaussian.
- a true class label e.g., a 3D semantic label
- FIG. 3G illustrates an example block diagram for learning semantic labels in 3D, according to some embodiments. More specifically, FIG. 3G illustrates a block diagram of learning semantic labels that are referenced at 314F of FIG. 3F.
- a rendering process may be invoked at 302G, and a number (N) of labels from one or more segmentation masks in 3D may be projected onto a 2D space at 304G to generate a number (N) of projected labels.
- the number (N) of projected labels may be decompressed at 306G into a number (N) of decompressed labels.
- decompressing a projected semantic label into a decompressed semantic label may be performed using a linear layer, a dense layer, or a fully connected (FC) layer where every input neuron or node is connected to every output neuron or node, and the linear layer maps an input to an output with a weight (W or a weight matrix) and a bias (b or a bias vector).
- the number (N) of decompressed labels may be expanded at 308G into the number plus one (N+1 ) labels, and these N+1 decompressed labels may be transformed into a probabilistic distribution of N+1 outcomes at 310G.
- a hierarchical part-whole structure of objects in an input image may be learned at 312G using a decoder based at least in part upon a frequency schedule. These embodiments learn a hierarchical part-whole structure of objects, rather than learning a semantic class per pixel as most of legacy approaches otherwise learn.
- a frequency schedule embeds one or more objects into different latent spaces (also referred to as embedding space or latent feature space) based at least in part upon the one or more respective frequencies where a latent space comprises an embedding of a set of items within a manifold where similar items are positioned closer to each other.
- a plurality of Gaussians may be clustered at 314G into one or more clusters based at least in part upon respective object identifiers. For example, a group of Gaussians corresponding to the same object may be clustered into the same cluster.
- a semantic vector may be embedded at 316G to each of the one or more clusters. In some embodiments, embedding a semantic vector to a cluster may be achieved by running a segmenter (e g., a contrastive language-image pre-training or CLIP model) on a plurality of masked image crops.
- a segmenter e g., a contrastive language-image pre-training or CLIP model
- An object may be selected at 318G at least by querying the embedded semantic vectors obtained from 316G using a user input with a textual model and/or speech-to-text model.
- the object may be provisioned by an extended reality device worn by a user in an extended reality session to facilitate user interactions with the rendered object at 320G, and the 3D representation of the object may be updated in real-time or nearly real-time to reflect the effects of the user interactions upon the object.
- the object is semantically disjoint from other objects or 3D representations generated by the extended reality device to allow such user interactions and manipulations.
- FIG. 4 illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
- an input pertaining to an environment from a sensor may be received at 402.
- the input may include a textual input (e.g., a string of text in natural language), a 2D floorplan of an environment, an image input (e.g., a photograph, a sequence of photographs or images captured by an image sensor, or a video captured by an image sensor, a voice input captured by a microphone, or any combination thereof) in some embodiments.
- the more information a model has about an object’s surface the better the underlying model will be in discovering the object’s shape and the way it interacts with lights.
- Semantics may be determined for the input at 404.
- semantics may be determined at least by parsing the input received at 402 with a large language model (LLM), a visual language model, or a diffusion model that produces hypotheses of complete meshes based at least in part upon some incomplete information pertaining to the 3D representation to be reconstructed.
- LLM large language model
- a diffusion model comprises a deep neural network that holds latent variables capable of learning the structure of a given image by removing its blur (i.e. , noise). After a model’s network is trained to “know” or “understand” the concept abstraction behind an input, the model can create new variations of that image. For example, by removing the noise from an image of a cat, the diffusion model “sees” a clean image of the cat, learns how the cat looks, and applies this knowledge to create new cat image variations.
- At least an incremental portion of the environment may be reconstructed (e.g., via inpainting and/or outpainting) and rendered at 406 at least by using one or more generative models based at least in part upon the semantics determined at 404.
- a 3D model with meshes may be provided for the incremental portion.
- the 3D model may include a diffusion model or a large language model which is a predictive model that receives an input and predicts what comes after the input.
- a diffusion model includes a forward process and a reverse process where the forward process (e.g., an auto-encoder) incrementally injects noise into a clean sample, and the reverse incrementally removes noise.
- a diffusion model When compared to a GAN (generative adversarial network), a diffusion model provides better distribution coverage of the input signals because GANs are usually biased towards what GANs reconstruct well (and thus may have poorer coverage of the input signals) yet may produce lower quality results.
- the issues with lower quality results involving diffusion models are addressed in some embodiments by using, for example, 3D Gaussian splatted neural radiance field techniques or the refinement network described herein while GANs may be reserved for training purposes, if any, in some embodiments.
- the incremental portion may be optionally updated at 408 at least by using a separate model in some embodiments.
- some embodiments may optionally use a generative adversarial network (GAN) for GAN losses as an additional, optional feature by using, for example, the discriminator of GAN to check to see if the generated environment appears correct in terms of styles, appearances, any correspondence, etc.
- GAN generative adversarial network
- a semantically integrated e.g., via embedding a semantic vector to a cluster of Gaussians as described herein
- manipulatable e.g., via user interaction
- photo-realistic 3D representation in the incremental portion may be modified in real-time or nearly real-time at 410 in response to a user interaction (e.g., an interaction by a user via an XR device).
- a SLAM simultaneous localization and mapping
- these embodiments address the shortcomings of SLAMs by further leveraging various neural 3D reconstruction methods to create semantic, high- quality digital twins of a given scene in real-time or nearly real-time. More particularly, SLAM primarily focuses on creating sparse maps for accurate camera localization but not so much in creating photorealistic digital assets.
- various neural 3D reconstruction methods described herein capture and reproduce highly detailed, photo-realistic 3D representations.
- various techniques described herein combine the benefits of SLAM (e.g., producing sparse maps for accurate camera localization) and the efficient, photo-realistic reconstruction of 3D representations of such neural 3D reconstruction methods to facilitate semantically disjoint, individually and independently manipulatable 3D representations of objects for various applications such as extended reality experiences with XR devices.
- FIG. 5A illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments. More specifically, FIG. 5A illustrates a high-level block diagram of a 3D Gaussian splatted neural radiance field technique that may be used in conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems.
- the keyframes e.g., 316E
- camera poses e.g., 318E
- depth information e.g., 320E
- semantically mapped or associated 3D digital scenes may be dynamically (e.g., in real-time or nearly realtime) reconstructed at 502.
- a lower-dimensional grouping of Gaussians may be generated at 504.
- a plurality of Gaussians may be clustered based at least in part upon one or more object identifiers to which the plurality of Gaussians pertains.
- the plurality of Gaussians is clustered into one or more lower-dimensional groups before semantic meanings, labels, or vectors are attached to or otherwise associated with any such Gaussians.
- semantic meanings, labels, or vectors may be attached to an individual Gaussian in some embodiments or to a lower-dimensional grouping of Gaussians in some other embodiments.
- 3D object representations generated by various techniques described herein may be semantically disjoint from each other so that a user wearing an XR device may interact with or manipulate such 3D object representations independently and individually.
- FIGS. 5B-5C illustrate a more detailed block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments.
- semantically mapped or associated 3D digital scenes may be dynamically (e.g., in real-time or nearly real-time) reconstructed at 502B.
- a SLAM system may be initialized at 504B as the backend to generate a sparse 3D point set for localization purposes.
- One or more non-keyframes may 1 be optionally embedded at 506B.
- One or more camera poses and/or depth data may be estimated at 508B for at least some of the one or more non-keyframes. At least these one or more non-keyframes may be used to augment a training data set at 51 OB for the 3D Gaussian splatted neural radiance field model to increase the number of training samples in some embodiments.
- respective camera poses and depth information may also be determined for the one or more non-keyframes and associated with or incorporated into the augmented training data set.
- the keyframes (which may be determined from a video or a sequence of images prior to the processing of the 3D Gaussian splatted neural radiance field model) may be analyzed at 512B with a segmentation model that generates a plurality of object segmentation masks at least by performing a segmentation process on a grid (e.g., a 32 x 32 point grid).
- a plurality of object masks (e.g., three or more) may be generated for each point in the point grid.
- a point in the point grid may be selected based on its uncertainty.
- a part-whole hierarchy of masks may be generated at least by analyzing the keyframes with a semantic segmentation model (which will be described in greater details below) at least by using a zero-shot, one-shot, or few-shot tracker neural network or multi-view geometry at 514B.
- a sparse 3D initial point set may be created at 516B based at least in part upon an inverse projection of the estimated keyframe poses, depth information, and/or depth covariance.
- points corresponding to lower uncertainty may be selected and unprojected into 3D using the camera pose or keyframe pose of the keyframe in which the point is identified, and/or depth information and selected as a part of the sparse 3D initial point set at 516B.
- the 3D Gaussian splatted neural radiance field model may be improved or optimized at 518B with semantic information. The 3D Gaussian splatted neural radiance field model is thus made semantically aware.
- a lower-dimensional semantic vector may be optionally attached to each Gaussian at 520B in some embodiments although an alternative process is to attach a semantic vector to a lower-dimensional grouping of a plurality of Gaussians in some other embodiments as described below with reference to 526B.
- the lower-dimensional semantic vectors may be optionally learned at 522B using a rendering process during training (e.g., a supervised training) with a loss (e.g., cross-entropy loss).
- the lower-dimensional vectors may be learned at least by projecting the semantic labels onto a 2D space with a linear layer (e.g., a fully connected layer connecting each input to each of a plurality of outputs) to decompress the 3D semantic label, further expanding the number of semantic labels to the number plus one semantic labels, and by performing a softmax function on the number plus one semantic labels into a probabilistic distribution of possible outcomes (a total of the number plus one possible outcomes).
- a linear layer e.g., a fully connected layer connecting each input to each of a plurality of outputs
- a plurality of Gaussians may be clustered at 524B into one or more lowerdimensional groupings based at least in part upon one or more object identifiers. For example, Gaussians corresponding to the same object identifier may be clustered into the same lower-dimensional grouping in some embodiments.
- a semantic vector determined may then be attached at 526B to each cluster of Gaussians at least by running a segmentation model (e.g., a contrastive language-image pre-training or CLIP model) on a plurality of masked image crops.
- a segmentation model e.g., a contrastive language-image pre-training or CLIP model
- a user input may be gathered at 528B to query the 3D representation.
- a user input may be a gesture, a movement or motion of a part of a user’s body (e.g., a user’s hand squeezing a virtual object, a user’s foot kicking a virtual object, etc.), a textual input, a voice input, and/or an input of an image although input of different formats may be processed with different encoders to generate respective embeddings.
- a textual input (e.g., by entering textual instructions in natural language) may be processed by a text encoder; an image input may be processed by an image encoder; and a voice input may be processed by a speech encoder or by a combination of a speech-to-text encoder and a textual encoder.
- Similarity e.g., cosine similarity
- this similarity is to correlate the user input with the pertinent portions of the 3D representation to more accurately determine and manipulate the 3D representation in response to the user input.
- One or more objects may be selected at 532B in response to the user input based at least in part upon the similarity determined at 530B.
- the one or more selected objects may be optionally highlighted at 534B with, for example, graphical, textual, or voice emphasis.
- the one or more objects selected in response to the user input may then be independently, individually manipulated at 536B.
- inpainting and/or outpainting may be performed at 538B for the 3D scene in response to manipulating the one or more selected objects based at least in part upon the contextual information and/or the user input. For example, a user may use the user input to move a selected virtual object from first location to a second location where the first location used to be occluded by the virtual object. Once the virtual object is moved out of the first location, the first location may be left with a void. [00136] In some embodiments where the first location belongs to the physical environment, the portion of the physical environment corresponding to the void may become visible by the user.
- the void may be filled with inpainting using techniques at 538B described herein.
- 3D representation of at least the area around the second location may be generated using outpainting techniques at 538B.
- FIG. 6A illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields without pre-training, according to some embodiments. More particularly, FIG. 6A illustrates an example segmentation block diagram that may be used to generate a plurality of masks. In these embodiments, an image 602 may be provided to an image encoder 604 to produce image embeddings 606.
- input data may be converted into a numerical format. This may involve creating a “bag of words” representation for text data, converting images into pixel values, or transforming graph data into a numerical matrix, etc.
- objects that come into an embedding model are output as embeddings, represented as vectors.
- a vector includes an array of numbers, where each number indicates where an object is along a specified dimension. The number of dimensions may reach a thousand or more depending on the input data’s complexity. The closer an embedding is to other embeddings in this n-dimensional space, the more similar they are.
- Distribution similarity may be determined by the length of the vector points from one object to the other (e.g., as measured by Euclidean, cosine or other).
- Embeddings may be used in various domains and applications due to their ability to transform high-dimensional and categorical data into continuous vector representations, capturing meaningful patterns, relationships and semantics. Below are a few reasons why embedding is used in data science. By mapping entities (words, images, nodes in a graph, voices, etc.) to vectors in a continuous space, embeddings may capture the semantic relationships and similarities, enabling models to understand and generalize better. An embedding layer is thus commonly used in neural network architectures to map categorical inputs to continuous vectors, facilitating backpropagation and optimization while meaningful relationships in the original input data are preserved.
- a masked autoencoder may be used for the image encoder 604 where the masked autoencoder masks random patches of the input image 602 and reconstructs the missing pixels.
- an asymmetric encoder-decoder architecture may be utilized, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image 602 from the latent representation and mask tokens.
- masking a higher proportion of the input image (e.g., 75%), may yield a nontrivial and meaningful self-supervisory task.
- a large random subset of image patches (e.g., 75%) may be masked out.
- the masked encoder may be applied to the small subset of visible patches.
- Mask tokens may be introduced after the masked encoder, and the full set of encoded patches and mask tokens may be processed by a small decoder that reconstructs the original image in pixels.
- the decoder may be discarded, and the encoder may be applied to uncorrupted images (e.g., a full sets of patches) for recognition tasks in some embodiments.
- the image embeddings 606 may be combined at 608 with one or more masks 612, once these one or more masks 612 are processed by a convolution network 610, and the combined result may be provided to a mask encoder 614.
- the mask decoder 614 may also receive various other embeddings from a prompt encoder 616.
- the prompt encoder may receive various prompts such as points 618, boxes 620, textual prompt 622, voice prompt 624, and/or image prompt 626, etc. and generate the respective embeddings therefor.
- Various prompts may be categorized into two sets - a sparse set including points, boxes, and/or texts, and a dense set including voice and/or image prompts in some embodiments.
- the sparse set may be represented by positional encodings summed with learned embeddings for each prompt type (e.g., point, box, text) while the dense set of prompts may be embedded using convolutions and summed element-wise with the image embedding 606.
- prompt type e.g., point, box, text
- dense set of prompts may be embedded using convolutions and summed element-wise with the image embedding 606.
- a textual prompt (e.g., by entering textual instructions in natural language) may be processed by a text encoder; an image prompt may be processed by an image encoder; and a voice prompt may be processed by a speech encoder or by a combination of a speech-to-text encoder and a textual encoder.
- the mask decoder 614 maps the image embedding 606, prompt embeddings produced by the prompt encoder(s) 616, and an output token to first compute confidence scores 628 respectively associated with corresponding masks. The masks corresponding to sufficiently high confidence scores may be determined to be valid masks
- FIG. 6B illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields with pre-training, according to some embodiments.
- an input image 602B may be identified.
- the input image includes a plurality of masked image patches or masked image tokens 602_1 B and a plurality of visible image patch or image tokens 602_2B.
- the plurality of visible image patch or image tokens 602_2B may be provided to an image encoder 604 whose output is then combined with the plurality of masked image patch or image tokens 602_1 B to form the image imbedding 606 with the respective positional information corresponding the plurality of image patches or tokens (both masked and visible image patches or tokens).
- the image imbedding 606 is provided to a decoder 650 that reconstructs, for pre-training purposes, the image embeddings without masks 652B that may then be used to reconstruct the original image from the reconstructed image embeddings 652B, also for training purposes.
- the reconstructed image 654B may be provided to a loss mechanism or a discriminator that determines the loss therebetween and uses the loss to fine tune various components (e.g., the encoder 604 and/or the decoder 650).
- the image embeddings 606 may be combined with one or more masks 612B after these one or more masks have been processed by a convolution network 610B.
- the combined result may be provided to a mask decoder 614B.
- the mask decoder 614B further receives prompt embeddings from the prompt encoder 616B to generate confidence a score for each of a plurality of masks 612B.
- the masks with sufficiently high confidence scores (e.g., above a threshold confidence score) may be determined to be value masks 630B.
- various prompts may be categorized into two sets - a sparse set including points 618B, boxes 620B, and/or texts 622B, and a dense set including voice 624B and/or image prompts 626B in some embodiments.
- the sparse set may be represented by positional encodings summed with learned embeddings for each prompt type (e.g., point, box, text) while the dense set of prompts may be embedded using convolutions and summed element-wise with the image embedding 606.
- prompt type e.g., point, box, text
- dense set of prompts may be embedded using convolutions and summed element-wise with the image embedding 606.
- FIG. 7A illustrates s simplified schematic diagram for a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments. More specifically, FIG. 7A illustrates a model 700 that jointly trains an image encoder, a text encoder, and a voice / speech encoder to predict the correct pairings of a batch of (image, text), (image, voice), or (image, text, voice) training examples.
- a learned text encoder synthesizes a zero-shot classifier by embedding the names or descriptions of a target dataset’s classes
- a learned voice / speech encoder synthesizes a zero-shot classifier by embedding the voice or speech of a target dataset’s classes.
- an image encoder 704 may process one or more images 702 into one or more corresponding sets of image embeddings 706 each having a plurality of values in a continuous space.
- a text encoder 710 may process one or more text inputs 708 into one or more corresponding sets of text embeddings 720 each having a plurality of values in a continuous space.
- speech encoder 716 may process one or more speech inputs 714 into one or more corresponding sets of speech embeddings 718 each having a plurality of values in a continuous space.
- the image embeddings 706, text embeddings 720, and speech encodings 718 may be used to form a multi-dimensional data structure 712 although a two-dimensional table-like data structure is shown in FIG. 7A for the ease of illustration.
- a speech encoder may be replaced by a speech-to- text model that transforms a speech input 71 into text input which may then be processed by the text encoder 710. These embodiments thus set aside the need for a separate speech encoder although at the additional expense of having a speech-to-text model.
- the various text inputs 708 and speech inputs 714 comprise natural language inputs to leverage the advantages of learning perception from supervision in natural language such as easier to scale than crowd-sourced labeling (e.g.., ImageNet) because natural language does not require annotations to be in a machine learning compatible format (e.g., a canonical one-of-N majority vote “gold label”). Further, learning from natural language does not simply learn a representation but also connects a specific representation to natural language that enables the more flexible zero-shot knowledge transfer.
- natural language such as easier to scale than crowd-sourced labeling (e.g.., ImageNet) because natural language does not require annotations to be in a machine learning compatible format (e.g., a canonical one-of-N majority vote “gold label”).
- learning from natural language does not simply learn a representation but also connects a specific representation to natural language that enables the more flexible zero-shot knowledge transfer.
- the network 700 pre-trains an image encoder (704), a text encoder (708), and a speech encoder (714) to predict which images were paired with which texts and/or speech in the dataset.
- the model uses this behavior to turn into a zero-shot classifier and converts all of a dataset’s classes into captions and/or speeches such as “a photo of a dog” and predict the class that is estimated to best pair with a given image.
- FIG. 7B illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments.
- the circuit 721 shows an image encoder 704 receives an image 702 and determines an image embedding 730 for the received image 702, and the image embedding includes a vector form having a plurality of values in a continuous space.
- the circuit 721 shows a text encoder 710 tasked with a task 724 for pairing an image with its description and a plurality of texts 722_1 , 722_2, 722_3, 722_4, ... , 722_n.
- the text encoder 710 determines the textual embedding 726 for each of the texts (722_1 , 722_2, 722_3, 722_4, ... , 722_n), each having a vector form having a plurality of values in a continuous space.
- the circuit 738 shows a speech encoder 716 tasked with a task 725 for pairing an image with its corresponding speech and a plurality of speech segments 736_1 , 736_2, 736_3, 736_4, ... , 736_n.
- the text encoder 710 determines the speech embedding 740 for each of the speech segments (736_1 , 736_2, 736_3, 736_4, ... , 736_n), each having a vector form having a plurality of values in a continuous space.
- the image encoder 704, the text encoder 710, and the speech encoder 716 may be jointly trained to determine the correct pairings of (image, text), (image, text, speech), and/or (image, text, speech). These three types of embeddings may be arranged in one or more multi-dimensional data structure. In practice, when a prompt is provided for one of the three inputs, the remaining two may be identified by querying the one or more multi-dimensional data structure to find the correct pairing(s) of (image, text), (image, text, speech), and/or (image, text, speech).
- a prompt of an image may be received at the image encoder 704, the image encoder determines the image embedding 732 for the received image 702 from a plurality of image embeddings 732.
- the text embedding and the speech embedding may then be used to query the one or more multi-dimensional data structure to identify the best match of text encoding and speech encoding for the prompted image 702.
- the text embedding and the speech embedding may be respectively processed to produce the textual description and the verbal description that may then be associated with or attached to the received image.
- FIG. 70 illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments. More specifically, FIG. 7C illustrates an example block diagram for an image encoder 700C that may be used in a zero-shot (or one-shot or few-shot) image classification model that learns by associating text and/or speech with images.
- an image encoder 700C may include an input stem 702C and four subsequent stages 704C (stage 1 ), 706C (stage 2), 708C (stage 3), and 7100 (stage 4) as well as the final output layer
- the input stem 702C reduces the input width and height and increases its channel size by using a network comprising an N x N convolution
- the input stem 702C may include 7 x 7 convolutions with an output channel of 64 with a stride of two (“2”) followed by a 3 x 3 max pooling layer with a stride of 2 in some embodiments. In these embodiments, the input stem 702C increases the input channel size to 64 and reduces the input width and height by four times.
- the second stage 706C and on begins with a down sampling layer 716C that is followed by several residual blocks (e.g., 718C and 720C).
- the down sampling layer 716C may have two paths where the first path includes three convolution layers 722C, 724C, and 726C, and the second path includes a single convolution layer 728C.
- the convolution layer 722C may include a kernel size of 1 x 1 with a stride of 2 (to half the input width and height) and a channel size of 512; the convolution layer 724C may include a kernel size of 3 x 3 with a channel size of 512.
- the convolution layer 726C may include a kernel size of 1 x 1 with a channel size of 2048 (which is four times larger than those of 724C and 722C); and the convolution layer 728C may include a kernel size of 1 x 1 with a stride of 2 and a channel size of 2048.
- the second path is to transform the input shape to be the output shape of the first path to that the outputs of the first path and the second path may be summed together to determine the output of the down sampling block 716C.
- a residual block is similar in architecture to the down sampling block 716C with the exception that the residual block has a stride of one (“1”).
- FIG. 7D illustrates some example modifications of the simplified schematic diagram illustrated in FIG.
- a down sampling layer 716C may include the architecture 716D that is identical to 716C described above with reference to FIG. 7C.
- a down sampling layer 716C may include the architecture 716D1 that also has a first path and a second path.
- the first path include a first convolution layer 730D having a kernel size of 1 x 1 , a second convolution layer 732D having a kernel size of 3 x 3 with a stride of 2, and a third convolution layer 734D having a kernel size of 1 x 1 .
- the second path includes a single convolution layer 736D having a kernel size of 1 x 1 with a stride of 2.
- the first convolution in the first path starting with 722C ignores three-quarters of the input feature map because it uses a kernel size 1 *1 with a stride of 2.
- 716D1 switches the strides size of the first two convolutions (730D and 732D) in the first path so no information is ignored.
- the second convolution 732D has a kernel size 3 x 3 with a stride of 2, the output shape of the first path remains unchanged.
- a down sampling layer 716C may include or be replaced with the architecture 716D2 that also has a first path and a second path.
- the first path include the same first convolution layer 730D having a kernel size of 1 x 1 , the same second convolution layer 732D having a kernel size of 3 x 3 with a stride of 2, and the same third convolution layer 734D having a kernel size of 1 x 1.
- the second path includes a pooling layer 738D having a kernel size of 2 x 2 with a stride of 2 followed by a convolution layer 740D having a kernel size of 1 x 1 with a stride of 2.
- adding a 2 x 2 pooling layer e.g., an average pooling layer
- a down sampling layer 716C may include or be replaced with the architecture 716D3 that has a single path including a first convolution layer 742D having a kernel size of 3 x 3 with a channel size of 32 and a stride of 2, a second convolution layer 744D having a kernel size of 3 x 3 a channel size of 32 with a stride of 2, and a third convolution layer 746D having a kernel size of 3 x 3 with a channel size of 64 and a stride of 2 followed by a pooling layer 748D having a kernel size of 3 x 3 with a stride of 2.
- the computational cost of a convolution is quadratic to the kernel width or height.
- a 7 x 7 convolution is 5.4 times more expensive than a 3 x 3 convolution. So this down sampling layer 716D3 replacing the 7 x 7 convolution in the input stem with three conservative 3 x 3 convolutions with the first and second convolutions have their output channel of 32 and a stride of 2, while the last convolution uses a 64-output channel.
- FIG. 7E illustrates a simplified block diagram for an example image encoder that may be employed in FIG. 7A or 7B, according to some embodiments. More specifically, FIG. 7E illustrates another example block diagram for an image encoder that may be used in a zero-shot (or one-shot or few-shot) image classification model that learns by associating text and/or speech with images.
- an image encoder may split an image into a plurality of patches (e.g., a plurality of fixed and/or variable size patches), linearly combine the plurality of patches, add position embeddings, and feed the resulting sequence of vectors in the resulting embeddings to a transformer encoder.
- an image 702E may be split into a plurality of patches 704E.
- an image may be split into a plurality of fixed-size patches while in some other embodiments, an image may be split into a plurality of variable-sized patches.
- the plurality of patches 704E may be provided to a projection module 706E which projects the plurality of patches into projected patches 708E at least by combining or otherwise associating the respective positions (in the original image 702E) with the plurality of patches 704E where the numbers in the projected patches 708E respectively indicate the positions of the projected patches in the original image 702E.
- an extra classification token 710E may be added to the projected patches 708E to enable the performance of classification.
- the extra classification token 710E and the projected patches 708E may be provided to a transformer encoder 712E.
- the extra classification token 710E and the projected patches 708E may be provided to a normalization layer 714E in the transformer encoder 712E that performs batch normalization or group normalization on the extra classification token 710E and the projected patches 708E.
- the normalized input may be forwarded to an attention layer 716E.
- the attention layer 716E includes a multi-head attention layer such as a multi-head self-attention layer.
- a multi-head attention layer includes a module of multiple attention mechanisms which runs through an attention mechanism several times in parallel. The independent attention outputs may then be concatenated and linearly transformed into the expected dimension. In some of these embodiments, multiple attention heads allows for attending to parts of the sequence differently (e.g. longer-term dependencies versus shorter-term dependencies).
- the output of the attention layer 716A may be combined with the extra classification token 710E and the projected patches 708E, which are also separately provided to the combiner.
- the output of the combiner may be forwarded to both a normalization layer 718E and another combiner which follows a separate multi-layer perceptron (MLP) 720E receiving the output of the normalization layer 718E.
- MLP multi-layer perceptron
- the output the MLP may be combined with the output of the previous combiner following the attention layer 716E to generate a combined output that is in turn forwarded to a classification 722E that produces a plurality of classes 724E.
- a multilayer perceptron comprises a modern feed-forward artificial neural network, including, for example, fully connected neurons with a nonlinear kind of activation function, organized in at least three layers, notable for being able to distinguish data that is not linearly separable.
- FIG. 7F illustrates a simplified block diagram for an example text encoder that may be employed in FIG. 7A or 7B, according to some embodiments.
- the text encoder 720 includes an encoding portion that receives an input 702F (e.g., textual input) at an input embedding module 704F that generates text embeddings for the input 702F.
- the text embeddings may be combined or associated with the position embeddings 706F at a combiner that forwards its combined output to an attention mechanism 708F as well as a normalization layer 71 OF that follows the attention mechanism 708F.
- the normalization layer 71 OF performs a batch normalization or group normalization and transmits its output to a position-wise feed forward layer and a normalization layer 714F that also receives the output of the position-wise feed forward layer 712F.
- the output of the normalization layer 714F is transmitted to an attention mechanism 726F in the decoding portion of the text encoder 720 that is described immediately below.
- the above encoding portion of the text encoder 720 may include a stack of N identical layers (e.g., six identical layers in some embodiments) where each layer has two sub-layers.
- the first sublayer includes a multi-head self-attention mechanism
- the second sublayer includes a simple, position-wise fully connected feed-forward network.
- Soe embodiments employ a residual connection around each of the two sublayers, followed by layer normalization.
- the output of each sublayer is LayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer itself.
- the text encoder 720 may also include the decoding portion that receives output 716F at an output embedding module 718F that generates embeddings for the output 716F. These embeddings, together with the positional embeddings generated by a positional encoding module 720F for the output 716F, may be combined at a combiner whose output is then provided to a masked attention mechanism 722F and a normalization layer 724F following and receiving the output of the masked attention mechanism 722F.
- the normalization layer 724F performs a batch normalization or group normalization to generate normalized output that may then be sent to the attention mechanism 726F as well as the normalization layer 728F following and receiving the output of the attention mechanism 726F in the decoding portion of the text encoder 720.
- the normalization layer 728F performs a batch normalization or group normalization to generate normalized output that may be provided to a position-wise feed forward layer 730F and a normalization layer 730F that follows and receives the output of the position-wise feed forward layer 730F.
- the normalization layer 732F performs a batch normalization or group normalization on the input to generate normalized output that may be provided to a linear layer 734F that is followed by a softmax layer 736F.
- the softmax layer 736F performs a softmax operation to transform values in a continuous space into probabilistic distributions of multiple possible outcomes 738F.
- the decoding portion may also include a stack of N identical layers (e.g., six identical layers).
- the decoding portion inserts a third sub-layer, which performs multi-head attention over the output of the encoder stack. Similar to the encoding portion, residual connections may be employed around each of the sub-layers, followed by layer normalization.
- the self-attention sub-layer in the decoder stack may be modified to prevent positions from attending to subsequent positions in some of these embodiments. This masking, combined with the fact that the output embeddings may be offset by one position, ensures that the predictions for position i may depend only on the known outputs at positions less than i.
- an attention mechanism performs an attention function which may be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are vectors.
- the output is computed as a weighted sum of the values, where the weight assigned to each value may be computed by a compatibility function of the query with the corresponding key.
- FIGS. 7G-7H illustrate an example attention mechanism that may be employed for a sequence transduction module in a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments.
- the example attention mechanism 700G may receive a query matrix 702G and a key matrix 704G at an inner product module 706G that performs an inner product on the query matrix 702G and the key matrix 704G.
- the query matrix 702G in encoding or decoding attention layers, may come from the previous decoder layer.
- the keys 704G and values 716G may come from the output of the encoding portion. This allows every position in the decoder to attend over all positions in the input sequence to mimic the typical encoder-decoder attention mechanisms in sequence-to-sequence models.
- an encoder includes one or more self-attention layers.
- a self-attention layer all of the keys, values, and queries may come from the same place (e.g., the output of the previous layer in the encoding portion). Each position in the encoder may attend to all positions in the previous layer of the encoding portion.
- a Self-attention layer in the decoding portion allows each position in the decoder to attend to all positions in the decoder up to and including that position.
- Leftward information flow in the decoder may need to be prevented to preserve the auto-regressive property, and this may be implemented inside of a scaled dot-product attention by masking out (e.g. , by setting to -°°) all values in the input of the softmax which correspond to illegal connections.
- the output of the inner product module 706G may be provided to an optional scaling module 708G that scales the output of the inner product module 706G (e.g., scale the output by the square root of the dimensionality of the queries or the keys, d_k) and/or an optional masking module 710G that optionally masks out values in the input to the normalization module 712G that correspond to illegal connections, if any.
- the normalization module 712G normalizes the output of the inner product module 706G (or that of the optional scaling module and/or that of the optional masking module 710G) into a probabilistic distribution of several possible outcomes in order to determine the weights on the values.
- the normalization module 712G may perform a softmax function that converts a vector of N real numbers into a probability distribution of N possible outcomes.
- the output of the normalization module 712G may be provided to another inner product module 714G that performs an inner product operation on the output of the normalization module 712G and the value matrix 716G.
- FIG. 7H illustrates a text encoder with multi-head attention, according to some embodiments.
- multiple example attention mechanisms 700G may be run in parallel where each example attention mechanism 700G receives projected query matrices 702G, projected key matrices 704G, and projected value matrices 716G.
- a query matrix 702G may be linearly projected to, for example, d_k dimensions 706H (e.g., 64 or 128);
- a key matrix 708G may be linearly projected to d_k dimensions 708H (e.g., 64 or 128), and a value matrix 716G may be linearly projected to d_v dimensions 71 OH.
- These projected matrices may be provided to respective example attention mechanisms 700G that perform the attention function in parallel to produce d_v dimensional output values.
- These d_v dimensional output values may be forwarded and concatenated by the concatenation module 712H which may then be projected again at 714H, resulting in the final values.
- FIG. 8 illustrates a simplified example of a wearable extended-reality (XR) device with a belt pack external to the XR glasses in some embodiments. More specifically, FIG. 8 illustrates a simplified example of a user-wearable VR (virtual reality) / AR (augmented reality) I MR (mixed reality) /XR (extended reality) system that includes an optical sub-system 802 and an external module 804 and may include multiple instances of personal augmented reality systems, for example a respective personal augmented reality system for a user. Any of the neural networks, module, processes, services, microservices, etc. described herein may be embedded in whole or in part in or on the wearable XR device.
- VR virtual reality
- AR augmented reality
- I MR mixed reality
- XR extended reality
- a neural network described herein as well as other peripherals may be embedded in the external module 804 alone, the optical sub-system 802 alone, or distributed between the external module 804 and the optical sub-system 802.
- the external module 804 may include a rechargeable battery that is operatively connected to the optical sub-system 802 in order to provide power to the optical sub-system 802.
- Some embodiments of the VR/AR/MR/XR system may comprise optical sub-system 802 that delivers virtual content to the user’s eyes as well as the external module 804 that performs a multitude of processing tasks to present the relevant virtual content to a user.
- the external module 804 may, for example, take the form of the belt pack, which can be convenience coupled to a belt or belt line of pants during use. Alternatively, the external module 804 may, for example, take the form of a personal digital assistant or smartphone type device.
- the external module 804 may include one or more processors, for example, one or more micro-controllers, microprocessors, graphical processing units, digital signal processors, application specific integrated circuits (ASICs), programmable gate arrays, programmable logic circuits, or other circuits either embodying logic or capable of executing logic embodied in instructions encoded in software or firmware.
- the external module 804 may include one or more non-transitory computer- or processor-readable media, for example volatile and/or nonvolatile memory, for instance read only memory (ROM), random access memory (RAM), static RAM, dynamic RAM, Flash memory, EEPROM, etc.
- the external module 804 may be communicatively coupled to the head worn component.
- the external module 804 may be communicatively tethered to the head worn component via one or more wires or optical fibers via a cable with appropriate connectors.
- the external module 804 and the optical sub-system 802 may communicate according to any of a variety of tethered protocols, for example UBS®, USB2®, USB3®, USB-C®, Ethernet®, Thunderbolt®, Lightning® protocols.
- the external module 804 may be wirelessly communicatively coupled to the head worn component.
- the external module 804 may be wirelessly communicatively coupled to the head worn component.
- the external module 804 may be wirelessly communicatively coupled to the head worn component.
- the external module 804 may be wirelessly communicatively coupled to the head worn component.
- the external module 804 may be wirelessly communicatively coupled to the head worn component.
- the external module 804 may be wirelessly communicatively coupled to the head worn component
- the optical sub-system 804 and the optical sub-system 802 may each include a transmitter, receiver or transceiver (collectively radio) and associated antenna to establish wireless communications there between.
- the radio and antenna(s) may take a variety of forms.
- the radio may be capable of short-range communications, and may employ a communications protocol such as BLUETOOTH®, WI-FI®, or some IEEE 802.11 compliant protocol (e.g., IEEE 802.11 n, IEEE 802.11 a/c).
- a communications protocol such as BLUETOOTH®, WI-FI®, or some IEEE 802.11 compliant protocol (e.g., IEEE 802.11 n, IEEE 802.11 a/c).
- any of the subcomponents described above for the external module 804 may be alternatively incorporated into the optical sub-system 802 while the external module 804 includes a battery to provide power to the optical subsystem 802 as well as the modules therein.
- FIG. 9 illustrates a computerized system on which a method for management of extended-reality systems or devices may be implemented.
- Computer system 900 includes a bus 906 or other communication module for communicating information, which interconnects subsystems and devices, such as processor 907, system memory 908 (e.g., RAM), static storage device 909 (e.g., ROM), disk drive 910 (e.g., magnetic or optical), communication interface 914 (e.g., modem or Ethernet card), display 911 (e.g., CRT or LCD), input device 912 (e.g., keyboard), and cursor control (not shown).
- processor 907 e.g., system memory 908 (e.g., RAM), static storage device 909 (e.g., ROM), disk drive 910 (e.g., magnetic or optical), communication interface 914 (e.g., modem or Ethernet card), display 911 (e.g., CRT or LCD), input device 912 (e.g., keyboard), and cursor control (not shown).
- the illustrative computing system 900 may include an Internet-based computing platform providing a shared pool of configurable computer processing resources (e.g., computer networks, servers, storage, applications, services, etc.) and data to other computers and devices in a ubiquitous, on-demand basis via the Internet.
- the computing system 900 may include or may be a part of a cloud computing platform in some embodiments.
- computer system 900 performs specific operations by one or more processor or processor cores 907 executing one or more sequences of one or more instructions contained in system memory 908. Such instructions may be read into system memory 908 from another computer readable/usable storage medium, such as static storage device 909 or disk drive 910.
- hard-wired circuitry may be used in place of or in combination with software instructions to implement the invention.
- embodiments of the invention are not limited to any specific combination of hardware circuitry and/or software.
- the term “logic” shall mean any combination of software or hardware that is used to implement all or part of the invention.
- Various actions or processes as described in the preceding paragraphs may be performed by using one or more processors, one or more processor cores, or combination thereof 907, where the one or more processors, one or more processor cores, or combination thereof executes one or more threads.
- various acts of determination, identification, synchronization, calculation of graphical coordinates, rendering, transforming, translating, rotating, generating software objects, placement, assignments, association, etc. may be performed by one or more processors, one or more processor cores, or combination thereof.
- Non-volatile media includes, for example, optical or magnetic disks, such as disk drive 910.
- Volatile media includes dynamic memory, such as system memory 908.
- Computer readable storage media includes, for example, electromechanical disk drives (such as a floppy disk, a flexible disk, or a hard disk), a flash-based, RAM-based (such as SRAM, DRAM, SDRAM, DDR, MRAM, etc.), or any other solid-state drives (SSD), magnetic tape, any other magnetic or magneto-optical medium, CD-ROM, any other optical medium, any other physical medium with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read.
- electromechanical disk drives such as a floppy disk, a flexible disk, or a hard disk
- RAM-based such as SRAM, DRAM, SDRAM, DDR, MRAM, etc.
- SSD solid-state drives
- magnetic tape any other magnetic or magneto-optical medium
- CD-ROM any other optical medium
- any other physical medium with patterns of holes RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other
- execution of the sequences of instructions to practice the invention is performed by a single computer system 900.
- two or more computer systems 900 coupled by communication link 915 may perform the sequence of instructions required to practice the invention in coordination with one another.
- Computer system 900 may transmit and receive messages, data, and instructions, including program (e.g., application code) through communication link 915 and communication interface 914. Received program code may be executed by processor 907 as it is received, and/or stored in disk drive 910, or other non-volatile storage for later execution.
- the computer system 900 operates in conjunction with a data storage system 931 , e.g., a data storage system 931 that includes a database 932 that is readily accessible by the computer system 900.
- the computer system 900 communicates with the data storage system 931 through a data interface 933.
- a data interface 933 which is coupled to the bus 906 (e.g., memory bus, system bus, data bus, etc.), transmits and receives electrical, electromagnetic or optical signals that include data streams representing various types of signal information, e.g., instructions, messages and data.
- the functions of the data interface 933 may be performed by the communication interface 914.
- FIG. 9 shows an example architecture 2500 for the electronics operatively coupled to an optics system or XR device in one or more embodiments.
- the optics system or XR device itself or an external device (e.g., a belt pack) coupled to the or XR device may include one or more printed circuit board components, for instance left (2502) and right (2504) printed circuit board assemblies (PCBA).
- the left PCBA 2502 includes most of the active electronics, while the right PCBA 604supports principally supports the display or projector elements.
- the right PCBA 2504 may include a number of projector driver structures which provide image information and control signals to image generation components.
- the right PCBA 2504 may carry a first or left projector driver structure 2506 and a second or right projector driver structure 2508.
- the first or left projector driver structure 2506 joins a first or left projector fiber 2510 and a set of signal lines (e.g., piezo driver wires).
- the second or right projector driver structure 2508 joins a second or right projector fiber 2512 and a set of signal lines (e.g., piezo driver wires).
- the first or left projector driver structure 2506 is communicatively coupled to a first or left image projector
- the second or right projector drive structure 2508 is communicatively coupled to the second or right image projector.
- the image projectors render virtual content to the left and right eyes (e.g. , retina) of the user via respective optical components, for instance waveguides and/or compensation lenses to alter the light associated with the virtual images.
- respective optical components for instance waveguides and/or compensation lenses to alter the light associated with the virtual images.
- the image projectors may, for example, include left and right projector assemblies.
- the projector assemblies may use a variety of different image forming or production technologies, for example, fiber scan projectors, liquid crystal displays (LCD), LCOS (Liquid Crystal On Silicon) displays, digital light processing (DLP) displays.
- a fiber scan projector images may be delivered along an optical fiber, to be projected therefrom via a tip of the optical fiber.
- the tip may be oriented to feed into the waveguide.
- the tip of the optical fiber may project images, which may be supported to flex or oscillate.
- a number of piezoelectric actuators may control an oscillation (e.g., frequency, amplitude) of the tip.
- the projector driver structures provide images to respective optical fiber and control signals to control the piezoelectric actuators, to project images to the user’s eyes.
- a button board connector 2514 may provide communicative and physical coupling to a button board 2516 which carries various user accessible buttons, keys, switches or other input devices.
- the right PCBA 2504 may include a right earphone or speaker connector 2518, to communicatively couple audio signals to a right earphone 2520 or speaker of the head worn component.
- the right PCBA 2504 may also include a right microphone connector 2522 to communicatively couple audio signals from a microphone of the head worn component.
- the right PCBA 2504 may further include a right occlusion driver connector 2524 to communicatively couple occlusion information to a right occlusion display 2526 of the head worn component.
- the right PCBA 2504 may also include a board-to-board connector to provide communications with the left PCBA 2502 via a board-to-board connector 2534 thereof.
- the right PCBA 2504 may be communicatively coupled to one or more right outward facing or world view cameras 2528 which are body or head worn, and optionally a right cameras visual indicator (e.g., LED) which illuminates to indicate to others when images are being captured.
- the right PCBA 2504 may be communicatively coupled to one or more right eye cameras 2532, carried by the head worn component, positioned and orientated to capture images of the right eye to allow tracking, detection, or monitoring of orientation and/or movement of the right eye.
- the right PCBA 2504 may optionally be communicatively coupled to one or more right eye illuminating sources 2530 (e g., LEDs), which as explained herein, illuminates the right eye with a pattern (e.g., temporal, spatial) of illumination to facilitate tracking, detection or monitoring of orientation and/or movement of the right eye.
- illuminating sources 2530 e g., LEDs
- the left PCBA 2502 may include a control subsystem, which may include one or more controllers (e.g., microcontroller, microprocessor, digital signal processor, graphical processing unit, central processing unit, application specific integrated circuit (ASIC), field programmable gate array (FPGA) 2540, and/or programmable logic unit (PLU)).
- the control system may include one or more non-transitory computer- or processor readable medium that stores executable logic or instructions and/or data or information.
- the non-transitory computer- or processor readable medium may take a variety of forms, for example volatile and nonvolatile forms, for instance read only memory (ROM), random access memory (RAM, DRAM, SD-RAM), flash memory, etc.
- the non- transitory computer or processor readable medium may be formed as one or more registers, for example of a microprocessor, FPGA or ASIC.
- the left PCBA 2502 may include a left earphone or speaker connector 2536, to communicatively couple audio signals to a left earphone or speaker 2538 of the head worn component.
- the left PCBA 2502 may include an audio signal amplifier (e.g., stereo amplifier) 2542, which is communicative coupled to the drive earphones or speakers.
- the left PCBA 2502 may also include a left microphone connector 2544 to communicatively couple audio signals from a microphone of the head worn component.
- the left PCBA 2502 may further include a left occlusion driver connector 2546 to communicatively couple occlusion information to a left occlusion display 2548 of the head worn component.
- the left PCBA 2502 may also include one or more sensors or transducers which detect, measure, capture or otherwise sense information about an ambient environment and/or about the user.
- an acceleration transducer 2550 e.g., three axis accelerometer
- a gyroscopic sensor 2552 may detect orientation and/or magnetic or compass heading or orientation.
- Other sensors or transducers may be similarly employed.
- the left PCBA 2502 may be communicatively coupled to one or more left outward facing or world view cameras 2554 which are body or head worn, and optionally a left cameras visual indicator (e.g., LED) 2556 which illuminates to indicate to others when images are being captured.
- a left cameras visual indicator e.g., LED
- the left PCBA may be communicatively coupled to one or more left eye cameras 2558, carried by the head worn component, positioned and orientated to capture images of the left eye to allow tracking, detection, or monitoring of orientation and/or movement of the left eye.
- the left PCBA 2502 may optionally be communicatively coupled to one or more left eye illuminating sources (e.g., LEDs) 2556, which as explained herein, illuminates the left eye with a pattern (e.g., temporal, spatial) of illumination to facilitate tracking, detection or monitoring of orientation and/or movement of the left eye.
- a pattern e.g., temporal, spatial
- the PCBAs 2502 and 2504 are communicatively coupled with the distinct computation component (e.g., belt pack) via one or more ports, connectors and/or paths.
- the left PCBA 2502 may include one or more communications ports or connectors to provide communications (e.g., bi-directional communications) with the belt pack.
- the one or more communications ports or connectors may also provide power from the belt pack to the left PCBA 2502.
- the left PCBA 2502 may include power conditioning circuitry 2580 (e.g., DC/DC power converter, input filter), electrically coupled to the communications port or connector and operable to condition (e.g., step up voltage, step down voltage, smooth current, reduce transients).
- the communications port or connector may, for example, take the form of a data and power connector or transceiver 2582 (e.g., Thunderbolt® port, USB® port).
- the right PCBA 2504 may include a port or connector to receive power from the belt pack.
- the image generation elements may receive power from a portable power source (e.g., chemical battery cells, primary or secondary battery cells, ultra-capacitor cells, fuel cells), which may, for example be located in the belt pack.
- the left PCBA 2502 includes most of the active electronics, while the right PCBA 2504 supports principally supports the display or projectors, and the associated piezo drive signals. Electrical and/or fiber optic connections are employed across a front, rear or top of the body or head worn component of the optics system or XR device. Both PCBAs 2502 and 2504 are communicatively (e.g., electrically, optically) coupled to the belt pack.
- the left PCBA 2502 includes the power subsystem and a highspeed communications subsystem.
- the right PCBA 2504 handles the fiber display piezo drive signals. In the illustrated embodiment, only the right PCBA 2504 needs to be optically connected to the belt pack. In other embodiments, both the right PCBA and the left PCBA may be connected to the belt pack.
- the electronics of the body or head worn component may employ other architectures.
- some implementations may use a fewer or greater number of PCBAs.
- various components or subsystems may be arranged differently than illustrated in FIG. 9.
- some of the components illustrated in FIG. 9 as residing on one PCBA may be located on the other PCBA, without loss of generality.
- An optics system or an XR device described herein may present virtual contents to a user so that the virtual contents may perceived as three-dimensional contents in some embodiments. In some other embodiments, an optics system or XR device may present virtual contents in a four- or five-dimensional lightfield (or light field) to a user.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Software Systems (AREA)
- Multimedia (AREA)
- Evolutionary Computation (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computing Systems (AREA)
- Artificial Intelligence (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Computer Graphics (AREA)
- Computer Hardware Design (AREA)
- Mathematical Physics (AREA)
- Data Mining & Analysis (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Computational Linguistics (AREA)
- Probability & Statistics with Applications (AREA)
- Algebra (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Biomedical Technology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Geometry (AREA)
- Processing Or Creating Images (AREA)
Abstract
An extended reality system that includes a wearable eyepiece presenting virtual contents to a user, a belt pack operatively coupled to the wearable eyepiece, a processor, and a non-transitory computer readable medium storing thereupon a sequence of instructions which, when executed by a model with the processor, causes the processor to estimate a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device, to generate a semantically annotated, manipulatable three-dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece, and to modify the semantically annotated, manipulatable 3D representation in real-time or nearly real-time in response to a user interaction.
Description
CONDITIONAL, PROBABILISTIC GENERATION OF 3D VIRTUAL ENVIRONMENT BASED ON INCOMPLETE INFORMATION AVAILABLE TO AN EXTENDED REALITY DEVICE
INCORPORATION BY REFERENCE
[001] This application claims priority to U.S. Provisional Application Ser. No. 63/666,058, filed on June 28, 2024, the contents of which are hereby expressly and fully incorporated by reference in its entirety.
COPYRIGHT NOTICE
[002] A portion of the disclosure of this patent document includes material, which is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.
BACKGROUND
[003] Modern computing and display technologies have facilitated the development of systems for so-called “virtual-reality” (VR), “augmented reality” (AR) experiences, “mixed-reality” (MR) experiences, and/or extended-reality (XR) experiences (hereinafter collectively referred to as “extended-reality” and/or “XR”), where digitally reproduced images or portions thereof are presented to a user in a manner where they seem to be, or may be perceived as, real. A VR scenario typically involves presentation of digital or virtual image information without transparency to other actual real-world visual input, whereas an AR or MR scenario typically involves presentation of digital or virtual
image information as an augmentation to visualization of the real world around the user such that the digital or virtual image (e.g., virtual content) may appear to be a part of the real world. However, MR may integrate the virtual content in a contextually meaningful way, whereas AR may not.
[004] Applications of extended-reality technologies have been expanding from, for example, gaming, military training, simulation-based training, etc. to productivity and content creation and management. An extended-reality system has the capabilities to create virtual objects that appear to be, or are perceived as, real. Such capabilities, when applied to the Internet technologies, may further expand and enhance the capability of the Internet as well as the user experiences so that using the web resources is no longer limited by the planar, two-dimensional representation of web pages.
[005] Neural radiance fields (NeRFs) are techniques that generate 3D representations of an object or scene from sparse two-dimensional (2D) images by using machine learning. These techniques involve encoding an object or scene into an artificial neural network, which predicts the light intensity -- or radiance -- at any point in the 2D image to generate novel 3D views from different angles. The NeRF model enables learning of novel view synthesis, scene geometry, and the reflectance properties of the scene. Gaussian splatting is a method for representing 3D scenes and rendering novel views. In Gaussian splatting, a 3D world is represented with a set of 3D points where each point is a 3D Gaussian with its own unique parameters that are fitted per scene such that renders of this scene match closely to the known dataset images. 3D Gaussian splatting is, therefore, analogous to triangle rasterization in computer graphics, which is used to draw many triangles on the screen. However, instead of drawing triangles, they
are Gaussian. Therefore, a Gaussian is described by parameters such as position, covariance measuring how a Gaussian stretches and/or scales, color such as Red, Green, and Blue, and alpha measuring the transparency of a Gaussian, etc.
[006] Nonetheless, conventional 3D Gaussian splatted radiance field methods directly embed language-based semantics into the Gaussians in their attempt to address the integration of segmentation and open vocabulary semantics into 3D Gaussian Radiance Fields. As a result, the 3D representations may not be semantically disjoint to allow manipulation (e.g., user interactions in an extended-reality application).
[007] Some legacy approaches uses simultaneous localization and mapping (SLAM) for constructing or updating a map of an environment while keeping track of the user’s location in the environment. Some of these legacy SLAM approaches have attempted to combine Gaussian spatting. Nonetheless, these legacy approaches either require RDG-D (red, blue, and green for color plus depth) data or are incompatible with real-time or nearly real-time applications.
[008] Therefore, there exists a need for methods, systems, and computer program products for extended-reality systems.
SUMMARY
[009] Disclosed are method(s), system(s), and article(s) of manufacture for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems in one or more embodiments. Some embodiments are directed at a method for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems.
[0010] Some embodiments are directed to an extended-reality system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems. The extended-reality system includes a wearable eyepiece that presents virtual contents to a user, a belt pack operatively coupled to the wearable eyepiece, a processor; and a non- transitory computer readable medium storing thereupon a sequence of instructions which, when executed by a model with the processor, causes the processor to estimate a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device, to generate a semantically annotated, manipulatable three-dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece, and to modify the semantically annotated, manipulatable 3D representation in real-time or nearly real-time in response to a user interaction.
[0011] In some of these embodiments where the processor is to estimate the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further includes receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; determining, from the video or the sequence of RGB images, the plurality of keyframes; and estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images.
[0012] In some of the immediately preceding embodiments, the set of acts executed by the model further includes performing a global bundle adjustment that
corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
[0013] In addition or in the alternative, the model further executes the set of acts to identify a non-keyframe from the video or the sequence of RGB images. A nonkeyframe pose may be determined for the non-keyframe, and a training data set may be augmented at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
[0014] In some embodiments, the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to generate a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
[0015] In some of the immediately preceding embodiments, the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to establish a semantic association among the plurality of object segmentation masks at least by using a zeroshot, one-shot, or few-shot neural network or a multi-view geometry. A three-dimensional (3D) point set may be generated based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
[0016] In some of the immediately preceding embodiments, the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to improve 3D Gaussian splatted neural radiance field at least by attaching a lower-dimensional semantic vector to a cluster
of one or more Gaussians for the semantically annotated, manipulatable three- dimensional (3D) representation.
[0017] Some embodiments are directed to a method for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems. In these embodiments, a plurality of keyframes, a plurality of keyframe poses, and depth data may be estimated or otherwise determined from a plurality of captures that is captured by at least the extended reality device. A semantically annotated, manipulatable three-dimensional (3D) representation may be generated in a physical environment for perception by the user wearing the wearable eyepiece. The semantically annotated, manipulatable 3D representation may then be modified in real-time or nearly real-time in response to a user interaction.
[0018] In some of these embodiments, a video or a sequence of RGB (Red, Green, Blue) images may be received for estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, wherein the video or the sequence of RGB images are not required to include depth information. The plurality of keyframes may be determined from the video or the sequence of RGB images. At least one keyframe of the plurality of keyframes may be estimated across the video or the sequence of RGB images. [0019] In at least some of the immediately preceding embodiments, a global bundle adjustment may be performed to correct a camera pose corresponding to at least one of the plurality of keyframe poses or to perform a loop closure for estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data.
[0020] In addition or in the alternative, a non-keyframe may be identified from the video or the sequence of RGB images for estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data. A non-keyframe pose may be determined for the non-keyframe. A training data set may be augmented at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
[0021] In some embodiments, generating the semantically annotated, manipulatable three-dimensional (3D) representation may include establishing a semantic association among the plurality of object segmentation masks at least by using a zero-shot, one-shot, or few-shot neural network or a multi-view geometry. A three- dimensional (3D) point set may be generated based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
[0022] In at least some of the immediately preceding embodiments, 3D Gaussian splatted neural radiance field may be improved at least by attaching a lower-dimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three-dimensional (3D) representation.
[0023] Some embodiments are directed at a hardware system that may be invoked to perform any of the methods, processes, or sub-processes disclosed herein. The hardware system may include or involve an extended-reality system having at least one processor or at least one processor core, which executes one or more threads of execution to perform any of the methods, processes, or sub-processes disclosed herein in some embodiments. The hardware system may further include one or more forms of
non-transitory machine-readable storage media or devices to temporarily or persistently store various types of data or information. Some exemplary modules or components of the hardware system may be found in the System Architecture Overview section below.
[0024] Some embodiments are directed at an article of manufacture that includes a non-transitory computer readable medium having stored thereupon a sequence of instructions which, when executed by at least one processor or at least one processor core, causes the at least one processor or the at least one processor core to perform any of the methods, processes, or sub-processes disclosed herein. Some exemplary forms of the non-transitory machine-readable storage media may also be found in the System Architecture Overview section below.
[0025] In these embodiments, the non-transitory computer readable medium stores thereupon a sequence of instructions which, when executed by a model with a processor, causes the processor to estimate a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device, to generate a semantically annotated, manipulatable three- dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece, and to modify the semantically annotated, manipulatable 3D representation in real-time or nearly real-time in response to a user interaction.
[0026] In some of these embodiments where the processor is to estimate the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further includes receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; determining, from the video or the sequence of RGB images,
the plurality of keyframes; and estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images.
[0027] In some of the immediately preceding embodiments, the set of acts executed by the model further includes performing a global bundle adjustment that corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
[0028] In addition or in the alternative, the model further executes the set of acts to identify a non-keyframe from the video or the sequence of RGB images. A nonkeyframe pose may be determined for the non-keyframe, and a training data set may be augmented at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
[0029] In some embodiments, the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to generate a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
[0030] In some of the immediately preceding embodiments, the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional (3D) representation may further execute the set of acts to establish a semantic association among the plurality of object segmentation masks at least by using a zeroshot, one-shot, or few-shot neural network or a multi-view geometry. A three-dimensional (3D) point set may be generated based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
[0031] In some of the immediately preceding embodiments, the model executing the set of acts to generate the semantically annotated, manipulatable three-dimensional
(3D) representation may further execute the set of acts to improve 3D Gaussian splatted neural radiance field at least by attaching a lower-dimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three- dimensional (3D) representation.
BRIEF DESCRIPTION OF THE DRAWINGS
[0032] This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the U.S. Patent and Trademark Office upon request and payment of the necessary fee.
[0033] The drawings illustrate the design and utility of various embodiments of the invention. It should be noted that the figures are not drawn to scale and that elements of similar structures or functions are represented by like reference numerals throughout the figures. In order to better appreciate how to obtain the above-recited and other advantages and objects of various embodiments of the invention, a more detailed description of the present inventions briefly described above will be rendered by reference to specific embodiments thereof, which are illustrated in the accompanying drawings. Understanding that these drawings depict only typical embodiments of the invention and are not therefore to be considered limiting of its scope, the invention will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:
[0034] FIG. 1 illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
[0035] FIG. 1A illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1 in some embodiments.
[0036] FIG. 1 B illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1 in some embodiments.
[0037] FIG. 10 illustrates a more detailed block diagram conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
[0038] FIG. 2A illustrates a high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
[0039] FIG. 2B illustrates another high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments.
[0040] FIG. 3A illustrates a simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual
environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments.
[0041] FIG. 3B illustrates another simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
[0042] FIG. 3C illustrates more details about a portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments.
[0043] FIG. 3D illustrates more details about another portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments.
[0044] FIG. 3E illustrates a simplified block diagram of a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
[0045] FIG. 3F illustrates a block diagram that continues from FIG. 3E and generates segmentation masks, according to some embodiments.
[0046] FIG. 3G illustrates an example block diagram for learning semantic labels in 3D, according to some embodiments.
[0047] FIG. 4 illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments.
[0048] FIG. 5A illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments.
[0049] FIGS. 5B-5C illustrate a more detailed block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments.
[0050] FIG. 6A illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields without pre-training, according to some embodiments.
[0051] FIG. 6B illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields with pre-training, according to some embodiments.
[0052] FIG. 7A illustrates s simplified schematic diagram for a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments.
[0053] FIG. 7B illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments.
[0054] FIG. 7C illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments.
[0055] FIG. 7D illustrates some example modifications of the simplified schematic diagram illustrated in FIG. 7C, according to some embodiments.
[0056] FIG. 7E illustrates a simplified block diagram for an example image encoder that may be employed in FIG. 7A or 7B, according to some embodiments.
[0057] FIG. 7F illustrates a simplified block diagram for an example text encoder that may be employed in FIG. 7A or 7B, according to some embodiments.
[0058] FIGS. 7G-7H illustrate an example attention mechanism that may be employed for a sequence transduction module in a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments.
[0059] FIG. 8 illustrates a simplified example of a wearable extended-reality (XR) device with a belt pack external to the XR glasses in some embodiments.
[0060] FIG. 9 illustrates a computerized system on which some of the methods described herein may be implemented.
[0061] FIG. 10 shows an example architecture 2500 for the electronics operatively coupled to an optics system or XR device in one or more embodiments.
DETAILED DESCRIPTION
[0062] In the following description, certain specific details are set forth in order to provide a thorough understanding of various disclosed embodiments. However, one skilled in the relevant art will recognize that embodiments may be practiced without one or more of these specific details, or with other methods, components, materials, etc. In other instances, well-known structures associated with computer systems, server computers, and/or communications networks have not been shown or described in detail to avoid unnecessarily obscuring descriptions of the embodiments.
[0063] It shall be noted that, unless the context requires otherwise, throughout the specification and claims which follow, the word “comprise” and variations thereof, such as, “comprises” and “comprising” are to be construed in an open, inclusive sense, that is as “including, but not limited to.”
[0064] It shall be further noted that Reference throughout this specification to “one embodiment” or “an embodiment” means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics, unless otherwise explicitly described as mutually exclusive of one another, may be readily combined in any suitable manner in one or more embodiments. Furthermore, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the content clearly dictates otherwise. It should also be noted that the term “or” is generally employed in its sense including “and/or” unless the content clearly dictates otherwise.
[0065] Various embodiments will now be described in detail with reference to the drawings, which are provided as illustrative examples of the invention so as to enable those skilled in the art to practice the invention. Notably, the figures and the examples below are not meant to limit the scope of the present invention. Where certain elements of the present invention may be partially or fully implemented using known components (or methods or processes), only those portions of such known components (or methods or processes) that are necessary for an understanding of the present invention will be described, and the detailed descriptions of other portions of such known components (or methods or processes) will be omitted so as not to obscure the invention. Various embodiments are directed to management of a virtual-reality (“VR”), augmented reality (“AR”), mixed-reality (“MR”), and/or extended reality (“XR”) system (collectively referred to as an “XR system” or extended-reality system) in various embodiments.
[0066] FIG. 1 illustrates a block diagram for learning semantic labels for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More specifically, FIG. 1 as well as other embodiments described herein illustrates an online reconstruction of instance-aware Gaussian splatted neural radiance field technique that may be used in conditional, probabilistic generation of 3D virtual environment in response to a prompt by a user wearing an extended-reality device in real-time or nearly real-time based at least in part upon incomplete information made available to the extended-reality device. In various embodiments, the aforementioned technique is instance aware in the sense that a virtual object so rendered is manipulatable and interactable by users using extended-reality devices. In these embodiments, an
instance may refer to a single Gaussian or a cluster or a grouping of a plurality of Gaussians that may be clustered or otherwise grouped into the same cluster or grouping which may then be associated with semantic meanings or a semantic label. For example, Gaussians that correspond to the same virtual object may be clustered into the same cluster or instance, and a semantic label may be associated with the cluster or instance. When a system renders the virtual object using rasterization techniques to “draw” the cluster, grouping, or instance of one or more Gaussians, the system is made aware of the cluster, grouping, or instance (as well as the semantics) and may thus render the cluster, grouping, or instance of one or more Gaussians as an undivided whole to make the rendered cluster, instance, or grouping manipulatable or interactable (e.g., in response to prompts).
[0067] In these embodiments, a plurality of keyframes, respective camera poses (may also be referred to as keyframe poses) for the plurality of keyframes, and/or depth information may be determined or estimated at 102 from a video or a sequence of images (e.g., photographs) that is acquired by or provided by (e.g., obtained from another XR device or server) an XR device worn by a user. Unlike conventional SLAM (simultaneous localization and mapping)-based techniques that require RBG-D (data pertaining to the three primary colors as well as depth) in the input images, these embodiments described herein do not require depth information or data from the input image in order to generate 3D representation for conditional, probabilistic generation of three-dimensional (3D) virtual environment.
[0068] With the plurality of keyframes, camera poses, and the depth data or information estimated or determined from the input video or the input sequence of images
at 102, a semantically annotated, editable 3D representation may be generated at 104 and placed in a physical environment as perceived by a user wearing the XR device. The
3D representation may be interacted upon by the user via the XR device, and various techniques described herein may modify the 3D representation at 106 to reflect the effect of the user interaction upon the 3D representation in real-time or nearly real-time. Nearly rea-time or near real-time near real time (NRT) pertains to the timeliness of data or information which has been delayed by the time required for electronic communication and automatic data processing and implies that there are no significant delays other than the delay caused by the time required for electronic communication and automatic data processing.
[0069] FIG. 1A illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1A in some embodiments. More particularly, FIG. 1A illustrates more details about estimating keyframes, camera poses, and depth information at 102 of FIG. 1 . In these embodiments, a video or a sequence of RGB images may be received at 102A wherein the video or the sequence of RGB images are not required to include depth information or data. The plurality of keyframes may be estimated or determined from the video or the sequence of RGB images. At least one keyframe of the plurality of keyframes may be estimated and tracked at 104A across the entire video or the entire sequence of RGB images or at least a portion thereof. In some embodiments, respective keyframe poses or the respective camera poses for the corresponding keyframes may be estimated or determined at 104A.
[0070] A global bundle adjustment may be performed globally (e.g., to the entire video, the entire sequence of RGB images, or a portion thereof) at 106A to correct one
or more camera poses and/or to perform loop closure in some embodiments. In some embodiments where any models described herein need to be trained, retrained, or validated with additional images, one or more non-keyframes in the video or the sequence may be embedded at 108A, and the one or more corresponding camera poses for the one or more embedded non-keyframes may be estimated at 110A. A training data set may be augmented at 112A at least by embedding the one or more non-keyframes and the one or more corresponding camera poses for the one or more non-keyframes into the training data set, wherein the training data set is used to train a model that performs the acts described above with reference to FIG. 1 .
[0071] FIG. 1 B illustrates more details about a portion of the high-level block diagram illustrated in FIG. 1A in some embodiments. More particularly, FIG. 1 B illustrates more details about generating a semantically annotated, editable 3D representation at 104 in FIG. 1A. In these embodiments, one or more object segmentation masks may be generated at 102B at least by analyzing the plurality of keyframes. In some embodiments where one or more non-keyframes have been embedded, one or more object segmentation masks may be generated at 102B at least by analyzing the plurality of keyframes as well as the one or more embedded non-keyframes. More details about object segmentation masks will be described below with reference to multiple drawing figures.
[0072] Semantic association may be established among the one or more object segmentation masks at least by using a zero-shot, one-shot, or few-shot tracker neural network or multi-view geometry at 104B. With the semantic association established, a 3D initial potin set may be generated at 106B based at least in part upon an inverse projection
of the estimated or determined camera poses, depth data, and/or the associated depth covariance (e.g., uncertainty). In some embodiments, the 3D initial point set may be generated by selecting points with low uncertainty (e.g., uncertainty below a certain threshold) and by reverse-projecting the selected points in 3D using the corresponding pose and depth information.
[0073] A lower dimensional semantic vector may be attached to a lower dimensional grouping of Gaussians at 108B to improve Gaussian splatted radiance field with semantics. In some embodiments, a semantic vector may include a compact representation of a true or predicted class label (hence in a lower dimensional space) for an object. In these embodiments, a low dimensional grouping of Gaussians may be generated before attaching semantic vectors or labels so as to form semantically disjointed objects. Unlike legacy approaches that directly embed language-based semantics onto the Gaussians, various embodiments generate a lower-dimensional grouping of Gaussians and then attach semantic meaning to each such low-dimensional grouping of Gaussians. As a result, 3D object representations generated by various techniques described herein may be semantically disjoint from each other so that a user wearing an XR device may interact with or manipulate such 3D object representations independently and individually.
[0074] In some embodiments, a 3D space may be defined as a set of Gaussians where each Gaussian is described by a set of parameters that are calculated or determined by machine learning. These parameters may include position, covariance measuring how a Gaussian stretches and/or scales, color such as Red, Green, and Blue, and alpha measuring the transparency of a Gaussian, etc. Gaussian rasterization may be
deemed analogous to triangle rasterization in legacy computer graphics that draw triangles on a display device whereas Gaussian rasterization draws Gaussians, instead of triangles.
[0075] Gaussian splatting is a rendering technique and includes a rasterization technique for real-time (or nearly real-time) 3D reconstruction and rendering of images taken from multiple points of view. For example, 3D Gaussian splatting represents a 3D scene as a large number of particles (Gaussians), and each 3D Gaussian is associated with its position, orientation, scale, opacity (or transparency), and color. To render a Gaussian, the Gaussian may be first converted into the 2D space and then organized for rendering. One method to render Gaussians may include creating a point cloud from one or more images; converting each point in the point cloud to a Gaussian by inferring or determining one or more of the aforementioned Gaussian parameters from the data or metadata associated with the point in the point cloud to enable rasterization; training a network to produce high-quality results by using, for example yet without limitation, stochastic gradient descent and by adjusting the Gaussian parameters according to the loss; and performing differentiable Gaussian rasterization by projecting the 2D Gaussian from the camera’s perspective (e.g., pose) and by repeated backward and forward combination for each Gaussian. It shall be noted that although various examples or embodiments refer to 3D Gaussian splatting for generation of 3D representations, NeRF (neural radiance field) may also be used while both NeRF and 3D Gaussian splatting address the shortcomings of photometry (e.g., unable to represent scenes that do not have contours such as the sky) although in some scenarios 3D Gaussian splatting may not require as much computing power to render the scene in real-time as NeRF does.
[0076] FIG. 1 C illustrates a more detailed block diagram conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More specifically, FIG. 1 C illustrates a block diagram for learning semantic labels in 3D. In these embodiments, a rendering process may be invoked at 102C, and a number (N) of labels from one or more segmentation masks in 3D may be projected onto a 2D space at 104C to generate a number (N) of projected labels.
[0077] The number (N) of projected labels may be decompressed at 106C into a number (N) of decompressed labels. In some of these embodiments, decompressing a projected semantic label into a decompressed semantic label may be performed using a linear layer, a dense layer, or a fully connected (FC) layer where every input neuron or node is connected to every output neuron or node, and the linear layer maps an input to an output with a weight (W or a weight matrix) and a bias (b or a bias vector).
[0078] The number (N) of decompressed labels may be expanded at 108C into the number plus one (N+1 ) labels, and these N+1 decompressed labels may be transformed into a probabilistic distribution of N+1 outcomes at 110C. In some embodiments, a hierarchical part-whole structure of objects in an input image may be learned at 112C using a decoder based at least in part upon a frequency schedule. These embodiments learn a hierarchical part-whole structure of objects, rather than learning a semantic class per pixel as most of legacy approaches otherwise learn.
[0079] Moreover, unlike conventional semantic label learning techniques, these embodiments learn the hierarchical part-whole structure of objects, instead of multiple semantic embeddings. In some embodiments, a frequency schedule embeds one or more
objects into different latent spaces (also referred to as embedding space or latent feature space) based at least in part upon the one or more respective frequencies where a latent space comprises an embedding of a set of items within a manifold where similar items are positioned closer to each other.
[0080] A plurality of Gaussians may be clustered at 114C into one or more clusters based at least in part upon respective object identifiers. For example, a group of Gaussians corresponding to the same object may be clustered into the same cluster. A semantic vector may be embedded at 116C to each of the one or more clusters. In some embodiments, embedding a semantic vector to a cluster may be achieved by running a segmenter (e.g., a contrastive language-image pre-training or CLIP model) on a plurality of masked image crops.
[0081] An object may be selected at 118C at least by querying the embedded semantic vectors obtained from 116C using a user input with a textual model and/or speech-to-text model. The object may be provisioned by an extended reality device worn by a user in an extended reality session to facilitate user interactions with the rendered object at 120C, and the 3D representation of the object may be updated in real-time or nearly real-time to reflect the effects of the user interactions upon the object. In these embodiments, the object is semantically disjoint from other objects or 3D representations generated by the extended reality device to allow such user interactions and manipulations.
[0082] FIG. 2A illustrates a high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems,
according to some embodiments. More particularly, FIG. 2A illustrates provisioning reconstruction of 3D representations of a scene using neural radiance field (NeRF). NeRF is one of several techniques that may be used in various embodiments described herein for reconstructing 3D scenes without contours to address the shortcomings of photometry that is unable to represent scenes that do not have contours. In these embodiments, a prompt or input may be received at 202A for reconstructing a 3D representation for interaction by a user using an extended reality device. For example, a user wearing an XR device may provide a prompt as a condition for generating a 3D representation for a scene or a portion thereof.
[0083] A prompt may include a textual input, a voice input, and/or an input of an image although different inputs may be processed with different encoders to generate respective embeddings. For example, a textual input (e.g., by entering textual instructions in natural language) may be processed by a text encoder; an image input may be processed by an image encoder; and a voice input may be processed by a speech encoder or by a combination of a speech-to-text encoder and a textual encoder.
[0084] A 3D image or 3D model may be generated at 204A using one or more captures in response to the prompt. In some embodiments, at least one of the one or more captures may be captured by using one or more sensors. For example, an XR device may invoke its image sensor(s) to capture one or more 2D captures (e.g., photographs or other types of images) of a scene for the generation of the 3D image or 3D model at 204A.
[0085] One or more holds may be identified and filled at 206A with in-painting in response to the prompt. For example, an area in the scene may be occluded by an object
so that the XR device cannot “see” the area due to occlusion. The XR device may invoke an in-painting module to predict what the occluded area may appear like and fill the area with predicted representation. In some of these embodiments, one or more areas may be beyond the visible range of the XR device, and the XR device may extrapolate the 3D image or 3D model at 208A with out-painting in response to the prompt. For example, a user wearing an XR device within a room having a wall where the XR device is thus unable to see what lies beyond the wall or outside the room. In this example, the XR device may invoke an out-painting module to expand the 3D image or 3D model beyond the wall or room by predicting what may lie beyond the wall or outside the room.
[0086] FIG. 2B illustrates another high-level block diagram for a portion of a process for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments. More particularly, FIG. 2B illustrates more details about neural radiance field that represents the light field as integrated radiance rays through the 3D space. In these embodiments, a plurality of 2D captures may be received at 202B where the plurality of 2D captures comprises color information such as the RGB data but not depth information.
[0087] These embodiments illustrated in FIG. 2D are in sharp contrast with conventional SLAM-based Gaussian splatting techniques that require RGB-D (color information in RGB plus the depth information) for reconstruction of 3D representations. A 2D capture may be deemed a five-dimensional (5D) input including its spatial location (x, y, z) and a viewing direction (in terms of theta for azimuth and phi for elevation). As described above, legacy approaches require six-dimensional (6D) input - the spatial
location including depth (x, y, z, d) and viewing direction (theta and phi). The volume density at a location includes the differential probability of a ray that terminates at an infinitesimal particle at that location in some embodiments.
[0088] An output including view-dependent emitted radiance may be generated at 204B in response to the aforementioned input at the spatial location (emitted color in terms of R, G, B) and the volume density (alpha measuring the transparency or opacity at the location). In some embodiments, the output view-dependent emitted radiance may be generated at 204B by compositing the color information and the volume density into the view-dependent emitted radiance. Depth information may be determined at206B from the 2D captures. A view may be synthesized at 208B at least by querying or sampling the aforementioned 5D coordinates in the input along a camera ray and further by using volume rendering techniques to project the queried output colors and the volume density information (alpha) to a specific pixel represented in the 3D representation of the view where the process illustrated in FIG. 2B represents the light field as integrated radiance along light rays through the 3D space.
[0089] In some embodiments, synthesizing and rendering a view includes estimating the expected color of a camera ray with near and far bounds, and the expected color may be expressed as an integral to represent the accumulated transmittance along the ray from the near bound to the far bound without hitting any other particles. In some of these embodiments, the expected color for the output may be represented as an integral for a camera ray traced through each pixel of a desired virtual camera for the 3D representation. The synthesized view may be expanded into a 3D expanded view at 21 OB at least by using the depth information determined at 206B from the 2D captures. In some
embodiments, the volume rendering techniques may be optionally improved at 212B at least by reducing a residual or loss between the 3D expanded view and a corresponding ground truth image during training of the model for generating the aforementioned 3D expanded views for 3D representation.
[0090] FIG. 3A illustrates a simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended- reality systems, according to some embodiments. More particularly, these embodiments illustrated in FIG. 3A comprise a model that may be invoked by various techniques described herein to fill a “hole” or “gap” (e.g., an occluded area in a scene). In these embodiments, the example network receives an input image 302A (e.g., a photograph captured by an XR device when a user wearing the XR device enters a scene) which may be turned into a masked input 304A.
[0091] The masked input 304A may be sent to a coarse network 306A for processing. In some embodiments, the coarse network 306A may comprise a plurality of dilated convolution layers 326A that may be used to provide sufficiently large receptive fields. The plurality of dilated convolution layers thus enables the coarse network to capture a wider context, without significant increase in computational costs. In some embodiments, the plurality of dilated convolution layers may include a first dilated convolution layer with a dilation rate of 2, a second dilated convolution layer with a dilation rate of 4, a third dilated convolution layer with a dilation rate of 8, and a fourth dilated convolution layer with a dilation rate of 16, each with 1x1 stride for 256 outputs each, in some embodiments.
[0092] In some embodiments, the refinement network 31 OA may include a plurality of dilated convolution layers 326A to expand the receptive field as well as a contextual attention mechanism. In some embodiments, in addition to the four dilated convolution layers 326A, the coarse network 306A may include, from the input, a first layer having a 5x5 kernel with one dilation and 1x1 stride, a second layer having a 3x3 kernel with one dilation and 2x2 stride, a third layer having a 3x3 kernel with one dilation and 1x1 stride, a fourth layer having a 3x3 kernel with one dilation and 2x2 stride, a fifth layer having a 3x3 kernel with one dilation and 1x1 stride, and a sixth layer having a 3x3 kernel with one dilation and 1x1 stride. The sixth layer feeds into the aforementioned four dilated convolution layers. The sixth layers receiving outputs from the four dilated convolution layers may be the aforementioned six layers arranged in the reverse order.
[0093] The coarse network 306A processes the masked input 304A to generate the coarse output 308A which may be forwarded to a refinement network 31 OA which the generates the inpainting results 312A to fill the “hole” or “gap” that is masked in 304A by a mask or patch (e.g., for the purposes of training the network). The inpainting result 312A may include the global image 318A (e.g., the masked input image 304A plus the inpainting result for the “hole” or “gap”) and a local image 314A (e.g., the local image for the inpainting result for the “hole” or “gap”). The global image 318A may be fed to a global discriminator 320A that determines the loss (e.g.,. cross-entropy loss) between the global image 318A and the corresponding ground truth, and the local inpainting result 314A may be processed by a local discriminator 316A to determine the loss (e.g., cross-entropy loss). The losses from the global discriminator 320A and the local discriminator 316A may be forwarded to the loss module 322A that determines the reconstruction loss and the
GAN (generative adversarial network) loss both of which may then be propagated to tune the coarse network output 308A and the refinement network output (the inpainting result) 312A.
[0094] FIG. 3B illustrates another simplified, example network architecture that may be used as a part of conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More particularly, FIG. 3B illustrates more details about the network illustrated in FIG. 3A. In these embodiments, an input 302B (e.g., a 3D input image or a 2.5D input image having 2D image with depth information) may be processed with a mask 304B (e.g., a binary mask) and a textual or voice input 324B into a masked output 306B, a global spatial loss 326B, and a local spatial loss 327B.
[0095] The masked output 306B may be provided to a dilated convolution network 308B, which may also receive the mask 304B, to increase the receptive field. The output of the dilated convolution network 308B may be provided to the coarse prediction module 31 OB that also receives the global spatial loss 326B, if available, to generate a coarse inpainting result 360B. The output including the masked output 306B and the coarse inpainting result 360B may be provided to a refinement network 312B which receives contextual attention 326B and generates a refined prediction 314B. The refined prediction 314B including the masked output 306B with a refined inpainting result 362B may be provided to a global critic 316B (e.g., the global discriminator 320A in FIG. 3A) to generate a global output 320B.
[0096] The refined inpainting result 362B may be provided to a local critic 318B (e.g., the local discriminator 316A in FIG. 3A) that generates a local output 322B. The global output 320B and the local output 322B may be provided to a loss module 332B that determines the reconstruction loss 321 B (e.g., a weighted Mean Squared Error (MSE) loss) and the GAN loss 323B. In some embodiments, the reconstruction loss 321 B is also to enhance training stability. The GAN loss 323B and the spatial loss 327B may be provided for training 328B that propagates the losses (e.g., by backward propagation) to fine turn the refinement network 312B in order to produce high-fidelity inpainting results.
[0097] The reconstruction loss 321 B and the spatial loss 326B may be provided to training 328B to fine tune the dilated convolution network 308B and the coarse prediction network 310B. In some embodiments, the global discriminator takes the entire image as input, while the local discriminator takes only a small region around the completed area as input. In some embodiments, both the global and local discriminators are trained to determine whether an image (e.g., the output image with inpainting result) is real or completed by the completion network (not shown), while the completion network is trained to fool both discriminator networks.
[0098] FIG. 3C illustrates more details about a portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments. More particularly, FIG. 3C illustrates more details about the coarse prediction (e.g., 310B in FIG. 3B or 306A in FIG. 3A). In these embodiments, the input 302B may be masked by a mask or patch 304B (e.g., a binary mask) to generate the masked input 306B which may be fed to a dilated convolution network 308B, and the input 302B may be further optionally supplemented with textual input 324B as similarly shown in FIG. 3B.
[0099] The dilated convolution network 308B may process the masked input 306B to generate an intermediate output that may be further processed by a coarse prediction network 31 OB to generate a coarse inpainting result or a patch 302C. The coarse prediction output including the masked input 306B and the patch 302C may be forwarded to a global critic 316B that generates the global output 320; and the coarse inpainting prediction or the patch 302C may be processed by a local critic 318B to generate a local output 322B.
[00100] Both the global output 320B and the local output 322B may be forwarded to a loss module 322B that generates the reconstruction loss 321 B and the GAN loss 323B. The reconstruction loss 321 B may be backward propagated 304C to the masked output 306B in order to train the dilated convolution network 308B and/or the coarse prediction network 310B. In some embodiments, the global critic 318B and the local critic 316B are included for training purposes but are not included for testing purposes.
[00101] FIG. 3D illustrates more details about another portion of the simplified, example network architecture illustrated in FIG. 3B, according to some embodiments. More particularly, FIG. 3D illustrates more details about a refinement network. In these embodiments, a first input 302D may be provided to an extraction module 304D that produces foreground input features 306D and background input features 308D. the background input features 308D may be further processed by a filter generation module 310D to generate one or more filters 312D for subsequent convolution. Both the foreground input features 306D and the one or more filters 312D may be provided to a convolution network 314D that generates a matching score 316D. For example, attention score may be determined at 316D for each pixel in the foreground input features 306D.
[00102] A vector-to-probabilistic distribution may be generated at 318D based at least in part upon the matching score determined at 316D. For example, a channel-wise softmax may be performed to compare and determine the attention score 320D for each pixel and the attention map 322D to transform the values into the probabilistic distribution of possible outcomes at 318D. Both the attention scores 320D and the attention map 322D may be provided to an attention fusion mechanism 324D that generates the attention score 326D. The attention fusion mechanism 324D performs a left-right propagation to compute an intermediate attention score followed by top-down propagation with a kernel (e.g., an identity matrix) of a certain size to compute the attention score 326D in some of these embodiments. The attention score 326D may be forwarded to a deconvolution layer 328D to generate the second output 330D which comprises the output (e.g., the refined prediction 314B) of the refinement network (e.g., 312B).
[00103] FIG. 3E illustrates a simplified block diagram of a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. More particularly, FIG. 3E illustrates a block diagram of a 3D Gaussian splatted radiance field module. In these embodiments, a video or a sequence of RGB images 304E (without depth information) may be provided to two simultaneous localization and mapping (SLAM) module 306E. In some of these embodiments where depth information is no longer required, the devices (e.g., an XR device) is also not required to include any depth sensors.
[00104] A SLAM module 306E processes the RGB images 304E to determine and track keyframes, to generate the depth data estimate 308E, to perform a global bundle adjustment operation 310E (e.g., to globally correct camera poses and/or to perform loop closure), to embed one or more non-keyframes in the video or the sequence of RGB images 312E, and/or to estimate camera poses for the respective keyframes 314E. The SLAM 306E thus generates a plurality of keyframes 316E, a plurality of camera poses 318E for the plurality of keyframes 316E, and depth information 320E. The respective outputs of these two SLAM modules 306E are provided to a 3D Gaussian splatted radiance field module 322E to generate one or more new scenes 324E (e.g., reconstruction of 3D representations) including one or more individually, independently manipulatable objects 326E by generating these one or more semantically disjointed objects. In some embodiments, instead of directly embedding language-based semantics onto the Gaussians, these embodiments first generate a low dimensional grouping of Gaussians (e.g., based on the object identifier to which the Gaussians correspond) into semantically disjoint objects before attaching semantic meaning to each low dimensional grouping.
[00105] The output of one of the two SLAM modules 306E may be provided to initialize the 3D scene 314E. Moreover, a plurality of 3D scenes 302E may be captured by one or more devices such as cameras, video cameras, XR devices, etc., and the plurality of 3D scenes may also be provided to the 3D Gaussian splatted radiance field module 322E for the generation or reconstruction of 3D representations 324E.
[00106] FIG. 3F illustrates a block diagram that continues from FIG. 3E and generates segmentation masks, according to some embodiments. The process illustrated
in FIG. 3F occurs after acquiring the keyframes, the camera poses, and the depth information described above with reference to FIG. 3E. In these embodiments, an image, a sequence of images, or a video may be received at 302F, and the image, the sequence, or the video may be analyzed at 304F using a 2D grid. In some embodiments, the image, the sequence of images, or the video may be acquired by a device such as an XR device worn by a user. In some embodiments, a grid may include, for example without limitation, a 32 x 32 grid by using a segmentation process. In some of these embodiments, the segmentation process returns a part-whole hierarchy of masks, rather than learning a semantic class per pixel as most of legacy approaches otherwise learn.
[00107] A plurality of masks may be determined at 306F for each point in the 2D grid, and semantic association may be established among the plurality of masks at 308F. In some embodiments, semantic association may be established at 308F by using a zeroshot, one-shot, or few-shot neural network, or multi-view geometry. One or more points may be selected at 31 OF from a plurality of points in the 2D grid based at least in part upon an inverse projection of the estimated or determined camera poses, depth data, and/or the associated depth covariance (e.g., uncertainty) some or all of which may be determined according to the block diagram illustrated in FIG. 3E. In some of these embodiments, points with low uncertainty below a threshold may be selected from the 2D grid at 391 F.
[00108] A sparse 3D point set may be generated at 312F for the selected points.
As described immediately above, the sparse 3D point set may be determined at 312F based at least in part upon an inverse projection of the estimated or determined camera poses, depth data, and/or the associated depth covariance (e.g., uncertainty) some or all
of which may be determined according to the block diagram illustrated in FIG. 3E. A selected point may be unprojected into 3D by using, for example, the corresponding camera pose, the depth information, and/or the depth covariance which measures how the point is stretched and/or scaled. The 3D Gaussian splatted neural radiance field may be improved or optimized at 314F at least by attaching a low-dimensional semantic vector representing a true class label (e.g., a 3D semantic label) to each cluster of Gaussians or to each Gaussian.
[00109] FIG. 3G illustrates an example block diagram for learning semantic labels in 3D, according to some embodiments. More specifically, FIG. 3G illustrates a block diagram of learning semantic labels that are referenced at 314F of FIG. 3F. In these embodiments, a rendering process may be invoked at 302G, and a number (N) of labels from one or more segmentation masks in 3D may be projected onto a 2D space at 304G to generate a number (N) of projected labels.
[00110] The number (N) of projected labels may be decompressed at 306G into a number (N) of decompressed labels. In some of these embodiments, decompressing a projected semantic label into a decompressed semantic label may be performed using a linear layer, a dense layer, or a fully connected (FC) layer where every input neuron or node is connected to every output neuron or node, and the linear layer maps an input to an output with a weight (W or a weight matrix) and a bias (b or a bias vector).
[00111] The number (N) of decompressed labels may be expanded at 308G into the number plus one (N+1 ) labels, and these N+1 decompressed labels may be transformed into a probabilistic distribution of N+1 outcomes at 310G. In some embodiments, a hierarchical part-whole structure of objects in an input image may be learned at 312G
using a decoder based at least in part upon a frequency schedule. These embodiments learn a hierarchical part-whole structure of objects, rather than learning a semantic class per pixel as most of legacy approaches otherwise learn.
[00112] Moreover, unlike conventional semantic label learning techniques, these embodiments learn the hierarchical part-whole structure of objects, instead of multiple semantic embeddings. In some embodiments, a frequency schedule embeds one or more objects into different latent spaces (also referred to as embedding space or latent feature space) based at least in part upon the one or more respective frequencies where a latent space comprises an embedding of a set of items within a manifold where similar items are positioned closer to each other.
[00113] A plurality of Gaussians may be clustered at 314G into one or more clusters based at least in part upon respective object identifiers. For example, a group of Gaussians corresponding to the same object may be clustered into the same cluster. A semantic vector may be embedded at 316G to each of the one or more clusters. In some embodiments, embedding a semantic vector to a cluster may be achieved by running a segmenter (e g., a contrastive language-image pre-training or CLIP model) on a plurality of masked image crops.
[00114] An object may be selected at 318G at least by querying the embedded semantic vectors obtained from 316G using a user input with a textual model and/or speech-to-text model. The object may be provisioned by an extended reality device worn by a user in an extended reality session to facilitate user interactions with the rendered object at 320G, and the 3D representation of the object may be updated in real-time or nearly real-time to reflect the effects of the user interactions upon the object. In these
embodiments, the object is semantically disjoint from other objects or 3D representations generated by the extended reality device to allow such user interactions and manipulations.
[00115] FIG. 4 illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems, according to some embodiments. In these embodiments, an input pertaining to an environment from a sensor may be received at 402. The input may include a textual input (e.g., a string of text in natural language), a 2D floorplan of an environment, an image input (e.g., a photograph, a sequence of photographs or images captured by an image sensor, or a video captured by an image sensor, a voice input captured by a microphone, or any combination thereof) in some embodiments. In some embodiments, the more information a model has about an object’s surface, the better the underlying model will be in discovering the object’s shape and the way it interacts with lights.
[00116] Semantics may be determined for the input at 404. In some embodiments, semantics may be determined at least by parsing the input received at 402 with a large language model (LLM), a visual language model, or a diffusion model that produces hypotheses of complete meshes based at least in part upon some incomplete information pertaining to the 3D representation to be reconstructed. In some of these embodiments, a diffusion model comprises a deep neural network that holds latent variables capable of learning the structure of a given image by removing its blur (i.e. , noise). After a model’s network is trained to “know” or “understand” the concept abstraction behind an input, the model can create new variations of that image. For example, by removing the noise from
an image of a cat, the diffusion model “sees” a clean image of the cat, learns how the cat looks, and applies this knowledge to create new cat image variations.
[00117] At least an incremental portion of the environment may be reconstructed (e.g., via inpainting and/or outpainting) and rendered at 406 at least by using one or more generative models based at least in part upon the semantics determined at 404. In some embodiments, a 3D model with meshes may be provided for the incremental portion. In some of these embodiments, the 3D model may include a diffusion model or a large language model which is a predictive model that receives an input and predicts what comes after the input. A diffusion model includes a forward process and a reverse process where the forward process (e.g., an auto-encoder) incrementally injects noise into a clean sample, and the reverse incrementally removes noise.
[00118] When compared to a GAN (generative adversarial network), a diffusion model provides better distribution coverage of the input signals because GANs are usually biased towards what GANs reconstruct well (and thus may have poorer coverage of the input signals) yet may produce lower quality results. The issues with lower quality results involving diffusion models are addressed in some embodiments by using, for example, 3D Gaussian splatted neural radiance field techniques or the refinement network described herein while GANs may be reserved for training purposes, if any, in some embodiments.
[00119] The incremental portion may be optionally updated at 408 at least by using a separate model in some embodiments. For example, some embodiments may optionally use a generative adversarial network (GAN) for GAN losses as an additional, optional feature by using, for example, the discriminator of GAN to check to see if the
generated environment appears correct in terms of styles, appearances, any correspondence, etc.
[00120] A semantically integrated (e.g., via embedding a semantic vector to a cluster of Gaussians as described herein), manipulatable (e.g., via user interaction), photo-realistic 3D representation in the incremental portion may be modified in real-time or nearly real-time at 410 in response to a user interaction (e.g., an interaction by a user via an XR device). In some embodiments where a SLAM (simultaneous localization and mapping) model is utilized, these embodiments address the shortcomings of SLAMs by further leveraging various neural 3D reconstruction methods to create semantic, high- quality digital twins of a given scene in real-time or nearly real-time. More particularly, SLAM primarily focuses on creating sparse maps for accurate camera localization but not so much in creating photorealistic digital assets.
[00121] In contrast, various neural 3D reconstruction methods described herein (e.g., various semantic segmentation models, diffusion models, large language models, neural radiance field, etc.) capture and reproduce highly detailed, photo-realistic 3D representations. In other words, various techniques described herein combine the benefits of SLAM (e.g., producing sparse maps for accurate camera localization) and the efficient, photo-realistic reconstruction of 3D representations of such neural 3D reconstruction methods to facilitate semantically disjoint, individually and independently manipulatable 3D representations of objects for various applications such as extended reality experiences with XR devices.
[00122] FIG. 5A illustrates a high-level block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon
incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments. More specifically, FIG. 5A illustrates a high-level block diagram of a 3D Gaussian splatted neural radiance field technique that may be used in conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems. As described above, the keyframes (e.g., 316E), camera poses (e.g., 318E), and depth information (e.g., 320E) have been determined prior to invoking the 3D Gaussian splatted neural radiance field model (e.g., 322E in FIG. 3E or that in FIG. 5) so that the aforementioned keyframes, camera poses, and depth information are available to the 3D Gaussian splatted neural radiance field model. In these embodiments, semantically mapped or associated 3D digital scenes, rather than or instead of static 3D meshes that are then composed of digital assets, may be dynamically (e.g., in real-time or nearly realtime) reconstructed at 502.
[00123] A lower-dimensional grouping of Gaussians may be generated at 504. For example, a plurality of Gaussians may be clustered based at least in part upon one or more object identifiers to which the plurality of Gaussians pertains. In some embodiments, the plurality of Gaussians is clustered into one or more lower-dimensional groups before semantic meanings, labels, or vectors are attached to or otherwise associated with any such Gaussians. At 506, semantic meanings, labels, or vectors may be attached to an individual Gaussian in some embodiments or to a lower-dimensional grouping of Gaussians in some other embodiments.
[00124] In these embodiments, 3D object representations generated by various techniques described herein may be semantically disjoint from each other so that a user
wearing an XR device may interact with or manipulate such 3D object representations independently and individually.
[00125] FIGS. 5B-5C illustrate a more detailed block diagram for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields, according to some embodiments. In these embodiments, semantically mapped or associated 3D digital scenes, rather than or instead of static 3D meshes that are then composed of digital assets, may be dynamically (e.g., in real-time or nearly real-time) reconstructed at 502B. A SLAM system may be initialized at 504B as the backend to generate a sparse 3D point set for localization purposes.
[00126] One or more non-keyframes (e.g., frames in the sequence of images or video that are not determined to be keyframes) may 1 be optionally embedded at 506B. One or more camera poses and/or depth data may be estimated at 508B for at least some of the one or more non-keyframes. At least these one or more non-keyframes may be used to augment a training data set at 51 OB for the 3D Gaussian splatted neural radiance field model to increase the number of training samples in some embodiments. In some of these embodiments, respective camera poses and depth information may also be determined for the one or more non-keyframes and associated with or incorporated into the augmented training data set.
[00127] The keyframes (which may be determined from a video or a sequence of images prior to the processing of the 3D Gaussian splatted neural radiance field model) may be analyzed at 512B with a segmentation model that generates a plurality of object segmentation masks at least by performing a segmentation process on a grid (e.g., a 32
x 32 point grid). In some of these embodiments, a plurality of object masks (e.g., three or more) may be generated for each point in the point grid. In addition or in the alternative, a point in the point grid may be selected based on its uncertainty.
[00128] In some other embodiments, rather than generating object segmentation masks by using a point grid, a part-whole hierarchy of masks may be generated at least by analyzing the keyframes with a semantic segmentation model (which will be described in greater details below) at least by using a zero-shot, one-shot, or few-shot tracker neural network or multi-view geometry at 514B. A sparse 3D initial point set may be created at 516B based at least in part upon an inverse projection of the estimated keyframe poses, depth information, and/or depth covariance. For example, points corresponding to lower uncertainty (e.g., uncertainty below a threshold) may be selected and unprojected into 3D using the camera pose or keyframe pose of the keyframe in which the point is identified, and/or depth information and selected as a part of the sparse 3D initial point set at 516B. [00129] The 3D Gaussian splatted neural radiance field model may be improved or optimized at 518B with semantic information. The 3D Gaussian splatted neural radiance field model is thus made semantically aware. For example, a lower-dimensional semantic vector may be optionally attached to each Gaussian at 520B in some embodiments although an alternative process is to attach a semantic vector to a lower-dimensional grouping of a plurality of Gaussians in some other embodiments as described below with reference to 526B. In some embodiments, the lower-dimensional semantic vectors may be optionally learned at 522B using a rendering process during training (e.g., a supervised training) with a loss (e.g., cross-entropy loss).
[00130] For example, the lower-dimensional vectors may be learned at least by projecting the semantic labels onto a 2D space with a linear layer (e.g., a fully connected layer connecting each input to each of a plurality of outputs) to decompress the 3D semantic label, further expanding the number of semantic labels to the number plus one semantic labels, and by performing a softmax function on the number plus one semantic labels into a probabilistic distribution of possible outcomes (a total of the number plus one possible outcomes).
[00131] A plurality of Gaussians may be clustered at 524B into one or more lowerdimensional groupings based at least in part upon one or more object identifiers. For example, Gaussians corresponding to the same object identifier may be clustered into the same lower-dimensional grouping in some embodiments. A semantic vector determined may then be attached at 526B to each cluster of Gaussians at least by running a segmentation model (e.g., a contrastive language-image pre-training or CLIP model) on a plurality of masked image crops.
[00132] A user input may be gathered at 528B to query the 3D representation. A user input may be a gesture, a movement or motion of a part of a user’s body (e.g., a user’s hand squeezing a virtual object, a user’s foot kicking a virtual object, etc.), a textual input, a voice input, and/or an input of an image although input of different formats may be processed with different encoders to generate respective embeddings. For example, a textual input (e.g., by entering textual instructions in natural language) may be processed by a text encoder; an image input may be processed by an image encoder; and a voice input may be processed by a speech encoder or by a combination of a speech-to-text encoder and a textual encoder.
[00133] Similarity (e.g., cosine similarity) among the attached or embedded semantic vectors or labels attached to a plurality of clusters of Gaussians and the user input may be determined at 530B. In some embodiments, this similarity is to correlate the user input with the pertinent portions of the 3D representation to more accurately determine and manipulate the 3D representation in response to the user input.
[00134] One or more objects (or portions thereof) may be selected at 532B in response to the user input based at least in part upon the similarity determined at 530B. The one or more selected objects may be optionally highlighted at 534B with, for example, graphical, textual, or voice emphasis. The one or more objects selected in response to the user input may then be independently, individually manipulated at 536B.
[00135] In some embodiments, inpainting and/or outpainting may be performed at 538B for the 3D scene in response to manipulating the one or more selected objects based at least in part upon the contextual information and/or the user input. For example, a user may use the user input to move a selected virtual object from first location to a second location where the first location used to be occluded by the virtual object. Once the virtual object is moved out of the first location, the first location may be left with a void. [00136] In some embodiments where the first location belongs to the physical environment, the portion of the physical environment corresponding to the void may become visible by the user. In some other embodiments where the first location is part of a 3D representation generated by, for example, an XR device using techniques described herein, the void may be filled with inpainting using techniques at 538B described herein. In some embodiments where the second location is beyond the 3D representation currently displayed to the user (e.g., the virtual object is moved to the second location
that has not been rendered visible to the user), 3D representation of at least the area around the second location may be generated using outpainting techniques at 538B.
[00137] FIG. 6A illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields without pre-training, according to some embodiments. More particularly, FIG. 6A illustrates an example segmentation block diagram that may be used to generate a plurality of masks. In these embodiments, an image 602 may be provided to an image encoder 604 to produce image embeddings 606.
[00138] Most machine learning algorithms take low-dimensional numerical data as inputs. Therefore, input data may be converted into a numerical format. This may involve creating a “bag of words” representation for text data, converting images into pixel values, or transforming graph data into a numerical matrix, etc. Moreover, objects that come into an embedding model are output as embeddings, represented as vectors. A vector includes an array of numbers, where each number indicates where an object is along a specified dimension. The number of dimensions may reach a thousand or more depending on the input data’s complexity. The closer an embedding is to other embeddings in this n-dimensional space, the more similar they are. Distribution similarity may be determined by the length of the vector points from one object to the other (e.g., as measured by Euclidean, cosine or other). Embeddings may be used in various domains and applications due to their ability to transform high-dimensional and categorical data into continuous vector representations, capturing meaningful patterns,
relationships and semantics. Below are a few reasons why embedding is used in data science. By mapping entities (words, images, nodes in a graph, voices, etc.) to vectors in a continuous space, embeddings may capture the semantic relationships and similarities, enabling models to understand and generalize better. An embedding layer is thus commonly used in neural network architectures to map categorical inputs to continuous vectors, facilitating backpropagation and optimization while meaningful relationships in the original input data are preserved.
[00139] In some embodiments, a masked autoencoder may be used for the image encoder 604 where the masked autoencoder masks random patches of the input image 602 and reconstructs the missing pixels. In some of these embodiments, an asymmetric encoder-decoder architecture may be utilized, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image 602 from the latent representation and mask tokens. In some of these embodiments, masking a higher proportion of the input image (e.g., 75%), may yield a nontrivial and meaningful self-supervisory task.
[00140] During pre-training, a large random subset of image patches (e.g., 75%) may be masked out. The masked encoder may be applied to the small subset of visible patches. Mask tokens may be introduced after the masked encoder, and the full set of encoded patches and mask tokens may be processed by a small decoder that reconstructs the original image in pixels. After pre-training, the decoder may be discarded, and the encoder may be applied to uncorrupted images (e.g., a full sets of patches) for recognition tasks in some embodiments.
[00141] The image embeddings 606 may be combined at 608 with one or more masks 612, once these one or more masks 612 are processed by a convolution network 610, and the combined result may be provided to a mask encoder 614. The mask decoder 614 may also receive various other embeddings from a prompt encoder 616. For example, the prompt encoder may receive various prompts such as points 618, boxes 620, textual prompt 622, voice prompt 624, and/or image prompt 626, etc. and generate the respective embeddings therefor. Various prompts may be categorized into two sets - a sparse set including points, boxes, and/or texts, and a dense set including voice and/or image prompts in some embodiments. The sparse set may be represented by positional encodings summed with learned embeddings for each prompt type (e.g., point, box, text) while the dense set of prompts may be embedded using convolutions and summed element-wise with the image embedding 606.
[00142] It shall be noted that different prompts may be processed with different encoders to generate respective embeddings. For example, a textual prompt (e.g., by entering textual instructions in natural language) may be processed by a text encoder; an image prompt may be processed by an image encoder; and a voice prompt may be processed by a speech encoder or by a combination of a speech-to-text encoder and a textual encoder. The mask decoder 614 maps the image embedding 606, prompt embeddings produced by the prompt encoder(s) 616, and an output token to first compute confidence scores 628 respectively associated with corresponding masks. The masks corresponding to sufficiently high confidence scores may be determined to be valid masks
630.
[00143] FIG. 6B illustrates a simplified, example network architecture of a process or system that may be employed as a part of the process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment based at least in part upon incomplete information available to extended-reality systems using 3D Gaussian radiance fields with pre-training, according to some embodiments. In these embodiments, an input image 602B may be identified. The input image includes a plurality of masked image patches or masked image tokens 602_1 B and a plurality of visible image patch or image tokens 602_2B. The plurality of visible image patch or image tokens 602_2B may be provided to an image encoder 604 whose output is then combined with the plurality of masked image patch or image tokens 602_1 B to form the image imbedding 606 with the respective positional information corresponding the plurality of image patches or tokens (both masked and visible image patches or tokens).
[00144] The image imbedding 606 is provided to a decoder 650 that reconstructs, for pre-training purposes, the image embeddings without masks 652B that may then be used to reconstruct the original image from the reconstructed image embeddings 652B, also for training purposes. For example, the reconstructed image 654B may be provided to a loss mechanism or a discriminator that determines the loss therebetween and uses the loss to fine tune various components (e.g., the encoder 604 and/or the decoder 650). [00145] The image embeddings 606 may be combined with one or more masks 612B after these one or more masks have been processed by a convolution network 610B. The combined result may be provided to a mask decoder 614B. The mask decoder 614B further receives prompt embeddings from the prompt encoder 616B to generate confidence a score for each of a plurality of masks 612B. The masks with sufficiently high
confidence scores (e.g., above a threshold confidence score) may be determined to be value masks 630B. As described above, various prompts may be categorized into two sets - a sparse set including points 618B, boxes 620B, and/or texts 622B, and a dense set including voice 624B and/or image prompts 626B in some embodiments. The sparse set may be represented by positional encodings summed with learned embeddings for each prompt type (e.g., point, box, text) while the dense set of prompts may be embedded using convolutions and summed element-wise with the image embedding 606.
[00146] FIG. 7A illustrates s simplified schematic diagram for a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments. More specifically, FIG. 7A illustrates a model 700 that jointly trains an image encoder, a text encoder, and a voice / speech encoder to predict the correct pairings of a batch of (image, text), (image, voice), or (image, text, voice) training examples. A learned text encoder synthesizes a zero-shot classifier by embedding the names or descriptions of a target dataset’s classes, and a learned voice / speech encoder synthesizes a zero-shot classifier by embedding the voice or speech of a target dataset’s classes.
[00147] In these embodiments, an image encoder 704 may process one or more images 702 into one or more corresponding sets of image embeddings 706 each having a plurality of values in a continuous space. Similarly, a text encoder 710 may process one or more text inputs 708 into one or more corresponding sets of text embeddings 720 each having a plurality of values in a continuous space. Further, speech encoder 716 may process one or more speech inputs 714 into one or more corresponding sets of speech
embeddings 718 each having a plurality of values in a continuous space. The image embeddings 706, text embeddings 720, and speech encodings 718 may be used to form a multi-dimensional data structure 712 although a two-dimensional table-like data structure is shown in FIG. 7A for the ease of illustration.
[00148] In some embodiments, a speech encoder may be replaced by a speech-to- text model that transforms a speech input 71 into text input which may then be processed by the text encoder 710. These embodiments thus set aside the need for a separate speech encoder although at the additional expense of having a speech-to-text model.
[00149] The various text inputs 708 and speech inputs 714 comprise natural language inputs to leverage the advantages of learning perception from supervision in natural language such as easier to scale than crowd-sourced labeling (e.g.., ImageNet) because natural language does not require annotations to be in a machine learning compatible format (e.g., a canonical one-of-N majority vote “gold label”). Further, learning from natural language does not simply learn a representation but also connects a specific representation to natural language that enables the more flexible zero-shot knowledge transfer.
[00150] Moreover, the network 700 pre-trains an image encoder (704), a text encoder (708), and a speech encoder (714) to predict which images were paired with which texts and/or speech in the dataset. The model then uses this behavior to turn into a zero-shot classifier and converts all of a dataset’s classes into captions and/or speeches such as “a photo of a dog” and predict the class that is estimated to best pair with a given image.
[00151] FIG. 7B illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments. In these embodiments, the circuit 721 shows an image encoder 704 receives an image 702 and determines an image embedding 730 for the received image 702, and the image embedding includes a vector form having a plurality of values in a continuous space.
[00152] The circuit 721 shows a text encoder 710 tasked with a task 724 for pairing an image with its description and a plurality of texts 722_1 , 722_2, 722_3, 722_4, ... , 722_n. The text encoder 710 determines the textual embedding 726 for each of the texts (722_1 , 722_2, 722_3, 722_4, ... , 722_n), each having a vector form having a plurality of values in a continuous space.
[00153] The circuit 738 shows a speech encoder 716 tasked with a task 725 for pairing an image with its corresponding speech and a plurality of speech segments 736_1 , 736_2, 736_3, 736_4, ... , 736_n. The text encoder 710 determines the speech embedding 740 for each of the speech segments (736_1 , 736_2, 736_3, 736_4, ... , 736_n), each having a vector form having a plurality of values in a continuous space.
[00154] The image encoder 704, the text encoder 710, and the speech encoder 716 may be jointly trained to determine the correct pairings of (image, text), (image, text, speech), and/or (image, text, speech). These three types of embeddings may be arranged in one or more multi-dimensional data structure. In practice, when a prompt is provided for one of the three inputs, the remaining two may be identified by querying the one or more multi-dimensional data structure to find the correct pairing(s) of (image, text), (image, text, speech), and/or (image, text, speech).
[00155] For example, a prompt of an image may be received at the image encoder 704, the image encoder determines the image embedding 732 for the received image 702 from a plurality of image embeddings 732. The text embedding and the speech embedding may then be used to query the one or more multi-dimensional data structure to identify the best match of text encoding and speech encoding for the prompted image 702. The text embedding and the speech embedding may be respectively processed to produce the textual description and the verbal description that may then be associated with or attached to the received image. It shall be noted that an image is referenced as the input in the aforementioned example where textual and speech embeddings and descriptions are subsequently determined, and that this example is presented for the sole purpose of demonstrating the functionality of these techniques but shall not be considered as limiting the scope of other embodiments or examples where the same process may receive a text prompt, a voice prompt, or a combination prompt of more than one type of the aforementioned types.
[00156] FIG. 70 illustrates more details about the simplified schematic diagram for the process or system illustrated in FIG. 7A, according some embodiments. More specifically, FIG. 7C illustrates an example block diagram for an image encoder 700C that may be used in a zero-shot (or one-shot or few-shot) image classification model that learns by associating text and/or speech with images. In these embodiments, an image encoder 700C may include an input stem 702C and four subsequent stages 704C (stage 1 ), 706C (stage 2), 708C (stage 3), and 7100 (stage 4) as well as the final output layer
712C.
[00157] In some embodiments, the input stem 702C reduces the input width and height and increases its channel size by using a network comprising an N x N convolution
716C with an output channel and a stride followed by a pooling layer 714C. For example, the input stem 702C may include 7 x 7 convolutions with an output channel of 64 with a stride of two (“2”) followed by a 3 x 3 max pooling layer with a stride of 2 in some embodiments. In these embodiments, the input stem 702C increases the input channel size to 64 and reduces the input width and height by four times.
[00158] Among the four stages 704C, 706C, 708C, and 710C, the second stage 706C and on begins with a down sampling layer 716C that is followed by several residual blocks (e.g., 718C and 720C). In some embodiments, the down sampling layer 716C may have two paths where the first path includes three convolution layers 722C, 724C, and 726C, and the second path includes a single convolution layer 728C. In some of these embodiments, the convolution layer 722C may include a kernel size of 1 x 1 with a stride of 2 (to half the input width and height) and a channel size of 512; the convolution layer 724C may include a kernel size of 3 x 3 with a channel size of 512.
[00159] The convolution layer 726C may include a kernel size of 1 x 1 with a channel size of 2048 (which is four times larger than those of 724C and 722C); and the convolution layer 728C may include a kernel size of 1 x 1 with a stride of 2 and a channel size of 2048. The second path is to transform the input shape to be the output shape of the first path to that the outputs of the first path and the second path may be summed together to determine the output of the down sampling block 716C. In some embodiments, a residual block is similar in architecture to the down sampling block 716C with the exception that the residual block has a stride of one (“1”).
[00160] FIG. 7D illustrates some example modifications of the simplified schematic diagram illustrated in FIG. 7C, according to some embodiments. In some embodiments, a down sampling layer 716C may include the architecture 716D that is identical to 716C described above with reference to FIG. 7C. In some embodiments, a down sampling layer 716C may include the architecture 716D1 that also has a first path and a second path. The first path include a first convolution layer 730D having a kernel size of 1 x 1 , a second convolution layer 732D having a kernel size of 3 x 3 with a stride of 2, and a third convolution layer 734D having a kernel size of 1 x 1 . The second path includes a single convolution layer 736D having a kernel size of 1 x 1 with a stride of 2.
[00161] In 716D, the first convolution in the first path starting with 722C ignores three-quarters of the input feature map because it uses a kernel size 1 *1 with a stride of 2. 716D1 switches the strides size of the first two convolutions (730D and 732D) in the first path so no information is ignored. Because the second convolution 732D has a kernel size 3 x 3 with a stride of 2, the output shape of the first path remains unchanged.
[00162] In some embodiments, a down sampling layer 716C may include or be replaced with the architecture 716D2 that also has a first path and a second path. The first path include the same first convolution layer 730D having a kernel size of 1 x 1 , the same second convolution layer 732D having a kernel size of 3 x 3 with a stride of 2, and the same third convolution layer 734D having a kernel size of 1 x 1. The second path includes a pooling layer 738D having a kernel size of 2 x 2 with a stride of 2 followed by a convolution layer 740D having a kernel size of 1 x 1 with a stride of 2.
[00163] In these embodiments, adding a 2 x 2 pooling layer (e.g., an average pooling layer) with a stride of 2 before the convolution 740D, whose stride is changed to
1 , works well in practice and impacts the computational cost little.
[00164] In some embodiments, a down sampling layer 716C may include or be replaced with the architecture 716D3 that has a single path including a first convolution layer 742D having a kernel size of 3 x 3 with a channel size of 32 and a stride of 2, a second convolution layer 744D having a kernel size of 3 x 3 a channel size of 32 with a stride of 2, and a third convolution layer 746D having a kernel size of 3 x 3 with a channel size of 64 and a stride of 2 followed by a pooling layer 748D having a kernel size of 3 x 3 with a stride of 2.
[00165] In these embodiments, the computational cost of a convolution is quadratic to the kernel width or height. A 7 x 7 convolution is 5.4 times more expensive than a 3 x 3 convolution. So this down sampling layer 716D3 replacing the 7 x 7 convolution in the input stem with three conservative 3 x 3 convolutions with the first and second convolutions have their output channel of 32 and a stride of 2, while the last convolution uses a 64-output channel.
[00166] FIG. 7E illustrates a simplified block diagram for an example image encoder that may be employed in FIG. 7A or 7B, according to some embodiments. More specifically, FIG. 7E illustrates another example block diagram for an image encoder that may be used in a zero-shot (or one-shot or few-shot) image classification model that learns by associating text and/or speech with images. In these embodiments, an image encoder may split an image into a plurality of patches (e.g., a plurality of fixed and/or variable size patches), linearly combine the plurality of patches, add position embeddings,
and feed the resulting sequence of vectors in the resulting embeddings to a transformer encoder.
[00167] More specifically, an image 702E may be split into a plurality of patches 704E. In some embodiments, an image may be split into a plurality of fixed-size patches while in some other embodiments, an image may be split into a plurality of variable-sized patches. The plurality of patches 704E may be provided to a projection module 706E which projects the plurality of patches into projected patches 708E at least by combining or otherwise associating the respective positions (in the original image 702E) with the plurality of patches 704E where the numbers in the projected patches 708E respectively indicate the positions of the projected patches in the original image 702E. In addition, an extra classification token 710E may be added to the projected patches 708E to enable the performance of classification.
[00168] The extra classification token 710E and the projected patches 708E may be provided to a transformer encoder 712E. For example, the extra classification token 710E and the projected patches 708E may be provided to a normalization layer 714E in the transformer encoder 712E that performs batch normalization or group normalization on the extra classification token 710E and the projected patches 708E. The normalized input may be forwarded to an attention layer 716E. in some embodiments, the attention layer 716E includes a multi-head attention layer such as a multi-head self-attention layer.
[00169] A multi-head attention layer includes a module of multiple attention mechanisms which runs through an attention mechanism several times in parallel. The independent attention outputs may then be concatenated and linearly transformed into the expected dimension. In some of these embodiments, multiple attention heads allows
for attending to parts of the sequence differently (e.g. longer-term dependencies versus shorter-term dependencies).
[00170] The output of the attention layer 716A may be combined with the extra classification token 710E and the projected patches 708E, which are also separately provided to the combiner. The output of the combiner may be forwarded to both a normalization layer 718E and another combiner which follows a separate multi-layer perceptron (MLP) 720E receiving the output of the normalization layer 718E. the output the MLP may be combined with the output of the previous combiner following the attention layer 716E to generate a combined output that is in turn forwarded to a classification 722E that produces a plurality of classes 724E.
[00171] In some embodiments, a multilayer perceptron (MLP) comprises a modern feed-forward artificial neural network, including, for example, fully connected neurons with a nonlinear kind of activation function, organized in at least three layers, notable for being able to distinguish data that is not linearly separable.
[00172] FIG. 7F illustrates a simplified block diagram for an example text encoder that may be employed in FIG. 7A or 7B, according to some embodiments. In these embodiments, the text encoder 720 includes an encoding portion that receives an input 702F (e.g., textual input) at an input embedding module 704F that generates text embeddings for the input 702F. The text embeddings may be combined or associated with the position embeddings 706F at a combiner that forwards its combined output to an attention mechanism 708F as well as a normalization layer 71 OF that follows the attention mechanism 708F.
[00173] The normalization layer 71 OF performs a batch normalization or group normalization and transmits its output to a position-wise feed forward layer and a normalization layer 714F that also receives the output of the position-wise feed forward layer 712F. The output of the normalization layer 714F is transmitted to an attention mechanism 726F in the decoding portion of the text encoder 720 that is described immediately below.
[00174] The above encoding portion of the text encoder 720 may include a stack of N identical layers (e.g., six identical layers in some embodiments) where each layer has two sub-layers. The first sublayer includes a multi-head self-attention mechanism, and the second sublayer includes a simple, position-wise fully connected feed-forward network. Soe embodiments employ a residual connection around each of the two sublayers, followed by layer normalization. In these embodiments, the output of each sublayer is LayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer itself. To facilitate these residual connections, all sub-layers in the model, as well as the embedding layers, produce outputs of dimension d model = 512 in some of these embodiments.
[00175] The text encoder 720 may also include the decoding portion that receives output 716F at an output embedding module 718F that generates embeddings for the output 716F. These embeddings, together with the positional embeddings generated by a positional encoding module 720F for the output 716F, may be combined at a combiner whose output is then provided to a masked attention mechanism 722F and a normalization layer 724F following and receiving the output of the masked attention mechanism 722F.
[00176] The normalization layer 724F performs a batch normalization or group normalization to generate normalized output that may then be sent to the attention mechanism 726F as well as the normalization layer 728F following and receiving the output of the attention mechanism 726F in the decoding portion of the text encoder 720. [00177] The normalization layer 728F performs a batch normalization or group normalization to generate normalized output that may be provided to a position-wise feed forward layer 730F and a normalization layer 730F that follows and receives the output of the position-wise feed forward layer 730F. The normalization layer 732F performs a batch normalization or group normalization on the input to generate normalized output that may be provided to a linear layer 734F that is followed by a softmax layer 736F. The softmax layer 736F performs a softmax operation to transform values in a continuous space into probabilistic distributions of multiple possible outcomes 738F.
[00178] In some of these embodiments, the decoding portion may also include a stack of N identical layers (e.g., six identical layers). In addition to the two sub-layers in each layer in the encoding portion, the decoding portion inserts a third sub-layer, which performs multi-head attention over the output of the encoder stack. Similar to the encoding portion, residual connections may be employed around each of the sub-layers, followed by layer normalization. The self-attention sub-layer in the decoder stack may be modified to prevent positions from attending to subsequent positions in some of these embodiments. This masking, combined with the fact that the output embeddings may be offset by one position, ensures that the predictions for position i may depend only on the known outputs at positions less than i.
[00179] In these embodiments, an attention mechanism performs an attention function which may be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value may be computed by a compatibility function of the query with the corresponding key.
[00180] FIGS. 7G-7H illustrate an example attention mechanism that may be employed for a sequence transduction module in a process or system for conditional, probabilistic generation of three-dimensional (3D) virtual environment using 3D Gaussian radiance fields based at least in part upon natural language queries or prompts, according to some embodiments. In these embodiments, the example attention mechanism 700G may receive a query matrix 702G and a key matrix 704G at an inner product module 706G that performs an inner product on the query matrix 702G and the key matrix 704G. [00181] The query matrix 702G, in encoding or decoding attention layers, may come from the previous decoder layer. The keys 704G and values 716G may come from the output of the encoding portion. This allows every position in the decoder to attend over all positions in the input sequence to mimic the typical encoder-decoder attention mechanisms in sequence-to-sequence models.
[00182] In some embodiments, an encoder includes one or more self-attention layers. In a self-attention layer, all of the keys, values, and queries may come from the same place (e.g., the output of the previous layer in the encoding portion). Each position in the encoder may attend to all positions in the previous layer of the encoding portion.
[00183] A Self-attention layer in the decoding portion allows each position in the decoder to attend to all positions in the decoder up to and including that position. Leftward
information flow in the decoder may need to be prevented to preserve the auto-regressive property, and this may be implemented inside of a scaled dot-product attention by masking out (e.g. , by setting to -°°) all values in the input of the softmax which correspond to illegal connections.
[00184] The output of the inner product module 706G may be provided to an optional scaling module 708G that scales the output of the inner product module 706G (e.g., scale the output by the square root of the dimensionality of the queries or the keys, d_k) and/or an optional masking module 710G that optionally masks out values in the input to the normalization module 712G that correspond to illegal connections, if any. The normalization module 712G normalizes the output of the inner product module 706G (or that of the optional scaling module and/or that of the optional masking module 710G) into a probabilistic distribution of several possible outcomes in order to determine the weights on the values. In some embodiments, the normalization module 712G may perform a softmax function that converts a vector of N real numbers into a probability distribution of N possible outcomes.
[00185] The output of the normalization module 712G may be provided to another inner product module 714G that performs an inner product operation on the output of the normalization module 712G and the value matrix 716G.
[00186] FIG. 7H illustrates a text encoder with multi-head attention, according to some embodiments. In these embodiments, multiple example attention mechanisms 700G may be run in parallel where each example attention mechanism 700G receives projected query matrices 702G, projected key matrices 704G, and projected value matrices 716G. More specifically, a query matrix 702G may be linearly projected to, for
example, d_k dimensions 706H (e.g., 64 or 128); a key matrix 708G may be linearly projected to d_k dimensions 708H (e.g., 64 or 128), and a value matrix 716G may be linearly projected to d_v dimensions 71 OH.
[00187] These projected matrices may be provided to respective example attention mechanisms 700G that perform the attention function in parallel to produce d_v dimensional output values. These d_v dimensional output values may be forwarded and concatenated by the concatenation module 712H which may then be projected again at 714H, resulting in the final values.
[00188] FIG. 8 illustrates a simplified example of a wearable extended-reality (XR) device with a belt pack external to the XR glasses in some embodiments. More specifically, FIG. 8 illustrates a simplified example of a user-wearable VR (virtual reality) / AR (augmented reality) I MR (mixed reality) /XR (extended reality) system that includes an optical sub-system 802 and an external module 804 and may include multiple instances of personal augmented reality systems, for example a respective personal augmented reality system for a user. Any of the neural networks, module, processes, services, microservices, etc. described herein may be embedded in whole or in part in or on the wearable XR device. For example, some or all of a neural network described herein as well as other peripherals (e.g., ToF or time-of-flight sensors) may be embedded in the external module 804 alone, the optical sub-system 802 alone, or distributed between the external module 804 and the optical sub-system 802. In some embodiments, the external module 804 may include a rechargeable battery that is operatively connected to the optical sub-system 802 in order to provide power to the optical sub-system 802.
[00189] Some embodiments of the VR/AR/MR/XR system may comprise optical sub-system 802 that delivers virtual content to the user’s eyes as well as the external module 804 that performs a multitude of processing tasks to present the relevant virtual content to a user. The external module 804 may, for example, take the form of the belt pack, which can be convenience coupled to a belt or belt line of pants during use. Alternatively, the external module 804 may, for example, take the form of a personal digital assistant or smartphone type device.
[00190] The external module 804 may include one or more processors, for example, one or more micro-controllers, microprocessors, graphical processing units, digital signal processors, application specific integrated circuits (ASICs), programmable gate arrays, programmable logic circuits, or other circuits either embodying logic or capable of executing logic embodied in instructions encoded in software or firmware. The external module 804 may include one or more non-transitory computer- or processor-readable media, for example volatile and/or nonvolatile memory, for instance read only memory (ROM), random access memory (RAM), static RAM, dynamic RAM, Flash memory, EEPROM, etc.
[00191] The external module 804 may be communicatively coupled to the head worn component. For example, the external module 804 may be communicatively tethered to the head worn component via one or more wires or optical fibers via a cable with appropriate connectors. The external module 804 and the optical sub-system 802 may communicate according to any of a variety of tethered protocols, for example UBS®, USB2®, USB3®, USB-C®, Ethernet®, Thunderbolt®, Lightning® protocols.
[00192] Alternatively or additionally, the external module 804 may be wirelessly communicatively coupled to the head worn component. For example, the external module
804 and the optical sub-system 802 may each include a transmitter, receiver or transceiver (collectively radio) and associated antenna to establish wireless communications there between. The radio and antenna(s) may take a variety of forms. For example, the radio may be capable of short-range communications, and may employ a communications protocol such as BLUETOOTH®, WI-FI®, or some IEEE 802.11 compliant protocol (e.g., IEEE 802.11 n, IEEE 802.11 a/c). Various other details of the processing sub-system and the optical sub-system are described in U.S. Pat. App. Ser. No. 14/707,000 filed on May 08, 2015 and entitled “EYE TRACKING SYSTEMS AND METHOD FOR AUGMENTED OR EXTENDED-REALITY”, the content of which is hereby expressly incorporated by reference in its entirety for all purposes.
[00193] It shall be noted that any of the subcomponents described above for the external module 804 may be alternatively incorporated into the optical sub-system 802 while the external module 804 includes a battery to provide power to the optical subsystem 802 as well as the modules therein.
[00194] COMPUTER SYSTEM ARCHITECTURE OVERVIEW
[00195] FIG. 9 illustrates a computerized system on which a method for management of extended-reality systems or devices may be implemented. Computer system 900 includes a bus 906 or other communication module for communicating information, which interconnects subsystems and devices, such as processor 907, system memory 908 (e.g., RAM), static storage device 909 (e.g., ROM), disk drive 910 (e.g., magnetic or optical), communication interface 914 (e.g., modem or Ethernet card),
display 911 (e.g., CRT or LCD), input device 912 (e.g., keyboard), and cursor control (not shown). The illustrative computing system 900 may include an Internet-based computing platform providing a shared pool of configurable computer processing resources (e.g., computer networks, servers, storage, applications, services, etc.) and data to other computers and devices in a ubiquitous, on-demand basis via the Internet. For example, the computing system 900 may include or may be a part of a cloud computing platform in some embodiments.
[00196] According to one embodiment, computer system 900 performs specific operations by one or more processor or processor cores 907 executing one or more sequences of one or more instructions contained in system memory 908. Such instructions may be read into system memory 908 from another computer readable/usable storage medium, such as static storage device 909 or disk drive 910. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions to implement the invention. Thus, embodiments of the invention are not limited to any specific combination of hardware circuitry and/or software. In one embodiment, the term “logic” shall mean any combination of software or hardware that is used to implement all or part of the invention.
[00197] Various actions or processes as described in the preceding paragraphs may be performed by using one or more processors, one or more processor cores, or combination thereof 907, where the one or more processors, one or more processor cores, or combination thereof executes one or more threads. For example, various acts of determination, identification, synchronization, calculation of graphical coordinates, rendering, transforming, translating, rotating, generating software objects, placement,
assignments, association, etc. may be performed by one or more processors, one or more processor cores, or combination thereof.
[00198] The term “computer readable storage medium” or “computer usable storage medium” as used herein refers to any non-transitory medium that participates in providing instructions to processor 907 for execution. Such a medium may take many forms, including but not limited to, non-volatile media and volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as disk drive 910. Volatile media includes dynamic memory, such as system memory 908. Common forms of computer readable storage media includes, for example, electromechanical disk drives (such as a floppy disk, a flexible disk, or a hard disk), a flash-based, RAM-based (such as SRAM, DRAM, SDRAM, DDR, MRAM, etc.), or any other solid-state drives (SSD), magnetic tape, any other magnetic or magneto-optical medium, CD-ROM, any other optical medium, any other physical medium with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read.
[00199] In an embodiment of the invention, execution of the sequences of instructions to practice the invention is performed by a single computer system 900. According to other embodiments, two or more computer systems 900 coupled by communication link 915 (e.g., LAN, PTSN, or wireless network) may perform the sequence of instructions required to practice the invention in coordination with one another.
[00200] Computer system 900 may transmit and receive messages, data, and instructions, including program (e.g., application code) through communication link 915
and communication interface 914. Received program code may be executed by processor 907 as it is received, and/or stored in disk drive 910, or other non-volatile storage for later execution. In an embodiment, the computer system 900 operates in conjunction with a data storage system 931 , e.g., a data storage system 931 that includes a database 932 that is readily accessible by the computer system 900. The computer system 900 communicates with the data storage system 931 through a data interface 933. A data interface 933, which is coupled to the bus 906 (e.g., memory bus, system bus, data bus, etc.), transmits and receives electrical, electromagnetic or optical signals that include data streams representing various types of signal information, e.g., instructions, messages and data. In embodiments of the invention, the functions of the data interface 933 may be performed by the communication interface 914.
[00201] FIG. 9 shows an example architecture 2500 for the electronics operatively coupled to an optics system or XR device in one or more embodiments. The optics system or XR device itself or an external device (e.g., a belt pack) coupled to the or XR device may include one or more printed circuit board components, for instance left (2502) and right (2504) printed circuit board assemblies (PCBA). As illustrated, the left PCBA 2502 includes most of the active electronics, while the right PCBA 604supports principally supports the display or projector elements.
[00202] The right PCBA 2504 may include a number of projector driver structures which provide image information and control signals to image generation components. For example, the right PCBA 2504 may carry a first or left projector driver structure 2506 and a second or right projector driver structure 2508. The first or left projector driver structure 2506 joins a first or left projector fiber 2510 and a set of signal lines (e.g., piezo
driver wires). The second or right projector driver structure 2508 joins a second or right projector fiber 2512 and a set of signal lines (e.g., piezo driver wires). The first or left projector driver structure 2506 is communicatively coupled to a first or left image projector, while the second or right projector drive structure 2508 is communicatively coupled to the second or right image projector.
[00203] In operation, the image projectors render virtual content to the left and right eyes (e.g. , retina) of the user via respective optical components, for instance waveguides and/or compensation lenses to alter the light associated with the virtual images.
[00204] The image projectors may, for example, include left and right projector assemblies. The projector assemblies may use a variety of different image forming or production technologies, for example, fiber scan projectors, liquid crystal displays (LCD), LCOS (Liquid Crystal On Silicon) displays, digital light processing (DLP) displays. Where a fiber scan projector is employed, images may be delivered along an optical fiber, to be projected therefrom via a tip of the optical fiber. The tip may be oriented to feed into the waveguide. The tip of the optical fiber may project images, which may be supported to flex or oscillate. A number of piezoelectric actuators may control an oscillation (e.g., frequency, amplitude) of the tip. The projector driver structures provide images to respective optical fiber and control signals to control the piezoelectric actuators, to project images to the user’s eyes.
[00205] Continuing with the right PCBA 2504, a button board connector 2514 may provide communicative and physical coupling to a button board 2516 which carries various user accessible buttons, keys, switches or other input devices. The right PCBA 2504 may include a right earphone or speaker connector 2518, to communicatively
couple audio signals to a right earphone 2520 or speaker of the head worn component. The right PCBA 2504 may also include a right microphone connector 2522 to communicatively couple audio signals from a microphone of the head worn component. The right PCBA 2504 may further include a right occlusion driver connector 2524 to communicatively couple occlusion information to a right occlusion display 2526 of the head worn component. The right PCBA 2504 may also include a board-to-board connector to provide communications with the left PCBA 2502 via a board-to-board connector 2534 thereof.
[00206] The right PCBA 2504 may be communicatively coupled to one or more right outward facing or world view cameras 2528 which are body or head worn, and optionally a right cameras visual indicator (e.g., LED) which illuminates to indicate to others when images are being captured. The right PCBA 2504 may be communicatively coupled to one or more right eye cameras 2532, carried by the head worn component, positioned and orientated to capture images of the right eye to allow tracking, detection, or monitoring of orientation and/or movement of the right eye. The right PCBA 2504 may optionally be communicatively coupled to one or more right eye illuminating sources 2530 (e g., LEDs), which as explained herein, illuminates the right eye with a pattern (e.g., temporal, spatial) of illumination to facilitate tracking, detection or monitoring of orientation and/or movement of the right eye.
[00207] The left PCBA 2502 may include a control subsystem, which may include one or more controllers (e.g., microcontroller, microprocessor, digital signal processor, graphical processing unit, central processing unit, application specific integrated circuit (ASIC), field programmable gate array (FPGA) 2540, and/or programmable logic unit
(PLU)). The control system may include one or more non-transitory computer- or processor readable medium that stores executable logic or instructions and/or data or information. The non-transitory computer- or processor readable medium may take a variety of forms, for example volatile and nonvolatile forms, for instance read only memory (ROM), random access memory (RAM, DRAM, SD-RAM), flash memory, etc. The non- transitory computer or processor readable medium may be formed as one or more registers, for example of a microprocessor, FPGA or ASIC.
[00208] The left PCBA 2502 may include a left earphone or speaker connector 2536, to communicatively couple audio signals to a left earphone or speaker 2538 of the head worn component. The left PCBA 2502 may include an audio signal amplifier (e.g., stereo amplifier) 2542, which is communicative coupled to the drive earphones or speakers. The left PCBA 2502 may also include a left microphone connector 2544 to communicatively couple audio signals from a microphone of the head worn component. The left PCBA 2502 may further include a left occlusion driver connector 2546 to communicatively couple occlusion information to a left occlusion display 2548 of the head worn component.
[00209] The left PCBA 2502 may also include one or more sensors or transducers which detect, measure, capture or otherwise sense information about an ambient environment and/or about the user. For example, an acceleration transducer 2550 (e.g., three axis accelerometer) may detect acceleration in three axes, thereby detecting movement. A gyroscopic sensor 2552 may detect orientation and/or magnetic or compass heading or orientation. Other sensors or transducers may be similarly employed.
[00210] The left PCBA 2502 may be communicatively coupled to one or more left outward facing or world view cameras 2554 which are body or head worn, and optionally a left cameras visual indicator (e.g., LED) 2556 which illuminates to indicate to others when images are being captured. The left PCBA may be communicatively coupled to one or more left eye cameras 2558, carried by the head worn component, positioned and orientated to capture images of the left eye to allow tracking, detection, or monitoring of orientation and/or movement of the left eye. The left PCBA 2502 may optionally be communicatively coupled to one or more left eye illuminating sources (e.g., LEDs) 2556, which as explained herein, illuminates the left eye with a pattern (e.g., temporal, spatial) of illumination to facilitate tracking, detection or monitoring of orientation and/or movement of the left eye.
[00211] The PCBAs 2502 and 2504 are communicatively coupled with the distinct computation component (e.g., belt pack) via one or more ports, connectors and/or paths. For example, the left PCBA 2502 may include one or more communications ports or connectors to provide communications (e.g., bi-directional communications) with the belt pack. The one or more communications ports or connectors may also provide power from the belt pack to the left PCBA 2502. The left PCBA 2502 may include power conditioning circuitry 2580 (e.g., DC/DC power converter, input filter), electrically coupled to the communications port or connector and operable to condition (e.g., step up voltage, step down voltage, smooth current, reduce transients).
[00212] The communications port or connector may, for example, take the form of a data and power connector or transceiver 2582 (e.g., Thunderbolt® port, USB® port). The right PCBA 2504 may include a port or connector to receive power from the belt pack.
The image generation elements may receive power from a portable power source (e.g., chemical battery cells, primary or secondary battery cells, ultra-capacitor cells, fuel cells), which may, for example be located in the belt pack.
[00213] As illustrated, the left PCBA 2502 includes most of the active electronics, while the right PCBA 2504 supports principally supports the display or projectors, and the associated piezo drive signals. Electrical and/or fiber optic connections are employed across a front, rear or top of the body or head worn component of the optics system or XR device. Both PCBAs 2502 and 2504 are communicatively (e.g., electrically, optically) coupled to the belt pack. The left PCBA 2502 includes the power subsystem and a highspeed communications subsystem. The right PCBA 2504 handles the fiber display piezo drive signals. In the illustrated embodiment, only the right PCBA 2504 needs to be optically connected to the belt pack. In other embodiments, both the right PCBA and the left PCBA may be connected to the belt pack.
[00214] While illustrated as employing two PCBAs 2502 and 2504, the electronics of the body or head worn component may employ other architectures. For example, some implementations may use a fewer or greater number of PCBAs. As another example, various components or subsystems may be arranged differently than illustrated in FIG. 9. For example, in some alternative embodiments some of the components illustrated in FIG. 9 as residing on one PCBA may be located on the other PCBA, without loss of generality.
[00215] An optics system or an XR device described herein may present virtual contents to a user so that the virtual contents may perceived as three-dimensional contents in some embodiments. In some other embodiments, an optics system or XR
device may present virtual contents in a four- or five-dimensional lightfield (or light field) to a user.
[00216] In the foregoing specification, the invention has been described with reference to specific embodiments thereof. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention. For example, the above-described process flows are described with reference to a particular ordering of process actions. However, the ordering of many of the described process actions may be changed without affecting the scope or operation of the invention. The specification and drawings are, accordingly, to be regarded in an illustrative rather than restrictive sense.
Claims
1 . An extended reality device, comprising: a wearable eyepiece that presents virtual contents to a user; a belt pack operatively coupled to the wearable eyepiece; a processor; and a non-transitory computer readable medium storing thereupon a sequence of instructions which, when executed by a model with the processor, causes the processor to perform a set of acts, the set of acts executed by the model comprising: estimating a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device; generating a semantically annotated, manipulatable three-dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece; and modifying the semantically annotated, manipulatable 3D representation in realtime or nearly real-time in response to a user interaction.
2. The extended reality device of claim 2, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further comprising:
receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; and determining, from the video or the sequence of RGB images, the plurality of keyframes; and estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images.
3. The extended reality device of claim 2, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further comprising: performing a global bundle adjustment that corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
4. The extended reality device of claim 2, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further comprising: identifying a non-keyframe from the video or the sequence of RGB images; determining a non-keyframe pose for the non-keyframe; and augmenting a training data set at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
5. The extended reality device of claim 1 , the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising generating the semantically annotated, manipulatable three-dimensional (3D) representation, the set of acts executed by the model further comprising: generating a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
6. The extended reality device of claim 5, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising generating the semantically annotated, manipulatable three-dimensional (3D) representation, the set of acts executed by the model further comprising: establishing a semantic association among the plurality of object segmentation masks at least by using a zero-shot, one-shot, or few-shot neural network or a multi-view geometry; and generating a three-dimensional (3D) point set based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
7. The extended reality device of claim 6, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising generating the semantically annotated, manipulatable three-dimensional (3D) representation, the set of acts executed by the model further comprising:
improving 3D Gaussian splatted neural radiance field at least by attaching a lowerdimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three-dimensional (3D) representation.
8. A method for an extended reality system, comprising: estimating a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device; generating a semantically annotated, manipulatable three-dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece; and modifying the semantically annotated, manipulatable 3D representation in realtime or nearly real-time in response to a user interaction.
9. The method of claim 8, estimating the plurality of keyframes, the plurality of keyframe poses further comprising: receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; and determining, from the video or the sequence of RGB images, the plurality of keyframes; estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images; and
performing a global bundle adjustment that corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
10. The method of claim 8, estimating the plurality of keyframes, the plurality of keyframe poses further comprising: identifying a non-keyframe from the video or the sequence of RGB images; determining a non-keyframe pose for the non-keyframe; and augmenting a training data set at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
11. The method of claim 8, generating the semantically annotated, manipulatable three-dimensional (3D) representation comprising: generating a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
12. The method of claim 11 , generating the semantically annotated, manipulatable three-dimensional (3D) representation comprising: establishing a semantic association among the plurality of object segmentation masks at least by using a zero-shot, one-shot, or few-shot neural network or a multi-view geometry; and generating a three-dimensional (3D) point set based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
13. The method of claim 12, generating the semantically annotated, manipulatable three-dimensional (3D) representation comprising:
improving 3D Gaussian splatted neural radiance field at least by attaching a lowerdimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three-dimensional (3D) representation.
14. A computer program product comprising a non-transitory computer readable medium storing thereupon a sequence of instructions which, when executed by a model with the processor, causes the processor to perform a set of acts, the set of acts executed by the model comprising: estimating a plurality of keyframes, a plurality of keyframe poses, and depth data from a plurality of captures that is captured by at least the extended reality device; generating a semantically annotated, manipulatable three-dimensional (3D) representation in a physical environment for perception by the user wearing the wearable eyepiece; and modifying the semantically annotated, manipulatable 3D representation in realtime or nearly real-time in response to a user interaction.
15. The computer program product of claim 14, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further comprising:
receiving a video or a sequence of RGB (Red, Green, Blue) images, wherein the video or the sequence of RGB images are not required to include depth information; and determining, from the video or the sequence of RGB images, the plurality of keyframes; and estimating at least one keyframe of the plurality of keyframes across the video or the sequence of RGB images.
16. The computer program product of claim 15, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further comprising: performing a global bundle adjustment that corrects a camera pose corresponding to at least one of the plurality of keyframe poses or performs a loop closure.
17. The computer program product of claim 15, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising estimating the plurality of keyframes, the plurality of keyframe poses, and the depth data, the set of acts executed by the model further comprising: identifying a non-keyframe from the video or the sequence of RGB images; determining a non-keyframe pose for the non-keyframe; and augmenting a training data set at least by embedding the non-keyframe and the non-keyframe pose into the training data set that is used to train the model.
18. The computer program product of claim 14, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising generating the semantically annotated, manipulatable three-dimensional (3D) representation, the set of acts executed by the model further comprising: generating a plurality of object segmentation masks at least by analyzing the plurality of keyframes or one or more non-keyframes.
19. The computer program product of claim 18, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising generating the semantically annotated, manipulatable three-dimensional (3D) representation, the set of acts executed by the model further comprising: establishing a semantic association among the plurality of object segmentation masks at least by using a zero-shot, one-shot, or few-shot neural network or a multi-view geometry; and generating a three-dimensional (3D) point set based at least in part upon an inverse projection of at least some of the plurality of keyframe poses, the depth data, or associated depth covariance pertaining to the depth data.
20. The computer program product of claim 19, the non-transitory computer readable medium further storing thereupon instructions which, when executed by the processor, cause the processor to perform the set of acts comprising generating the semantically annotated, manipulatable three-dimensional (3D) representation, the set of acts executed by the model further comprising:
improving 3D Gaussian splatted neural radiance field at least by attaching a lowerdimensional semantic vector to a cluster of one or more Gaussians for the semantically annotated, manipulatable three-dimensional (3D) representation.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463666058P | 2024-06-28 | 2024-06-28 | |
| US63/666,058 | 2024-06-28 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026006752A1 true WO2026006752A1 (en) | 2026-01-02 |
Family
ID=98222824
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/035725 Pending WO2026006752A1 (en) | 2024-06-28 | 2025-06-27 | Conditional, probabilistic generation of 3d virtual environment based on incomplete information available to an extended reality device |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026006752A1 (en) |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20190026942A1 (en) * | 2017-07-18 | 2019-01-24 | Sony Corporation | Robust mesh tracking and fusion by using part-based key frames and priori model |
| US20220139057A1 (en) * | 2019-06-14 | 2022-05-05 | Magic Leap, Inc. | Scalable three-dimensional object recognition in a cross reality system |
| US20230082420A1 (en) * | 2021-09-13 | 2023-03-16 | Qualcomm Incorporated | Display of digital media content on physical surface |
| CN117671108A (en) * | 2023-12-07 | 2024-03-08 | 浙江大学 | A dynamic human body modeling method based on three-dimensional Gaussian |
| CN118196306A (en) * | 2024-05-15 | 2024-06-14 | 广东工业大学 | 3D modeling and reconstruction system, method and device based on point cloud information and Gaussian cloud |
| US20240203138A1 (en) * | 2020-03-04 | 2024-06-20 | Magic Leap, Inc. | Systems and methods for efficient floorplan generation from 3d scans of indoor scenes |
-
2025
- 2025-06-27 WO PCT/US2025/035725 patent/WO2026006752A1/en active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20190026942A1 (en) * | 2017-07-18 | 2019-01-24 | Sony Corporation | Robust mesh tracking and fusion by using part-based key frames and priori model |
| US20220139057A1 (en) * | 2019-06-14 | 2022-05-05 | Magic Leap, Inc. | Scalable three-dimensional object recognition in a cross reality system |
| US20240203138A1 (en) * | 2020-03-04 | 2024-06-20 | Magic Leap, Inc. | Systems and methods for efficient floorplan generation from 3d scans of indoor scenes |
| US20230082420A1 (en) * | 2021-09-13 | 2023-03-16 | Qualcomm Incorporated | Display of digital media content on physical surface |
| CN117671108A (en) * | 2023-12-07 | 2024-03-08 | 浙江大学 | A dynamic human body modeling method based on three-dimensional Gaussian |
| CN118196306A (en) * | 2024-05-15 | 2024-06-14 | 广东工业大学 | 3D modeling and reconstruction system, method and device based on point cloud information and Gaussian cloud |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11900518B2 (en) | Interactive systems and methods | |
| CN111882026B (en) | Optimizing unsupervised generative adversarial networks via latent space regularization | |
| IL305340A (en) | High resolution neural processing | |
| CN111754596A (en) | Editing model generation, face image editing method, device, equipment and medium | |
| US12217395B2 (en) | Exemplar-based object appearance transfer driven by correspondence | |
| CN116721334A (en) | Training methods, devices, equipment and storage media for image generation models | |
| CN113850182B (en) | Action recognition method based on DAMR_3DNet | |
| CN119836650B9 (en) | User authentication based on 3D facial modeling using partial facial images | |
| CN118365509B (en) | A facial image generation method and related device | |
| CN118015142B (en) | Face image processing method, device, computer equipment and storage medium | |
| CN120068923A (en) | Digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization | |
| Habib et al. | Exploring progress in Text-to-Image synthesis: an In-Depth survey on the evolution of generative adversarial networks | |
| CN121263824A (en) | Planar mesh reconstruction using images from multiple camera poses | |
| WO2026006752A1 (en) | Conditional, probabilistic generation of 3d virtual environment based on incomplete information available to an extended reality device | |
| US20250095259A1 (en) | Avatar animation with general pretrained facial movement encoding | |
| CN115631285B (en) | Face rendering method, device, equipment and storage medium based on unified driving | |
| CN119922393A (en) | Customize motion and appearance in video generation | |
| CN113821338A (en) | Image transformation method and device | |
| US12602875B2 (en) | Technique for three dimensional (3D) human model parsing | |
| CN121211378B (en) | Digital smell generation method and system based on multi-mode feature fusion | |
| US20250285230A1 (en) | Conditional and marginal model based frame generation | |
| CN121074897B (en) | Multimodal intent recognition methods, apparatus, computer devices, and storage media | |
| US20240169633A1 (en) | Interactive systems and methods | |
| US20260080250A1 (en) | Heavy-tailed diffusion models | |
| CN121767393A (en) | Dynamic novel view reconstruction based on stream re-matching |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25827659 Country of ref document: EP Kind code of ref document: A1 |