WO2025211547A1 - Method and system for reconstructing multi-dimensional extended reality scene - Google Patents

Method and system for reconstructing multi-dimensional extended reality scene

Info

Publication number
WO2025211547A1
WO2025211547A1 PCT/KR2025/000786 KR2025000786W WO2025211547A1 WO 2025211547 A1 WO2025211547 A1 WO 2025211547A1 KR 2025000786 W KR2025000786 W KR 2025000786W WO 2025211547 A1 WO2025211547 A1 WO 2025211547A1
Authority
WO
WIPO (PCT)
Prior art keywords
scene
user
capture
dimensional
capturing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2025/000786
Other languages
French (fr)
Inventor
Midhun Sreekumar MENON
Viswanath VEERA
Jayesh Mundayadan KOROTH
Rajath C ARALIKATTI
Srinidhi NAGARAJA RAO
Lakshmi Priya Muraleedharan
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Samsung Electronics Co Ltd
Original Assignee
Samsung Electronics Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Samsung Electronics Co Ltd filed Critical Samsung Electronics Co Ltd
Publication of WO2025211547A1 publication Critical patent/WO2025211547A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T19/00Manipulating three-dimensional [3D] models or images for computer graphics
    • G06T19/006Mixed reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/048Interaction techniques based on graphical user interfaces [GUI]
    • G06F3/0481Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance
    • G06F3/04815Interaction with a metaphor-based environment or interaction object displayed as three-dimensional [3D], e.g. changing the user viewpoint with respect to the environment or object
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T15/00Three-dimensional [3D] image rendering
    • G06T15/50Lighting effects
    • G06T15/506Illumination models
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T19/00Manipulating three-dimensional [3D] models or images for computer graphics
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T19/00Manipulating three-dimensional [3D] models or images for computer graphics
    • G06T19/20Editing of three-dimensional [3D] images, e.g. changing shapes or colours, aligning objects or positioning parts
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/54Extraction of image or video features relating to texture
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/20Scenes; Scene-specific elements in augmented reality scenes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/70Labelling scene content, e.g. deriving syntactic or semantic representations

Definitions

  • Embodiments disclosed herein relate to Extended Reality (XR) systems, and more particularly to methods and systems for providing guidance to a user to capture a scene for reconstruction using an XR device.
  • XR Extended Reality
  • Extended Reality is rapidly expanding, enabling users in different locations to share an immersive experience of the same physical space through scene reconstruction.
  • the 3D XR experience involves spatial, acoustic, illumination, and appearance information components, which require advanced techniques to help users capture rich data for high-quality reconstruction.
  • Modelling scene illumination is crucial for enhancing the realism and immersiveness of an XR experience. It includes various elements (such as, but not limited to, shadows, depth perception, virtual objects, and mood lighting) which contribute to a more convincing and engaging environment. For instance, in a Virtual Reality (VR) party, dynamic lighting that changes based on the music being played can significantly enhance the atmosphere and make the experience more enjoyable and immersive for participants. Properly modelled lighting helps in creating believable and lifelike virtual scenes that closely mimic real-world interactions with light, thereby improving the overall XR experience.
  • VR Virtual Reality
  • modelling scene acoustics is vital for achieving spatial audio, which allows users to identify the direction and distance of sounds within the XR environment, thereby enhancing realism.
  • Spatial audio adds a layer of depth to the experience, making the experience feel more authentic and immersive.
  • accurately modelled acoustics enable users to pinpoint where a sound is coming from, (for example, a conversation in a virtual meeting or the direction of footsteps in a VR game). This auditory information complements the visual cues, creating a more cohesive and lifelike experience.
  • VST Visual See Through
  • the embodiments herein provide a method and system for reconstructing a multi-dimensional extended reality (XR) scene, the method comprising, obtaining, by an XR device, information corresponding to at least one of: semantics of a scene, geometry of the scene and acoustic of the scene using a sensor data, while a user moves within the scene during the capturing of the scene, determining, by the XR device, a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, wherein the multi-dimensional scene comprises at least one of: texture characteristics of the scene, spatial characteristics of the scene, acoustics characteristics of the scene, and illumination characteristics of the scene, generating, by the XR device, a user guidance for an assisted scene capturing using the scene type and the at least one capture threshold, wherein the user guidance includes at least one of: a path to be followed by the user and at least one action to be performed by the user along the path, and capturing, by the XR
  • the embodiments herein provide a method that may perform at least one of: indicating a gaze fixation point for the user to focus an imaging device, indicating a walking speed depending on at least one detail in a part of the scene, and augmenting a captured data for at least one optimized texture, optimized acoustics, optimized illumination, and optimized material reconstruction in the scene.
  • the embodiments herein provide a method that may perform at least one of: listening experience at a given position in the multi-dimensional XR scene.
  • the embodiments herein provide a method that may capture the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene and the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance comprises: processing a tag associated with an image, a depth of the image, a microphone response signal, a material classification associated with the scene, a sound generation time, lighting information, and acoustic information; and capturing, by the XR device, the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene and the acoustic characteristics of the scene based on the processing.
  • the embodiments herein provide a method and system to determine the multi-dimensional scene by generating a map upon initiating a traverse through an environment, marking an area in which user intended to travel in the environment and usage in the map, estimating a ground plane, a user height and a pose region in the environment, and estimating an occlusion blind spot by ray propagation from a pose region in the environment.
  • the embodiments herein provide a method and system wherein the user guidance comprises at least one of: modifying a lighting in the scene by at least one of: turning ON light, turning OFF light, adjusting a brightness of the light, and changing a colour of the light to capture a model different light source on scene lighting.
  • FIG. 1 depicts a scene reconstruction method, according to existing arts
  • FIGS. 2A-2E depict flow diagrams for providing live guidance to a user for capturing data for performing joint spatial-acoustic-illumination for XR scene reconstruction, according to embodiments as disclosed herein;
  • FIGS. 3A-3G depicts an example user journey, wherein the user uses joint spatial-acoustic-illumination for XR scene reconstruction, according to embodiments as disclosed herein;
  • FIG. 4 depicts hardware component of the XR Device, according to embodiments as disclosed herein;
  • FIG. 5 depicts an example user scenario of a virtual experience, according to embodiments as disclosed herein;
  • FIG. 7 depicts an example user scenario of productivity, according to embodiments as disclosed herein.
  • FIG. 8 depicts an example user scenario of architecture and design, according to embodiments as disclosed herein;
  • FIG. 9A-9B depicts the difference between an example spatial video and a video captured using joint spatial-acoustic-illumination XR scene reconstruction, according to embodiments as disclosed herein;
  • FIG. 10 depicts an example of real-time guidance and feedback for capturing media using the joint spatial-acoustic-illumination XR scene reconstruction., according to embodiments as disclosed herein;
  • FIGS. 11A-11B depicts example illumination and acoustics of a reconstructed scene that can be modified using the joint spatial-acoustic-illumination XR scene reconstruction, according to embodiments as disclosed herein;
  • FIGS. 12A-12B depicts an example flowchart with live user assistance, according to embodiments as disclosed herein;
  • FIGS. 13A-13B depicts an example flowchart with XR scene reconstruction, according to embodiments as disclosed herein;
  • FIG. 14A and 14B depict an example surface with higher texture richness alongside a surface with lower texture richness, according to embodiments as disclosed herein
  • FIG. 15 depicts an example of recreation and high-fidelity spatial reconstruction with a spatial and pose coverage, according to embodiments as disclosed herein;
  • FIG. 16 depicts an example view of direction of a light source for modelling and recreating illumination experience(s), according to embodiments as disclosed herein;
  • FIG. 17 is a flowchart depicting a method for providing user guidance for reconstructing a multi-dimensional XR scene, according to embodiments as disclosed herein;
  • FIG. 18 is a flowchart depicting a method for reconstructing a multi-dimensional XR scene, according to embodiments as disclosed herein.
  • Embodiments herein may be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which may be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and/or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by a firmware.
  • the circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like.
  • circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block.
  • a processor e.g., one or more programmed microprocessors and associated circuitry
  • Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure.
  • the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
  • An embodiment according to the present disclosure may disclose methods and systems for guiding a user to capture a scene (e.g., multi-dimensional extended reality (XR) scene or the like) for reconstruction using an extended reality (XR) device.
  • a scene e.g., multi-dimensional extended reality (XR) scene or the like
  • XR extended reality
  • an embodiment according to the present disclosure may provide real-time feedback and active guidance to a user (when the user is capturing data) to ensure accurate data collection for capturing multi-dimensional scene information, while simultaneously delivering high-quality scene reconstruction in less time.
  • the gathered data including pose-tagged HDR images, depth images, microphone responses, and material classifications, are sent to the high-dimensional joint scene reconstruction module 240.
  • the high-dimensional joint scene module 240 combines spatial, acoustic, illumination, and material properties to build a comprehensive 3D representation of the scene, either on the device or in the cloud.
  • the user guided capture stage 242 may comprise a user action planner, a user gaze fixation planner, and/or a path and speed planner.
  • the user action planner may recommend an action plan to the user 232.
  • the user gaze fixation planner may recommend one or more fixation points to the user 232.
  • the path and speed planner may provide an optimal path and speed for capture to the user 232.
  • the guided capture stage 242 may provide a guidance comprise at least one of the recommended action plan, the recommended fixation points, or the optimal path and speed for capture to the user 232.
  • the user guided capture stage 242 may suggest or recommend one or more actions to the user 232.
  • the one or more actions may include, but are not limited to, placing audio source(s) and/or audio receiver(s) (e.g., a phone, buds, and/or a speaker) at indicated position(s) along the guided path for richer acoustic reconstruction.
  • the one or more actions may include, but are not limited to, modifying a lighting in the scene by turning on or off one or more light sources, modifying brightness and/or color of a light source for better capture, and/or modeling how each light source influence lighting of the scene.
  • FIG. 2E depicts the process of performing multi-dimensional Joint Scene Reconstruction.
  • a method according to the present disclosure involves verifying the capture region marking using XR inputs from various sources such as cameras 206, IMUs 201, ToF sensors 202, mics 203, and eye-tracking cameras 205 illustrated in FIG. 2A.
  • a scene understanding module e.g., the scene understanding module 238 of FIG. 2A
  • the coarse mesh depiction is a representation of a scene in 3D, which provides good surface understanding.
  • the mesh density or number of points in the mesh is lower giving a large-scale understanding. Though the coarse mesh depiction covers the mesh density or number of points in the mesh is lower and might miss on minute details, but will be generated rapidly and with lesser sensor data due to lower mesh resolution for dynamic tasks like obstacle avoidance and path planning.
  • a method includes determining one or more capture settings, such as, but not limited to, texture, acoustics, illumination accuracy, and scene type, which are selected by the user and sent to the user guided capture stage 242.
  • a typical listening experience i.e., acoustics
  • the user guided capture stage 242 can send one or more parameters to a high-dimensional joint scene reconstruction module 240, which performs spatial, acoustic, illumination, and material property reconstructions either on-device or in the cloud.
  • Examples of the parameters can be, but not limited to, pose-tagged HDR images, depth images, microphone responses, material classification, sound generation time, and so on.
  • the Pose-tagged HDR images capture both high-dynamic-range visuals and the camera's position, aiding 3D scene reconstruction, whereas the depth images map object distances, essential for 3D modeling. The depth information is crucial for building a 3D model of the scene.
  • the audio data captured by microphones, for example, the microphone responses include the sound characteristics to recreate the acoustic environment.
  • Material classification involves identifying and categorizing the materials present in a scene, helping in creating a more accurate reconstruction of the scene.
  • the specific time taken at which a sound is produced within a scene helps in synchronizing the sound with the visual and spatial data, ensuring that the reconstructed scene accurately reflects the original environment.
  • the various actions mentioned above may be performed in the order presented, in a different order or simultaneously. Additionally or alternatively, the method provide real-time guidance during the capture process. This allows users to make informed adjustments on-the-fly, thereby enhancing the overall quality and precision of the XR experience.
  • Embodiments herein can assess the quality of input data (prior to reconstruction), thereby enhancing the accuracy and relevance of captured data.
  • embodiments herein provide dynamic, real-time guidance tailored to the user's context during the capture process.
  • Embodiments herein use a personalized feedback mechanism to ensure that the input data is optimized for the specific needs of various user personas, for example, users involved in social media, architecture, music, and so on.
  • Embodiments herein can employ default and custom profiles to address the unique requirements of different users, enhancing the versatility and applicability of the method across diverse fields and applications.
  • Embodiments herein can improve the user experience by delivering context-aware, actionable insights at the moment of data capture. Further, in some embodiments, some actions listed in FIG. 2A-2E may be omitted.
  • Users can enhance their audio experience by strategically placing audio sources and receivers (such as, but not limited to, phones, earbuds, and speakers), at one or more specified locations along a suggested path to achieve a richer acoustic reconstruction. Additionally, the users can optimize the lighting within the scene by adjusting existing light sources (such as, by turning them on or off, or by altering their brightness and color settings) to accurately capture and model the influence of each light source on the overall scene lighting. These adjustments will significantly improve both the auditory and visual quality of the scene, creating a more immersive and detailed environment.
  • audio sources and receivers such as, but not limited to, phones, earbuds, and speakers
  • Embodiments herein enhance the comprehensiveness and quality of the data collected, thereby facilitating superior XR reconstruction of various scene attributes (such as, but not limited to, geometry, textures, acoustics, and illumination).
  • scene attributes such as, but not limited to, geometry, textures, acoustics, and illumination.
  • FIGs. 3A-3G depict an example user journey using the proposed method of joint spatial-acoustic-illumination XR scene reconstruction.
  • a user 30 uses a XR device 300.
  • the XR device 300 is illustrated as a pair of handy controllers, a configuration of the XR device 300 is not limited thereto.
  • the XR device 300 may include a display device (e.g., a head-mounted display apparatus, an augmented reality (AR)/XR helmet, or an AR/VR glasses) which provides a user interface to the user 30.
  • the XR device 300 may utilize one or more algorithms for tracking one or more gestures of the user 30 without any handy controller.
  • FIG. 3A illustrates an example image indicating that a user 30 of a XR device 300 has accessed an application for XR scene reconstruction. While the user 30 uses the XR device 300, the XR device 300 may provide a user interface 302 comprising one or more icons representing corresponding functionality.
  • the user interface 302 may comprise at least one of: an icon representing current time, an icon representing a face (or an icon) of the user 30, an icon representing one or more wireless connections of the XR device 300, an icon corresponding to a functionality representing a list of applications installed on the XR device 300, an icon corresponding to a functionality for representing a list of contacts stored on the XR device 300, an icon corresponding to a functionality for representing a list of notifications occurred in the XR device 300, an icon corresponding to a functionality for sharing one or more XR scenes generated by the XR device 300, or an icon corresponding to a functionality for modifying settings of the XR device 300.
  • the XR device 300 may provide an application list user interface 304.
  • the application list user interface 304 may comprise one or more icons respectively corresponding to one or more applications of the XR device 300.
  • the XR device may further representing a line user interface 306 which indicates tracked intention of the user 30.
  • the line user interface 306 may represented based on tracked, by the XR device 300, gesture of the user 30.
  • the application for joint spatial-acoustic-illumination XR scene reconstruction may be opened.
  • the user 30 can be provided with an interface that prompts them to begin the process.
  • FIG. 3B illustrates an example image, wherein the application for XR scene reconstruction provides instructions via a guidance user interface 310.
  • the application may provide the guidance user interface 310 which guides the user 30 to walk around the scene and mark the capture region. For example, via the guidance user interface 310, the application instructs the user 30 to walk around the environment and mark the boundaries of the capture region.
  • FIG. 3C illustrates an example image indicating that the user 30 wearing the XR device 300 confirms that the capture region is marked correctly on the ground plane.
  • the XR device 300 may comprise a head-mounted wearable device 300c which has an imaging device (for example, a camera).
  • FIG. 3D illustrates an example image indicating an estimated path 312 for optimal capture that has been highlighted for the user 30 to follow.
  • the XR device may display a user interface 314 for asking the user 30 to confirm the estimated path 312 (e.g., the capture region).
  • the application may re-estimate the capture region.
  • the gaze fixation point is indicated for the user to focus an imaging device (e.g., a camera of the XR device 300) on. Additionally or alternatively, the walking speed is indicated depending on the details in different parts of the scene. Other actions to augment the richness of the captured data for better textures, acoustics, illumination, material reconstruction are suggested.
  • an imaging device e.g., a camera of the XR device 300
  • FIGs. 3E and 3F illustrate an example image indicating that the user 30 is guided to walk along a highlighted path 316, and to perform an instructed action for optimization of image capturing.
  • the XR device 300c may guide the user 30 to gaze a certain fixed point (e.g., a point 320) in the environment while walking along the highlighted path 326.
  • the application continuously captures spatial data, while simultaneously collecting information on the scene's acoustic properties and illumination conditions.
  • the application may reconstruct the XR scene based on the collected information, and indicates the user 30 that data associated with the scene is successfully captured by using an user interface 322 (as depicted in the example depicted in FIG. 3G). As the user 30 walks along this highlighted path 316, the user 30 is guided to perform specific action(s) to ensure optimal data capture, including focusing the camera on designated gaze fixation points and adjusting walking speed based on scene details.
  • the application can suggest one or more additional actions to enhance data richness, such as, but not limited to, fixating on points to capture detailed lighting information.
  • the user 30 can be prompted to focus on a specific point (e.g., the point 320 illustrated in FIG. 3F) to gather more light source details.
  • the application can indicate that the user 30 has correctly marked the capture region on the ground plane.
  • the application analyzes the gathered data, and extracts detailed scene semantics and geometric information from the gathered data. Thereafter, the scene semantics and geometry information of the marked region is collected, and scene type and capture accuracy thresholds are determined (for example, one or more thresholds for texture, acoustics, and/or illumination accuracy).
  • the application determines the scene type and sets specific capture accuracy thresholds for texture, acoustics, and illumination.
  • the captured data including pose-tagged HDR images, depth images, microphone response signals, material classifications, signature sound generation times, and other lighting and acoustic information, can be then processed to create a reconstruction. This processing can be performed either offline (e.g., by the XR device 300) or on the cloud, ensuring comprehensive and high-quality data for accurate scene reconstruction.
  • This comprehensive dataset ensures that the reconstructed XR scene accurately reflects the real-world environment, providing a highly immersive and realistic experience.
  • FIG. 4 depicts hardware component of the XR Device 230 comprises of processor 230a, a scene reconstruction controller 230b and a memory 230c.
  • the XR device 230 may exclude at least one of these components or may further include at least one other component.
  • the processor 230a includes one or more processing devices or processing circuitry, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs).
  • the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processor unit (GPU).
  • the processor 230a is able to perform control on at least one of the other components of the XR device 230.
  • the scene reconstruction controller (230b) is coupled with the processor (230a) and the memory (230c).
  • the scene reconstruction controller (230b) is configured to obtain information corresponding to at least one of: the semantics of the scene, the geometry of the scene, or acoustic of the scene using the sensor data, while a user moves within the scene during capturing the scene.
  • the scene reconstruction controller 230b is configured to activate user marking of the capture region if the capture region is verified, determine a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, generate a user guidance for an assisted scene capturing using the scene type and the at least one capture thresholds, wherein the user guidance includes at least one of: a path to be followed by the user or at least one action to be performed by the user along the path, and capture the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene, or the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance.
  • the memory 230c may store commands or data related to at least one other component of the XR device 230.
  • the memory 230c can include volatile memory (e.g., a random-access memory (RAM)) and/or non-volatile memory (e.g., a flash memory or a solid-state drive (SSD)).
  • the memory 230c may comprise one or more storage medium which store(s) one or more instructions. The one or more instructions may cause, when executed by the processor 230a and/or the scene reconstruction controller 230b individually or collectively, the XR device 230 to perform any combinations of operations described herein.
  • FIG. 5 depicts an example virtual experience.
  • Embodiments herein can provide the user with an immersive virtual concert experience that closely simulates the sensations of attending a live event in person.
  • Embodiments herein can replicate the intricate details of concert lighting and sound systems, thereby ensuring that the user can experience the atmosphere and ambience of a live performance with remarkable fidelity.
  • Embodiments herein can capture the visual and auditory nuances of a concert, and can also extend the experience beyond geographical limitations, thereby allowing friends and loved ones to join in and enjoy the event together, regardless of their location. This enhanced social interaction amplifies the enjoyment and creates a shared, memorable experience, bridging the gap between physical presence and virtual participation.
  • FIG. 6 depicts an example virtual experience.
  • the user utters the following: "I want to share my experience of this museum exhibit with my friends in an immersive way?".
  • Embodiments herein can capture a high-fidelity reconstruction of scene in 3D. The user can share this with friends who can view the exhibit from any viewpoint in an immersive way.
  • FIG. 7 depicts an example virtual experience.
  • the user utters the following: "I want to have a meeting, but my team works in a hybrid mode across multiple offices, and it is hard to communicate effectively with just video meetings".
  • Embodiments herein offer the potential to create a high-dimensional reconstruction of a meeting room, meticulously modeling not just the physical layout of the meeting room, but also the acoustics and lighting in the meeting room, thereby ensuring a highly immersive virtual environment.
  • Embodiments herein can capture and replication of in-person interactions accurately, including subtle aspects such as tone of voice, gestures, and body language of the participants in the meeting, all presented in a 1:1 scale.
  • embodiments herein effectively bridge the communication gap, allowing for interactions that closely mimic face-to-face meetings. This level of detail enhances the realism of virtual engagements, making them more effective and engaging compared to traditional virtual communication platforms.
  • embodiments herein can significantly enhance the architectural design process by enabling precise scaled 3D reconstructions, which are crucial for accurate design, acoustic analysis, and illumination modelling.
  • Embodiments herein allow architects to experiment with various materials and design styles in a virtual environment, offering a comprehensive perspective from any vantage point within the reconstructed scene. Additionally, embodiments herein facilitates real-time collaboration by allowing on-site workers to update and share the current progress of a project with the architect remotely, ensuring continuous and effective communication even when the architect is off-site. This seamless integration of technology not only optimizes design accuracy and efficiency, but also enhances the overall management and execution of architectural projects.
  • FIG. 10 depicts an example scenario, wherein real-time guidance and feedback is used for capturing the scene.
  • the XR device may start capturing data from surroundings.
  • the XR device may suggest following a path and performing one or more actions (e.g., capturing one or more points nearby with the XR device) to the user. While the user follows the suggestion, the XR device may collect data from the surroundings, and reconstruct the scene based on the collected data.
  • FIGs. 11A-11B depict the illumination and acoustics of the reconstructed scene that can be modified using the proposed joint spatial-acoustic-illumination XR scene reconstruction.
  • a jointed XR experience may be reconstructed, and provided to a user as illustrated in FIG. 11B.
  • the jointed XR experience may include one or spatial characteristics (e.g., one or more objects included in the environment), one or more illumination characteristics (e.g., one or more lightings 1102), and one or more acoustic characteristics (e.g., one or more sound elements 1104).
  • One or more characteristics included in the joint XR experience may be customized.
  • the one or more lightings 1102 and/or the one or more sound elements 1104 may be customized (or adjusted) by the user.
  • spatial video viewers may be limited to the perspective from which the video was initially recorded, typically bound to the positions of the original camera(s). Accordingly, users may be confined to a fixed viewpoint, limiting their ability to explore the scene dynamically.
  • Spatial video technology offers the potential for users to explore a scene from various angles and distances, providing a more immersive and flexible experience. In an embodiment, this dynamic exploration is facilitated by a calibrated stereo camera setup, which records the scene from multiple viewpoints.
  • a method according to an embodiment of the present disclosure revolutionizes this approach by eliminating the need for such complex equipment. It allows for capturing spatial video with just a single camera, thus simplifying the recording process while still enabling users to interact with and view the scene from different perspectives. This advancement opens up new possibilities for more accessible and versatile spatial video applications.
  • a method according to an embodiment of the present disclosure delivers real-time feedback and active guidance to users, enhancing their ability to capture high-quality XR experiences.
  • By integrating high-dimensional capture within a single framework it enables users to record spatial richness, including pose, acoustic, and illumination elements.
  • Accurate audio recreation is fundamental to creating a realistic and immersive XR experience, as it enhances spatial awareness, facilitates social interaction, fosters emotional connections, and provides a nuanced understanding of the environment.
  • precise scene lighting modelling is vital for lifelike rendering, affecting the perception of depth, shape, and texture, thereby ensuring a believable experience. This approach allows users dynamic control over the scene, making applications like XR design and digital twins more practical and effective.
  • FIGs. 12A & 12B depicts a flow chart with live user assistance flow.
  • FIGs. 12A & 12B depicts a system 200 for providing user guidance for reconstructing a multidimensional XR scene using a server connected to a XR Device 230.
  • the XR device 230 may comprise one or more modules illustrated in FIGs. 12A and 12B.
  • the XR Device 230 is configured to receive sensor data from a plurality of sensors 201-205 of FIG. 2A(on a user initiating the capture of a scene), derive information on the scene's semantics and geometry using the received data, and activate a user marking out capture region module 236 of FIG. 2A (on the region being successfully verified).
  • the user marking out capture region module 236 creates a coarse semantic scene mesh based on the derived information, while optionally allowing the user to mark a capture region.
  • the XR Device 230 determines the scene type and capture accuracy thresholds using a pre-trained model, wherein the pre-trained model uses texture, spatial, acoustics, and/or illumination characteristics from a metric semantic localization & mapping module 210, a material type classifier module 212, and a shadow region classifier & reconstruction module 214.
  • the XR Device 230 using the pre-trained model generates user guidance for a second round of scene capturing, which includes a path and actions for the user. During this second round, the user captures the multidimensional scene data following the provided guidance.
  • FIGs. 13A & 13B depicts a system 400 detailed flow-chart with an XR scene reconstruction flow.
  • the XR device 230 may comprise one or more modules illustrated in FIGs. 13A and 13B.
  • the XR Device 230 is configured to create a coarse semantic scene mesh using the derived information, determine the scene type and capture thresholds for multidimensional data, determine if a capture region is marked on a ground plane, and highlight an optimal capture path.
  • the capture threshold defines the threshold value for quality of reconstruction metrics of the scene. If the metrics or reconstruction quality scores from scene reconstruction algorithms improve above the threshold, the scene is deemed to be good at having reconstructed the environment with accuracy as expected by user through the threshold.
  • the thresholds are set at the beginning, based on user history, environment, user inputs, and user preference.
  • the reconstruction quality in each of these modalities or dimensions depends on these threshold values.
  • the capture path guides the user along this path, provides real-time visual feedback for reconstructing the scene (wherein the scene encompasses spatial, acoustic, illumination, and material appearance characteristics), and offers user assistance for creating multi-dimensional XR scenes (which integrates integrating texture, spatial, audio, visual, and illumination experiences).
  • the XR Device 230 is configured to evaluate data quality, provide guidance based on user profiles and scenes, and/or support both default and custom user personas tailored to the user and/or application.
  • a scene segmentation module 222 creates a coarse scene mesh by estimating the semantics and geometry of a scene, while the user is marking a capture region.
  • the pre-trained model determines the scene type and sets capture thresholds for multi-dimensional scene data, including texture, spatial, acoustic, and illumination characteristics.
  • a ground plane confirmation module verifies if the capture region is correctly marked on the ground plane.
  • a path highlighting module (e.g., a path planner module 218 of FIG. 12B) highlights an optimal capture path for the user.
  • a user guidance module 220 of FIG. 12B instructs the user to follow this path and perform necessary actions.
  • a data processing module processes the captured multi-dimensional scene data to reconstruct an XR scene, while a real-time feedback module provides visual feedback for scene reconstruction.
  • a live user assistance module offers real-time support for creating the multi-dimensional XR scene.
  • An evaluation and guidance module benchmarks reconstruction input data quality and provides live guidance based on user profiles and scenes.
  • a user persona support module 211 of FIG. 12A accommodates default and custom user personas 213 of FIG. 12A based on different user needs.
  • the XR Device 230 can locate the audio sources within a room using Direction of Arrival (DOA) algorithms combined with visual cues extracted from recorded video or images.
  • DOA Direction of Arrival
  • the XR Device 230 can further estimate one or more location related parameters (such as, but not limited to, dimensions and absorption values) from recorded video or images.
  • the XR Device 230 can perform object segmentation for identifying material types and their properties, which can be used as initial conditions. Subsequently, the XR Device 230 can fine-tune these parameters through analysis of recorded audio response spectra.
  • Modelling and recreating the listening experience using the proposed method includes estimating the Room Impulse Response (RIR), and user guidance for recording the RIR.
  • the RIR is the transfer function between the sound source and the microphone.
  • T60 and space dimensions can be used for estimating the RIR, wherein T60 is the time taken for the sound to decay by 60dB. T60 can be different at different locations in the location.
  • the XR Device 230 can provide guidance to the user to record RIRs at probable locations in the space where the listener is located (for example, in front of the TV, chairs, etc.)
  • the XR Device 230 confirms if the capture region is on a ground plane, and in step 1808, the XR Device 230 highlights an estimated path for the user to follow for effective scene capture. In step 1810, as the user navigates this path, the XR Device 230 provides guidance to ensure accurate scene data collection. In step 1810, the XR Device 230 processes the captured data to reconstruct the XR scene, which incorporates spatial, acoustic, illumination, and material appearance characteristics into a single framework, offering real-time visual feedback to achieve high-quality, rich XR scene capture with customizable and realistic lighting, acoustics, and virtual object manipulation in step 1812.
  • the XR Device 230 provides user assistance for creating multi-dimensional XR scenes by integrating texture, spatial, audio, visual, and illumination experiences within a unified framework. It also serves as a benchmark to evaluate the quality of reconstruction input data, offering user guidance tailored to individual profiles and scenes while supporting both default and custom user personas to cater to diverse user needs.
  • the various actions in method 1800 may be performed in the order presented, in a different order or simultaneously. Further, in some embodiments, some actions listed in FIG. 18 may be omitted.
  • Embodiments herein disclose can enhance the collected data to improve the XR reconstruction of a scene by refining its geometry, textures, acoustics, and lighting.
  • the embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the network elements.
  • the elements include blocks which can be at least one of a hardware device, or a combination of hardware device and software module.
  • the embodiments disclosed herein describe a method for user guidance for reconstructing a multi-dimensional extended reality (XR) scene using an XR Device and providing real-time feedback and active guidance to the user when capturing data. Therefore, it is understood that the scope of the protection is extended to such a program and in addition to a computer readable means having a message therein, such computer readable storage means (e.g., computer-readable storage medium) contain (or store) program code means for implementation of one or more steps of the method, when the program runs on a server or mobile device or any suitable programmable device.
  • XR extended reality
  • the method is implemented in at least one embodiment through or together with a software program written in e.g., Very high speed integrated circuit Hardware Description Language (VHDL) another programming language, or implemented by one or more VHDL or several software modules being executed on at least one hardware device.
  • VHDL Very high speed integrated circuit Hardware Description Language
  • the hardware device can be any kind of portable device that can be programmed.
  • the device may also include means which could be e.g., hardware means like e.g., an ASIC, or a combination of hardware and software means, e.g., an ASIC and an FPGA, or at least one microprocessor and at least one memory with software modules located therein.
  • the method embodiments described herein could be implemented partly in hardware and partly in software.
  • the invention may be implemented on different hardware devices, e.g., using a plurality of CPUs.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Graphics (AREA)
  • Human Computer Interaction (AREA)
  • Computer Hardware Design (AREA)
  • Software Systems (AREA)
  • Multimedia (AREA)
  • Architecture (AREA)
  • Computational Linguistics (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

Embodiments herein disclose methods and systems for reconstructing a multi-dimensional extended reality (XR) scene. A method, performed by an XR Device, may obtain information corresponding to at least one of: semantics of a scene, geometry of the scene, or acoustic of the scene using a sensor data, determining a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, and offering real-time feedback and active guidance to user. The method generates a user guidance for an assisted scene capturing using the scene type and the at least one capture threshold, to capture the multi-dimensional scene using the generated user guidance.

Description

METHOD AND SYSTEM FOR RECONSTRUCTING MULTI-DIMENSIONAL EXTENDED REALITY SCENE
Embodiments disclosed herein relate to Extended Reality (XR) systems, and more particularly to methods and systems for providing guidance to a user to capture a scene for reconstruction using an XR device.
Extended Reality (XR) is rapidly expanding, enabling users in different locations to share an immersive experience of the same physical space through scene reconstruction. The 3D XR experience involves spatial, acoustic, illumination, and appearance information components, which require advanced techniques to help users capture rich data for high-quality reconstruction.
Modelling scene illumination is crucial for enhancing the realism and immersiveness of an XR experience. It includes various elements (such as, but not limited to, shadows, depth perception, virtual objects, and mood lighting) which contribute to a more convincing and engaging environment. For instance, in a Virtual Reality (VR) party, dynamic lighting that changes based on the music being played can significantly enhance the atmosphere and make the experience more enjoyable and immersive for participants. Properly modelled lighting helps in creating believable and lifelike virtual scenes that closely mimic real-world interactions with light, thereby improving the overall XR experience.
Similarly, modelling scene acoustics is vital for achieving spatial audio, which allows users to identify the direction and distance of sounds within the XR environment, thereby enhancing realism. Spatial audio adds a layer of depth to the experience, making the experience feel more authentic and immersive. For example, accurately modelled acoustics enable users to pinpoint where a sound is coming from, (for example, a conversation in a virtual meeting or the direction of footsteps in a VR game). This auditory information complements the visual cues, creating a more cohesive and lifelike experience.
Visual See Through (VST) technology represents a significant advancement in extending computing into the 3D realm. VST offers unprecedented opportunities for creating and sharing media and experiences in a new dimension. Equipped with a wide array of sensors for environmental perception, VST opens up possibilities that were previously unattainable. These sensors allow for detailed capture and interaction with the 3D environment, facilitating advanced applications (such as, but not limited to, multi-dimensional XR scene reconstruction). This, in turn, necessitates sophisticated user guidance methods to optimally capture and utilize data from various sensors, ensuring a seamless and enriched XR experience.
Accordingly, the embodiments herein provide a method and system for reconstructing a multi-dimensional extended reality (XR) scene, the method comprising, obtaining, by an XR device, information corresponding to at least one of: semantics of a scene, geometry of the scene and acoustic of the scene using a sensor data, while a user moves within the scene during the capturing of the scene, determining, by the XR device, a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, wherein the multi-dimensional scene comprises at least one of: texture characteristics of the scene, spatial characteristics of the scene, acoustics characteristics of the scene, and illumination characteristics of the scene, generating, by the XR device, a user guidance for an assisted scene capturing using the scene type and the at least one capture threshold, wherein the user guidance includes at least one of: a path to be followed by the user and at least one action to be performed by the user along the path, and capturing, by the XR device, the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene and the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance.
Accordingly, the embodiments herein provide a method and system for obtaining information corresponding to at least one of: the semantics of the scene, the geometry of the scene and the acoustic of the scene using the sensor data comprises, receiving the sensor data from at least one sensor in the XR device when the user of the XR device initiates capturing of the scene, and obtaining information corresponding to at least one of: the semantics of the scene, the geometry of the scene and the acoustic of the scene using the sensor data.
Accordingly, the embodiments herein provide a method that may perform at least one of: indicating a gaze fixation point for the user to focus an imaging device, indicating a walking speed depending on at least one detail in a part of the scene, and augmenting a captured data for at least one optimized texture, optimized acoustics, optimized illumination, and optimized material reconstruction in the scene.
Accordingly, the embodiments herein provide a method that may perform at least one of: listening experience at a given position in the multi-dimensional XR scene.
Accordingly, the embodiments herein provide a method that may capture the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene and the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance comprises: processing a tag associated with an image, a depth of the image, a microphone response signal, a material classification associated with the scene, a sound generation time, lighting information, and acoustic information; and capturing, by the XR device, the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene and the acoustic characteristics of the scene based on the processing.
Accordingly, the embodiments herein provide a method and system to determine the multi-dimensional scene by generating a map upon initiating a traverse through an environment, marking an area in which user intended to travel in the environment and usage in the map, estimating a ground plane, a user height and a pose region in the environment, and estimating an occlusion blind spot by ray propagation from a pose region in the environment.
Accordingly, the embodiments herein provide a method and system wherein the user guidance comprises at least one of: modifying a lighting in the scene by at least one of: turning ON light, turning OFF light, adjusting a brightness of the light, and changing a colour of the light to capture a model different light source on scene lighting.
These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating at least one embodiment and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
Embodiments herein are illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the following illustratory drawings. Embodiments herein are illustrated by way of examples in the accompanying drawings, and in which:
FIG. 1 depicts a scene reconstruction method, according to existing arts;
FIGS. 2A-2E depict flow diagrams for providing live guidance to a user for capturing data for performing joint spatial-acoustic-illumination for XR scene reconstruction, according to embodiments as disclosed herein;
FIGS. 3A-3G depicts an example user journey, wherein the user uses joint spatial-acoustic-illumination for XR scene reconstruction, according to embodiments as disclosed herein;
FIG. 4 depicts hardware component of the XR Device, according to embodiments as disclosed herein;
FIG. 5 depicts an example user scenario of a virtual experience, according to embodiments as disclosed herein;
FIG. 6 depicts an example user scenario of a social media, according to embodiments as disclosed herein;
FIG. 7 depicts an example user scenario of productivity, according to embodiments as disclosed herein.
FIG. 8 depicts an example user scenario of architecture and design, according to embodiments as disclosed herein;
FIG. 9A-9B depicts the difference between an example spatial video and a video captured using joint spatial-acoustic-illumination XR scene reconstruction, according to embodiments as disclosed herein;
FIG. 10 depicts an example of real-time guidance and feedback for capturing media using the joint spatial-acoustic-illumination XR scene reconstruction., according to embodiments as disclosed herein;
FIGS. 11A-11B depicts example illumination and acoustics of a reconstructed scene that can be modified using the joint spatial-acoustic-illumination XR scene reconstruction, according to embodiments as disclosed herein;
FIGS. 12A-12B depicts an example flowchart with live user assistance, according to embodiments as disclosed herein;
FIGS. 13A-13B depicts an example flowchart with XR scene reconstruction, according to embodiments as disclosed herein;
FIG. 14A and 14B depict an example surface with higher texture richness alongside a surface with lower texture richness, according to embodiments as disclosed herein
FIG. 15 depicts an example of recreation and high-fidelity spatial reconstruction with a spatial and pose coverage, according to embodiments as disclosed herein;
FIG. 16 depicts an example view of direction of a light source for modelling and recreating illumination experience(s), according to embodiments as disclosed herein;
FIG. 17 is a flowchart depicting a method for providing user guidance for reconstructing a multi-dimensional XR scene, according to embodiments as disclosed herein; and
FIG. 18 is a flowchart depicting a method for reconstructing a multi-dimensional XR scene, according to embodiments as disclosed herein.
The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
For the purposes of interpreting this specification, the definitions (as defined herein) will apply and whenever appropriate the terms used in singular will also include the plural and vice versa. It is to be understood that the terminology used herein is for the purposes of describing particular embodiments only and is not intended to be limiting. The terms "comprising", "having" and "including" are to be construed as open-ended terms unless otherwise noted.
The words/phrases "exemplary", "example", "illustration", "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.," , "i.e.," are merely used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein using the words/phrases "exemplary", "example", "illustration", "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.," , "i.e.," is not necessarily to be construed as preferred or advantageous over other embodiments.
Embodiments herein may be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which may be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and/or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by a firmware. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
It should be noted that elements in the drawings are illustrated for the purposes of this description and ease of understanding and may not have necessarily been drawn to scale. For example, the flowcharts/sequence diagrams illustrate the method in terms of the steps required for understanding aspects of the embodiments as disclosed herein. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. Furthermore, in terms of the system, one or more components/modules which comprise the system may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any modifications, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings and the corresponding description. Usage of words such as first, second, third etc., to describe components/elements/steps is for the purposes of this description and should not be construed as sequential ordering/placement/occurrence unless specified otherwise.
An embodiment according to the present disclosure may disclose methods and systems for guiding a user to capture a scene (e.g., multi-dimensional extended reality (XR) scene or the like) for reconstruction using an extended reality (XR) device.
Additionally or alternatively, an embodiment according to the present disclosure may provide real-time feedback and active guidance to a user (when the user is capturing data) to ensure accurate data collection for capturing multi-dimensional scene information, while simultaneously delivering high-quality scene reconstruction in less time.
Additionally or alternatively, an embodiment according to the present disclosure may provide methods and systems for capturing multi-dimensional scene information (such as, but not limited to, spatial, acoustics, appearance and illumination information, and so on) in a single framework.
Additionally or alternatively, an embodiment according to the present disclosure may provide methods and systems for customizing a scene through one or more of scene relighting, and acoustic remodeling.
Additionally or alternatively, an embodiment according to the present disclosure may generate a user guidance for an assisted scene capturing using the scene type and the at least one capture threshold. The user guidance includes at least one of: a path to be followed by the user and at least one action to be performed by the user along the path.
Additionally or alternatively, an embodiment according to the present disclosure may capture the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene and the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance.
These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating at least one embodiment and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
The embodiments herein provide a user guidance for reconstructing a multidimensional XR scene by XR Device offering real-time feedback and active guidance to user to capture high quality XR experiences, where high-dimensional capture allows to capture spatial richness, pose richness, acoustic richness and illumination richness all in a single framework.
Referring now to the drawings, and more particularly to FIGS. 1 through 18, providing a method and system for reconstructing a multi-dimensional extended reality (XR) scene, where similar reference characters denote corresponding features consistently throughout the figures.
FIG. 1 depicts an exemplary scene reconstruction method. In step 102, the input frames are received from the extended reality (XR) scenes. In step 104, an initial capture of the existing scene is performed. In step 106, the initial capture of the existing scene is converted to 3D structure. In step 108, he quality of the converted 3D structure is analyzed, when the user views the capture structure. In step 110, if the quality of the converted 3D structure is good, the captured frames are exported as 3D scene. In step 112, if the quality of the converted 3D structure is of poor quality, receiving input frames is started again.
According to the embodiment of the present disclosure, the method for reconstructing (or capturing) XR scenes may comprise obtainment of real-time feedback for users, and/or validation or profiling tools or proxy metrics for evaluating data capture quality prior to scene reconstruction. The existence of proxy metrics results in a process that saves users' time. Additionally or alternatively, the method may adequately address the reconstruction of acoustic properties, illumination, and material details, leading to optimal results in these aspects. Additionally or alternatively, the method for capturing XR scenes may comprise obtaining user feedback, making it easy for users to gauge the richness of the captured data or the expected quality of the reconstruction. As a result, non-experts' experience to capture high-quality XR scenes may be improved. For example, there may be no need to wait for post-processing to determine if the reconstruction is satisfactory. Even though the quality is poor, users may not need to recapture the scene, thereby saving time, and decreasing costs. Additionally or alternatively, the method may capture various aspects of a scene, thereby enhancing the immersive experience. For example, information on illumination and acoustic elements may be used to convincingly customize reconstructed scenes.
FIGs. 2A-2E depict flow diagrams for providing live user guidance for enhancing data-capture for joint spatial-acoustic-illumination XR scene reconstruction. FIG. 2A depicts the process of providing XR Inputs to a user guided stage involves using various extended reality (XR) inputs―such as data from cameras 206, inertial measurement units (IMUs) 201, time-of-flight (ToF) sensors 202, microphones 203, and eye-tracking cameras 205―to verify and refine the marked capture region in a scene. For example, cameras 206 may proviced processed high dynamic range (HDR) images. The IMU 201 may provide filtered IMU readings. The ToF sensor 202 may provide processed dense depth. The mics 203 may provide conditioned voice signals. The eye-tracking cameras 206 may provide information for gaze tracking.
The XR inputs may be processed and used to verify (or identify, or determine) capture region marking. If the capture region is verified to be marked ('True'), a user marking out capture region module 236 may activate the capture region. The processed input data may be transferred to a user marking out capture region module 236. If the capture region is not marked or fails to be verified ('False'), the user marking out capture region module 236 may not activate the captured region. Alternatively or additionally, the user 232 may provide user-marked out capture region to a user guided capture stage 242.
The method according to an embodiment of the present disclosure comprises obtaining information corresponding to at least one of: the semantics of the scene, the geometry of the scene and the acoustic of the scene using the sensor data. The method according to an embodiment of the present disclosure comprises receiving the sensor data from at least one sensor in the XR device when the user of the XR device initiates capturing of the scene. The method according to an embodiment of the present disclosure comprises obtaining information corresponding to at least one of: the semantics of the scene, the geometry of the scene and the acoustic of the scene using the sensor data. The method according to an embodiment of the present disclosure comprises performing at least one of: listening experience at a given position in the multi-dimensional XR scene.
Additionally or alternatively, the method may comprise capturing the multi-dimensional scene representing at least one of: the spatial characteristics of the scene, the illumination characteristics of the scene, or the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance. Additionally or alternatively, the method comprises: processing at least one of: a tag associated with an image, a depth of the image, a microphone response signal, a material classification associated with the scene, a sound generation time, lighting information, or acoustic information. Additionally or alternatively, the method comprises capturing the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene, or the acoustic characteristics of the scene based on the processing. A scene understanding module 238 processes these inputs to estimate the geometry and semantics, creating a coarse mesh of the scene. The user then selects specific capture settings like texture, acoustics, illumination, and scene type, incorporating illumination and material property modeling to simulate lighting effects and material textures realistically. These settings guide the user guided capture stage 242, which sends relevant parameters for multi-dimensional scene reconstruction to a high-dimensional joint scene reconstruction module 240 illustrated in FIG. 2B, ensuring an accurate and immersive representation of the environment. The multi-dimensional scene is basically a scene which includes at least one of: 3D scene information, reconstruction, segmentation of scene, illumination details of scene including light sources, color, reflectance or bidirectional reflectance distribution function (BRDF) characteristics, acoustics of the room, acoustics properties of objects in scene, or acoustic response function of the scene. According to an embodiment of the present disclosure, a real-time user assistance for data capture may be provided for ensuring that users can efficiently gather and integrate data during their interactions, while using quality and richness metrics to evaluate the effectiveness and fidelity of the XR data capture process. For example, the user guided capture stage 242 may provide an optimal path and speed for capture, recommended fixation points, and/or recommended action plan to the user 232.
FIG. 2B depicts the process of performing multi-dimensional Joint Scene Reconstruction. The FIG. 2B depicts creating a detailed and accurate 3D model of a scene by integrating various data types by addressing critical aspects of scene representation, acoustic accuracy, and user interaction in dynamic and realistic virtual environments. This process begins with incorporating pose richness, which captures the precise positioning and movement within the scene, as well as detailed illumination and material attributes, and verifying the marked capture region using inputs from devices like cameras, IMUs, ToF sensors, microphones, and eye-tracking cameras. Multi-dimensional scenes are determined by generating a map upon initiating a traverse through an environment and marking an area in which user intended to travel in the environment and usage in the map estimating a ground plane, a user height and a pose region in the environment and estimating an occlusion blind spot by ray propagation from a pose region in the environment. The semantics of the scene is obtained by using a trained object detection and segmentation module, and the geometry of the scene is obtained by using Simultaneous Localization And Mapping (SLAM) technique to formulate a coarse mesh depiction of the scene wherein the SLAM technique is used to understand the 3D scene to build a map and localize camera/sensor cluster in the map at the same time. The XR device directs an object in the scene, to trigger an audio device associated with the XR device to emit an acoustic signal and receive an acoustic response, and activate a light source associated with the XR device to capture the scene and its objects under varying lighting condition to guide a user head pose by asking the user to look with a visual center pointed at a moving overlaid visual cue in XR device. A scene understanding module 238 illustrated in FIG. 2A then estimates the scene's geometry and meaning, forming a basic mesh. Users select capture settings such as texture, acoustics, and illumination, which guide the capture stage. The gathered data, including pose-tagged HDR images, depth images, microphone responses, and material classifications, are sent to the high-dimensional joint scene reconstruction module 240. The high-dimensional joint scene module 240 combines spatial, acoustic, illumination, and material properties to build a comprehensive 3D representation of the scene, either on the device or in the cloud.
FIG. 2C depicts a process of providing XR Inputs to a user-guided stage. The process involves using data from various sources like cameras 206, IMUs 201, ToF sensors 202, microphones 203, and eye-tracking cameras 205 to accurately mark and understand a capture region. This data helps a scene understanding module (e.g., user marking out capture region module 236) to estimate the geometry and semantics of the scene, creating a rough 3D model or coarse mesh. The user then selects specific user marking out the capture region module 236, such as simulation localization and mapping (SLAM) module 231, Segmentation module 222 and Classification Module 212, and coarse mesher 233. The user marking out capture region module 236 uses the learned user profile/user defined custom settings 235 to mark out the capture region and accurately reconstruct the scene's spatial, acoustic, and material properties, either on the device or in the cloud, ensuring a realistic and detailed scene reconstruction. In an embodiment, the SLAM module 231,the segmentation module 222, classification module 212, and coarse mesher 233 may be included to the scene understanding module 238 illustrated in FIG. 2A. The classification module 212 may perform object detection. The classification module 212 and segmentation module 222 are a trained module to identify and localize objects in images and/or scenes, and for instance segmentation in pixel space or object segmentation. SLAM is a method used for 3D scene understanding that builds a map and localize cameras and sensors in the map at the same time. SLAM module 231 allows to map out the environments. The trained model is a neural network which has been trained on labelled data using different techniques in machine learning to predict output for different new unseen inputs to minimize the error in prediction over the known datasets and tested for accuracy on unseen data. Additionally or alternatively, the trained model can also be deterministic algorithm like SLAM module 231 which is a series of steps to generate a 3D reconstruction map and localization of camera/sensor cluster on the map. The capture threshold comprises at least one of: a spatial quality threshold, an acoustic threshold, an illumination threshold, or a material appearance threshold. Each threshold value included in the capture threshold may define the quality of reconstruction metrics for the scene, such that if the reconstruction quality scores from scene reconstruction algorithms exceed the threshold values, the scene is deemed to be accurately reconstructed according to the user's expected level of accuracy. The capture threshold comprises predefined criteria using pre-trained model for the optimal conditions for capturing the texture, spatial, acoustics, and illumination characteristics of the scene. Additionally or alternatively, the thresholds are determined based on the semantics and geometry of the scene using the received sensor data. The scene geometry comprises at least one of three-dimensional structure of the scene. Additionally or alternatively, the three-dimensional structure comprises at least one of: a shape of an object in the scene, a distance between the objects in the scene, or relative positions of the objects in the scene. The texture characteristics of the scene comprises surface details of objects in the scene. The spatial characteristics of the scene comprise arrangement of objects in the scene.
Additionally or alternatively, the acoustic characteristics of the scene comprises sound properties, including echoes and/or ambient noise during the scene capture. Additionally or alternatively, illumination characteristics of the scene comprise a scene capturing lighting condition. The trained model is a neural network which has been trained on labelled data using different techniques in machine learning to predict output for different new unseen inputs to minimize the error in prediction over the known datasets and tested for accuracy on unseen data. Additionally or alternatively, the trained model can be deterministic algorithm like SLAM module 231 which is a series of steps for generation of a 3D reconstruction map and localization of camera/sensor cluster on the map for various metrics for specific and guided actions. The deterministic algorithm may effectively create a layer of abstraction that simplifies the process of achieving a high-quality XR experience, even for individuals without specialized knowledge. A method according to the present disclosure enables an application to meticulously capture and reconstruct physical spaces and objects with high fidelity and multi-dimensional detail. For example, the method may provide live capture guidance by tracking and analyzing one or more key metrics (such as, but not limited to, pose coverage (which ensures accurate spatial positioning); audio impulse coverage (which measures the effectiveness of sound capture); and texture detail score (which assesses the clarity and resolution of visual elements)).
FIG. 2D depicts the process of user guided capture stage based on the user-marked out capture region. The XR system estimates the scene's geometry and key details to create a rough 3D model (e.g., coarse mech of the scene). The user selects capture settings (e.g., texture, acoustics, illumination), which guide the next stage. The selected capture settings or user-marked out capture region may be provided to the user guided capture stage 242. The SLAM module 231 may provide scene understanding to the user guided capture stage 242. The segmentation module 222 and the classification module 212 may provide scene geometry and semantics to the user guided capture stage 242. The coarse mesher 233 may provide a coarse mesh of the scene to the user guided capture stage 242. The model to determine capture settings may provide one or more capture thresholds (e.g., a texture threshold, an acoustic threshold, an illumination threshold, and/or an accuracy threshold) to the user guided capture stage 242.
The user guided capture stage 242 may comprise a user action planner, a user gaze fixation planner, and/or a path and speed planner. The user action planner may recommend an action plan to the user 232. The user gaze fixation planner may recommend one or more fixation points to the user 232. The path and speed planner may provide an optimal path and speed for capture to the user 232. The guided capture stage 242 may provide a guidance comprise at least one of the recommended action plan, the recommended fixation points, or the optimal path and speed for capture to the user 232.
The user guided capture stage 242 may suggest or recommend one or more actions to the user 232. The one or more actions may include, but are not limited to, placing audio source(s) and/or audio receiver(s) (e.g., a phone, buds, and/or a speaker) at indicated position(s) along the guided path for richer acoustic reconstruction. The one or more actions may include, but are not limited to, modifying a lighting in the scene by turning on or off one or more light sources, modifying brightness and/or color of a light source for better capture, and/or modeling how each light source influence lighting of the scene.
FIG. 2E depicts the process of performing multi-dimensional Joint Scene Reconstruction. A method according to the present disclosure involves verifying the capture region marking using XR inputs from various sources such as cameras 206, IMUs 201, ToF sensors 202, mics 203, and eye-tracking cameras 205 illustrated in FIG. 2A. If a region is to be marked, a scene understanding module (e.g., the scene understanding module 238 of FIG. 2A) estimates the scene's geometry and semantics to create a coarse mesh. The coarse mesh depiction is a representation of a scene in 3D, which provides good surface understanding. The mesh density or number of points in the mesh is lower giving a large-scale understanding. Though the coarse mesh depiction covers the mesh density or number of points in the mesh is lower and might miss on minute details, but will be generated rapidly and with lesser sensor data due to lower mesh resolution for dynamic tasks like obstacle avoidance and path planning.
A method according to the present disclosure includes determining one or more capture settings, such as, but not limited to, texture, acoustics, illumination accuracy, and scene type, which are selected by the user and sent to the user guided capture stage 242. A typical listening experience (i.e., acoustics) at a given position in a location can be decided by the distance between the audio sources and the ear (or microphone); and/or dimension(s) of the location and the sound absorption coefficient of the walls and floor of the location (if the given audio scene corresponds to a location). The user guided capture stage 242 can send one or more parameters to a high-dimensional joint scene reconstruction module 240, which performs spatial, acoustic, illumination, and material property reconstructions either on-device or in the cloud.
Examples of the parameters can be, but not limited to, pose-tagged HDR images, depth images, microphone responses, material classification, sound generation time, and so on. The Pose-tagged HDR images capture both high-dynamic-range visuals and the camera's position, aiding 3D scene reconstruction, whereas the depth images map object distances, essential for 3D modeling. The depth information is crucial for building a 3D model of the scene. The audio data captured by microphones, for example, the microphone responses include the sound characteristics to recreate the acoustic environment. Material classification involves identifying and categorizing the materials present in a scene, helping in creating a more accurate reconstruction of the scene. The specific time taken at which a sound is produced within a scene, for example, the signature sound generation time helps in synchronizing the sound with the visual and spatial data, ensuring that the reconstructed scene accurately reflects the original environment. The various actions mentioned above may be performed in the order presented, in a different order or simultaneously. Additionally or alternatively, the method provide real-time guidance during the capture process. This allows users to make informed adjustments on-the-fly, thereby enhancing the overall quality and precision of the XR experience.
Embodiments herein can assess the quality of input data (prior to reconstruction), thereby enhancing the accuracy and relevance of captured data. By using a learned user profile and scene-specific quality settings, embodiments herein provide dynamic, real-time guidance tailored to the user's context during the capture process.
Embodiments herein use a personalized feedback mechanism to ensure that the input data is optimized for the specific needs of various user personas, for example, users involved in social media, architecture, music, and so on. Embodiments herein can employ default and custom profiles to address the unique requirements of different users, enhancing the versatility and applicability of the method across diverse fields and applications. Embodiments herein can improve the user experience by delivering context-aware, actionable insights at the moment of data capture. Further, in some embodiments, some actions listed in FIG. 2A-2E may be omitted.
Users can enhance their audio experience by strategically placing audio sources and receivers (such as, but not limited to, phones, earbuds, and speakers), at one or more specified locations along a suggested path to achieve a richer acoustic reconstruction. Additionally, the users can optimize the lighting within the scene by adjusting existing light sources (such as, by turning them on or off, or by altering their brightness and color settings) to accurately capture and model the influence of each light source on the overall scene lighting. These adjustments will significantly improve both the auditory and visual quality of the scene, creating a more immersive and detailed environment.
Embodiments herein enhance the comprehensiveness and quality of the data collected, thereby facilitating superior XR reconstruction of various scene attributes (such as, but not limited to, geometry, textures, acoustics, and illumination). By improving the fidelity and detail of the captured data, the user can achieve a more accurate and realistic representations of physical environments in XR, leading to more immersive and believable user experiences. This comprehensive data augmentation ensures that every aspect of the scene, from the spatial arrangement and material properties to the sound dynamics and lighting conditions, is meticulously reconstructed to reflect real-world conditions with high precision.
FIGs. 3A-3G depict an example user journey using the proposed method of joint spatial-acoustic-illumination XR scene reconstruction. In an embodiment illustrated in FIGs. 3A-3G, a user 30 uses a XR device 300. Although the XR device 300 is illustrated as a pair of handy controllers, a configuration of the XR device 300 is not limited thereto. For example, the XR device 300 may include a display device (e.g., a head-mounted display apparatus, an augmented reality (AR)/XR helmet, or an AR/VR glasses) which provides a user interface to the user 30. In an embodiment, the XR device 300 may utilize one or more algorithms for tracking one or more gestures of the user 30 without any handy controller.
FIG. 3A illustrates an example image indicating that a user 30 of a XR device 300 has accessed an application for XR scene reconstruction. While the user 30 uses the XR device 300, the XR device 300 may provide a user interface 302 comprising one or more icons representing corresponding functionality. For example, the user interface 302 may comprise at least one of: an icon representing current time, an icon representing a face (or an icon) of the user 30, an icon representing one or more wireless connections of the XR device 300, an icon corresponding to a functionality representing a list of applications installed on the XR device 300, an icon corresponding to a functionality for representing a list of contacts stored on the XR device 300, an icon corresponding to a functionality for representing a list of notifications occurred in the XR device 300, an icon corresponding to a functionality for sharing one or more XR scenes generated by the XR device 300, or an icon corresponding to a functionality for modifying settings of the XR device 300.
In response to an input of the user 30 for icon corresponding to the functionality representing a list of applications installed on the XR device 300, the XR device 300 may provide an application list user interface 304. The application list user interface 304 may comprise one or more icons respectively corresponding to one or more applications of the XR device 300. The XR device may further representing a line user interface 306 which indicates tracked intention of the user 30. The line user interface 306 may represented based on tracked, by the XR device 300, gesture of the user 30. In response to an input for an icon 308, the application for joint spatial-acoustic-illumination XR scene reconstruction may be opened. Moreover, in the depicted example, when the user 30 opens the application for joint spatial-acoustic-illumination XR scene reconstruction, the user 30 can be provided with an interface that prompts them to begin the process.
FIG. 3B illustrates an example image, wherein the application for XR scene reconstruction provides instructions via a guidance user interface 310. The application may provide the guidance user interface 310 which guides the user 30 to walk around the scene and mark the capture region. For example, via the guidance user interface 310, the application instructs the user 30 to walk around the environment and mark the boundaries of the capture region.
FIG. 3C illustrates an example image indicating that the user 30 wearing the XR device 300 confirms that the capture region is marked correctly on the ground plane. In an embodiment illustrated in FIG. 3C, the XR device 300 may comprise a head-mounted wearable device 300c which has an imaging device (for example, a camera). FIG. 3D illustrates an example image indicating an estimated path 312 for optimal capture that has been highlighted for the user 30 to follow. The XR device may display a user interface 314 for asking the user 30 to confirm the estimated path 312 (e.g., the capture region). In response to an input for 'back' icon from the user 30, the application may re-estimate the capture region. In an embodiment of the present disclosure, the gaze fixation point is indicated for the user to focus an imaging device (e.g., a camera of the XR device 300) on. Additionally or alternatively, the walking speed is indicated depending on the details in different parts of the scene. Other actions to augment the richness of the captured data for better textures, acoustics, illumination, material reconstruction are suggested.
For example, in the image on the left, the person is asked to fixate on the point marked to capture more details about the light source. FIGs. 3E and 3F illustrate an example image indicating that the user 30 is guided to walk along a highlighted path 316, and to perform an instructed action for optimization of image capturing. For example, in FIG. 3F, the XR device 300c may guide the user 30 to gaze a certain fixed point (e.g., a point 320) in the environment while walking along the highlighted path 326. As the user 30 moves through the space, the application continuously captures spatial data, while simultaneously collecting information on the scene's acoustic properties and illumination conditions. The application may reconstruct the XR scene based on the collected information, and indicates the user 30 that data associated with the scene is successfully captured by using an user interface 322 (as depicted in the example depicted in FIG. 3G). As the user 30 walks along this highlighted path 316, the user 30 is guided to perform specific action(s) to ensure optimal data capture, including focusing the camera on designated gaze fixation points and adjusting walking speed based on scene details.
The application can suggest one or more additional actions to enhance data richness, such as, but not limited to, fixating on points to capture detailed lighting information. For example, the user 30 can be prompted to focus on a specific point (e.g., the point 320 illustrated in FIG. 3F) to gather more light source details. Furthermore, the application can indicate that the user 30 has correctly marked the capture region on the ground plane. Upon the user 30 completing the walk-around, the application analyzes the gathered data, and extracts detailed scene semantics and geometric information from the gathered data. Thereafter, the scene semantics and geometry information of the marked region is collected, and scene type and capture accuracy thresholds are determined (for example, one or more thresholds for texture, acoustics, and/or illumination accuracy). The application then determines the scene type and sets specific capture accuracy thresholds for texture, acoustics, and illumination. The captured data, including pose-tagged HDR images, depth images, microphone response signals, material classifications, signature sound generation times, and other lighting and acoustic information, can be then processed to create a reconstruction. This processing can be performed either offline (e.g., by the XR device 300) or on the cloud, ensuring comprehensive and high-quality data for accurate scene reconstruction.
This comprehensive dataset ensures that the reconstructed XR scene accurately reflects the real-world environment, providing a highly immersive and realistic experience.
FIG. 4 depicts hardware component of the XR Device 230 comprises of processor 230a, a scene reconstruction controller 230b and a memory 230c. In some embodiments, the XR device 230 may exclude at least one of these components or may further include at least one other component.
The processor 230a includes one or more processing devices or processing circuitry, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processor unit (GPU). The processor 230a is able to perform control on at least one of the other components of the XR device 230.
The scene reconstruction controller (230b) is coupled with the processor (230a) and the memory (230c). The scene reconstruction controller (230b) is configured to obtain information corresponding to at least one of: the semantics of the scene, the geometry of the scene, or acoustic of the scene using the sensor data, while a user moves within the scene during capturing the scene. The scene reconstruction controller 230b is configured to activate user marking of the capture region if the capture region is verified, determine a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, generate a user guidance for an assisted scene capturing using the scene type and the at least one capture thresholds, wherein the user guidance includes at least one of: a path to be followed by the user or at least one action to be performed by the user along the path, and capture the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene, or the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance.
The memory 230c may store commands or data related to at least one other component of the XR device 230. The memory 230c can include volatile memory (e.g., a random-access memory (RAM)) and/or non-volatile memory (e.g., a flash memory or a solid-state drive (SSD)). According to embodiments of the present disclosure, the memory 230c may comprise one or more storage medium which store(s) one or more instructions. The one or more instructions may cause, when executed by the processor 230a and/or the scene reconstruction controller 230b individually or collectively, the XR device 230 to perform any combinations of operations described herein.
FIG. 5 depicts an example virtual experience. In an example scenario, consider that the user utters the following: "I wish I could attend this concert, but I couldn't get my hands on the tickets?". Embodiments herein can provide the user with an immersive virtual concert experience that closely simulates the sensations of attending a live event in person. Embodiments herein can replicate the intricate details of concert lighting and sound systems, thereby ensuring that the user can experience the atmosphere and ambiance of a live performance with remarkable fidelity. Embodiments herein can capture the visual and auditory nuances of a concert, and can also extend the experience beyond geographical limitations, thereby allowing friends and loved ones to join in and enjoy the event together, regardless of their location. This enhanced social interaction amplifies the enjoyment and creates a shared, memorable experience, bridging the gap between physical presence and virtual participation.
FIG. 6 depicts an example virtual experience. In an example scenario, consider that the user utters the following: "I want to share my experience of this museum exhibit with my friends in an immersive way?". Embodiments herein can capture a high-fidelity reconstruction of scene in 3D. The user can share this with friends who can view the exhibit from any viewpoint in an immersive way.
FIG. 7 depicts an example virtual experience. In an example scenario, consider that the user utters the following: "I want to have a meeting, but my team works in a hybrid mode across multiple offices, and it is hard to communicate effectively with just video meetings". Embodiments herein offer the potential to create a high-dimensional reconstruction of a meeting room, meticulously modeling not just the physical layout of the meeting room, but also the acoustics and lighting in the meeting room, thereby ensuring a highly immersive virtual environment. Embodiments herein can capture and replication of in-person interactions accurately, including subtle aspects such as tone of voice, gestures, and body language of the participants in the meeting, all presented in a 1:1 scale. By integrating these elements into the virtual reconstruction, embodiments herein effectively bridge the communication gap, allowing for interactions that closely mimic face-to-face meetings. This level of detail enhances the realism of virtual engagements, making them more effective and engaging compared to traditional virtual communication platforms.
FIG. 8 depicts an example virtual experience. In an example scenario, consider that the user utters the following: "when designing or remodeling houses, it is time consuming to design with computer software as I do not get a good mental image of the scale and structure of a building. Also, I need to be on-site at any moment to review construction progress". Embodiments herein can facilitate an architect to obtain a proper scaled 3D reconstruction, which can help with design and additionally with acoustic and illumination modelling. This enables the architect to try out different materials, layouts, and styles, and experience the design from any location within the reconstructed scene. Also, on-site workers can share the current progress of a project in the 3D reconstruction, when the architect is off-site. Furthermore, embodiments herein can significantly enhance the architectural design process by enabling precise scaled 3D reconstructions, which are crucial for accurate design, acoustic analysis, and illumination modelling. Embodiments herein allow architects to experiment with various materials and design styles in a virtual environment, offering a comprehensive perspective from any vantage point within the reconstructed scene. Additionally, embodiments herein facilitates real-time collaboration by allowing on-site workers to update and share the current progress of a project with the architect remotely, ensuring continuous and effective communication even when the architect is off-site. This seamless integration of technology not only optimizes design accuracy and efficiency, but also enhances the overall management and execution of architectural projects.
FIGS. 9A-9B is an example scenario, depicting the difference between a spatial video and proposed joint spatial-acoustic-illumination XR scene reconstruction. FIG. 9A depicts a spatial video frame viewpoint that is restricted to the stereo camera's capture position. FIG. 9B depicts a video frame of the scene using the proposed joint spatial-acoustic-illumination XR scene reconstruction, where the scene can be viewed from any viewpoint at any distance.
FIG. 10 depicts an example scenario, wherein real-time guidance and feedback is used for capturing the scene. For reconstruction of an XR scene, the XR device may start capturing data from surroundings. The XR device may suggest following a path and performing one or more actions (e.g., capturing one or more points nearby with the XR device) to the user. While the user follows the suggestion, the XR device may collect data from the surroundings, and reconstruct the scene based on the collected data.
FIGs. 11A-11B depict the illumination and acoustics of the reconstructed scene that can be modified using the proposed joint spatial-acoustic-illumination XR scene reconstruction. From an environment illustrated in FIG. 11A, a jointed XR experience may be reconstructed, and provided to a user as illustrated in FIG. 11B. The jointed XR experience may include one or spatial characteristics (e.g., one or more objects included in the environment), one or more illumination characteristics (e.g., one or more lightings 1102), and one or more acoustic characteristics (e.g., one or more sound elements 1104). One or more characteristics included in the joint XR experience may be customized. For example, the one or more lightings 1102 and/or the one or more sound elements 1104 may be customized (or adjusted) by the user.
In spatial video, viewers may be limited to the perspective from which the video was initially recorded, typically bound to the positions of the original camera(s). Accordingly, users may be confined to a fixed viewpoint, limiting their ability to explore the scene dynamically. Spatial video technology offers the potential for users to explore a scene from various angles and distances, providing a more immersive and flexible experience. In an embodiment, this dynamic exploration is facilitated by a calibrated stereo camera setup, which records the scene from multiple viewpoints. A method according to an embodiment of the present disclosure revolutionizes this approach by eliminating the need for such complex equipment. It allows for capturing spatial video with just a single camera, thus simplifying the recording process while still enabling users to interact with and view the scene from different perspectives. This advancement opens up new possibilities for more accessible and versatile spatial video applications.
A method according to an embodiment of the present disclosure delivers real-time feedback and active guidance to users, enhancing their ability to capture high-quality XR experiences. By integrating high-dimensional capture within a single framework, it enables users to record spatial richness, including pose, acoustic, and illumination elements. Accurate audio recreation is fundamental to creating a realistic and immersive XR experience, as it enhances spatial awareness, facilitates social interaction, fosters emotional connections, and provides a nuanced understanding of the environment. Similarly, precise scene lighting modelling is vital for lifelike rendering, affecting the perception of depth, shape, and texture, thereby ensuring a believable experience. This approach allows users dynamic control over the scene, making applications like XR design and digital twins more practical and effective.
FIGs. 12A & 12B depicts a flow chart with live user assistance flow. FIGs. 12A & 12B depicts a system 200 for providing user guidance for reconstructing a multidimensional XR scene using a server connected to a XR Device 230. The XR device 230 may comprise one or more modules illustrated in FIGs. 12A and 12B. The XR Device 230 is configured to receive sensor data from a plurality of sensors 201-205 of FIG. 2A(on a user initiating the capture of a scene), derive information on the scene's semantics and geometry using the received data, and activate a user marking out capture region module 236 of FIG. 2A (on the region being successfully verified). The user marking out capture region module 236 creates a coarse semantic scene mesh based on the derived information, while optionally allowing the user to mark a capture region. The XR Device 230 determines the scene type and capture accuracy thresholds using a pre-trained model, wherein the pre-trained model uses texture, spatial, acoustics, and/or illumination characteristics from a metric semantic localization & mapping module 210, a material type classifier module 212, and a shadow region classifier & reconstruction module 214. The XR Device 230 using the pre-trained model, generates user guidance for a second round of scene capturing, which includes a path and actions for the user. During this second round, the user captures the multidimensional scene data following the provided guidance.
FIGs. 13A & 13B depicts a system 400 detailed flow-chart with an XR scene reconstruction flow. The XR device 230 may comprise one or more modules illustrated in FIGs. 13A and 13B. The XR Device 230 is configured to create a coarse semantic scene mesh using the derived information, determine the scene type and capture thresholds for multidimensional data, determine if a capture region is marked on a ground plane, and highlight an optimal capture path. The capture threshold defines the threshold value for quality of reconstruction metrics of the scene. If the metrics or reconstruction quality scores from scene reconstruction algorithms improve above the threshold, the scene is deemed to be good at having reconstructed the environment with accuracy as expected by user through the threshold. The thresholds are set at the beginning, based on user history, environment, user inputs, and user preference. The reconstruction quality in each of these modalities or dimensions depends on these threshold values. The capture path guides the user along this path, provides real-time visual feedback for reconstructing the scene (wherein the scene encompasses spatial, acoustic, illumination, and material appearance characteristics), and offers user assistance for creating multi-dimensional XR scenes (which integrates integrating texture, spatial, audio, visual, and illumination experiences). The XR Device 230 is configured to evaluate data quality, provide guidance based on user profiles and scenes, and/or support both default and custom user personas tailored to the user and/or application. A scene segmentation module 222 creates a coarse scene mesh by estimating the semantics and geometry of a scene, while the user is marking a capture region. The pre-trained model determines the scene type and sets capture thresholds for multi-dimensional scene data, including texture, spatial, acoustic, and illumination characteristics. A ground plane confirmation module verifies if the capture region is correctly marked on the ground plane. A path highlighting module (e.g., a path planner module 218 of FIG. 12B) highlights an optimal capture path for the user. A user guidance module 220 of FIG. 12B instructs the user to follow this path and perform necessary actions. A data processing module processes the captured multi-dimensional scene data to reconstruct an XR scene, while a real-time feedback module provides visual feedback for scene reconstruction. A live user assistance module offers real-time support for creating the multi-dimensional XR scene. An evaluation and guidance module benchmarks reconstruction input data quality and provides live guidance based on user profiles and scenes. A user persona support module 211 of FIG. 12A accommodates default and custom user personas 213 of FIG. 12A based on different user needs.
The XR Device 230 can locate the audio sources within a room using Direction of Arrival (DOA) algorithms combined with visual cues extracted from recorded video or images. The XR Device 230 can further estimate one or more location related parameters (such as, but not limited to, dimensions and absorption values) from recorded video or images. The XR Device 230 can perform object segmentation for identifying material types and their properties, which can be used as initial conditions. Subsequently, the XR Device 230 can fine-tune these parameters through analysis of recorded audio response spectra. The XR Device 230 can derive dimensions and geometry of the location from a combination of sparse Simultaneous Localization And Mapping (SLAM) maps and preliminary dense view synthesis techniques using Neural Radiance Fields (NERF) or Signed Distance Fields (SDF). The XR Device 230 can isolate the audio information using sound source separation algorithms, and the XR Device 230 can apply de-reverberation algorithms to these isolated signals to mitigate the influence of location acoustics on each component.
Modelling and recreating the listening experience using the proposed method includes estimating the Room Impulse Response (RIR), and user guidance for recording the RIR. The RIR is the transfer function between the sound source and the microphone. T60 and space dimensions can be used for estimating the RIR, wherein T60 is the time taken for the sound to decay by 60dB. T60 can be different at different locations in the location. The XR Device 230 can provide guidance to the user to record RIRs at probable locations in the space where the listener is located (for example, in front of the TV, chairs, etc.)
The XR Device 230 can generate RIRs for each audio information based on the estimated source positions and room parameters. The XR Device 230 can apply these RIRs to the individual audio sources, and combine the results to recreate the audio experience at a new listening point, effectively simulating the sound as if heard from this new location.
Recreating high-fidelity spatial reconstructions involves media capture techniques that ensure the accurate representation of textures and occlusions in 3D models, which includes verifying the texture richness and managing occlusions effectively.
FIGS. 14A & 14B depict examples of surfaces with varying levels of texture richness, where one surface displays a higher level of detail compared to another surface with a lesser texture richness. Consequently, embodiments herein can be capable of identifying regions with varying texture richness, which can be achieved by analyzing texture granularity in the captured data, which can be computed using histograms of edge data in the media. For capturing media with enhanced texture richness, embodiments herein can actively guide the users during the capture process, directing them to areas/surfaces requiring closer inspection or additional focus. By doing so, the XR device 230 can provide effective guidance to improve the fidelity of the spatial reconstruction.
To detect occluded real-world objects within a scene, the XR Device 230 can provide guidance to a user to conduct a comprehensive 360-degree view and execute a minimal walk-through of the space. This preliminary exploration can help the XR Device 230 initialize and estimate how objects obscure one another (if any). In an embodiment herein, the XR Device 230 can assess the occlusion coverage using a ray hit test, which involves calculating the intersection of rays with objects and identifying triangles within the field of view that are obscured by one or more other objects. The XR Device 230 can estimate the occlusion at various points on a 3D grid, based on the direction of view of the user. By integrating these occlusion estimates with additional matrices, the XR Device 230 can create and visualize a detailed representation of the occlusion. Embodiments herein ensure that the occlusion analysis captures the complexities of how objects in the scene obstruct each other from different perspectives.
The spatial reconstruction process involves creating a detailed and accurate 3D representation of a space using a mesh grid system that records the quality of coverage at various points, as illustrated in FIG. 15. Embodiments herein also disclose a user interface (UI) to help visualize the extent of coverage within the 3D space through a miniature map. In an embodiment herein, the miniature map features a heat map that visually represents the density of coverage. In an embodiment herein, the heat map can be colour coded, wherein colors are used to indicate the number of observed points; for example, red signifies areas with extensive coverage and multiple viewpoints, yellow highlights regions with less coverage, and so on. Embodiments herein can determine a coverage score, wherein the coverage score(s) can be encoded into the vertices of the 3D grid structure, thereby allowing the heat map to be dynamically rendered from any camera angle, providing an intuitive and interactive way to assess spatial reconstruction quality. In an embodiment, the position of each vertex of the 3D grid structure can be optimized. The coverage score is determined by analyzing the 3D representation of a space using a mesh grid system. The system records coverage quality at various points. The coverage score is encoded into the vertices of the 3D grid, enabling a heat map that visually represents coverage density. Colors like red and yellow indicate areas with different coverage levels. The heat map can be dynamically rendered from any angle, helping assess spatial reconstruction quality.
As depicted in FIG. 16, a method according to an embodiment of the present disclosure provides modelling and recreating illumination within the location based on the direction of the light source. The quality of capture of the media can be heavily dependent on the light source direction. For example, the quality of capture of the media will be better in the direction facing the light source, and the quality of capture of the media will be worse in the direction facing away from the light source. To compensate for the worse quality of capture, the user needs to improve the lighting or make more number of captures from the direction facing away from the light source. For example, an application has a lighting estimation application programming interface (API) which provides detailed information on the lighting of a scene, using cues, such as, but not limited to, shadows, ambient light, shading, specular highlights, reflections, and so on. Embodiments herein can determine a quality metric associated with a viewing direction from a weighted representation of the main direction of the direction of the light and ambient spherical harmonics.
Thus, a user can capture an accurate acoustic reconstruction of a scene with the guidance of the proposed method, without the need to hire professional help and equipment. Hence, the factors mentioned above such as 3D experience, several new sensors, requirement for sophisticated guidance methods makes this method most suitable for the VST when compared to traditional 2D devices like mobiles or computers.
FIG. 17 is a flow chart depicting a method (1700) for providing user guidance for reconstructing a multi-dimensional XR scene according to an embodiment of the present disclosure. The one or more steps in the method 1700 may be performed by the XR device 230. In step 1702, the XR Device 230 (e.g., a head-mounted display (HMD) device) receives data from one or more sensors included in the XR device 230, when the user is performing an initial scene capture. In step 1704, the sensor data can be used to derive information corresponding to semantics and geometry of the scene during the first round of the scene capture. In step 1706, a coarse semantic scene mesh can be created using the derived information while optionally marking a capture region. In step 1708, the scene type and capture accuracy thresholds can be determined using a pre-trained model, with multi-dimensional scene data including texture, spatial, acoustics, and illumination characteristics. In step 1710, user guidance can be generated for a second scene capture based on the scene type and capture thresholds, which includes a path for the user to follow. In step 1712 the multi-dimensional scene data can be captured during the second round using the generated user guidance. The various actions in method 1700 may be performed in the order presented, in a different order or simultaneously. Further, in some embodiments, some actions listed in FIG. 17 may be omitted.
FIG. 18 is a flow chart depicting a method (1800) for reconstructing a multi-dimensional XR scene. The XR Device 230 can reconstruct a multi-dimensional extended reality (XR) scene. In step 1802, the XR Device 230 creates a coarse semantic scene mesh by utilizing derived information related to the semantics and geometry of the scene as the user 232 optionally marks a capture region. In step 1804, the XR Device 230 determines the scene type and establishes one or more capture accuracy thresholds using a pre-trained model to capture multi-dimensional scene data, including texture, spatial, acoustics, and illumination characteristics. In step 1806, the XR Device 230 confirms if the capture region is on a ground plane, and in step 1808, the XR Device 230 highlights an estimated path for the user to follow for effective scene capture. In step 1810, as the user navigates this path, the XR Device 230 provides guidance to ensure accurate scene data collection. In step 1810, the XR Device 230 processes the captured data to reconstruct the XR scene, which incorporates spatial, acoustic, illumination, and material appearance characteristics into a single framework, offering real-time visual feedback to achieve high-quality, rich XR scene capture with customizable and realistic lighting, acoustics, and virtual object manipulation in step 1812. Furthermore, in step 1814, the XR Device 230 provides user assistance for creating multi-dimensional XR scenes by integrating texture, spatial, audio, visual, and illumination experiences within a unified framework. It also serves as a benchmark to evaluate the quality of reconstruction input data, offering user guidance tailored to individual profiles and scenes while supporting both default and custom user personas to cater to diverse user needs. The various actions in method 1800 may be performed in the order presented, in a different order or simultaneously. Further, in some embodiments, some actions listed in FIG. 18 may be omitted.
Embodiments herein disclose can enhance the collected data to improve the XR reconstruction of a scene by refining its geometry, textures, acoustics, and lighting. Embodiments herein process and enrich the raw data to create a more accurate and immersive virtual or augmented representation of the real-world environment, ensuring that the reconstructed scene closely resembles the physical space in terms of visual details, sound quality, and lighting effects.
The embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the network elements. The elements include blocks which can be at least one of a hardware device, or a combination of hardware device and software module.
The embodiments disclosed herein describe a method for user guidance for reconstructing a multi-dimensional extended reality (XR) scene using an XR Device and providing real-time feedback and active guidance to the user when capturing data. Therefore, it is understood that the scope of the protection is extended to such a program and in addition to a computer readable means having a message therein, such computer readable storage means (e.g., computer-readable storage medium) contain (or store) program code means for implementation of one or more steps of the method, when the program runs on a server or mobile device or any suitable programmable device. The method is implemented in at least one embodiment through or together with a software program written in e.g., Very high speed integrated circuit Hardware Description Language (VHDL) another programming language, or implemented by one or more VHDL or several software modules being executed on at least one hardware device. The hardware device can be any kind of portable device that can be programmed. The device may also include means which could be e.g., hardware means like e.g., an ASIC, or a combination of hardware and software means, e.g., an ASIC and an FPGA, or at least one microprocessor and at least one memory with software modules located therein. The method embodiments described herein could be implemented partly in hardware and partly in software. Alternatively, the invention may be implemented on different hardware devices, e.g., using a plurality of CPUs.
The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and/or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of embodiments and examples, those skilled in the art will recognize that the embodiments and examples disclosed herein can be practiced with modification within the scope of the embodiments as described herein.

Claims (15)

  1. A method for reconstructing a multi-dimensional extended reality (XR) scene, the method comprising:
    obtaining, by an XR device (230), information corresponding to at least one of: semantics of a scene, geometry of the scene, or an acoustic of the scene, using a sensor data, while a user of the XR device (230) moves within the scene during capturing the scene;
    determining, by the XR device (230), a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, wherein the multi-dimensional scene comprises at least one of: texture characteristics of the scene, spatial characteristics of the scene, acoustics characteristics of the scene, or illumination characteristics of the scene;
    generating, by the XR device (230), a user guidance for an assisted scene capturing using the scene type and the at least one capture threshold, wherein the user guidance includes at least one of: a path to be followed by the user or at least one action to be performed by the user along the path; and
    capturing, by the XR device (230), the multi-dimensional scene representing at least one of: the spatial characteristics of the scene, the illumination characteristics of the scene, or the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance.
  2. The method as claimed in claim 1, wherein obtaining the information comprises:
    obtaining, by the XR device (230), the sensor data from at least one sensor in the XR device (230) in response to the user of the XR device (230) initiating capturing the scene.
  3. The method as claimed in claim 1 or claim 2, wherein the method further comprises performing, by the XR device (230), at least one of: indicating a gaze fixation point for the user to focus, indicating a walking speed depending on at least one detail in a part of the scene, or augmenting a captured data for at least one of optimized texture, optimized acoustics, optimized illumination, or optimized material reconstruction in the scene.
  4. The method as claimed in any one of claims 1 to claims 3, wherein capturing, by the XR device (230), the multi-dimensional scene comprises:
    processing at least one of: a tag associated with an image, a depth of the image, a microphone response signal, a material classification associated with the scene, a sound generation time, lighting information, or acoustic information; and
    capturing, by the XR device (230), the multi-dimensional scene based on a result of the processing.
  5. The method as claimed in any one of claims 1 to claims 4, wherein the multi-dimensional scene is determined by:
    generating a map upon initiating a traverse through an environment;
    marking, in the map, an area in which the user intended to travel in the environment and usage;
    estimating a ground plane, a user height and a pose region in the environment; and
    estimating an occlusion blind spot by ray propagation from a pose region in the environment.
  6. The method as claimed in any one of claims 1 to claims 5, wherein the semantics of the scene is obtained by using a trained object detection and segmentation module, and wherein the geometry of the scene is obtained by using Simultaneous Localization And Mapping (SLAM) technique to formulate a coarse mesh depiction of the scene.
  7. The method as claimed in any one of claims 1 to 6, wherein the user guidance comprises modifying a lighting in the scene by at least one of: turning on a light, turning off the light, adjusting a brightness of the light, or changing a colour of the light to model scene lighting with different light source.
  8. The method as claimed in any one of claims 1 to 7, wherein the capture threshold comprises predefined criteria using pre-trained model for optimal conditions for capturing the texture, spatial, acoustics, and illumination characteristics of the scene wherein each threshold included in the capture threshold determined based on the semantics and geometry of the scene using the received sensor data.
  9. The method as claimed in any one of claims 1 to 8, wherein the geometry of the scene comprises at least one of three-dimensional structure of the scene, wherein the three-dimensional structure comprises at least one of: a shape of an object in the scene, a distance between the objects in the scene, or relative positions of the objects in the scene.
  10. A XR device (230), comprises,
    at least one processor (230a) including processing circuitry; and
    a memory (230b) comprising one or more storage media storing one or more instructions,
    wherein, when executed by the at least one processor individually or collectively, the one or more instructions cause the XR device (230) to:
    obtain information corresponding to at least one of: semantics of a scene, geometry of a scene, or an acoustic of a scene, using a sensor data, while a user moves within the scene during the capturing of the scene;
    determine a scene type and at least one capture threshold for capturing a multi-dimensional scene using a trained model based on the obtained information, wherein the multi-dimensional scene comprises at least one of: texture characteristics of the scene, spatial characteristics of the scene, acoustics characteristics of the scene, or illumination characteristics of the scene;
    generate a user guidance for an assisted scene capturing using the scene type and the at least one capture thresholds, wherein the user guidance includes at least one of: a path to be followed by the user and at least one action to be performed by the user along the path; and
    capture the multi-dimensional scene representing at least one of the spatial characteristics of the scene, the illumination characteristics of the scene, or the acoustic characteristics of the scene during the assisted scene capture using the generated user guidance.
  11. The XR device (230) as claimed in claim 10, wherein, when executed by the at least one processor individually or collectively, the one or more instructions further cause the XR device (230) to:
    obtain the sensor data from at least one sensor in the XR device (230) when the user of the XR device (230) initiates capturing of the scene.
  12. The XR device (230) as claimed in claim 10 or claim 11, wherein, when executed by the at least one processor individually or collectively, the one or more instructions further cause the XR device (230) to perform at least of: indicating a gaze fixation point for the user to focus, to indicate a walking speed depending on at least one detail in a part of the scene, or augmenting a captured data for at least one optimized texture, optimized acoustics, optimized illumination, or optimized material reconstruction in the scene.
  13. The XR device (230) as claimed in any one of claims 10 to 12, wherein, when executed by the at least one processor individually or collectively, the one or more instructions further cause the XR device (230) to:
    process at least one of: a tag associated with an image, a depth of the image, a microphone response signal, a material classification associated with the scene, a sound generation time, lighting information, or acoustic information; and
    capture the multi-dimensional scene based on a result of the processing.
  14. The XR device (230) as claimed in any one of claims 10 to 13, wherein, to determine the multi-dimensional scene, when executed by the at least one processor individually or collectively, the one or more instructions further cause the XR device (230) to:
    generate a map upon initiating a traverse through an environment;
    mark, in the map, an area in which the user intended to travel in the environment and usage;
    estimate a ground plane, a user height and a pose region in the environment; and
    estimate an occlusion blind spot by ray propagation from a pose region in the environment.
  15. The XR device (230) as claimed in any one of claims 10 to 14, wherein, when executed by the at least one processor individually or collectively, the one or more instructions further cause the XR device (230) to obtain semantics of the scene by using a trained object detection and segmentation module, and to obtain the geometry of the scene by using Simultaneous Localization And Mapping (SLAM) technique to formulate a coarse mesh depiction of the scene.
PCT/KR2025/000786 2024-04-04 2025-01-14 Method and system for reconstructing multi-dimensional extended reality scene Pending WO2025211547A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
IN202441028031 2024-04-04
IN202441028031 2024-12-30

Publications (1)

Publication Number Publication Date
WO2025211547A1 true WO2025211547A1 (en) 2025-10-09

Family

ID=97269159

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2025/000786 Pending WO2025211547A1 (en) 2024-04-04 2025-01-14 Method and system for reconstructing multi-dimensional extended reality scene

Country Status (1)

Country Link
WO (1) WO2025211547A1 (en)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20160234432A1 (en) * 2013-10-28 2016-08-11 Olympus Corporation Image processing apparatus and image processing method
KR101867051B1 (en) * 2011-12-16 2018-06-14 삼성전자주식회사 Image pickup apparatus, method for providing composition of pickup and computer-readable recording medium
US20190349562A1 (en) * 2017-02-14 2019-11-14 Samsung Electronics Co., Ltd. Method for providing interface for acquiring image of subject, and electronic device
CN111754569A (en) * 2020-06-28 2020-10-09 中国银行股份有限公司 Method, device and system for alarming number of people in room, electronic equipment and storage medium
KR20240012449A (en) * 2021-07-13 2024-01-29 엘지전자 주식회사 Route guidance device and route guidance system based on augmented reality and mixed reality

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101867051B1 (en) * 2011-12-16 2018-06-14 삼성전자주식회사 Image pickup apparatus, method for providing composition of pickup and computer-readable recording medium
US20160234432A1 (en) * 2013-10-28 2016-08-11 Olympus Corporation Image processing apparatus and image processing method
US20190349562A1 (en) * 2017-02-14 2019-11-14 Samsung Electronics Co., Ltd. Method for providing interface for acquiring image of subject, and electronic device
CN111754569A (en) * 2020-06-28 2020-10-09 中国银行股份有限公司 Method, device and system for alarming number of people in room, electronic equipment and storage medium
KR20240012449A (en) * 2021-07-13 2024-01-29 엘지전자 주식회사 Route guidance device and route guidance system based on augmented reality and mixed reality

Similar Documents

Publication Publication Date Title
Yang et al. Audio augmented reality: A systematic review of technologies, applications, and future research directions
TWI813098B (en) Neural blending for novel view synthesis
CN102413414B (en) System and method for high-precision 3-dimensional audio for augmented reality
JP6377082B2 (en) Providing a remote immersive experience using a mirror metaphor
TWI647593B (en) System and method for providing simulated environment
KR20220125358A (en) Systems, methods and media for displaying real-time visualizations of physical environments in artificial reality
Garg et al. Geometry-aware multi-task learning for binaural audio generation from video
Geronazzo et al. Applying a single-notch metric to image-guided head-related transfer function selection for improved vertical localization
US20230401789A1 (en) Methods and systems for unified rendering of light and sound content for a simulated 3d environment
Garg et al. Visually-guided audio spatialization in video with geometry-aware multi-task learning
Kim et al. Immersive audio-visual scene reproduction using semantic scene reconstruction from 360 cameras
CN118138789A (en) A digital human live broadcast method, device, equipment, medium and program product
Privitera et al. On the effect of user tracking on perceived source positions in mobile audio augmented reality
CN111881807A (en) VR conference control system and method based on face modeling and expression tracking
US12087090B2 (en) Information processing system and information processing method
WO2025211547A1 (en) Method and system for reconstructing multi-dimensional extended reality scene
CN117292094B (en) Digitalized application method and system for performance theatre in karst cave
US20240119619A1 (en) Deep aperture
Siegl et al. An augmented reality human–computer interface for object localization in a cognitive vision system
Chang et al. Applying deep learning and building information modeling to indoor positioning based on sound
Thery et al. Impact of the visual rendering system on subjective auralization assessment in VR
US12462508B1 (en) User representation based on an anchored recording
Córdova-Esparza et al. Telepresence system based on simulated holographic display
Henson We’re In This Together: Embodied Interaction, Affect, and Design Methods in Asymmetric, Co-Located, Co-Present Mixed Reality
Menzer Preliminary study on integrating 3d audio with 2d game engines for immersive storytelling

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25782853

Country of ref document: EP

Kind code of ref document: A1