WO2026019672A1 - Systems and methods for providing intelligent embodied interactive agents with spatial understanding - Google Patents

Systems and methods for providing intelligent embodied interactive agents with spatial understanding

Info

Publication number
WO2026019672A1
WO2026019672A1 PCT/US2025/037380 US2025037380W WO2026019672A1 WO 2026019672 A1 WO2026019672 A1 WO 2026019672A1 US 2025037380 W US2025037380 W US 2025037380W WO 2026019672 A1 WO2026019672 A1 WO 2026019672A1
Authority
WO
WIPO (PCT)
Prior art keywords
language model
user
augmented reality
embodied
reality headset
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2025/037380
Other languages
French (fr)
Inventor
Pranav DESHPANDE
Mengyu CHEN
Elvir Azanli
Monica LANDERLANDER
Kristine W MA
Joseph W Ligman
Marco Pistoia
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
JPMorgan Chase Bank NA
Original Assignee
JPMorgan Chase Bank NA
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by JPMorgan Chase Bank NA filed Critical JPMorgan Chase Bank NA
Publication of WO2026019672A1 publication Critical patent/WO2026019672A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T19/00Manipulating three-dimensional [3D] models or images for computer graphics
    • G06T19/006Mixed reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/16Sound input; Sound output
    • G06F3/167Audio in a user interface, e.g. using voice commands for navigating, audio feedback
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/20Three-dimensional [3D] animation
    • G06T13/40Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/20Scenes; Scene-specific elements in augmented reality scenes
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/18Speech classification or search using natural language modelling
    • G10L15/183Speech classification or search using natural language modelling using context dependencies, e.g. language models
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F2203/00Indexing scheme relating to G06F3/00 - G06F3/048
    • G06F2203/038Indexing scheme relating to G06F3/038
    • G06F2203/0381Multimodal input, i.e. interface arrangements enabling the user to issue commands by simultaneous use of input devices of different nature, e.g. voice plus gesture on digitizer
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/226Procedures used during a speech recognition process, e.g. man-machine dialogue using non-speech characteristics

Definitions

  • Embodiments relate to systems and methods for providing intelligent embodied interactive agents with spatial understanding.
  • a method may include: (1) receiving, by a conversational artificial intelligence engine and from an augmented reality headset worn by a user, a query that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset, wherein the query may include audio of a user utterance and images or video captured by a camera of the augmented reality headset of what the user is seeing; (2) generating, by the conversational artificial intelligence engine, a prompt for a large language model based on the user utterance and the images or video; (3) providing, by the conversational artificial intelligence engine, the prompt to the large language model; (4) receiving, by the conversational artificial intelligence engine, an output of the large language model, wherein the output may include text and gestures for the embodied interactive agent; (5) generating, by the conversational artificial intelligence engine, animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and (6)
  • the method may also include receiving, by the conversational artificial intelligence engine, user location information from the augmented reality headset, wherein the prompt may be further based on the user location information.
  • the method may also include inferring, by the conversational artificial intelligence engine, a task goal associated with the query, wherein the inference may be based on a user interaction history, environment object labels, and user location information.
  • the conversational artificial intelligence engine generates a text prompt for the large language model based on the user utterance and an image prompt for a visual language model, and the large language model and the visual language model return outputs.
  • the large language model may include a multi-modal large language model.
  • the large language model further outputs an identification of a document to provide to the augmented reality headset.
  • the display in the augmented reality headset displays the animations for the embodied interactive agent, and a speaker in the augmented reality headset outputs the speech for the embodied interactive agent.
  • a system may include: an augmented reality headset comprising a camera, a microphone, a display, and a speaker, wherein the augmented reality headset may be configured to be worn by a user; and a multi-modal conversational platform comprising a conversational artificial intelligence engine that may be configured to receive, from the augmented reality headset, a query that is made to an embodied interactive agent that is displayed by the display, wherein the query may include audio of a user utterance and images or video captured by the camera of what the user is seeing; to generate a prompt for a large language model based on the user utterance and the images or video; to provide the prompt to the large language model, to receive an output of the large language model, wherein the output may include text and gestures for the embodied interactive agent, to generate animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and to output the animations and the speech to the augmented reality headset.
  • the display of the augmented reality headset comprising a camera,
  • the conversational artificial intelligence engine may be further configured to receive user location information from the augmented reality headset, and the prompt may be further based on the user location information.
  • the conversational artificial intelligence engine may be further configured to infer a task goal associated with the query, wherein the inference may be based on a user interaction histoiy, environment object labels, and user location information.
  • the conversational artificial intelligence engine may be further configured to generate a text prompt for the large language model based on the user utterance, and an image prompt for a visual language model, and the large language model and the visual language model return outputs.
  • the large language model may include a multi-modal large language model.
  • the large language model further outputs an identification of a document to provide to the augmented reality headset.
  • a non-transitory computer readable storage medium may include instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving, from an augmented reality headset worn by a user, a query that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset, wherein the query may include audio of a user utterance and images or video captured by a camera of the augmented reality headset of what the user is seeing; generating a prompt for a large language model based on the user utterance and the images or video; providing the prompt to the large language model; receiving an output of the large language model, wherein the output may include text and gestures for the embodied interactive agent; generating animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and outputting the animations and the speech to the augmented reality headset.
  • the non-transitoiy computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving user location information from the augmented reality headset, wherein the prompt may be further based on the user location information.
  • the non-transitory computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: inferring a task goal associated with the query, wherein the inference may be based on a user interaction history, enviromnent object labels, and user location information.
  • the non-transitory computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: generating, a text prompt for the large language model based on the user utterance and an image prompt for a visual language model, and receiving outputs from the large language model and the visual language model.
  • the large language model may include a multi-modal large language model.
  • the non-transitory computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving from the large language model further, an identification of a document to provide to the augmented reality headset.
  • the display in the augmented reality headset displays the animations for the embodied interactive agent, and a speaker in the augmented reality headset outputs the speech for the embodied interactive agent.
  • Figure 1 illustrates a system for providing intelligent embodied interactive agents with spatial understanding according to an embodiment
  • Figure 2 illustrates an exemplary architecture for a conversational Al engine according to an embodiment
  • Figure 3 illustrates an exemplary architecture for a conversational Al engine according to another embodiment
  • Figure 4 illustrates an exemplary architecture for a conversational Al engine according to another embodiment
  • Figure 5 illustrates a system for providing intelligent embodied interactive agents with spatial understanding according to an embodiment
  • Figure 6 depicts an exemplary computing system for implementing aspects of the present disclosure.
  • Embodiments relate to systems and methods for providing intelligent embodied interactive agents with spatial understanding.
  • Embodiments may provide a system with an embodied agent (e.g., a virtual avatar, a 3D avatar, etc.) that can have an intelligent conversation with the user, backed by multiple APIs for fetching data and also capable of making API calls to alter data elsewhere.
  • the conversational backend may include a large language model and text- to- speech and speech- to-text models.
  • the agents are capable of spatial understanding as they are integrated with vision-language understanding for understanding the scene. Hence, not only can the agent interact with digital computing APIs and fetch/edit data; it can also offer an understanding of the user’s scene view from the headset and compute upon the input data.
  • Embodiments may perform task inference based on the user’s input command, and the scene along with the knowledge available to it.
  • the embodied agent has spatial scene awareness as perceived from the mounted cameras on the user’s headset.
  • Task inference may be based on spatial and auditory awareness and user commands.
  • Spatial and auditoiy awareness are part of the scene understanding process, using the mounted cameras images and feeding them into a visual language model or a multi-modal large language model.
  • These models generally do an object recognition, segmentation and localization, as well as text recognition (if recognized any) to extract the necessary information from the user's surrounding environment. They also associate individual objects with their possible states (trained knowledge in the model) and identify the most possible event that is happening with these objects or human activities. With this level of spatial awareness, the user can ask for suggestions of how to deal with certain situations (e.g., find out where some objects are located, or search for more specific information about a given object).
  • the large language capability behind the visual language model reasons about some suggestions for what the user asked (e.g., what should I do to better organize the room layout, how much to renovate this room).
  • Such knowledge can either come from a training/fine-tuning databases (e.g., public database, the Internet, government documents, etc.), from Internet searches perfoimed by the Al engine, etc.
  • System 100 may include headset 110, which may be worn by a user.
  • Headset 110 may include, for example, display 112, speakers or other audio output 114, microphone 118, and camera 116.
  • Display 112 may display imagery for a user inside of headset 110, including an embodied agent, imagery from a virtual area, pass-through imagery from the surrounding area (e.g., augmented reality), documents, etc.
  • Speakers 114 may output audio, such as sounds, text, etc.
  • Camera 116 may capture a view of what the user is seeing and may be mounted with a forward-looking orientation. This may include objects in the surrounding area, gestures made by the user’s appendages, etc. Imagery from camera 116 may be provided to video streaming client 142 in multi-modal conversational platform 130. [0036] Images from camera 116 may also be provided to user behavior tracking module 140.
  • Microphone 118 may capture audible utterances from the user, as well as ambient sounds in the surrounding area. Audio captured by microphone 118 may be provided to user behavior tracking module 140 in multi-modal conversational platform 130. User behavior tracking module 140 may also receive, for example, gestures, interaction history, user texts, user location information, images from camera 116, etc.
  • Multi-modal conversational platform 130 may be provided with conversational artificial intelligence (Al) engine 132.
  • Conversational Al engine 132 may generate voice output, text output, and gesture output for an embodied agent, such as an avatar, a virtual person, etc.
  • Conversational Al engine 132 may also identify documents or other data from document database 120 to provide to headset 110.
  • Conversational Al engine 132 may have several architectures, which will be discussed below in conjunction with Figures 2, 3, and 4.
  • Conversational Al engine 132 may receive, as inputs, information on the user’s situation, such as a context of the interaction with the user, the user’s location (e.g., in a bank, at a merchant location, at home, etc.), etc. This information may be used as an input to a task inference algorithm to assist the language model in determining the information that is to be provided to the end user.
  • information on the user’s situation such as a context of the interaction with the user, the user’s location (e.g., in a bank, at a merchant location, at home, etc.), etc.
  • This information may be used as an input to a task inference algorithm to assist the language model in determining the information that is to be provided to the end user.
  • Conversational Al engine 132 may output gestures for the embodied agent to body animation synthesizer 134 and facial animation synthesizer 136.
  • Body animation synthesizer 134 and facial animation synthesizer 136 may generate commands for animating the body and facial expressions for the embodied agent.
  • Voice and animation synchronizer 138 may receive the commands from body animation synthesizer 134 and facial animation synthesizer 136, as well as audio from conversational Al engine 132, so that the visual animation is synchronized with the audio.
  • the output of voice and animation synchronizer 138 may be provided to headset 110, where the animation may be displayed on display 112, and the audio output by speakers 114.
  • Conversational Al engine 132 may output natural language responses to text builder 137, which may provide text output to headset 110 for display by display 112.
  • Multi-modal conversational platform 130 may also identify documents in document database 120 to be displayed on display 112.
  • User prompt and image handler 144 may receive the output of user behavior tracking module 140 and video streaming client 142 and may identify user prompts, questions, gestures, etc., as well as an identification of what the user is looking at. For example, user prompt and image handler 144 may determine whether the user is looking at the embodied agent, data being displayed on display 112, or something else.
  • FIG. 2 illustrates an exemplary architecture for a conversational Al engine according to an embodiment.
  • Conversational Al engine 132 may receive the output of user prompt and image handler 144, which may include speech, text, and images.
  • Text prompt engine 214 may receive audio and/or text and may generate a text prompt, such as “What is the user asking for by this utterance: [text].” This may be provided to large language model 220, which may return a response to the prompt to Al engine 230.
  • text prompt engine 214 may collect and build text entries from the user and may connect some preset questions with the input to prepare the final prompts that are sent to LLM 220 to generate appropriate responses.
  • a prompt sentence may be “User input: “ + “I want to know what my card balance is.” + “Instruction: construct your answer to this user input in j son format, and put your answers in an order of, ‘response:’, ‘emotion:’, ‘gesture:’, ‘user task:’, ‘action call’) so that LLM 220 may output a query data structure that can be parsed and converted into corresponding avatar control function calls.
  • Image prompt engine 216 may receive streaming video from the camera on headset 110, and may generate an image prompt, such as “what is the user looking at” with the image or video. This prompt may be provided to visual language model 224, which may be similar to a large language model, but may interpret images or video instead of text.
  • Al engine 230 may include text-to-speech model 232, which may convert text to speech for the audio response, natural language response model 234, which may be the response that is to be spoken by the avatar and may be sent to text builder 137 which may generate text for display as a subtitle, body animation inference model 236, which may infer what a proper avatar animation should be (e.g., sending the spoken sentence into a lip-sync animation generator which returns a queiy of facial animation blendshape weights, and a gesture cue (e.g., wave hand, stretch) to trigger the body animation on the avatar, etc.) and task inference/document retrieval 238, which may infer the task goal for the user.
  • text-to-speech model 232 which may convert text to speech for the audio response
  • natural language response model 234 which may be the response that is to be spoken by the avatar and may be sent to text builder 137 which may generate text for display as a subtitle
  • user interaction history e.g., a history of queries from the user and responses
  • enviromnent object labels e.g., labels for objects identified in the images or video, such as “chair,” “couch,” “car,” etc.
  • current user location information e.g., in a furniture store, at a car dealership, etc.
  • user hand gestures, etc. may be used as a text prompt by text prompt engine 214, and LLM 220 may generate a corresponding estimation of what the user’s task goal.
  • the output of text-to-speech model 232 may be output to voice and animation synchronizer 138
  • the output of natural language response e.g., text output
  • the output of body animation inference model 236 may be output to body animation synthesizer 134 and facial animation synthesizer 136
  • the output of task inference/document retrieval 238 may be provided to retrieve a document from document database 120.
  • the outputs of modules 120, 134, 136, 137, and 138 may then be provided to headset 110 as illustrated in Figure 1.
  • Figure 3 illustrates an exemplary architecture for a conversational Al engine according to another embodiment.
  • conversational Al engine 132 may receive the output of user prompt and image handler 144, which may include speech, text, and images.
  • multi-modal large language model 302 may be provided to respond to the prompts from text prompt engine 214 and image prompt engine 216.
  • the output of multi-modal large language model 320 may be provided to Al engine 230 for processing as described with regard to Figure 2.
  • multi-modal large language model 302 may receive one or more prompts involving multiple modalities - such as a text modality and an image modality - and may output text responses with the additional capability of outputting gestures.
  • Figure 4 illustrates an exemplary architecture for a conversational Al engine according to another embodiment.
  • conversational Al engine 132 may receive the output of user prompt and image handler 144, which may include speech, text, and images, and may further provide an audio prompt to multi-modal large language model 320 using audio prompt engine 412.
  • multi-modal large language model 320 may output audio directly to voice synchronizer 138. In another embodiment, multimodal large language model 320 may output audio directly to headset 110.
  • Figure 5 illustrates a system for providing intelligent embodied interactive agents with spatial understanding according to an embodiment.
  • a user may wear a headset.
  • the headset may be a virtual reality headset, an augmented reality headset, etc.
  • the headset may include a display to display a virtual reality or augmented reality view, one or more speakers to output audio, a microphone to capture the user’s audio (e.g., speech), and one or more camera that may capture what the user is looking at.
  • the camera may also capture gestures made by the user, such as the user pointing or gesturing at something.
  • the headset or another connected device may identify a location of the user, such as in a store, at a bank, at home, etc.
  • step 510 the user may initiate an interaction with an Al agent.
  • an Al agent For example, while in augmented reality mode, the user may look at a furniture set, and may ask the Al agent about the cost and financing for the furniture set.
  • the conversational Al engine may receive the audio of user’s query, video from user’s headset, and the user location information.
  • the video may include what the user is looking at (e.g., an object, such as a furniture set).
  • the conversational Al engine may generate one or more prompts for a large language model and/or a visual language model, or a multi-modal large language model, based on the audio and video received from the headset, the location information, and any inferred task information.
  • the conversational Al engine may convert the audio from the user headset to text, and a text prompt engine may generate a prompt for a large language model or multi-modal language model using the text.
  • An image prompt engine may generate a prompt for a visual language model, or a multi-modal language model based on the image(s) or video received from the user headset.
  • Additional information such as user location information, inferred task information, etc. may also be provided to the large language model and/or the visual language model, or the multi-modal large language model.
  • the conversational Al engine may receive output(s) from the large language model and/or the visual language model, or the multi-modal large language model in response to the queries.
  • the large language model and/or the visual language model, or the multimodal large language model may return text for a voice output, a text output, a gesture output, and a document output.
  • an Al engine may generate speech, facial animations, body animations, and retrieves information, such as documents, to present to the headset based on the output(s) of the large language model and/or the visual language model, or the multi-modal large language model.
  • the Al engine may use a text to speech model to generate audio of a text output, a natural language response model to generate text in a natural language format to be displayed to the user (e.g., as a subtitle), a body animation inference model that may provide output to a body animation synthesizer and a facial animation synthesizer to generate animation for the Al agent, and a task inference/document retrieval model that may retrieve a document to present to the user from a document database.
  • a text to speech model to generate audio of a text output
  • a natural language response model to generate text in a natural language format to be displayed to the user (e.g., as a subtitle)
  • a body animation inference model that may provide output to a body animation synthesizer and a facial animation synthesizer to generate animation for the Al agent
  • a task inference/document retrieval model that may retrieve a document to present to the user from a document database.
  • the animations and the speech of the Al agent may be synchronized using a voice and animation synchronizer.
  • the Al agent may be animated with facial and body animations and the animation may be presented in the headset display.
  • step 540 the audio of the text (i.e., speech) may be output to the speakers in the headset.
  • the documentary information may be presented in headset display.
  • the conversational Al engine may retrieve documentary information using APIs and may present those on the display.
  • Figure 6 depicts an exemplary computing system for implementing aspects of the present disclosure.
  • Figure 6 depicts exemplary computing device 600.
  • Computing device 600 may represent the system components described herein.
  • Computing device 600 may include processor 605 that may be coupled to memory 610.
  • Memory 610 may include volatile memory.
  • Processor 605 may execute computer-executable program code stored in memory 610, such as software programs 615.
  • Software programs 615 may include one or more of the logical steps disclosed herein as a programmatic instruction, which may be executed by processor 605.
  • Memory 610 may also include data repository 620, which may be nonvolatile memory for data persistence.
  • Processor 605 and memory 610 may be coupled by bus 630.
  • Bus 630 may also be coupled to one or more network interface connectors 640, such as wired network interface 642 or wireless network interface 644.
  • Computing device 600 may also have user interface components, such as a screen for displaying graphical user interfaces and receiving input from the user, a mouse, a keyboard and/or other input/output components (not shown).
  • Embodiments of the system or portions of the system may be in the form of a “processing machine,” such as a general-purpose computer, for example.
  • the tenn “processing machine” is to be understood to include at least one processor that uses at least one memory.
  • the at least one memory stores a set of instructions.
  • the instructions may be either permanently or temporarily stored in the memory or memories of the processing machine.
  • the processor executes the instructions that are stored in the memory or memories in order to process data.
  • the set of instructions may include various instructions that perform a particular task or tasks, such as those tasks described above. Such a set of instructions for performing a particular task may be characterized as a program, software program, or simply software.
  • the processing machine may be a specialized processor.
  • the processing machine may be a cloudbased processing machine, a physical processing machine, or combinations thereof.
  • the processing machine executes the instructions that ar e stored in the memory or memories to process data.
  • This processing of data may be in response to commands by a user or users of the processing machine, in response to previous processing, in response to a request by another processing machine and/or any other input, for example.
  • the processing machine used to implement embodiments may be a general-purpose computer.
  • the processing machine described above may also utilize any of a wide variety of other technologies including a special purpose computer, a computer system including, for example, a microcomputer, mini-computer or mainframe, a programmed microprocessor, a micro-controller, a peripheral integrated circuit element, a CSIC (Customer Specific Integrated Circuit) or ASIC (Application Specific Integrated Circuit) or other integrated circuit, a logic circuit, a digital signal processor, a programmable logic device such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), PLA (Programmable Logic Array), or PAL (Programmable Array Logic), or any other device or arrangement of devices that is capable of implementing the steps of the processes disclosed herein.
  • a programmable logic device such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), PLA (Programmable Logic Array), or PAL
  • the processing machine used to implement embodiments may utilize a suitable operating system.
  • each of the processors and/or the memories of the processing machine may be located in geographically distinct locations and connected so as to communicate in any suitable manner.
  • each of the processor and/or the memory may be composed of different physical pieces of equipment. Accordingly, it is not necessary that the processor be one single piece of equipment in one location and that the memory be another single piece of equipment in another location. That is, it is contemplated that the processor may be two pieces of equipment in two different physical locations. The two distinct pieces of equipment may be connected in any suitable manner. Additionally, the memory may include two or more portions of memory in two or more physical locations.
  • processing is performed by various components and various memories.
  • processing perfoimed by two distinct components as described above in accordance with a further embodiment, may be performed by a single component.
  • processing performed by one distinct component as described above may be performed by two distinct components.
  • the memory storage performed by two distinct memory portions as described above may be performed by a single memory portion. Further, the memory storage performed by one distinct memory portion as described above may be performed by two memory portions.
  • various technologies may be used to provide communication between the various processors and/or memories, as well as to allow the processors and/or the memories to communicate with any other entity; i.e., so as to obtain further instructions or to access and use remote memory stores, for example.
  • Such technologies used to provide such communication might include a network, the Internet, Intranet, Extranet, a LAN, an Ethernet, wireless communication via cell tower or satellite, or any client server system that provides communication, for example.
  • Such communications technologies may use any suitable protocol such as TCP/IP, UDP, or OSI, for example.
  • a set of instructions may be used in the processing of embodiments.
  • the set of instructions may be in the fonn of a program or software.
  • the software may be in the form of system software or application software, for example.
  • the software might also be in the form of a collection of separate programs, a program module within a larger program, or a portion of a program module, for example.
  • the software used might also include modular programming in the form of object-oriented programming. The software tells the processing machine what to do with the data being processed.
  • the instructions or set of instructions used in the implementation and operation of embodiments may be in a suitable form such that the processing machine may read the instructions.
  • the instructions that form a program may be in the form of a suitable programming language, which is converted to machine language or object code to allow the processor or processors to read the instructions. That is, written lines of programming code or source code, in a particular programming language, are converted to machine language using a compiler, assembler or interpreter.
  • the machine language is binary coded machine instructions that are specific to a particular type of processing machine, i.e., to a particular type of computer, for example. The computer understands the machine language.
  • Any suitable programming language may be used in accordance with the various embodiments.
  • the instructions and/or data used in the practice of embodiments may utilize any compression or enciyption technique or algorithm, as may be desired.
  • An encryption module might be used to encrypt data.
  • files or other data may be decrypted using a suitable decryption module, for example.
  • the embodiments may illustratively be embodied in the form of a processing machine, including a computer or computer system, for example, that includes at least one memory.
  • the set of instructions i.e., the software for example, that enables the computer operating system to perform the operations described above may be contained on any of a wide variety of media or medium, as desired.
  • the data that is processed by the set of instructions might also be contained on any of a wide variety of media or medium. That is, the particular medium, i.e., the memory in the processing machine, utilized to hold the set of instructions and/or the data used in embodiments may take on any of a variety of physical forms or transmissions, for example.
  • the medium may be in the form of a compact disc, a DVD, an integrated circuit, a hard disk, a floppy disk, an optical disc, a magnetic tape, a RAM, a ROM, a PROM, an EPROM, a wire, a cable, a fiber, a communications channel, a satellite transmission, a memory card, a SIM card, or other remote transmission, as well as any other medium or source of data that may be read by the processors.
  • the memory or memories used in the processing machine that implements embodiments may be in any of a wide variety of forms to allow the memory to hold instructions, data, or other information, as is desired.
  • the memory might be in the form of a database to hold data.
  • the database might use any desired arrangement of files such as a flat file arrangement or a relational database arrangement, for example.
  • a user interface includes any hardware, software, or combination of hardware and software used by the processing machine that allows a user to interact with the processing machine.
  • a user interface may be in the form of a dialogue screen for example.
  • a user interface may also include any of a mouse, touch screen, keyboard, keypad, voice reader, voice recognizer, dialogue screen, menu box, list, checkbox, toggle switch, a pushbutton or any other device that allows a user to receive information regarding the operation of the processing machine as it processes a set of instructions and/or provides the processing machine with infonnation.
  • the user interface is any device that provides communication between a user and a processing machine.
  • the information provided by the user to the processing machine through the user interface may be in the form of a command, a selection of data, or some other input, for example.
  • a user interface is utilized by the processing machine that perfonns a set of instructions such that the processing machine processes data for a user.
  • the user interface is typically used by the processing machine for interacting with a user either to convey information or receive information from the user.
  • the user interface might interact, i.e., convey and receive information, with another processing machine, rather than a human user. Accordingly, the other processing machine might be characterized as a user.
  • a user interface utilized in the system and method may interact partially with another processing machine or processing machines, while also interacting partially with a human user.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Software Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • Computer Hardware Design (AREA)
  • Computer Graphics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • General Health & Medical Sciences (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

A method may include: a conversational artificial intelligence engine receiving from an augmented reality headset worn by a user, a query including audio and images/video that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset; the conversational artificial intelligence engine generating a prompt for a large language model based on the user utterance and the images or video; the conversational artificial intelligence engine providing the prompt to the large language model; the conversational artificial intelligence engine receiving an output of the large language model, wherein the output comprises text and gestures for the embodied interactive agent; the conversational artificial intelligence engine generating animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and the conversational artificial intelligence engine outputting the animations and the speech to the augmented reality headset.

Description

SYSTEMS AND METHODS FOR PROVIDING INTELLIGENT EMBODIED INTERACTIVE AGENTS WITH SPATIAL UNDERSTANDING
BACKGROUND OF THE INVENTION
1. Field of the Invention
[0001] Embodiments relate to systems and methods for providing intelligent embodied interactive agents with spatial understanding.
2. Description of the Related Art
[0002] The introduction of mixed reality devices and headsets has led to the creation of a new paradigm of computing -spatial computing. Spatial computing offers the user with an opportunity to interact with the environment and the computer in various ways. There is, however, a lack of interactive agents with spatial understanding in the mixed reality and spatial computing space.
SUMMARY OF THE INVENTION
[0003] Systems and methods for providing intelligent embodied interactive agents with spatial understanding are disclosed. According to one embodiment, a method may include: (1) receiving, by a conversational artificial intelligence engine and from an augmented reality headset worn by a user, a query that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset, wherein the query may include audio of a user utterance and images or video captured by a camera of the augmented reality headset of what the user is seeing; (2) generating, by the conversational artificial intelligence engine, a prompt for a large language model based on the user utterance and the images or video; (3) providing, by the conversational artificial intelligence engine, the prompt to the large language model; (4) receiving, by the conversational artificial intelligence engine, an output of the large language model, wherein the output may include text and gestures for the embodied interactive agent; (5) generating, by the conversational artificial intelligence engine, animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and (6) outputting, by the conversational artificial intelligence engine, the animations and the speech to the augmented reality headset.
[0004] In one embodiment, the method may also include receiving, by the conversational artificial intelligence engine, user location information from the augmented reality headset, wherein the prompt may be further based on the user location information.
[0005] In one embodiment, the method may also include inferring, by the conversational artificial intelligence engine, a task goal associated with the query, wherein the inference may be based on a user interaction history, environment object labels, and user location information.
[0006] In one embodiment, the conversational artificial intelligence engine generates a text prompt for the large language model based on the user utterance and an image prompt for a visual language model, and the large language model and the visual language model return outputs.
[0007] In one embodiment, the large language model may include a multi-modal large language model.
[0008] In one embodiment, the large language model further outputs an identification of a document to provide to the augmented reality headset. [0009] In one embodiment, the display in the augmented reality headset displays the animations for the embodied interactive agent, and a speaker in the augmented reality headset outputs the speech for the embodied interactive agent.
[0010] According to another embodiment, a system may include: an augmented reality headset comprising a camera, a microphone, a display, and a speaker, wherein the augmented reality headset may be configured to be worn by a user; and a multi-modal conversational platform comprising a conversational artificial intelligence engine that may be configured to receive, from the augmented reality headset, a query that is made to an embodied interactive agent that is displayed by the display, wherein the query may include audio of a user utterance and images or video captured by the camera of what the user is seeing; to generate a prompt for a large language model based on the user utterance and the images or video; to provide the prompt to the large language model, to receive an output of the large language model, wherein the output may include text and gestures for the embodied interactive agent, to generate animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and to output the animations and the speech to the augmented reality headset. The display of the augmented reality headset is configured to display the animations of the embodied interactive agent, and the speaker is configured to output the speech of the embodied interactive agent.
[0011] In one embodiment, the conversational artificial intelligence engine may be further configured to receive user location information from the augmented reality headset, and the prompt may be further based on the user location information. [0012] In one embodiment, the conversational artificial intelligence engine may be further configured to infer a task goal associated with the query, wherein the inference may be based on a user interaction histoiy, environment object labels, and user location information.
[0013] In one embodiment, the conversational artificial intelligence engine may be further configured to generate a text prompt for the large language model based on the user utterance, and an image prompt for a visual language model, and the large language model and the visual language model return outputs.
[0014] In one embodiment, the large language model may include a multi-modal large language model.
[0015] In one embodiment, the large language model further outputs an identification of a document to provide to the augmented reality headset.
[0016] According to another embodiment, a non-transitory computer readable storage medium may include instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving, from an augmented reality headset worn by a user, a query that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset, wherein the query may include audio of a user utterance and images or video captured by a camera of the augmented reality headset of what the user is seeing; generating a prompt for a large language model based on the user utterance and the images or video; providing the prompt to the large language model; receiving an output of the large language model, wherein the output may include text and gestures for the embodied interactive agent; generating animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and outputting the animations and the speech to the augmented reality headset.
[0017] In one embodiment, the non-transitoiy computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving user location information from the augmented reality headset, wherein the prompt may be further based on the user location information.
[0018] In one embodiment, the non-transitory computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: inferring a task goal associated with the query, wherein the inference may be based on a user interaction history, enviromnent object labels, and user location information.
[0019] In one embodiment, the non-transitory computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: generating, a text prompt for the large language model based on the user utterance and an image prompt for a visual language model, and receiving outputs from the large language model and the visual language model.
[0020] In one embodiment, the large language model may include a multi-modal large language model.
[0021] In one embodiment, the non-transitory computer readable storage medium may also include instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving from the large language model further, an identification of a document to provide to the augmented reality headset.
[0022] In one embodiment, the display in the augmented reality headset displays the animations for the embodied interactive agent, and a speaker in the augmented reality headset outputs the speech for the embodied interactive agent.
BRIEF DESCRIPTION OF THE DRAWINGS
[0023] For a more complete understanding of the present invention, the objects and advantages thereof, reference is now made to the following descriptions taken in connection with the accompanying drawings in which:
[0024] Figure 1 illustrates a system for providing intelligent embodied interactive agents with spatial understanding according to an embodiment;
[0025] Figure 2 illustrates an exemplary architecture for a conversational Al engine according to an embodiment;
[0026] Figure 3 illustrates an exemplary architecture for a conversational Al engine according to another embodiment;
[0027] Figure 4 illustrates an exemplary architecture for a conversational Al engine according to another embodiment;
[0028] Figure 5 illustrates a system for providing intelligent embodied interactive agents with spatial understanding according to an embodiment; and [0029] Figure 6 depicts an exemplary computing system for implementing aspects of the present disclosure.
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0030] Embodiments relate to systems and methods for providing intelligent embodied interactive agents with spatial understanding.
[0031] Embodiments may provide a system with an embodied agent (e.g., a virtual avatar, a 3D avatar, etc.) that can have an intelligent conversation with the user, backed by multiple APIs for fetching data and also capable of making API calls to alter data elsewhere. The conversational backend may include a large language model and text- to- speech and speech- to-text models. Moreover, the agents are capable of spatial understanding as they are integrated with vision-language understanding for understanding the scene. Hence, not only can the agent interact with digital computing APIs and fetch/edit data; it can also offer an understanding of the user’s scene view from the headset and compute upon the input data. Embodiments may perform task inference based on the user’s input command, and the scene along with the knowledge available to it.
[0032] In embodiments, the embodied agent has spatial scene awareness as perceived from the mounted cameras on the user’s headset. Task inference may be based on spatial and auditory awareness and user commands.
[0033] Spatial and auditoiy awareness are part of the scene understanding process, using the mounted cameras images and feeding them into a visual language model or a multi-modal large language model. These models generally do an object recognition, segmentation and localization, as well as text recognition (if recognized any) to extract the necessary information from the user's surrounding environment. They also associate individual objects with their possible states (trained knowledge in the model) and identify the most possible event that is happening with these objects or human activities. With this level of spatial awareness, the user can ask for suggestions of how to deal with certain situations (e.g., find out where some objects are located, or search for more specific information about a given object). The large language capability behind the visual language model reasons about some suggestions for what the user asked (e.g., what should I do to better organize the room layout, how much to renovate this room). Such knowledge can either come from a training/fine-tuning databases (e.g., public database, the Internet, government documents, etc.), from Internet searches perfoimed by the Al engine, etc.
[0034] Referring to Figure 1, a system for providing intelligent embodied interactive agents with spatial understanding is provided according to an embodiment. System 100 may include headset 110, which may be worn by a user. Headset 110 may include, for example, display 112, speakers or other audio output 114, microphone 118, and camera 116. Display 112 may display imagery for a user inside of headset 110, including an embodied agent, imagery from a virtual area, pass-through imagery from the surrounding area (e.g., augmented reality), documents, etc. Speakers 114 may output audio, such as sounds, text, etc.
[0035] Camera 116 may capture a view of what the user is seeing and may be mounted with a forward-looking orientation. This may include objects in the surrounding area, gestures made by the user’s appendages, etc. Imagery from camera 116 may be provided to video streaming client 142 in multi-modal conversational platform 130. [0036] Images from camera 116 may also be provided to user behavior tracking module 140.
[0037] Microphone 118 may capture audible utterances from the user, as well as ambient sounds in the surrounding area. Audio captured by microphone 118 may be provided to user behavior tracking module 140 in multi-modal conversational platform 130. User behavior tracking module 140 may also receive, for example, gestures, interaction history, user texts, user location information, images from camera 116, etc.
[0038] Multi-modal conversational platform 130 may be provided with conversational artificial intelligence (Al) engine 132. Conversational Al engine 132 may generate voice output, text output, and gesture output for an embodied agent, such as an avatar, a virtual person, etc. Conversational Al engine 132 may also identify documents or other data from document database 120 to provide to headset 110. Conversational Al engine 132 may have several architectures, which will be discussed below in conjunction with Figures 2, 3, and 4.
[0039] Conversational Al engine 132 may receive, as inputs, information on the user’s situation, such as a context of the interaction with the user, the user’s location (e.g., in a bank, at a merchant location, at home, etc.), etc. This information may be used as an input to a task inference algorithm to assist the language model in determining the information that is to be provided to the end user.
[0040] Conversational Al engine 132 may output gestures for the embodied agent to body animation synthesizer 134 and facial animation synthesizer 136. Body animation synthesizer 134 and facial animation synthesizer 136 may generate commands for animating the body and facial expressions for the embodied agent.
[0041] Voice and animation synchronizer 138 may receive the commands from body animation synthesizer 134 and facial animation synthesizer 136, as well as audio from conversational Al engine 132, so that the visual animation is synchronized with the audio. The output of voice and animation synchronizer 138 may be provided to headset 110, where the animation may be displayed on display 112, and the audio output by speakers 114.
[0042] Conversational Al engine 132 may output natural language responses to text builder 137, which may provide text output to headset 110 for display by display 112.
[0043] Multi-modal conversational platform 130 may also identify documents in document database 120 to be displayed on display 112.
[0044] User prompt and image handler 144 may receive the output of user behavior tracking module 140 and video streaming client 142 and may identify user prompts, questions, gestures, etc., as well as an identification of what the user is looking at. For example, user prompt and image handler 144 may determine whether the user is looking at the embodied agent, data being displayed on display 112, or something else.
[0045] Figure 2 illustrates an exemplary architecture for a conversational Al engine according to an embodiment. Conversational Al engine 132 may receive the output of user prompt and image handler 144, which may include speech, text, and images.
[0046] Text prompt engine 214 may receive audio and/or text and may generate a text prompt, such as “What is the user asking for by this utterance: [text].” This may be provided to large language model 220, which may return a response to the prompt to Al engine 230.
[0047] In one embodiment, text prompt engine 214 may collect and build text entries from the user and may connect some preset questions with the input to prepare the final prompts that are sent to LLM 220 to generate appropriate responses.
[0048] For example, a prompt sentence may be “User input: “ + “I want to know what my card balance is.” + “Instruction: construct your answer to this user input in j son format, and put your answers in an order of, ‘response:’, ‘emotion:’, ‘gesture:’, ‘user task:’, ‘action call’) so that LLM 220 may output a query data structure that can be parsed and converted into corresponding avatar control function calls.
[0049] Image prompt engine 216 may receive streaming video from the camera on headset 110, and may generate an image prompt, such as “what is the user looking at” with the image or video. This prompt may be provided to visual language model 224, which may be similar to a large language model, but may interpret images or video instead of text.
[0050] The outputs of large language model 220 and visual language model 224 may be provided as inputs to Al engine 230. Al engine 230 may include text-to-speech model 232, which may convert text to speech for the audio response, natural language response model 234, which may be the response that is to be spoken by the avatar and may be sent to text builder 137 which may generate text for display as a subtitle, body animation inference model 236, which may infer what a proper avatar animation should be (e.g., sending the spoken sentence into a lip-sync animation generator which returns a queiy of facial animation blendshape weights, and a gesture cue (e.g., wave hand, stretch) to trigger the body animation on the avatar, etc.) and task inference/document retrieval 238, which may infer the task goal for the user. For example, user interaction history (e.g., a history of queries from the user and responses), enviromnent object labels (e.g., labels for objects identified in the images or video, such as “chair,” “couch,” “car,” etc.), current user location information (e.g., in a furniture store, at a car dealership, etc.), user hand gestures, etc. may be used as a text prompt by text prompt engine 214, and LLM 220 may generate a corresponding estimation of what the user’s task goal.
[0051] In one embodiment, the output of text-to-speech model 232 (e.g., a voice output) may be output to voice and animation synchronizer 138, the output of natural language response (e.g., text output) may be output to text builder 137, the output of body animation inference model 236 may be output to body animation synthesizer 134 and facial animation synthesizer 136, and the output of task inference/document retrieval 238 may be provided to retrieve a document from document database 120. The outputs of modules 120, 134, 136, 137, and 138 may then be provided to headset 110 as illustrated in Figure 1.
[0052] Figure 3 illustrates an exemplary architecture for a conversational Al engine according to another embodiment. In Figure 3, conversational Al engine 132 may receive the output of user prompt and image handler 144, which may include speech, text, and images.
[0053] Instead of large language model 220 and visual language model 224, a single model, multi-modal large language model 302 may be provided to respond to the prompts from text prompt engine 214 and image prompt engine 216. The output of multi-modal large language model 320 may be provided to Al engine 230 for processing as described with regard to Figure 2.
[0054] In one embodiment, multi-modal large language model 302 may receive one or more prompts involving multiple modalities - such as a text modality and an image modality - and may output text responses with the additional capability of outputting gestures.
[0055] Figure 4 illustrates an exemplary architecture for a conversational Al engine according to another embodiment. In Figure 4, conversational Al engine 132 may receive the output of user prompt and image handler 144, which may include speech, text, and images, and may further provide an audio prompt to multi-modal large language model 320 using audio prompt engine 412.
[0056] In addition, multi-modal large language model 320 may output audio directly to voice synchronizer 138. In another embodiment, multimodal large language model 320 may output audio directly to headset 110.
[0057] Figure 5 illustrates a system for providing intelligent embodied interactive agents with spatial understanding according to an embodiment.
[0058] In step 505, a user may wear a headset. The headset may be a virtual reality headset, an augmented reality headset, etc. The headset may include a display to display a virtual reality or augmented reality view, one or more speakers to output audio, a microphone to capture the user’s audio (e.g., speech), and one or more camera that may capture what the user is looking at.
[0059] The camera may also capture gestures made by the user, such as the user pointing or gesturing at something. [0060] The headset or another connected device may identify a location of the user, such as in a store, at a bank, at home, etc.
[0061] In step 510, the user may initiate an interaction with an Al agent. For example, while in augmented reality mode, the user may look at a furniture set, and may ask the Al agent about the cost and financing for the furniture set.
[0062] In step 515, the conversational Al engine may receive the audio of user’s query, video from user’s headset, and the user location information. The video may include what the user is looking at (e.g., an object, such as a furniture set).
[0063] In step 520, the conversational Al engine may generate one or more prompts for a large language model and/or a visual language model, or a multi-modal large language model, based on the audio and video received from the headset, the location information, and any inferred task information. For example, the conversational Al engine may convert the audio from the user headset to text, and a text prompt engine may generate a prompt for a large language model or multi-modal language model using the text. An image prompt engine may generate a prompt for a visual language model, or a multi-modal language model based on the image(s) or video received from the user headset.
[0064] Additional information, such as user location information, inferred task information, etc. may also be provided to the large language model and/or the visual language model, or the multi-modal large language model.
[0065] In step 525, the conversational Al engine may receive output(s) from the large language model and/or the visual language model, or the multi-modal large language model in response to the queries. For example, the large language model and/or the visual language model, or the multimodal large language model may return text for a voice output, a text output, a gesture output, and a document output.
[0066] In step 530, an Al engine may generate speech, facial animations, body animations, and retrieves information, such as documents, to present to the headset based on the output(s) of the large language model and/or the visual language model, or the multi-modal large language model.
[0067] For example, the Al engine may use a text to speech model to generate audio of a text output, a natural language response model to generate text in a natural language format to be displayed to the user (e.g., as a subtitle), a body animation inference model that may provide output to a body animation synthesizer and a facial animation synthesizer to generate animation for the Al agent, and a task inference/document retrieval model that may retrieve a document to present to the user from a document database.
[0068] In one embodiment, the animations and the speech of the Al agent may be synchronized using a voice and animation synchronizer.
[0069] In step 535, the Al agent may be animated with facial and body animations and the animation may be presented in the headset display.
[0070] In step 540, the audio of the text (i.e., speech) may be output to the speakers in the headset.
[0071] In step 545, the documentary information may be presented in headset display. For example, the conversational Al engine may retrieve documentary information using APIs and may present those on the display.
[0072] Figure 6 depicts an exemplary computing system for implementing aspects of the present disclosure. Figure 6 depicts exemplary computing device 600. Computing device 600 may represent the system components described herein. Computing device 600 may include processor 605 that may be coupled to memory 610. Memory 610 may include volatile memory. Processor 605 may execute computer-executable program code stored in memory 610, such as software programs 615. Software programs 615 may include one or more of the logical steps disclosed herein as a programmatic instruction, which may be executed by processor 605. Memory 610 may also include data repository 620, which may be nonvolatile memory for data persistence. Processor 605 and memory 610 may be coupled by bus 630. Bus 630 may also be coupled to one or more network interface connectors 640, such as wired network interface 642 or wireless network interface 644. Computing device 600 may also have user interface components, such as a screen for displaying graphical user interfaces and receiving input from the user, a mouse, a keyboard and/or other input/output components (not shown).
[0073] Hereinafter, general aspects of implementation of the systems and methods of embodiments will be described.
[0074] Embodiments of the system or portions of the system may be in the form of a “processing machine,” such as a general-purpose computer, for example. As used herein, the tenn “processing machine” is to be understood to include at least one processor that uses at least one memory. The at least one memory stores a set of instructions. The instructions may be either permanently or temporarily stored in the memory or memories of the processing machine. The processor executes the instructions that are stored in the memory or memories in order to process data. The set of instructions may include various instructions that perform a particular task or tasks, such as those tasks described above. Such a set of instructions for performing a particular task may be characterized as a program, software program, or simply software.
[0075] In one embodiment, the processing machine may be a specialized processor.
[0076] In one embodiment, the processing machine may be a cloudbased processing machine, a physical processing machine, or combinations thereof.
[0077] As noted above, the processing machine executes the instructions that ar e stored in the memory or memories to process data. This processing of data may be in response to commands by a user or users of the processing machine, in response to previous processing, in response to a request by another processing machine and/or any other input, for example.
[0078] As noted above, the processing machine used to implement embodiments may be a general-purpose computer. However, the processing machine described above may also utilize any of a wide variety of other technologies including a special purpose computer, a computer system including, for example, a microcomputer, mini-computer or mainframe, a programmed microprocessor, a micro-controller, a peripheral integrated circuit element, a CSIC (Customer Specific Integrated Circuit) or ASIC (Application Specific Integrated Circuit) or other integrated circuit, a logic circuit, a digital signal processor, a programmable logic device such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), PLA (Programmable Logic Array), or PAL (Programmable Array Logic), or any other device or arrangement of devices that is capable of implementing the steps of the processes disclosed herein.
[0079] The processing machine used to implement embodiments may utilize a suitable operating system.
[0080] It is appreciated that in order to practice the method of the embodiments as described above, it is not necessary that the processors and/or the memories of the processing machine be physically located in the same geographical place. That is, each of the processors and the memories used by the processing machine may be located in geographically distinct locations and connected so as to communicate in any suitable manner. Additionally, it is appreciated that each of the processor and/or the memory may be composed of different physical pieces of equipment. Accordingly, it is not necessary that the processor be one single piece of equipment in one location and that the memory be another single piece of equipment in another location. That is, it is contemplated that the processor may be two pieces of equipment in two different physical locations. The two distinct pieces of equipment may be connected in any suitable manner. Additionally, the memory may include two or more portions of memory in two or more physical locations.
[0081] To explain further, processing, as described above, is performed by various components and various memories. However, it is appreciated that the processing perfoimed by two distinct components as described above, in accordance with a further embodiment, may be performed by a single component. Further, the processing performed by one distinct component as described above may be performed by two distinct components.
[0082] In a similar manner, the memory storage performed by two distinct memory portions as described above, in accordance with a further embodiment, may be performed by a single memory portion. Further, the memory storage performed by one distinct memory portion as described above may be performed by two memory portions.
[0083] Further, various technologies may be used to provide communication between the various processors and/or memories, as well as to allow the processors and/or the memories to communicate with any other entity; i.e., so as to obtain further instructions or to access and use remote memory stores, for example. Such technologies used to provide such communication might include a network, the Internet, Intranet, Extranet, a LAN, an Ethernet, wireless communication via cell tower or satellite, or any client server system that provides communication, for example. Such communications technologies may use any suitable protocol such as TCP/IP, UDP, or OSI, for example.
[0084] As described above, a set of instructions may be used in the processing of embodiments. The set of instructions may be in the fonn of a program or software. The software may be in the form of system software or application software, for example. The software might also be in the form of a collection of separate programs, a program module within a larger program, or a portion of a program module, for example. The software used might also include modular programming in the form of object-oriented programming. The software tells the processing machine what to do with the data being processed.
[0085] Further, it is appreciated that the instructions or set of instructions used in the implementation and operation of embodiments may be in a suitable form such that the processing machine may read the instructions. For example, the instructions that form a program may be in the form of a suitable programming language, which is converted to machine language or object code to allow the processor or processors to read the instructions. That is, written lines of programming code or source code, in a particular programming language, are converted to machine language using a compiler, assembler or interpreter. The machine language is binary coded machine instructions that are specific to a particular type of processing machine, i.e., to a particular type of computer, for example. The computer understands the machine language.
[0086] Any suitable programming language may be used in accordance with the various embodiments. Also, the instructions and/or data used in the practice of embodiments may utilize any compression or enciyption technique or algorithm, as may be desired. An encryption module might be used to encrypt data. Further, files or other data may be decrypted using a suitable decryption module, for example.
[0087] As described above, the embodiments may illustratively be embodied in the form of a processing machine, including a computer or computer system, for example, that includes at least one memory. It is to be appreciated that the set of instructions, i.e., the software for example, that enables the computer operating system to perform the operations described above may be contained on any of a wide variety of media or medium, as desired. Further, the data that is processed by the set of instructions might also be contained on any of a wide variety of media or medium. That is, the particular medium, i.e., the memory in the processing machine, utilized to hold the set of instructions and/or the data used in embodiments may take on any of a variety of physical forms or transmissions, for example. Illustratively, the medium may be in the form of a compact disc, a DVD, an integrated circuit, a hard disk, a floppy disk, an optical disc, a magnetic tape, a RAM, a ROM, a PROM, an EPROM, a wire, a cable, a fiber, a communications channel, a satellite transmission, a memory card, a SIM card, or other remote transmission, as well as any other medium or source of data that may be read by the processors.
[0088] Further, the memory or memories used in the processing machine that implements embodiments may be in any of a wide variety of forms to allow the memory to hold instructions, data, or other information, as is desired. Thus, the memory might be in the form of a database to hold data. The database might use any desired arrangement of files such as a flat file arrangement or a relational database arrangement, for example.
[0089] In the systems and methods, a variety of “user interfaces” may be utilized to allow a user to interface with the processing machine or machines that are used to implement embodiments. As used herein, a user interface includes any hardware, software, or combination of hardware and software used by the processing machine that allows a user to interact with the processing machine. A user interface may be in the form of a dialogue screen for example. A user interface may also include any of a mouse, touch screen, keyboard, keypad, voice reader, voice recognizer, dialogue screen, menu box, list, checkbox, toggle switch, a pushbutton or any other device that allows a user to receive information regarding the operation of the processing machine as it processes a set of instructions and/or provides the processing machine with infonnation. Accordingly, the user interface is any device that provides communication between a user and a processing machine. The information provided by the user to the processing machine through the user interface may be in the form of a command, a selection of data, or some other input, for example.
[0090] As discussed above, a user interface is utilized by the processing machine that perfonns a set of instructions such that the processing machine processes data for a user. The user interface is typically used by the processing machine for interacting with a user either to convey information or receive information from the user. However, it should be appreciated that in accordance with some embodiments of the system and method, it is not necessary that a human user actually interact with a user interface used by the processing machine. Rather, it is also contemplated that the user interface might interact, i.e., convey and receive information, with another processing machine, rather than a human user. Accordingly, the other processing machine might be characterized as a user. Further, it is contemplated that a user interface utilized in the system and method may interact partially with another processing machine or processing machines, while also interacting partially with a human user.
[0091] It will be readily understood by those persons skilled in the art that embodiments are susceptible to broad utility and application. Many embodiments and adaptations of the present invention other than those herein described, as well as many variations, modifications and equivalent arrangements, will be apparent from or reasonably suggested by the foregoing description thereof, without departing from the substance or scope.
[0092] Accordingly, while the embodiments of the present invention have been described here in detail in relation to its exemplary embodiments, it is to be understood that this disclosure is only illustrative and exemplary of the present invention and is made to provide an enabling disclosure of the invention. Accordingly, the foregoing disclosure is not intended to be construed or to limit the present invention or otherwise to exclude any other such embodiments, adaptations, variations, modifications or equivalent arrangements.

Claims

CLAIMS What is claimed is:
1. A method, comprising: receiving, by a conversational artificial intelligence engine and from an augmented reality headset worn by a user, a query that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset, wherein the query comprises audio of a user utterance and images or video captured by a camera of the augmented reality headset of what the user is seeing; generating, by the conversational artificial intelligence engine, a prompt for a large language model based on the user utterance and the images or video; providing, by the conversational artificial intelligence engine, the prompt to the large language model; receiving, by the conversational artificial intelligence engine, an output of the large language model, wherein the output comprises text and gestures for the embodied interactive agent; generating, by the conversational artificial intelligence engine, animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and outputting, by the conversational artificial intelligence engine, the animations and the speech to the augmented reality headset.
2. The method of claim 1, further comprising: receiving, by the conversational artificial intelligence engine, user location information from the augmented reality headset, wherein the prompt is further based on the user location information.
3. The method of claim 1, further comprising: inferring, by the conversational artificial intelligence engine, a task goal associated with the query, wherein the inference is based on a user interaction histoiy, environment object labels, and user location information.
4. The method of claim 1, wherein the conversational artificial intelligence engine generates a text prompt for the large language model based on the user utterance and an image prompt for a visual language model, and the large language model and the visual language model return outputs.
5. The method of claim 1, wherein the large language model comprises a multi-modal large language model.
6. The method of claim 1, wherein the large language model further outputs an identification of a document to provide to the augmented reality headset.
7. The method of claim 1, wherein the display in the augmented reality headset displays the animations for the embodied interactive agent, and a speaker in the augmented reality headset outputs the speech for the embodied interactive agent.
8. A system, comprising: an augmented reality headset comprising a camera, a microphone, a display, and a speaker, wherein the augmented reality headset is configured to be worn by a user; and a multi-modal conversational platform comprising a conversational artificial intelligence engine that is configured to receive, from the augmented reality headset, a query that is made to an embodied interactive agent that is displayed by the display, wherein the query comprises audio of a user utterance and images or video captured by the camera of what the user is seeing; to generate a prompt for a large language model based on the user utterance and the images or video; to provide the prompt to the large language model, to receive an output of the large language model, wherein the output comprises text and gestures for the embodied interactive agent, to generate animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and to output the animations and the speech to the augmented reality headset; wherein the display of the augmented reality headset is configured to display the animations of the embodied interactive agent, and the speaker is configured to output the speech of the embodied interactive agent.
9. The system of claim 8, wherein the conversational artificial intelligence engine is further configured to receive user location information from the augmented reality headset, and the prompt is further based on the user location information.
10. The system of claim 8, wherein the conversational artificial intelligence engine is further configured to infer a task goal associated with the query, wherein the inference is based on a user interaction history, environment object labels, and user location information.
11. The system of claim 8, wherein the conversational artificial intelligence engine is further configured to generate a text prompt for the large language model based on the user utterance, and an image prompt for a visual language model, and the large language model and the visual language model return outputs.
12. The system of claim 8, wherein the large language model comprises a multi-modal large language model.
13. The system of claim 8, wherein the large language model further outputs an identification of a document to provide to the augmented reality headset.
14. A non -transitory computer readable storage medium, including instructions stored thereon, which when read and executed by one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving, from an augmented reality headset worn by a user, a query that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset, wherein the query comprises audio of a user utterance and images or video captured by a camera of the augmented reality headset of what the user is seeing; generating a prompt for a large language model based on the user utterance and the images or video; providing the prompt to the large language model; receiving an output of the large language model, wherein the output comprises text and gestures for the embodied interactive agent; generating animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and outputting the animations and the speech to the augmented reality headset.
15. The non- transitory computer readable storage medium of claim
14, further comprising instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving user location information from the augmented reality headset, wherein the prompt is further based on the user location information.
16. The non-transitory computer readable storage medium of claim 14, further comprising instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: inferring a task goal associated with the query, wherein the inference is based on a user interaction history, environment object labels, and user location information.
17. The non-transitory computer readable storage medium of claim 14, further comprising instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: generating, a text prompt for the large language model based on the user utterance and an image prompt for a visual language model, and receiving outputs from the large language model and the visual language model.
18. The non-transitory computer readable storage medium of claim 14, wherein the large language model comprises a multi-modal large language model.
19. The non-transitory computer readable storage medium of claim 14, further comprising instructions stored thereon, which when read and executed by the one or more computer processors, cause the one or more computer processors to perform steps comprising: receiving from the large language model further, an identification of a document to provide to the augmented reality headset.
20. The non-transitory computer readable storage medium of claim 14, wherein the display in the augmented reality headset displays the animations for the embodied interactive agent, and a speaker in the augmented reality headset outputs the speech for the embodied interactive agent.
PCT/US2025/037380 2024-07-19 2025-07-11 Systems and methods for providing intelligent embodied interactive agents with spatial understanding Pending WO2026019672A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US18/778,292 2024-07-19
US18/778,292 US20260024286A1 (en) 2024-07-19 2024-07-19 Systems and methods for providing intelligent embodied interactive agents with spatial understanding

Publications (1)

Publication Number Publication Date
WO2026019672A1 true WO2026019672A1 (en) 2026-01-22

Family

ID=96917809

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2025/037380 Pending WO2026019672A1 (en) 2024-07-19 2025-07-11 Systems and methods for providing intelligent embodied interactive agents with spatial understanding

Country Status (2)

Country Link
US (1) US20260024286A1 (en)
WO (1) WO2026019672A1 (en)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2021003471A1 (en) * 2019-07-03 2021-01-07 DMAI, Inc. System and method for adaptive dialogue management across real and augmented reality
US11645479B1 (en) * 2019-11-07 2023-05-09 Kino High Coursey Method for AI language self-improvement agent using language modeling and tree search techniques
US20240045704A1 (en) * 2022-07-29 2024-02-08 Meta Platforms, Inc. Dynamically Morphing Virtual Assistant Avatars for Assistant Systems
US20240095491A1 (en) * 2023-12-01 2024-03-21 Quantiphi, Inc. Method and system for personalized multimodal response generation through virtual agents

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2021003471A1 (en) * 2019-07-03 2021-01-07 DMAI, Inc. System and method for adaptive dialogue management across real and augmented reality
US11645479B1 (en) * 2019-11-07 2023-05-09 Kino High Coursey Method for AI language self-improvement agent using language modeling and tree search techniques
US20240045704A1 (en) * 2022-07-29 2024-02-08 Meta Platforms, Inc. Dynamically Morphing Virtual Assistant Avatars for Assistant Systems
US20240095491A1 (en) * 2023-12-01 2024-03-21 Quantiphi, Inc. Method and system for personalized multimodal response generation through virtual agents

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
GOOGLE: "Project Astra: Our vision for the future of AI assistants", 14 May 2024 (2024-05-14), XP093327884, Retrieved from the Internet <URL:https://www.youtube.com/watch?v=nXVvvRhiGjI> [retrieved on 20251028] *
YANG ZHENGYUAN ET AL: "MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action", ARXIV.ORG, 20 March 2023 (2023-03-20), XP093328439, Retrieved from the Internet <URL:https://arxiv.org/abs/2303.11381> [retrieved on 20251028] *

Also Published As

Publication number Publication date
US20260024286A1 (en) 2026-01-22

Similar Documents

Publication Publication Date Title
US12367640B2 (en) Virtual role-based multimodal interaction method, apparatus and system, storage medium, and terminal
US12225325B2 (en) Method, apparatus, electronic device, computer-readable storage medium, and computer program product for video communication
CN110519636B (en) Voice information playing method and device, computer equipment and storage medium
US6526395B1 (en) Application of personality models and interaction with synthetic characters in a computing system
US20210407504A1 (en) Generation and operation of artificial intelligence based conversation systems
JP2023552854A (en) Human-computer interaction methods, devices, systems, electronic devices, computer-readable media and programs
CN113316078B (en) Data processing method and device, computer equipment and storage medium
CN113067953A (en) Customer service method, system, device, server and storage medium
WO2018006375A1 (en) Interaction method and system for virtual robot, and robot
US20240096093A1 (en) Ai-driven augmented reality mentoring and collaboration
CN116737883A (en) Human-computer interaction methods, devices, equipment and storage media
CN116610777A (en) Conversational AI platform with extracted questions and answers
CN118658200A (en) Image processing method, electronic device, storage medium and computer program product
CN112637692B (en) Interaction method, device and equipment
TW202418138A (en) Language data processing system and method and computer program product suitable for processing the oral report content of a user
US20260024286A1 (en) Systems and methods for providing intelligent embodied interactive agents with spatial understanding
US20250150647A1 (en) Method for preventing false triggering of live broadcasting risk control, computer device, and product
CN120542470A (en) Intelligent agent adaptive decision-making method and device based on multimodal semantic alignment
CN120688648A (en) An information processing method
CN117373455B (en) Audio and video generation method, device, equipment and storage medium
CN118802403A (en) A smart service method, device and equipment applied to Internet of Things devices
US20240320519A1 (en) Systems and methods for providing a digital human in a virtual environment
CN111783928A (en) Animal interaction method, device, equipment and medium
CN120883250A (en) Programs, information processing devices and information processing methods
CN117649469A (en) Interaction method of artificial intelligent screen, artificial intelligent screen and storage medium

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25760445

Country of ref document: EP

Kind code of ref document: A1