WO2025255236A1 - Robot navigation utilizing visual language model - Google Patents
Robot navigation utilizing visual language modelInfo
- Publication number
- WO2025255236A1 WO2025255236A1 PCT/US2025/032275 US2025032275W WO2025255236A1 WO 2025255236 A1 WO2025255236 A1 WO 2025255236A1 US 2025032275 W US2025032275 W US 2025032275W WO 2025255236 A1 WO2025255236 A1 WO 2025255236A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- goal
- robot
- request
- environment
- image frames
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05D—SYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
- G05D1/00—Control of position, course, altitude or attitude of land, water, air or space vehicles, e.g. using automatic pilots
- G05D1/20—Control system inputs
- G05D1/22—Command input arrangements
- G05D1/228—Command input arrangements located on-board unmanned vehicles
- G05D1/2285—Command input arrangements located on-board unmanned vehicles using voice or gesture commands
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05D—SYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
- G05D1/00—Control of position, course, altitude or attitude of land, water, air or space vehicles, e.g. using automatic pilots
- G05D1/20—Control system inputs
- G05D1/24—Arrangements for determining position or orientation
- G05D1/246—Arrangements for determining position or orientation using environment maps, e.g. simultaneous localisation and mapping [SLAM]
- G05D1/2469—Arrangements for determining position or orientation using environment maps, e.g. simultaneous localisation and mapping [SLAM] using a topologic or simplified map
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/10—Terrestrial scenes
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01C—MEASURING DISTANCES, LEVELS OR BEARINGS; SURVEYING; NAVIGATION; GYROSCOPIC INSTRUMENTS; PHOTOGRAMMETRY OR VIDEOGRAMMETRY
- G01C21/00—Navigation; Navigational instruments not provided for in groups G01C1/00 - G01C19/00
- G01C21/20—Instruments for performing navigational calculations
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05D—SYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
- G05D2105/00—Specific applications of the controlled vehicles
- G05D2105/30—Specific applications of the controlled vehicles for social or care-giving applications
- G05D2105/31—Specific applications of the controlled vehicles for social or care-giving applications for attending to humans or animals, e.g. in health care environments
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05D—SYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
- G05D2107/00—Specific environments of the controlled vehicles
- G05D2107/60—Open buildings, e.g. offices, hospitals, shopping areas or universities
- G05D2107/63—Offices, universities or schools
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05D—SYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
- G05D2109/00—Types of controlled vehicles
- G05D2109/10—Land vehicles
Definitions
- VLMs visual language models
- long context VLMs such as those that have context windows of greater than 500,000 tokens, 1 million tokens, 1.5 million tokens, or other quantity of tokens.
- VLMs may not be readily adaptable to utilization in various robotic tasks, such as robotic navigation.
- VLMs can produce only textual outputs, while navigation tasks can require outputting continuous coordinate actions.
- Implementations disclosed herein relate to utilizing a vision/visual language model (VLM) in robotic navigation. Some of those implementations process, utilizing the VLM, a request and prior tour image frames, to generate output that reflects a goal image frame, of the prior tour image frames, that corresponds to the instruction.
- the request can be a user request that is directed to a robot in an environment, such as a multimodal user request.
- the multimodal user request can include natural language from the user and image(s) of the user and/or of content created by the user.
- the prior tour image frames capture at least some portions of the environment of the robot and are captured prior to receiving the request.
- the prior tour image frames can be from video previously captured by camera(s) of the robot, or an additional robot, while a human navigates (e.g., through remote operation or kinesthetic guiding) the robot or the additional robot through the environment.
- the goal image frame that is reflected by the output generated utilizing the VLM, can be used to identify a corresponding goal waypoint.
- the goal waypoint can correspond to a pose from which the goal image frame was captured and can be determined and associated with the goal image frame prior to receiving the request.
- waypoints can be previously associated with prior tour image frames utilizing structure-from-motion and/or other technique(s).
- a lower-level policy can then be utilized in processing, at each step, the determined goal waypoint and a corresponding observation, to produce waypoint actions for the robot to execute to navigate from a current waypoint (waypoint of the robot when the request was received) to the goal waypoint in the environment.
- the lower-level policy can be an image-based goal reaching policy such as one that utilizes structure-from-motion and a topological graph constructed based on the prior tour image frames.
- Some implementations disclosed herein are directed to receiving a request that is directed to at least one robot in an environment, the request including natural language and/or one or more images.
- the request and prior tour image frames for the environment are processed, using a VLM, to generate VLM output.
- the prior tour image frames for the environment captured, prior to receiving the request, throughout at least a portion of the environment.
- Processing the prior tour image frames for the environment, using the VLM and along with the request is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment.
- a goal image frame, of the prior tour image frames, that is responsive to the request is determined.
- the at least one robot is caused to navigate, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
- a processing system Upon receiving this request, a processing system, utilizing a vision language model (VLM), processes the request alongside the vast collection of prior tour image frames. Through such processing, the system generates VLM output that indicates a specific "goal image frame" from the prior tour images, which it determines to best represent the large conference room with the glass wall. This goal image frame directly corresponds to a "goal waypoint" in the office environment (e.g., based on the stored associated waypoint data). In response to determining this goal image frame and its associated goal waypoint, the system then causes the robot to navigate autonomously through the office building, using its mapping and localization systems, to reach the identified goal waypoint.
- VLM vision language model
- a system determines the goal waypoint that corresponds to a particular goal image frame and then transmits that goal waypoint to the robot. For instance, if the goal image frame is a picture of a kitchen counter, the system would identify the precise location (waypoint) where that picture was taken and send those coordinates to the robot. Alternatively, in some implementations, the system may transmit the goal image frame directly to the robot, and the robot itself determines the corresponding goal waypoint. This could involve the robot analyzing the received image and using its own internal mapping capabilities to deduce the exact location it needs to reach.
- a goal waypoint corresponds to a particular goal image frame because it was previously determined that the goal image frame was captured from that specific goal waypoint. For example, during a pre-recorded tour of an environment, if an image was taken at a specific corner of a room, that image would be linked to the waypoint representing that corner. When the system later identifies that image as a goal, it can utilize the corresponding linked waypoint.
- a robot when a robot navigates to a goal waypoint, it uses a topological graph.
- This graph can include various points (vertices), where each point represents a specific waypoint.
- Each of these points can be linked to a corresponding image frame from a prior tour of the environment, or to specific visual characteristics (image features) that were derived from those image frames.
- the topological graph might have a vertex for the living room doorway, which is linked to an image of that doorway captured during a previous walkthrough.
- the robot might use a technique known as structure-from-motion, which relies on the topological graph, to help it move to the goal waypoint.
- a system further processes a natural language representation of each prior tour image frame.
- This processing can use a VLM along with the request and the tour image frames.
- each image could have a generated text description, like "a view of the breakroom with a coffee machine” or "the main hallway leading to the elevators.”
- the VLM is used to process these natural language descriptions in addition to the images themselves.
- the system might determine the goal image frame by first identifying a natural language representation that matches the request. For instance, if the request is "Go to the coffee machine," the system might find the natural language representation "a view of the breakroom with a coffee machine” and then identify the corresponding image frame as the goal.
- the VLM might be instructed to output the natural language representation for the tour image frame that best matches the received request.
- the system also processes narrative language descriptions that are associated with some or all of the prior tour image frames. These descriptions are provided by a human and are captured prior to receiving a request. For example, a person might have walked through an office building and narrated observations like "This is where we keep the shared printer" while a tour video was being recorded. This narrative could then be linked to the specific image frames captured at that location. When a request is received, these human-provided narratives can be processed, using a VLM and along with the request and the tour images, to help identify the most relevant goal image frame.
- the prior tour image frames are captured by cameras that are part of the robot itself.
- a service robot might routinely traverse its operational environment, recording video footage that later becomes the set of prior tour image frames.
- a received request is a user request.
- This request can include natural language, which might be based on spoken or typed input from a user. For example, a user might verbally tell the robot, "Go to the supply closet," or type the same instruction.
- the request might also include one or more images, such as an image of the user, an object the user is holding, or even a drawing made by the user to indicate a desired location.
- These user requests can be captured through input devices on the robot, such as microphones for spoken commands or cameras for visual input.
- the request is sent by a robot located in the environment and received over one or more networks.
- a robot in a warehouse might send a request for navigation assistance to a central server via the internet.
- the processing of the request and the tour image frames occurs on one or more remote servers that are not physically located within the robot's immediate environment.
- a system captures a user request using input devices on a robot in a particular environment.
- This user request can include natural language provided by a human, such as a spoken command and/or one or more images of the human or their surroundings.
- the system arranges for the user request and prior tour image frames of the environment to be processed using a VLM. This processing helps to identify a goal image frame from the prior tour images that is relevant to the request.
- the prior tour image frames were captured beforehand, covering at least a portion of the environment.
- the robot receives an indication of a goal waypoint in the environment that corresponds to the identified goal image frame. Upon receiving this indication, the robot then navigates to that goal waypoint.
- navigating the robot to the goal waypoint involves identifying a specific "goal vertex" within a topological graph that corresponds to the goal waypoint.
- This topological graph contains various points (vertices), each representing a waypoint, and each linked to a corresponding prior tour image frame or features derived from those frames. For instance, if the goal waypoint is "the corner by the water cooler," the system would identify the specific vertex in the graph that represents that corner.
- the robot then uses structure-from-motion, based on this topological graph and using the goal vertex as its target, to move to the waypoint.
- the robot identifies a sequence of connected vertices in the topological graph that form a path from its current location to the goal waypoint, and then uses this sequence to navigate.
- the goal waypoint can correspond to the goal image frame because it was previously determined that the goal image frame was originally captured from that specific location.
- a natural language representation for each of the prior tour image frames is processed using the VLM, along with the user request and the prior tour image frames themselves.
- a tour image has a textual description like "area near the main entrance,” this description is used to help match the user's request, such as "Go to the entrance.”
- the system might process instructions that direct the VLM to output the natural language representation for the tour image frame that best matches the user's request, which then helps to identify the goal image.
- a corresponding narrative language description for one or more of the prior tour image frames is also processed using the VLM when determining the goal image.
- These narrative descriptions are provided by humans, captured before the request is received, and are associated with their respective prior tour image frames. For example, a human might have provided a narrative like "This is the conference room where we hold large meetings" while recording a tour, and this narrative would be used in processing using by the VLM in conjunction with the user request to find the relevant image.
- the prior tour image frames are captured by cameras that are part of the robot itself. For example, a robot might autonomously perform a mapping tour, recording visual data that is then used to create these image frames.
- the user request includes natural language, which is based on either spoken or typed input from a human. For example, a human might simply say "Find my keys" or type "Go to the break room.”
- the user request can additionally or alternatively include one or more images. For instance, a human might show the robot a picture of an object they are looking for.
- the input devices used to capture the user request include microphones, cameras, or both. This allows for multimodal input from the human user.
- the robot transmits the user request to one or more remote servers. For example, a robot might send a user's verbal command to a cloud server for processing. In some specific examples, the robot also transmits an indication of itself or its environment along with the user request to these servers. The servers then use this indication to identify the correct set of prior tour image frames relevant to that specific robot or environment for processing.
- one or more of the processors that perform these operations are integrated directly into the robot. This could allow for faster, more localized processing of the request and navigation commands.
- a system receives a request that is directed to a robot in an environment, and this request includes natural language and/or one or more images.
- the system processes both the request and prior tour image frames of the environment using a VLM.
- the tour image frames were captured beforehand, covering at least a portion of the environment.
- the processing of these image frames with the VLM, along with the request, is performed because the images cover the environment and the request is specifically for a robot within that environment.
- the system Based on the output generated using the VLM, the system identifies a goal image frame from the prior tour images that is relevant to the request. Then, in response to identifying this goal image frame, the robot is navigated in the environment to a specific goal waypoint that corresponds to that goal image frame.
- a system receives a user request directed to a robot in an environment, which includes natural language and/or one or more images.
- the system processes this user request and prior tour image frames of the environment using a vision language model (VLM) to generate VLM output.
- VLM vision language model
- the prior tour image frames were captured beforehand throughout at least a portion of the environment. This processing is based on the tour image frames covering the environment and the request being for a robot within that environment.
- VLM output Based on the VLM output, a goal image frame from the prior tour images is determined to be responsive to the user request. Subsequently, based on processing this goal image frame and the user request, a specific response to the user request is determined.
- the system then causes the robot to render this response in a way that is responsive to the user's request.
- the user request includes natural language
- the determined response to the user request is also a natural language response.
- This response can be audibly rendered by the robot using its speakers. For example, if a user asks, "Where is the nearest exit?", the robot might audibly reply, "The nearest exit is to your left, past the green door.”
- the determination of the response involves using the same vision language model (VLM) or an additional VLM for the processing.
- VLM vision language model
- the determination of the response to the user request involves deciding to cause the robot to render a response instead of causing the robot to navigate to a goal waypoint. For example, if the request is "Describe this object" while showing an image of an object, the system might determine that an audible description from the robot is the appropriate response, rather than instructing the robot to navigate to the object's location.
- FIG. 1 is a diagram illustrating an example of processing of a multimodal user instruction for robot navigation.
- FIG. 2 is a flowchart illustrating an example method of processing a request and prior tour image frames, using a VLM, to determine a goal image frame that is responsive to the request and causing a robot to navigate to a goal waypoint that corresponds to the goal image frame.
- FIG. 3 is a flowchart illustrating an example method of causing a request and prior tour image frames to be processed, using a VLM, to determine a goal image frame, receiving, in response to the causing, an indication of a goal waypoint that corresponds to the goal image frame, and navigating a robot to the goal waypoint.
- FIG. 4 is a flowchart illustrating an example method of processing a request and prior tour image frames, using a VLM, to determine a goal image frame that is responsive to the request, determining a response to the request based on processing the request and the goal image frame, and causing a robot to render the response.
- FIG. 5 illustrates an example robotic system according to some implementations.
- FIG. 6 illustrates an example computing system according to some implementations.
- MINT Multimodal Instruction Navigation with demonstration Tours
- VLMs Vision Language Models
- Mobility VLA Hierarchical Vision-Language-Action
- the high-level policy incorporates a long-context VLM that receives the demonstration tour video and the multimodal user instruction as input to identify a goal frame within the tour video. Subsequent to this, a low- level policy utilizes the identified goal frame and a topological graph, which is constructed offline, to generate robot actions at each timestep.
- Mobility VLA has been evaluated in an 836m 2 real-world environment. This evaluation indicates that Mobility VLA achieves high end-to-end success rates on previously unaddressed multimodal instructions, such as "Where should I return this?" when a plastic bin is held by a user.
- Robot navigation has undergone significant advancements.
- Object goal and Vision Language Navigation represent a notable progression in robot usability, enabling the use of open-vocabulary language to define navigation goals, such as "Go to the couch.”
- a further advancement is proposed by extending the natural language space of ObjNav and VLN into the multimodal domain. This implies that a robot is configured to accept natural language and/or image instructions concurrently. For instance, a person unfamiliar with a building may inquire, "Where should I return this?" while holding a plastic bin. In such a scenario, the robot guides the user to a designated shelf for returning the bin, based on both verbal and visual context.
- This category of navigation tasks is termed Multimodal Instruction Navigation (MIN).
- MIN Multimodal Instruction Navigation
- Multimodal Instruction Navigation is a task category encompassing environment exploration and instruction-guided navigation.
- exploration may be circumvented by employing a demonstration tour video that comprehensively traverses the environment.
- a demonstration tour offers several advantages. For example, collection is easy as a user may teleoperate the robot or utilize a portable electronic device, such as a smartphone, to record a video while traversing the environment.
- Various exploration algorithms can additionally or alternatively be employed to generate such a tour.
- such a tour aligns with common user practices. For instance, upon acquiring a new robotic device for a home environment, a user may naturally introduce the robotic device to the home environment and verbally indicate locations of interest during the tour.
- MINT a category of tasks referred to as MINT. This category leverages demonstration tours and focuses on the fulfillment of multimodal user instructions.
- VLMs Vision-Language Models
- MINT Vision-Language Models
- the utilization of VLMs in isolation for MINT presents certain challenges.
- the quantity of input images that can be processed by many VLMs is constrained by context-length limitations. This limitation can significantly impede the accuracy of environmental understanding in expansive environments.
- the execution of MINT tasks necessitates the determination of robot actions. Queries formulated to elicit such robot actions are typically outside the distribution of data used for VLM training or pre-training. Consequently, the performance of zero-shot navigation may be less than optimal.
- a Mobility Vision-Language-Action (VLA) navigation policy is provided for solving MINT.
- This hierarchical policy combines environment understanding and common sense reasoning capabilities of long-context Vision-Language Models (VLMs) with a robust low- level navigation policy based on topological graphs.
- VLMs Vision-Language Models
- a high-level VLM is utilized to process a demonstration tour video and a multimodal user instruction to generate output that is utilized to identify a goal frame within the tour video.
- a classical low-level policy employs the goal frame and a topological graph, which is constructed offline from the tour frames to generate robot actions, such as waypoints, at each timestep.
- the application of long-context VLMs addresses issues related to the fidelity of environment understanding.
- the topological graph facilitates bridging a gap between a VLM's training distribution and the specific robot actions for solving MINT.
- Mobility VLA was evaluated in a real-world office environment spanning 836m 2 and a home-like environment.
- the evaluation results indicate that implementations of Mobility VLA achieved success rates of 86% and 90% (26% and 60% higher than baseline methods) in these environments, respectively, on MINT tasks that were previously considered infeasible. These tasks involved complex reasoning, such as interpreting instructions like "I want to store something out of sight from the public eye. Where should I go?" and multimodal user instructions.
- the ease with which users can interact with a robot utilizing this approach was shown to be significantly advanced. For instance, a user could record a narrated video walkthrough in a home environment using a smartphone and subsequently inquire, "Where did I leave my coaster?"
- a low-level controller utilized in various implementations incorporates a visual Simultaneous Localization and Mapping (SLAM) algorithm, COLMAP, and a Model Predictive Control (MPC) method to track desired waypoints obtained from high-level VLMs.
- SLAM visual Simultaneous Localization and Mapping
- COLMAP COLMAP
- MPC Model Predictive Control
- object and image goal navigation techniques utilize rich input modalities. These modalities can include object categories, natural language instructions, dialogue, goal image conditions, and multimodal inputs that combine language and images.
- a MINT task involves providing a demonstration tour video and a multimodal user instruction as input.
- a robot operates to navigate to one or more specified goal locations in order to satisfy the user's instruction.
- the multimodal user instruction can be just a text instruction d G str (e.g., "Where can I find a ladder?"), or both text and image instructions I G ]R> HXWX3 (e.g., "Where can I get something to clean this?" + The robot sees the user pointing to a dirty whiteboard).
- Implementations aim to produce a navigation policy (a I O, F, N, d, I), where O G J ⁇ HXWX3 j s t e ro b ot ' s curr ent camera observation.
- the policy emits an embodimentagnostic waypoint action a G IR 3 representing longitudinal translation (Ax), lateral translation (A-r .), and rotation along the vertical axis (A0), all in the robot-centric frame.
- the robot has an embodiment-specific mechanism to execute waypoint actions.
- Mobility VLA A hierarchical navigation policy, referred to herein referred to as Mobility VLA, can incorporate both online and offline processing components.
- a topological graph G can be generated from the demonstration tour (N, F).
- the high-level policy takes the demonstration tour and the multimodal user instruction (d,l) to find the navigation goal frame index g, which is an integer corresponding to a specific frame of the tour.
- the lower-level policy utilizes the topological graph, the current camera observation (O) and g to produce a waypoint action (a) for the robot to execute at each timestep. This can be represented by the following two equations.
- h and I are the high and low-level policies.
- Mobility VLA utilizes a demonstration tour of an environment to solve MINT tasks. This tour can be provided by a human user via teleoperation, or by recording a video on a portable electronic device, such as a smartphone, while traversing the environment.
- Each vertex corresponds to a corresponding frame from the demonstration tour video, which is represented by a set of frames F and an optional set of narratives N.
- a structure-from- motion pipeline such as COLMAP, can be employed to determine an approximate 6-Degree- of-Freedom camera pose for each frame. This camera pose information is subsequently stored in the corresponding vertex.
- a directed edge is added to the topological graph G if a target vertex is positioned substantially in front of a source vertex and is within a predefined distance. Specifically, an edge is added if the target vertex is less than 90 degrees away from the source vertex's pose and is within 2 meters of the source vertex.
- Th is topological graph approach offers a simplified alternative to traditional navigation pipelines, which typically involve mapping the environment, identifying traversable areas, and then constructing a Probabilistic Roadmap (PRM).
- PRM Probabilistic Roadmap
- the topological graph captures the general connectivity of the environment based on the trajectory of the demonstration tour, thereby streamlining the environmental representation.
- a high-level policy leverages the common sense reasoning capabilities of VLMs to identify a navigation goal from the demonstration tour. This goal is determined such that it satisfies a wide range of multimodal, colloquial, and often ambiguous user instructions.
- a prompt P is prepared, which includes interleaving text and images.
- the prompt P can be defined as a function of the tour frames F, the narratives N, the text instruction d, and the image instruction I, denoted as P(F, N, d, I).
- An illustrative example of the prompt P, formulated for the multimodal user instruction "Where should I return this?" can be as follows:
- VLM output that reflects an integer goal frame index, denoted as g.
- g an integer goal frame index
- a low-level policy takes over and produces a waypoint action at every timestep (Eq. 1).
- Algorithm 1 An example low-level policy is demonstrated by Algorithm 1:
- Input goal frame index g, offline-constructed topological graph G.
- a real-time hierarchical visual localization system may be utilized to estimate the pose of a robot and an initial vertex for navigation.
- This localization system is configured to identify a plurality of candidate frames that are nearest to a current camera observation based on a global descriptor, and subsequently to compute the pose of the robot through a Perspective-n-Point (PnP) algorithm.
- PnP Perspective-n-Point
- a shortest path on the topological graph is identified between the initial vertex and a goal vertex.
- the goal vertex corresponds to a desired goal frame.
- a low-level policy is configured to generate a waypoint action. This waypoint action may include positional and orientational data of a subsequent vertex in the identified shortest path, relative to the current robot pose.
- RQ1 Does Mobility VLA perform well in MINT in the real world?
- RQ2 Does Mobility VLA outperform alternatives thanks to the use of long-context VLM?
- RQ3 Is the topological graph necessary? Can VLMs produce actions directly?
- the environments of the experiments is an office environment occupied by humans. This environment encompasses approximately 836 square meters and includes various items such as shelves, desks, and chairs.
- the robot of the experiments is a wheel-based mobile manipulator that employs a Model Predictive Control (MPC)-based algorithm to execute waypoint actions, identified as positional and orientational data within the robot's local frame while avoiding obstacles.
- MPC Model Predictive Control
- a demonstration tour is collected by teleoperating the robot with a gamepad. All corridors within the environment are traversed twice from opposing directions. The resulting tour has a duration of approximately 16 minutes, corresponding to 948 image frames captured at 1 Hertz.
- a selection of user instructions can be employed. For example, five user instructions may be randomly selected per category. The policy's performance can then be evaluated from a plurality of random starting poses such as four random starting poses, each located at a distance of at least 20 meters from each other. A long-context multimodal VLM, such as Gemini 1.5 Pro may be utilized for this purpose.
- Table 2 illustrates that a Mobility VLA policy exhibits a high end-to-end navigation success rate across most user instruction categories. This includes categories such as Reasoning-Required and Multimodal instructions, which were previously considered infeasible for such systems.
- the policy can also demonstrate a reasonable Success Rate weighted Path Length (SPL), indicating that the topological graph, as a navigation aid, does not impose a substantial penalty on path length.
- SPL Success Rate weighted Path Length
- the Mobility VLA policy has been observed to successfully incorporate personalization narratives from a demonstration tour. For instance, the policy correctly navigated to different locations in response to substantially similar instructions from distinct users.
- An example includes navigation to a frame at a timestamp of 7:14 when presented with the instruction "I'm Lewis, take me to a temp desk please,” and navigation to a frame at a timestamp of 5:28 when presented with the instruction "Hi robot, I'm visiting, can you take me to a temp desk?".
- Table 2 Mobility VLA end-to-end navigation Success Rate (SR) and SPL of various user instruction types in the real Office environment.
- SR Mobility VLA end-to-end navigation Success Rate
- Table 2 additionally indicates the robustness of a Mobility VLA policy's low-level goal reaching exhibiting a 100% success rate in real-world scenarios. This robustness is observed even when the demonstration tour was recorded several months prior to the experiments, during which time many objects, furniture arrangements, and lighting conditions in the environment had undergone changes.
- simulations can be utilized to scale evaluation numbers.
- a high-fidelity simulation reconstruction of an office environment may be created using techniques such as Neural Radiance Fields (NeRF).
- NeRF Neural Radiance Fields
- the Mobility VLA policy can then be evaluated against 20 language-instructed tasks, with 50 random starting poses allocated per task.
- Such an experiment yields a high-level goal finding success rate of 90% and a low-level goal reaching success rate of 100%, culminating in a total of 900 successful end-to-end executions.
- Table 8 End-to-end navigation Success Rate (SR) and SPL of various user instruction types in the simulated Office environment.
- a proof-of- concept experiment can be conducted in a real home-like environment. Rather than providing a robot with a teleoperated tour, a smartphone can be utilized to record a demonstration tour. Subsequently, the Mobility VLA policy can be evaluated end-to-end with, for example, 4 Reasoning-Required and 1 Small Object user instructions, each from 4 random starting positions. This evaluation results in a 100% success rate with an SPL of 0.87. This outcome indicates that a Mobility VLA policy performs well regardless of the specific environment. Additionally, it highlights the ease of deployment, as a user may simply utilize a smartphone to record a tour of an environment, upload the recorded tour to a robot, and then immediately commence providing instructions.
- Table 3 illustrates how well alternative methods perform compared to Mobility VLA.
- Mobility VLA is compared with CLIP-based retrieval and Text-Only Mobility VLA.
- CLIP-based retrieval the high-level goal finding module of NLMap is reproduced by adopting OWL-ViT for region proposal and CLIP for sub-regions and full-images embeddings extraction for tour frames.
- Goal frame retrieval is then performed using CLIP embeddings of the instruction language and image.
- Text-Only Mobility VLA the multimodal demonstration tour is captioned by a VLM frame-by-frame to form a "text tour”.
- An LLM uses the text tour to produce the goal frame index.
- Table 3 shows that high-level goal finding success rates of Mobility VLA are significantly higher than comparison methods. Given the 100% low-level success rate, this high-level goal finding success rates are representative of end-to-end success rates.
- Table 3 High-level goal finding Success Rates of Mobility VLA compared to baselines [0077] Feeding a full demonstration tour of a large environment into non-long-context VLMs is challenging since each image requires hundreds-of-token budgets.
- One solution for reducing input tokens number is reducing the tour video frame rate, at the cost of intermediate frames loss.
- Table 4 shows that the high-level goal finding success rate decreases as the tour frame rate decreases. This can be due to a lower frame rate tour sometimes missing the navigation target frame.
- comparing various VLMs only Gemini 1.5 Pro currentlyvyields satisfactory success rate thanks to its long IM token context-length.
- Table 4 High-level goal finding Success Rates with regards to various user instruction types (presented in the order of Reasoning Free (RF), Reasoning Required (RR), Small Objects (SO), MultiModal (MM)) as a function of VLM models (column) and multimodal demonstration tour Frames Per Second (FPS) (row). All VLMs were queried in June 2024.
- RF Reasoning Free
- RR Reasoning Required
- SO Small Objects
- MM MultiModal
- Mobility VLA uses a hierarchical architecture to harness long-context VLM's reasoning capability and uses a topological graph to produce waypoint actions. Such a topological graph can be beneficial for navigation success.
- Table 5 shows the end-to-end performance of Mobility VLA in simulation compared to prompting the VLM to output waypoint actions directly.
- the 0% end-to-end success rate shows that Gemini 1.5 Pro is incapable of navigating the robot zero-shot without the topological graph.
- Gemini almost always outputs the "move forward" waypoint action regardless of the current camera observation.
- the current Gemini 1.5 API requires the upload of all 948 tour images at every inference call, resulting in a prohibitively expensive 26s perstep running time for the robot to move just lm.
- Mobility VLA's high- level VLM spends 10-30s to find a goal index and then the robot navigates to the goal using the low-level topological graph results in a highly robust and efficient (0.19s per step) system for solving MINT.
- COLMAP a structure-from-motion pipeline to estimate the pose of the robot for each frame in the tour (i.e. reference images), 3D point landmarks in the environment, and their corresponding 2D projections across all reference images (i.e. 2D-3D correspondences).
- the poses are used to build a fully connected topological graph.
- the tour frames F, 3D landmarks, and 2D features are used in some implementations of a real-time hierarchical localizer.
- the method is hierarchical since it divides localization of the observed image O into two steps: a global search to determine a set of candidate reference images close to O followed by local feature matching and pose estimation.
- the candidate set C £ p o f k-nearest (w.r.t. the I 2 -norm of a global image descriptor [54]) tour frames to O is determined.
- 2D features in O are matched to the 2D features of each frame in C.
- correspondences between 2D features in O and 3D landmarks observed in the tour are established.
- the pose of O is computed by solving the corresponding Perspective-n-Point problem.
- the pose with the most inlier 2D-3D correspondences is selected as T o .
- Table 7 High-level goal finding Success Rates of multimodal user instructions as a function of VLM models and instruction representations (columns) and tour modalities (row).
- MM Instructions columns the robot's current camera observation is fed directly into the VLMs.
- Text Instructions columns the camera observation is captioned by Gemini 1.5 Pro and the caption text is then fed into the VLMs.
- the text tour was captioned w/ Gemini 1.5 Pro.
- FIG. 1 is a diagram illustrating an example of processing of a multimodal user instruction for robot navigation.
- FIG. 1 illustrates an overall system architecture 100 for robot navigation, representing a hierarchical approach that combines high-level understanding with low-level motion control.
- the system processes multimodal inputs to facilitate effective navigation.
- the multimodal user instruction and a demonstration tour video of the environment are used by a long-context VLM (high-level policy) to identify the goal frame in the video.
- VLM high-level policy
- the low-level policy uses the goal frame and an offline generated topological map (from the tour video using structure-from-motion) to compute a robot action at every timestep
- FIG. 1 depicts a multimodal user instruction 101 and a demonstration tour video 102 as primary inputs to a high level goal finding with long-context VLM module 120.
- the high level goal finding with long-context VLM module 120 utilizes a VLM to process the multimodal user instruction 101 and the demonstration tour video 102.
- the module 120 can generate a prompt that includes the multimodal user instruction 101 and the demonstration tour video 102 and process the prompt using the VLM.
- a navigation goal 103 is reflected in VLM output generated from such processing.
- the navigation goal 103 can correspond to a specific frame or location within the demonstration tour.
- the architecture 100 further includes an offline component for generating a topological graph 104 from the demonstration tour video 102 through a structure from motion module 110.
- the topological graph 104 provides a simplified representation of the environment's connectivity and can be utilized for efficient low-level navigation.
- a localization module 124 utilizes observation data 106 (e.g., current images and/or other state data captured by sensor(s) of robot 115) from the robot 115 and the topological graph 104 to determine the robot's current pose within the environment.
- the pose is provided to path finding module 122
- the navigation goal 103 identified by the module 120, is fed into a path finding module 122.
- the path finding module 122 utilizing the topological graph 104 and real-time robot pose information from localization module 124, computes a waypoint action 105.
- the waypoint action 105 represents the specific movement commands (e.g., change in x, y, and theta coordinates) required for the robot 115 to navigate towards the designated goal.
- the waypoint action 105 is then executed by the robot 115.
- This hierarchical design allows the system to leverage the powerful reasoning capabilities of VLMs for goal interpretation while relying on a robust, graph-based approach for precise and efficient movement.
- FIG. 2 is a flowchart illustrating an example method 200 of processing a request and prior tour image frames, using a vision language model (VLM), to determine a goal image frame that is responsive to the request and causing a robot to navigate to a goal waypoint that corresponds to the goal image frame.
- VLM vision language model
- a request is received that is directed to at least one robot in an environment.
- This request can include natural language, as shown in sub-block 252A, and/or one or more images, as shown in sub-block 252B.
- the user might say, "Go to the kitchen and grab the red apple," which is a natural language input, and/or show the robot a picture of a red apple, which is an image input.
- a vision language model (VLM) is used to process the request and prior tour image frames for the environment to generate VLM output.
- VLM vision language model
- These prior tour image frames are captured throughout at least a portion of the environment before the request is received. For example, if the robot previously toured the house and captured images of every room, including the kitchen, these prior tour images would be processed by the VLM along with the user's request. The processing considers that the images are from the environment and the request is directed to the robot operating within that environment.
- a goal image frame is determined from the prior tour image frames that is responsive to the request. For example, the VLM, having processed the request "Go to the kitchen and grab the red apple" and the tour images, identifies a specific image frame of the kitchen showing a red apple as the goal image frame.
- the at least one robot in response to determining the goal image frame, is caused to navigate, in the environment, to a goal waypoint that corresponds to the goal image frame.
- a system may determine the goal waypoint based on the goal image frame and then transmit that goal waypoint to the robot.
- the system identifies the precise GPS coordinates or internal map location (waypoint) of that counter and sends those coordinates to the robot.
- the system may transmit the goal image frame directly to the robot, and the robot itself determines the corresponding goal waypoint based on its internal mapping capabilities and the received image. The robot then proceeds to navigate to that determined goal waypoint, thereby completing the user's request to retrieve the red apple.
- FIG. 3 is a flowchart illustrating an example method 300 of causing a request and prior tour image frames to be processed, utilizing a vision language model (VLM), to determine a goal image frame, receiving, in response to the causing, an indication of a goal waypoint that corresponds to the goal image frame, and navigating a robot to the goal waypoint.
- VLM vision language model
- a user request is captured via one or more input devices of the robot in the factory environment.
- This request can include natural language, as shown in sub-block 352A, and/or one or more images, as shown in sub-block 352B.
- the worker might say, "Go to the wrench rack” (natural language), and/or show the robot a picture of a wrench (image).
- the robot's onboard microphone or camera captures this request.
- the request and prior tour image frames for the environment are caused to be processed, using a VLM, to determine a goal image frame from the prior tour image frames that is responsive to the request.
- VLM determines, for example, that a particular image frame of the wrench rack best matches the worker's request.
- the robot receives an indication of a goal waypoint in the environment that corresponds to the determined goal image frame. For instance, the central server sends back to the robot the precise coordinates of the wrench rack, which were previously associated with the identified goal image frame.
- the robot in response to receiving the indication of the goal waypoint, the robot navigates to that goal waypoint. As shown in sub-block 358A, this navigation can involve identifying a "goal vertex" in a topological graph based on the received goal waypoint indication. This topological graph contains a network of interconnected waypoints, each linked to a prior tour image frame or derived image features. The robot then utilizes this goal vertex, along with its internal localization and navigation systems, to calculate a path and move autonomously to the wrench rack.
- FIG. 4 is a flowchart illustrating an example method 400 for a robot to respond to a user request, which may involve rendering a response instead of navigating to a location. For example, consider a scenario where a robot is in a museum, and a user asks about a specific exhibit.
- a request is received that is directed to a robot in an environment, such as the museum.
- This request can include natural language, as shown in block 452A, and/or one or more images, as shown in block 452B.
- a user might point the robot's camera at a painting and verbally ask, "What can you tell me about this?” (natural language and image input).
- a vision language model (VLM) is used to process the request and prior tour image frames for the environment to generate VLM output. These prior tour image frames were captured during an earlier detailed tour of the museum, covering all exhibits.
- the VLM processes the user's question and the image of the painting alongside these stored tour images.
- a goal image frame is determined from the prior tour image frames that is responsive to the request.
- the VLM identifies a specific prior tour image frame of the painting the user is pointing to as the goal image frame.
- a response to the request is determined. For example, the VLM further processes the goal image (the painting) and the question ("What can you tell me about this?") to generate a descriptive text about the painting's history, artist, and significance.
- the robot is caused to render this determined response responsive to the user's request. As shown in block 460A, this could involve determining to cause the robot to render the response, such as audibly reciting the painting's information through its speakers, in lieu of causing the robot to navigate to a goal waypoint that corresponds to the goal image frame. This illustrates that the system can decide the most appropriate response type, which may not always be navigation.
- FIG. 5 schematically depicts an example architecture of a robot 520.
- the robot 520 includes a robot control system 560, one or more operational components 540a- 540n, and one or more sensors 542a-542m.
- the sensors 542a-542m may include, for example, vision sensors, light sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, and so forth. While sensors 542a-m are depicted as being integral with robot 520, this is not meant to be limiting. In some implementations, sensors 542a-m may be located external to robot 520, e.g., as standalone units.
- Operational components 540a-540n may include, for example, one or more end effectors and/or one or more servo motors or other actuators to effectuate movement of one or more components of the robot.
- the robot 520 may have multiple degrees of freedom and each of the actuators may control the actuation of the robot 520 within one or more of the degrees of freedom responsive to the control commands.
- the term actuator encompasses a mechanical or electrical device that creates motion (e.g., a motor), in addition to any driver(s) that may be associated with the actuator and that translate received control commands into one or more signals for driving the actuator.
- providing a control command to an actuator may comprise providing the control command to a driver that translates the control command into appropriate signals for driving an electrical or mechanical device to create desired motion.
- the robot control system 560 may be implemented in one or more processors, such as a CPU, GPU, and/or other controller(s) of the robot 520.
- the robot 520 may comprise a "brain box" that may include all or aspects of the control system 560.
- the brain box may provide real time bursts of data to the operational components 540a-n, with each of the real time bursts comprising a set of one or more control commands that dictate, inter alia, the parameters of motion (if any) for each of one or more of the operational components 540a-n.
- the robot control system 560 may perform one or more aspects of method(s) described herein.
- control system 560 in controlling a robot during performance of a robotic task, can be generated based on considering different poses and/or different time stamp image(s) generated according to techniques described herein.
- control system 560 is illustrated in FIG. 5 as an integral part of the robot 520, in some implementations, all or aspects of the control system 560 may be implemented in a component that is separate from, but in communication with, robot 520.
- all or aspects of control system 560 may be implemented on one or more computing devices that are in wired and/or wireless communication with the robot 520, such as computing device 610.
- FIG. 6 is a block diagram of an example computer system 610.
- Computer system 610 typically includes at least one processor 614 which communicates with a number of peripheral devices via bus subsystem 612. These peripheral devices may include a storage subsystem 624, including, for example, a memory subsystem 625 and a file storage subsystem 626, user interface output devices 620, user interface input devices 622, and a network interface subsystem 616. The input and output devices allow user interaction with computer system 610.
- Network interface subsystem 616 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
- User interface input devices 622 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and/or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computer system 610 or onto a communication network.
- User interface output devices 620 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices.
- the display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image.
- CTR cathode ray tube
- LCD liquid crystal display
- projection device or some other mechanism for creating a visible image.
- the display subsystem may also provide non-visual display such as via audio output devices.
- output device is intended to include all possible types of devices and ways to output information from computer system 610 to the user or to another machine or computer system.
- Storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein.
- the storage subsystem 624 may include the logic to perform selected aspects of method 200, method 300, method 400 and/or to implement one or more aspects of robot 500.
- Memory 625 used in the storage subsystem 624 can include a number of memories including a main randomaccess memory (RAM) 630 for storage of instructions and data during program execution and a read only memory (ROM) 632 in which fixed instructions are stored.
- a file storage subsystem 626 can provide persistent storage for program and data files, and may include a hard disk drive, a CD-ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 626 in the storage subsystem 624, or in other machines accessible by the processor(s) 614.
- Bus subsystem 612 provides a mechanism for letting the various components and subsystems of computer system 610 communicate with each other as intended. Although bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
- Computer system 610 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 610 depicted in FIG. 6 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer system 610 are possible having more or fewer components than the computer system depicted in FIG. 6.
- a VLM utilized herein can be a sequence-to- sequence based machine learning models capable of processing vision data and textual data and generating generative vision data, generative audio data, generative textual data, and/or other forms of generative data.
- Some non-limiting examples of these sequence-to- sequence based machine learning models capable that are capable of generating one or more forms of the generative data noted above include transformer-based machine learning models (e.g., encoder-decoder transformer models, encoder-only transformer models, decoder-only transformer models, etc. that optionally employ an attention mechanism or some other form of memory), diffusion-based machine learning models, recurrent neural network-based machine learning models, generative adversarial networkbased machine learning models, etc.
- Some particular non-limiting examples of these sequence-to-sequence based machine learning models include the Gemini family of models (e.g., Gemini Pro 1.5).
- a method implemented by processor(s) includes receiving a request that is directed to at least one robot in an environment.
- the request includes natural language and/or one or more images.
- the method further includes, in response to receiving the request, processing, using a vision language model (VLM), the request and prior tour image frames for the environment, to generate VLM output.
- VLM vision language model
- the prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment. Processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment.
- the method further includes determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request.
- the method further includes, in response to determining the goal image frame, causing the at least one robot to navigate, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
- causing the at least one robot to navigate to the goal waypoint includes determining, based on the goal image frame, the goal waypoint that corresponds to the goal image frame and transmitting, to the at least one robot, the goal waypoint.
- causing the at least one robot to navigate to the goal waypoint includes transmitting, to the at least one robot, the goal image frame.
- the robot determines, based on the transmitted goal image frame, the goal waypoint that corresponds to the goal image frame.
- the goal waypoint corresponds to the goal image frame based on a prior determination that the goal image frame was captured from the goal waypoint.
- the prior determination that the goal image frame was captured from the goal waypoint can be based on performing structure from motion based on the prior tour image frames.
- the prior tour image frames are captured by one or more cameras of the at least one robot.
- the at least one robot in navigating to the goal waypoint, utilizes a topological graph, such as a topological graph that includes vertexes that are each for a corresponding waypoint and that are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames.
- the at least one robot in navigating to the goal waypoint, utilizes structure-from-motion based on the topological graph.
- the topological graph is created utilizing structure-from-motion and based on the tour image frames.
- the at least one robot in navigating to the goal waypoint, utilizes a topological graph that includes vertexes that are each for a corresponding waypoint and that are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames.
- causing the at least one robot to navigate to the goal waypoint includes transmitting, to the at least one robot, data that specifies a vertex, of the vertexes, that corresponds to the goal image frame. For example, the data could specify an index value for the vertex.
- the method further includes processing, using the VLM and along with the request and the prior tour image frames, a corresponding natural language representation of each of the prior tour image frames.
- determining, based on the VLM output, the goal image frame that is responsive to the request includes determining, based on the VLM output, a given one of the corresponding natural language representations and determining the goal image frame based on the given one of the corresponding natural language representations being for the goal image frame.
- the method further includes processing, using the VLM and along with the request, the prior tour image frames, and the corresponding natural language representations, instructions to output the corresponding natural language representation for one of the prior tour image frames that matches the request.
- the method further includes processing, using the VLM and along with the request and the prior tour image frames, a corresponding narrative language description for one or more of the prior tour image frames.
- the corresponding narrative language descriptions can be human provided, are captured prior to receiving the request, and are each associated with a corresponding one of the prior tour image frames.
- a subset of the prior tour image frames can each be associated with a respective corresponding narrative language description.
- a given prior tour image frame can have an associated narrative language description based on a human providing the narrative language description at or near a time that the given prior tour image frame was captured, or based on a human later explicitly annotating (e.g., via a user interface) the narrative language description for the given prior tour image frame.
- the prior tour image frames are captured by one or more cameras of the at least one robot.
- the request is a user request.
- the request includes the natural language and the natural language is based on spoken or typed input provided by a user.
- the user request can include audio data that captures a spoken utterance of the user and the audio data can be processed utilizing the VLM and/or a transcription, generated based on the audio data, can be generated and processed utilizing the VLM.
- the request includes the one or more images.
- the one or more images can capture a user, can capture a drawing made by a user, or can capture an automatically generated image generated by an image generation model based on input(s) of a user (e.g., based on natural language input of a user).
- the user request is captured via one or more input devices of the at least one robot, such as one or more microphones and/or one or more cameras.
- the request is transmitted by the at least one robot in the environment and is received via one or more networks.
- one or more of the processors are in one or more servers that are not in the environment.
- a method implemented by processor(s) includes capturing a user request via one or more input devices of a robot in an environment.
- the user request includes natural language provided by a human and/or one or more images, such as current image(s) captured in the environment.
- the method further includes, in response to capturing the user request, causing the user request and prior tour image frames for the environment to be processed, using a vision language model (VLM), in determining a goal image frame, of the prior tour image frames, that is responsive to the request.
- VLM vision language model
- the prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment.
- the method further includes receiving, in response to the causing, an indication of a goal waypoint, in the environment, that corresponds to the determined goal image frame.
- the method further includes, in response to receiving the indication of the goal waypoint, navigating the robot to the goal waypoint.
- navigating the robot to the goal waypoint includes identifying, based on the indication of the goal waypoint and in a topological graph that includes a plurality of vertexes, a goal vertex that corresponds to the goal waypoint.
- the indication of the goal waypoint can directly or indirectly identify the goal vertex.
- each of the vertexes of the topological graph are for a corresponding waypoint and are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames.
- navigating the robot to the goal waypoint includes using structure-from-motion that is based on the topological graph and that utilizes the goal vertex as a goal.
- navigating the robot to the goal waypoint includes identifying a sequence of the vertexes that are connected in the topological graph and that define a path between a current waypoint of the robot and the goal waypoint, and utilizing the sequence of the vertexes in navigating to the goal waypoint.
- the goal waypoint corresponds to the goal image frame based on a prior determination that the goal image frame was captured from the goal waypoint.
- a corresponding natural language representation of each of the prior tour image frames is processed, using the VLM and along with the user request and the prior tour image frames, in determining the goal image.
- instructions, to output the corresponding natural language representation for one of the prior tour image frames that matches the user request are processed, using the VLM and along with the user request, the prior tour image frames, and the corresponding natural language representations, in determining the goal image.
- a corresponding narrative language description for one or more of the prior tour image frames is processed, using the VLM and along with the user request and the prior tour image frames, in determining the goal image.
- the corresponding narrative language descriptions are human provided, are captured prior to receiving the request, and are each associated with a corresponding one of the prior tour image frames.
- the prior tour image frames are captured by one or more cameras of the robot.
- the user request includes the natural language and the natural language is based on spoken or typed input provided by the human.
- the user request includes the one or more images.
- the one or more input devices include one or more microphones and/or one or more cameras.
- causing the user request and prior tour image frames for the environment to be processed, using the VLM, in determining the goal image that is responsive to the request includes transmitting the user request from the robot to one or more remote servers.
- the request can be transmitted over one or more networks to one or more edge or cloud servers and optionally utilizing an application programming interface (API).
- causing the user request and prior tour image frames for the environment to be processed, using the VLM, in determining the goal image that is responsive to the request includes transmitting, to the one or more servers and with the user request, an indication of the robot and/or of the environment.
- the one or more servers utilize the indication of the robot and/or the environment in identifying the prior tour image frames.
- a first request from a first robot in a first environment can be provided a first identifier that is used by the server(s) to identify first prior tour image frames, from a prior tour of the first environment, for processing along with the first request using the VLM.
- a second request from a second robot in a second environment can be provided with a second identifier that is used by the server(s) to identify second prior tour image frames, for a prior tour of the second environment, for processing along with the second request using the VLM.
- one or more of the processors are integrated as part of the robot.
- a method implemented by processor(s) includes receiving a request that is directed to a robot in an environment and that includes natural language and/or one or more images.
- the method further includes, in response to receiving the request, processing, using a vision language model (VLM), the request and prior tour image frames for the environment, to generate VLM output.
- VLM vision language model
- the prior tour image frames for the environment are captured, prior to the request, throughout at least a portion of the environment. Processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the robot in the environment.
- the method further includes determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request.
- the method further includes, in response to determining the goal image frame, navigating the robot, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
- a method implemented by processor(s) includes receiving a user request that is directed to a robot in an environment and that includes natural language and/or one or more images.
- the method further includes, in response to receiving the user request, processing, using a vision language model (VLM), the user request and prior tour image frames for the environment, to generate VLM output.
- VLM vision language model
- the prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment.
- Processing the prior tour image frames for the environment, using the VLM and along with the user request is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment.
- the method further includes determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the user request.
- the method further includes determining, based on processing the goal image frame and the user request, a response to the user request.
- the method further includes causing the robot to render the response responsive to the user request.
- the user request includes the natural language and/or the response to the user request is a natural language response that is audibly rendered via one or more speakers of the robot.
- determining, based on processing the goal image frame and the user request, the response includes processing the goal image frame and the user request using the VLM or an additional VLM.
- determining, based on processing the goal image frame and the user request, the response includes determining to cause the robot to render the response responsive to the user request in lieu of causing, responsive to the user request, the robot to navigate to a goal waypoint that corresponds to the goal image frame.
- Some implementations includes a system having memory storing instructions and one or more processors (e.g., graphic processing unit(s), tensor processing unit(s), and/or central processing unit(s)) operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
- Some implementations include at least one transitory or non-transitory computer-readable medium including instructions that, in response to execution by one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
- Some implementations includes a robot having actuators, vision component(s) (e.g., RGB camera, RGB-D camera, LIDAR, and/or other vision component(s)) memory storing instructions, and one or more processors operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
- vision component(s) e.g., RGB camera, RGB-D camera, LIDAR, and/or other vision component(s)
- memory e.g., a robot having actuators, vision component(s) (e.g., RGB camera, RGB-D camera, LIDAR, and/or other vision component(s)) memory storing instructions, and one or more processors operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
- Some implementations includes a system having memory storing instructions and one or more processors operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
- Some implementations include at least one transitory or non-transitory computer-readable medium including instructions that, in response to execution by one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
Landscapes
- Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Remote Sensing (AREA)
- Automation & Control Theory (AREA)
- Multimedia (AREA)
- General Health & Medical Sciences (AREA)
- Radar, Positioning & Navigation (AREA)
- Aviation & Aerospace Engineering (AREA)
- Artificial Intelligence (AREA)
- Health & Medical Sciences (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Engineering & Computer Science (AREA)
- Software Systems (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Manipulator (AREA)
- Image Analysis (AREA)
Abstract
Implementations relate to a method for robotic navigation that involves receiving a multimodal request directed to a robot in an environment. A vision language model (VLM) is used to process the request and previously captured tour image frames from the environment. Based on the VLM output, a goal image frame responsive to the request is determined from the tour image frames. The robot is then caused to navigate to a goal waypoint in the environment that corresponds to the determined goal image frame.
Description
ROBOT NAVIGATION UTILIZING VISUAL LANGUAGE MODEL
Background
[0001] Recent advances of visual language models (VLMs) have shown impressive capabilities across various reasoning and generation tasks. Such VLMs include long context VLMs such as those that have context windows of greater than 500,000 tokens, 1 million tokens, 1.5 million tokens, or other quantity of tokens.
[0002] However, such VLMs may not be readily adaptable to utilization in various robotic tasks, such as robotic navigation. For example, such VLMs can produce only textual outputs, while navigation tasks can require outputting continuous coordinate actions.
Summary
[0003] Implementations disclosed herein relate to utilizing a vision/visual language model (VLM) in robotic navigation. Some of those implementations process, utilizing the VLM, a request and prior tour image frames, to generate output that reflects a goal image frame, of the prior tour image frames, that corresponds to the instruction. The request can be a user request that is directed to a robot in an environment, such as a multimodal user request. For example, the multimodal user request can include natural language from the user and image(s) of the user and/or of content created by the user. The prior tour image frames capture at least some portions of the environment of the robot and are captured prior to receiving the request. For example, the prior tour image frames can be from video previously captured by camera(s) of the robot, or an additional robot, while a human navigates (e.g., through remote operation or kinesthetic guiding) the robot or the additional robot through the environment. The goal image frame, that is reflected by the output generated utilizing the VLM, can be used to identify a corresponding goal waypoint. For example, the goal waypoint can correspond to a pose from which the goal image frame was captured and can be determined and associated with the goal image frame prior to receiving the request. For instance, waypoints can be previously associated with prior tour image frames utilizing structure-from-motion and/or other technique(s).
[0004] A lower-level policy can then be utilized in processing, at each step, the determined goal waypoint and a corresponding observation, to produce waypoint actions for the robot to execute to navigate from a current waypoint (waypoint of the robot when the request was received) to the goal waypoint in the environment. For example, the lower-level policy can be an image-based goal reaching policy such as one that utilizes structure-from-motion and a topological graph constructed based on the prior tour image frames.
[0005] Some implementations disclosed herein are directed to receiving a request that is directed to at least one robot in an environment, the request including natural language and/or one or more images. In response to receiving the request, the request and prior tour image frames for the environment are processed, using a VLM, to generate VLM output. The prior tour image frames for the environment captured, prior to receiving the request, throughout at least a portion of the environment. Processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment. Based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request is determined. In response to determining the goal image frame, the at least one robot is caused to navigate, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
[0006] As a non-limiting example of some implementations disclosed herein, consider a scenario where a service robot operates within a large office building. Prior to receiving any requests, the robot has conducted a comprehensive "tour" of the building, capturing thousands of image frames as it traversed various hallways, meeting rooms, and common areas. These "prior tour image frames" are stored and associated with corresponding waypoint data, indicating the robot's pose (position and orientation) at the time each image was captured. A user, seeking to return a forgotten document to a specific conference room, initiates a request to the robot. This request can include natural language, such as "Take this document to the large conference room with the glass wall". Upon receiving this request, a processing system, utilizing a vision language model (VLM), processes the request
alongside the vast collection of prior tour image frames. Through such processing, the system generates VLM output that indicates a specific "goal image frame" from the prior tour images, which it determines to best represent the large conference room with the glass wall. This goal image frame directly corresponds to a "goal waypoint" in the office environment (e.g., based on the stored associated waypoint data). In response to determining this goal image frame and its associated goal waypoint, the system then causes the robot to navigate autonomously through the office building, using its mapping and localization systems, to reach the identified goal waypoint.
[0007] In some implementations, to cause a robot to navigate to a goal waypoint, a system determines the goal waypoint that corresponds to a particular goal image frame and then transmits that goal waypoint to the robot. For instance, if the goal image frame is a picture of a kitchen counter, the system would identify the precise location (waypoint) where that picture was taken and send those coordinates to the robot. Alternatively, in some implementations, the system may transmit the goal image frame directly to the robot, and the robot itself determines the corresponding goal waypoint. This could involve the robot analyzing the received image and using its own internal mapping capabilities to deduce the exact location it needs to reach.
[0008] In some implementations, a goal waypoint corresponds to a particular goal image frame because it was previously determined that the goal image frame was captured from that specific goal waypoint. For example, during a pre-recorded tour of an environment, if an image was taken at a specific corner of a room, that image would be linked to the waypoint representing that corner. When the system later identifies that image as a goal, it can utilize the corresponding linked waypoint.
[0009] In some implementations, when a robot navigates to a goal waypoint, it uses a topological graph. This graph can include various points (vertices), where each point represents a specific waypoint. Each of these points can be linked to a corresponding image frame from a prior tour of the environment, or to specific visual characteristics (image features) that were derived from those image frames. For example, if a robot is navigating in a house, the topological graph might have a vertex for the living room doorway, which is
linked to an image of that doorway captured during a previous walkthrough. In some particular examples, the robot might use a technique known as structure-from-motion, which relies on the topological graph, to help it move to the goal waypoint.
[0010] In some implementations, a system further processes a natural language representation of each prior tour image frame. This processing can use a VLM along with the request and the tour image frames. For example, if the system has thousands of tour images, each image could have a generated text description, like "a view of the breakroom with a coffee machine" or "the main hallway leading to the elevators." When a request comes in, the VLM is used to process these natural language descriptions in addition to the images themselves. In some specific examples, the system might determine the goal image frame by first identifying a natural language representation that matches the request. For instance, if the request is "Go to the coffee machine," the system might find the natural language representation "a view of the breakroom with a coffee machine" and then identify the corresponding image frame as the goal. In some further examples, the VLM might be instructed to output the natural language representation for the tour image frame that best matches the received request.
[0011] In some implementations, the system also processes narrative language descriptions that are associated with some or all of the prior tour image frames. These descriptions are provided by a human and are captured prior to receiving a request. For example, a person might have walked through an office building and narrated observations like "This is where we keep the shared printer" while a tour video was being recorded. This narrative could then be linked to the specific image frames captured at that location. When a request is received, these human-provided narratives can be processed, using a VLM and along with the request and the tour images, to help identify the most relevant goal image frame.
[0012] In some implementations, the prior tour image frames are captured by cameras that are part of the robot itself. For example, a service robot might routinely traverse its operational environment, recording video footage that later becomes the set of prior tour image frames.
[0013] In some implementations, a received request is a user request. This request can include natural language, which might be based on spoken or typed input from a user. For example, a user might verbally tell the robot, "Go to the supply closet," or type the same instruction. In some cases, the request might also include one or more images, such as an image of the user, an object the user is holding, or even a drawing made by the user to indicate a desired location. These user requests can be captured through input devices on the robot, such as microphones for spoken commands or cameras for visual input.
[0014] In some implementations, the request is sent by a robot located in the environment and received over one or more networks. For example, a robot in a warehouse might send a request for navigation assistance to a central server via the internet. In some of these cases, the processing of the request and the tour image frames occurs on one or more remote servers that are not physically located within the robot's immediate environment.
[0015] In some implementations, a system captures a user request using input devices on a robot in a particular environment. This user request can include natural language provided by a human, such as a spoken command and/or one or more images of the human or their surroundings. In response to capturing this request, the system arranges for the user request and prior tour image frames of the environment to be processed using a VLM. This processing helps to identify a goal image frame from the prior tour images that is relevant to the request. The prior tour image frames were captured beforehand, covering at least a portion of the environment. After this processing, the robot receives an indication of a goal waypoint in the environment that corresponds to the identified goal image frame. Upon receiving this indication, the robot then navigates to that goal waypoint.
[0016] In some implementations, navigating the robot to the goal waypoint involves identifying a specific "goal vertex" within a topological graph that corresponds to the goal waypoint. This topological graph contains various points (vertices), each representing a waypoint, and each linked to a corresponding prior tour image frame or features derived from those frames. For instance, if the goal waypoint is "the corner by the water cooler," the system would identify the specific vertex in the graph that represents that corner. In some examples, the robot then uses structure-from-motion, based on this topological graph
and using the goal vertex as its target, to move to the waypoint. In other examples, the robot identifies a sequence of connected vertices in the topological graph that form a path from its current location to the goal waypoint, and then uses this sequence to navigate. The goal waypoint can correspond to the goal image frame because it was previously determined that the goal image frame was originally captured from that specific location. [0017] In some implementations, when determining the goal image, a natural language representation for each of the prior tour image frames is processed using the VLM, along with the user request and the prior tour image frames themselves. For example, if a tour image has a textual description like "area near the main entrance," this description is used to help match the user's request, such as "Go to the entrance." In some further examples, the system might process instructions that direct the VLM to output the natural language representation for the tour image frame that best matches the user's request, which then helps to identify the goal image.
[0018] In some implementations, a corresponding narrative language description for one or more of the prior tour image frames is also processed using the VLM when determining the goal image. These narrative descriptions are provided by humans, captured before the request is received, and are associated with their respective prior tour image frames. For example, a human might have provided a narrative like "This is the conference room where we hold large meetings" while recording a tour, and this narrative would be used in processing using by the VLM in conjunction with the user request to find the relevant image. [0019] In some implementations, the prior tour image frames are captured by cameras that are part of the robot itself. For example, a robot might autonomously perform a mapping tour, recording visual data that is then used to create these image frames.
[0020] In some implementations, the user request includes natural language, which is based on either spoken or typed input from a human. For example, a human might simply say "Find my keys" or type "Go to the break room." In some implementations, the user request can additionally or alternatively include one or more images. For instance, a human might show the robot a picture of an object they are looking for. In some implementations, the
input devices used to capture the user request include microphones, cameras, or both. This allows for multimodal input from the human user.
[0021] In some implementations, to facilitate the processing of the user request and the prior tour image frames using the VLM for determining the goal image, the robot transmits the user request to one or more remote servers. For example, a robot might send a user's verbal command to a cloud server for processing. In some specific examples, the robot also transmits an indication of itself or its environment along with the user request to these servers. The servers then use this indication to identify the correct set of prior tour image frames relevant to that specific robot or environment for processing.
[0022] In some implementations, one or more of the processors that perform these operations are integrated directly into the robot. This could allow for faster, more localized processing of the request and navigation commands.
[0023] In some implementations, a system receives a request that is directed to a robot in an environment, and this request includes natural language and/or one or more images. In response to receiving the request, the system processes both the request and prior tour image frames of the environment using a VLM. The tour image frames were captured beforehand, covering at least a portion of the environment. The processing of these image frames with the VLM, along with the request, is performed because the images cover the environment and the request is specifically for a robot within that environment. Based on the output generated using the VLM, the system identifies a goal image frame from the prior tour images that is relevant to the request. Then, in response to identifying this goal image frame, the robot is navigated in the environment to a specific goal waypoint that corresponds to that goal image frame.
[0024] In some implementations, a system receives a user request directed to a robot in an environment, which includes natural language and/or one or more images. In response, the system processes this user request and prior tour image frames of the environment using a vision language model (VLM) to generate VLM output. The prior tour image frames were captured beforehand throughout at least a portion of the environment. This processing is based on the tour image frames covering the environment and the request being for a robot
within that environment. Based on the VLM output, a goal image frame from the prior tour images is determined to be responsive to the user request. Subsequently, based on processing this goal image frame and the user request, a specific response to the user request is determined. The system then causes the robot to render this response in a way that is responsive to the user's request.
[0025] In some implementations, the user request includes natural language, and the determined response to the user request is also a natural language response. This response can be audibly rendered by the robot using its speakers. For example, if a user asks, "Where is the nearest exit?", the robot might audibly reply, "The nearest exit is to your left, past the green door."
[0026] In some implementations, the determination of the response, based on processing the goal image frame and the user request, involves using the same vision language model (VLM) or an additional VLM for the processing.
[0027] In some implementations, the determination of the response to the user request involves deciding to cause the robot to render a response instead of causing the robot to navigate to a goal waypoint. For example, if the request is "Describe this object" while showing an image of an object, the system might determine that an audible description from the robot is the appropriate response, rather than instructing the robot to navigate to the object's location.
[0028] The preceding is presented as a non-limiting overview of only some implementations disclosed herein.
Brief Description of the Drawings
[0029] FIG. 1 is a diagram illustrating an example of processing of a multimodal user instruction for robot navigation.
[0030] FIG. 2 is a flowchart illustrating an example method of processing a request and prior tour image frames, using a VLM, to determine a goal image frame that is responsive to the request and causing a robot to navigate to a goal waypoint that corresponds to the goal image frame.
[0031] FIG. 3 is a flowchart illustrating an example method of causing a request and prior tour image frames to be processed, using a VLM, to determine a goal image frame, receiving, in response to the causing, an indication of a goal waypoint that corresponds to the goal image frame, and navigating a robot to the goal waypoint.
[0032] FIG. 4 is a flowchart illustrating an example method of processing a request and prior tour image frames, using a VLM, to determine a goal image frame that is responsive to the request, determining a response to the request based on processing the request and the goal image frame, and causing a robot to render the response.
[0033] FIG. 5 illustrates an example robotic system according to some implementations.
[0034] FIG. 6 illustrates an example computing system according to some implementations.
Detailed Description
[0035] Prior to turning to the Figures, a non-limiting description of some example aspects of the disclosure is provided.
[0036] A long-standing objective in navigation research involves the development of an intelligent agent capable of interpreting multimodal instructions, including natural language and images, to execute effective navigation. To achieve this, a category of navigation tasks referred to as Multimodal Instruction Navigation with demonstration Tours (MINT) is considered. In MINT, environmental context is provided through a previously recorded demonstration video.
[0037] Recent advancements in Vision Language Models (VLMs) have indicated a promising direction for achieving this objective, as VLMs exhibit capabilities in perceiving and reasoning about multimodal inputs. However, VLMs are typically trained to predict textual output, and their optimal utilization in navigation remains an area of ongoing research. [0038] To address MINT, a hierarchical Vision-Language-Action (VLA) navigation policy, denoted as Mobility VLA, is provided. This policy combines the environmental understanding and common-sense reasoning power of long-context VLMs with a robust low-level navigation policy based on topological graphs. The high-level policy incorporates a long-context VLM that receives the demonstration tour video and the multimodal user instruction as input to identify a goal frame within the tour video. Subsequent to this, a low-
level policy utilizes the identified goal frame and a topological graph, which is constructed offline, to generate robot actions at each timestep.
[0039] The Mobility VLA has been evaluated in an 836m2 real-world environment. This evaluation indicates that Mobility VLA achieves high end-to-end success rates on previously unaddressed multimodal instructions, such as "Where should I return this?" when a plastic bin is held by a user.
[0040] Robot navigation has undergone significant advancements. Early systems relied on users specifying physical coordinates in pre-mapped environments. Object goal and Vision Language Navigation (ObjNav and VLN) represent a notable progression in robot usability, enabling the use of open-vocabulary language to define navigation goals, such as "Go to the couch." To enhance the utility and pervasiveness of robots in daily life, a further advancement is proposed by extending the natural language space of ObjNav and VLN into the multimodal domain. This implies that a robot is configured to accept natural language and/or image instructions concurrently. For instance, a person unfamiliar with a building may inquire, "Where should I return this?" while holding a plastic bin. In such a scenario, the robot guides the user to a designated shelf for returning the bin, based on both verbal and visual context. This category of navigation tasks is termed Multimodal Instruction Navigation (MIN).
[0041] Multimodal Instruction Navigation (MIN) is a task category encompassing environment exploration and instruction-guided navigation. However, in many scenarios, exploration may be circumvented by employing a demonstration tour video that comprehensively traverses the environment. A demonstration tour offers several advantages. For example, collection is easy as a user may teleoperate the robot or utilize a portable electronic device, such as a smartphone, to record a video while traversing the environment. Various exploration algorithms can additionally or alternatively be employed to generate such a tour. As another example, such a tour aligns with common user practices. For instance, upon acquiring a new robotic device for a home environment, a user may naturally introduce the robotic device to the home environment and verbally indicate locations of interest during the tour. As yet another example, under certain circumstances,
restricting the motion of the robotic device to a pre-defined zone may be desirable for safety and/or privacy considerations. Accordingly, this disclosure introduces and examines a category of tasks referred to as MINT. This category leverages demonstration tours and focuses on the fulfillment of multimodal user instructions.
[0042] Recently, large Vision-Language Models (VLMs) have demonstrated substantial potential in addressing MINT due to their capabilities in language and image comprehension, as well as common-sense reasoning, all of which are pertinent for MINT. However, the utilization of VLMs in isolation for MINT presents certain challenges. For example, the quantity of input images that can be processed by many VLMs is constrained by context-length limitations. This limitation can significantly impede the accuracy of environmental understanding in expansive environments. As another example, the execution of MINT tasks necessitates the determination of robot actions. Queries formulated to elicit such robot actions are typically outside the distribution of data used for VLM training or pre-training. Consequently, the performance of zero-shot navigation may be less than optimal.
[0043] A Mobility Vision-Language-Action (VLA) navigation policy is provided for solving MINT. This hierarchical policy combines environment understanding and common sense reasoning capabilities of long-context Vision-Language Models (VLMs) with a robust low- level navigation policy based on topological graphs. Specifically, a high-level VLM is utilized to process a demonstration tour video and a multimodal user instruction to generate output that is utilized to identify a goal frame within the tour video. Subsequent to this identification, a classical low-level policy employs the goal frame and a topological graph, which is constructed offline from the tour frames to generate robot actions, such as waypoints, at each timestep. The application of long-context VLMs addresses issues related to the fidelity of environment understanding. Furthermore, the topological graph facilitates bridging a gap between a VLM's training distribution and the specific robot actions for solving MINT.
[0044] Mobility VLA was evaluated in a real-world office environment spanning 836m2 and a home-like environment. The evaluation results indicate that implementations of Mobility
VLA achieved success rates of 86% and 90% (26% and 60% higher than baseline methods) in these environments, respectively, on MINT tasks that were previously considered infeasible. These tasks involved complex reasoning, such as interpreting instructions like "I want to store something out of sight from the public eye. Where should I go?" and multimodal user instructions. Furthermore, the ease with which users can interact with a robot utilizing this approach was shown to be significantly advanced. For instance, a user could record a narrated video walkthrough in a home environment using a smartphone and subsequently inquire, "Where did I leave my coaster?"
[0045] The development of this technology encompasses several advancements. A new paradigm for robot navigation, Multimodal Instruction Navigation (MIN), and its variant, Multimodal Instruction Navigation with demonstration Tours (MINT), have been developed. These paradigms enhance the helpfulness and intuitive use of robots. Mobility VLA is provided as a solution for MINT, combining long-context Vision Language Models (VLMs) and topological maps. This combination has resulted in a notable improvement in the naturalness of human-robot interaction and a substantial increase in robot usability.
[0046] Regarding conventional navigation, such methods generally emphasize point-to- point robot movement, where goals are specified using metric coordinates. These systems commonly rely on pre-built or dynamically generated maps and implement path-planning algorithms, such as D* to generate fine-grained navigation commands (e.g., commands for twist drive velocity) to achieve collision-free movement. Similar to prior systems, a low-level controller utilized in various implementations incorporates a visual Simultaneous Localization and Mapping (SLAM) algorithm, COLMAP, and a Model Predictive Control (MPC) method to track desired waypoints obtained from high-level VLMs. This MPC method determines a control input that minimizes a quadratic cost function subject to linear dynamics.
[0047] In contrast to conventional navigation methods that typically exhibit robust behavior but do not leverage semantically meaningful information for specifying navigation targets, object and image goal navigation techniques utilize rich input modalities. These modalities
can include object categories, natural language instructions, dialogue, goal image conditions, and multimodal inputs that combine language and images.
[0048] Most of these approaches involve an active exploration phase because the robot operates without prior knowledge of the environment. In contrast, implementations disclosed herein leverage environment priors provided in the form of a previously collected video tour. Further, some of those implementations have an absence of explicit semantic scene representation graphs, instead relying on the capabilities of VLMs to process raw videos.
[0049] A MINT task involves providing a demonstration tour video and a multimodal user instruction as input. A robot operates to navigate to one or more specified goal locations in order to satisfy the user's instruction.
[0050] Under this setting, the demonstration tour video includes a sequence of first-person view image frames F = {ft \ft G
x W x 3, i = 1, 2, ..., k} taken during a tour of the environment, where k is the number of frames in the video. In addition, optional natural language narratives can be added to certain frames N = |n7 G str,j G [1, 2, ..., fc] . The
multimodal user instruction can be just a text instruction d G str (e.g., "Where can I find a ladder?"), or both text and image instructions I G ]R>HXWX3 (e.g., "Where can I get something to clean this?" + The robot sees the user pointing to a dirty whiteboard).
[0051] Implementations aim to produce a navigation policy (a I O, F, N, d, I), where O G J^HXWX3 js t e robot's current camera observation. The policy emits an embodimentagnostic waypoint action a G IR3 representing longitudinal translation (Ax), lateral translation (A-r .), and rotation along the vertical axis (A0), all in the robot-centric frame. We assume that the robot has an embodiment-specific mechanism to execute waypoint actions.
[0052] A hierarchical navigation policy, referred to herein referred to as Mobility VLA, can incorporate both online and offline processing components.
[0053] In the offline phase, a topological graph G can be generated from the demonstration tour (N, F). Online, the high-level policy takes the demonstration tour and the multimodal
user instruction (d,l) to find the navigation goal frame index g, which is an integer corresponding to a specific frame of the tour. Next, the lower-level policy utilizes the topological graph, the current camera observation (O) and g to produce a waypoint action (a) for the robot to execute at each timestep. This can be represented by the following two equations.
[0054] = h F (1)
[0055] n (a 10, (2)
[0056] where h and I are the high and low-level policies.
[0057] Mobility VLA utilizes a demonstration tour of an environment to solve MINT tasks. This tour can be provided by a human user via teleoperation, or by recording a video on a portable electronic device, such as a smartphone, while traversing the environment.
[0058] The Mobility VLA system constructs a topological graph, designated G = (V, E). The graph G comprises a set of vertices V and a set of edges E, denoted as G = (V, E). Each vertex corresponds to a corresponding frame from the demonstration tour video, which is represented by a set of frames F and an optional set of narratives N. A structure-from- motion pipeline, such as COLMAP, can be employed to determine an approximate 6-Degree- of-Freedom camera pose for each frame. This camera pose information is subsequently stored in the corresponding vertex. A directed edge is added to the topological graph G if a target vertex is positioned substantially in front of a source vertex and is within a predefined distance. Specifically, an edge is added if the target vertex is less than 90 degrees away from the source vertex's pose and is within 2 meters of the source vertex.
[0059] Th is topological graph approach offers a simplified alternative to traditional navigation pipelines, which typically involve mapping the environment, identifying traversable areas, and then constructing a Probabilistic Roadmap (PRM). The topological graph captures the general connectivity of the environment based on the trajectory of the demonstration tour, thereby streamlining the environmental representation.
[0060] During online execution, a high-level policy leverages the common sense reasoning capabilities of VLMs to identify a navigation goal from the demonstration tour. This goal is determined such that it satisfies a wide range of multimodal, colloquial, and often ambiguous user instructions. To facilitate this, a prompt P is prepared, which includes
interleaving text and images. The prompt P can be defined as a function of the tour frames F, the narratives N, the text instruction d, and the image instruction I, denoted as P(F, N, d, I). An illustrative example of the prompt P, formulated for the multimodal user instruction "Where should I return this?" can be as follows:
[0061] You are a robot operating in a building and your task is to respond to the user command about going to a specific location by finding the closest frame in the tour video to navigate to. These frames are from the tour of the building last year. [Frame 1 Image fl]. Frame 1. [Frame narrative nl] ... [Frame k Image fk]. Frame k. [Frame narrative nk]. This image is what you see now. You may or may not see the user in this image. [Image Instruction I], The user says: Where should I return this? How would you respond? Can you find the closest frame?
[0062] The processing utilizing the VLM results in VLM output that reflects an integer goal frame index, denoted as g. Once the goal frame index g is identified by the high-level policy, a low-level policy takes over and produces a waypoint action at every timestep (Eq. 1). [0063] An example low-level policy is demonstrated by Algorithm 1:
Algorithm 1 Low-level Goal Reaching Policy
1: Input: goal frame index g, offline-constructed topological graph G.
2:
3: while timestep < maximum steps do
4: Get new camera observation image O
5: Get start vertex vs and robot pose T by localizing O in G
6: if == it, then
7: Navigation goal reached, break
8: end if
9: Compute
...,vg , the shortest path between vs and vg.
10: Compute waypoint action a from the relative pose between T and
11: Execute a on robot
12: end while
[0064] At each timestep, a real-time hierarchical visual localization system may be utilized to estimate the pose of a robot and an initial vertex for navigation. This localization system is configured to identify a plurality of candidate frames that are nearest to a current camera
observation based on a global descriptor, and subsequently to compute the pose of the robot through a Perspective-n-Point (PnP) algorithm. Following this, a shortest path on the topological graph is identified between the initial vertex and a goal vertex. The goal vertex corresponds to a desired goal frame. Finally, a low-level policy is configured to generate a waypoint action. This waypoint action may include positional and orientational data of a subsequent vertex in the identified shortest path, relative to the current robot pose.
[0065] To demonstrate the performance of the Mobility VLA and to gain further insights into its design, experiments were configured to address the following research questions (RQs):
[0066] RQ1: Does Mobility VLA perform well in MINT in the real world?
[0067] RQ2: Does Mobility VLA outperform alternatives thanks to the use of long-context VLM?
[0068] RQ3: Is the topological graph necessary? Can VLMs produce actions directly? [0069] The environments of the experiments is an office environment occupied by humans. This environment encompasses approximately 836 square meters and includes various items such as shelves, desks, and chairs. The robot of the experiments is a wheel-based mobile manipulator that employs a Model Predictive Control (MPC)-based algorithm to execute waypoint actions, identified as positional and orientational data within the robot's local frame while avoiding obstacles. A demonstration tour is collected by teleoperating the robot with a gamepad. All corridors within the environment are traversed twice from opposing directions. The resulting tour has a duration of approximately 16 minutes, corresponding to 948 image frames captured at 1 Hertz. During the tour, narrative elements are added to specific frames to facilitate personalized navigation. For instance, the narrative "Temp desk for everyone" is associated with a frame at 5:28, and "Lewis' desk" is associated with a frame at 7:14. For the experiments, a collection of 57 user instructions was obtained through crowd-sourcing and categorized into four types: 20 Reasoning-Free (RF), 15 Reasoning-Required (RR), 12 Small Objects (SO), and 10 Multimodal (MM) instructions. A notable characteristic of Reasoning-Required instructions is that they do not explicitly mention the specific object or location to which the robot is directed. Furthermore, the
destination of Multimodal instructions is substantially difficult to infer without the inclusion of an image modality in the instruction. Prior systems are not configured for or evaluated against these two categories of tasks, which represent a key differentiation between MINT and conventional Object Navigation (ObjNav) and Vision Language Navigation (VLN) approaches.
[0070] To evaluate the performance of a VLA navigation policy, such as Mobility VLA, in a real-world MINT environment, a selection of user instructions can be employed. For example, five user instructions may be randomly selected per category. The policy's performance can then be evaluated from a plurality of random starting poses such as four random starting poses, each located at a distance of at least 20 meters from each other. A long-context multimodal VLM, such as Gemini 1.5 Pro may be utilized for this purpose. [0071]Table 2 illustrates that a Mobility VLA policy exhibits a high end-to-end navigation success rate across most user instruction categories. This includes categories such as Reasoning-Required and Multimodal instructions, which were previously considered infeasible for such systems. However, a comparatively lower success rate may be observed in the Small Object category. This outcome can be attributed to limitations in the resolution of the demonstration tour video. The policy can also demonstrate a reasonable Success Rate weighted Path Length (SPL), indicating that the topological graph, as a navigation aid, does not impose a substantial penalty on path length. Furthermore, the Mobility VLA policy has been observed to successfully incorporate personalization narratives from a demonstration tour. For instance, the policy correctly navigated to different locations in response to substantially similar instructions from distinct users. An example includes navigation to a frame at a timestamp of 7:14 when presented with the instruction "I'm Lewis, take me to a temp desk please," and navigation to a frame at a timestamp of 5:28 when presented with the instruction "Hi robot, I'm visiting, can you take me to a temp desk?".
[0072]Table 2 is presented below.
Reasoning-Free Reasoning-Required Small Objects Multimodal
Goal Finding SR 80% 80% 40% 85%
Goal Reaching SR 100% 100% 100% 100% End-to-end SR 80% 80% 40% 85%
SPL 0.59 0.69 0.38 0.64
Table 2: Mobility VLA end-to-end navigation Success Rate (SR) and SPL of various user instruction types in the real Office environment.
[0073]Table 2 additionally indicates the robustness of a Mobility VLA policy's low-level goal reaching exhibiting a 100% success rate in real-world scenarios. This robustness is observed even when the demonstration tour was recorded several months prior to the experiments, during which time many objects, furniture arrangements, and lighting conditions in the environment had undergone changes.
[0074] To further investigate end-to-end performance, simulations can be utilized to scale evaluation numbers. Specifically, a high-fidelity simulation reconstruction of an office environment may be created using techniques such as Neural Radiance Fields (NeRF). The Mobility VLA policy can then be evaluated against 20 language-instructed tasks, with 50 random starting poses allocated per task. Such an experiment yields a high-level goal finding success rate of 90% and a low-level goal reaching success rate of 100%, culminating in a total of 900 successful end-to-end executions. Comprehensive results are shown in Table 8.
Reasoning-Free Reasoning Required
High-Level Goal Finding SR 90% 90%
Low-Level Goal Reaching SR 100% 100%
End-to-end SR 90% 90%
SPL 0.83 0.84
Table 8: End-to-end navigation Success Rate (SR) and SPL of various user instruction types in the simulated Office environment.
[0075] To demonstrate the generality and ease of use of a Mobility VLA policy, a proof-of- concept experiment can be conducted in a real home-like environment. Rather than
providing a robot with a teleoperated tour, a smartphone can be utilized to record a demonstration tour. Subsequently, the Mobility VLA policy can be evaluated end-to-end with, for example, 4 Reasoning-Required and 1 Small Object user instructions, each from 4 random starting positions. This evaluation results in a 100% success rate with an SPL of 0.87. This outcome indicates that a Mobility VLA policy performs well regardless of the specific environment. Additionally, it highlights the ease of deployment, as a user may simply utilize a smartphone to record a tour of an environment, upload the recorded tour to a robot, and then immediately commence providing instructions.
[0076]Table 3 illustrates how well alternative methods perform compared to Mobility VLA. Concretely, Mobility VLA is compared with CLIP-based retrieval and Text-Only Mobility VLA. With CLIP-based retrieval, the high-level goal finding module of NLMap is reproduced by adopting OWL-ViT for region proposal and CLIP for sub-regions and full-images embeddings extraction for tour frames. Goal frame retrieval is then performed using CLIP embeddings of the instruction language and image. With Text-Only Mobility VLA, the multimodal demonstration tour is captioned by a VLM frame-by-frame to form a "text tour". An LLM then uses the text tour to produce the goal frame index. Table 3 shows that high-level goal finding success rates of Mobility VLA are significantly higher than comparison methods. Given the 100% low-level success rate, this high-level goal finding success rates are representative of end-to-end success rates.
Success Rates Reasoning- Reasoning- Small Multimodal
Free Required Objects
CLIP-based retrieval 35% 33% 25% 20%
Text Only Mobility VLA 70% 60% 50% 30%
Mobility VLA (Ours) 95% 86% 42% 90%
Table 3: High-level goal finding Success Rates of Mobility VLA compared to baselines [0077] Feeding a full demonstration tour of a large environment into non-long-context VLMs is challenging since each image requires hundreds-of-token budgets. One solution for reducing input tokens number is reducing the tour video frame rate, at the cost of intermediate frames loss. Table 4 shows that the high-level goal finding success rate
decreases as the tour frame rate decreases. This can be due to a lower frame rate tour sometimes missing the navigation target frame. In addition, comparing various VLMs, only Gemini 1.5 Pro currentlyvyields satisfactory success rate thanks to its long IM token context-length.
Frame GPT-4V | GPT-4o | Gemini 1.5 PRO
Rate RF RR SO MM RF RR SO MM RF RR SO MM
0.2 FPS 60% 53% 17% 30% 75% 40% 25% 50% 95% 67% 36% 60%
1 FPS Exceeds token limit Exceeds token limit 95% 86% 42% 90%
Table 4: High-level goal finding Success Rates with regards to various user instruction types (presented in the order of Reasoning Free (RF), Reasoning Required (RR), Small Objects (SO), MultiModal (MM)) as a function of VLM models (column) and multimodal demonstration tour Frames Per Second (FPS) (row). All VLMs were queried in June 2024.
[0078] As one selected qualitative comparison example for high-level goal finding of all candidates approaches, consider the multimodal instruction of "I want more of this." and a picture of several soft drink cans on a desk. Mobility VLA correctly identified the frame containing the refrigerator which it should lead the user to. On the other hand, CLIP-based retrieval finds a region in which a water bottle and some stuff are on a desk to be most similar to the full instruction image, given it is hard to extract "what the user want" from the instruction image using Owl-ViT. GPT-4o incorrectly attempts to find the frame closest to the instruction image, while GPT-4V refuses to give a frame number since it was unable to find a frame where beverages are. Lastly, the Text only approach cannot understand whether "this" refers to the Coke cans or the office setting, since it relies only on caption of the instruction image.
[0079] Mobility VLA uses a hierarchical architecture to harness long-context VLM's reasoning capability and uses a topological graph to produce waypoint actions. Such a topological graph can be beneficial for navigation success. Table 5 shows the end-to-end
performance of Mobility VLA in simulation compared to prompting the VLM to output waypoint actions directly. The 0% end-to-end success rate shows that Gemini 1.5 Pro is incapable of navigating the robot zero-shot without the topological graph. Empirically, we found that Gemini almost always outputs the "move forward" waypoint action regardless of the current camera observation. In addition, the current Gemini 1.5 API requires the upload of all 948 tour images at every inference call, resulting in a prohibitively expensive 26s perstep running time for the robot to move just lm. On the other hand, Mobility VLA's high- level VLM spends 10-30s to find a goal index and then the robot navigates to the goal using the low-level topological graph results in a highly robust and efficient (0.19s per step) system for solving MINT.
Direct Waypoint Goal Index Output +
Action Input Topological Graph
Success Rate 0% 90%
SPL - 0.84
Per-step inference Time 25.90+8.36s 0.19+0.047s
Table 5: End-to-end navigation Success Rate and SPL as a function of VLM (Gemini 1.5 Pro) output format in the simulated Office environment.
[0080] Various implementations use COLMAP, a structure-from-motion pipeline to estimate the pose of the robot for each frame in the tour (i.e. reference images), 3D point landmarks in the environment, and their corresponding 2D projections across all reference images (i.e. 2D-3D correspondences). The poses are used to build a fully connected topological graph.
[0081]The tour frames F, 3D landmarks, and 2D features are used in some implementations of a real-time hierarchical localizer. The method is hierarchical since it divides localization of the observed image O into two steps: a global search to determine a set of candidate reference images close to O followed by local feature matching and pose estimation.
[0082] In the global search, the candidate set C £ p of k-nearest (w.r.t. the I2 -norm of a global image descriptor [54]) tour frames to O is determined. 2D features in O are matched to the 2D features of each frame in C. Using the pre-computed 2D-3D correspondences,
correspondences between 2D features in O and 3D landmarks observed in the tour are established.
[0083] Given the set of 2D-3D correspondences for each frame in C, the pose of O is computed by solving the corresponding Perspective-n-Point problem. The pose with the most inlier 2D-3D correspondences is selected as To.
[0084] When To is used to determine the closest vertex on G, the scale-ambiguity characteristic of monocular structure-from-motion systems is inconsequential to the high- level goal-finding policy. However, when computing the waypoint action for low-level navigation (see Algorithm 1), the scale factor is utilized to generate metrically accurate actions.
[0085] Implementations also investigate if strictly multimodal user instructions (instructions that are nearly impossible to answer without the image) can be answered by the text modality alone. To this end, the image part of the multimodal user instructions is replaced with its caption. Table 7 shows the high-level goal reaching success rate of such setup in the Text Instruction columns compared to feeding VLMs the image (MM Instruction column). Table 7 shows that the success rate is much higher when multimodal demo tour and image instructions are fed to the VLM (lower right corner).
Success Rates GPT-4o GPT-4o Gemini 1.5 Pro Gemini 1.5 Pro
Text Instruction MM Instruction Text Instruction MM Instruction
Text Tour 0.10 0.10 0.20 0.20
Multimodal Tour Exceeds token limit Exceeds token limit 0.40 0.90
Table 7: High-level goal finding Success Rates of multimodal user instructions as a function of VLM models and instruction representations (columns) and tour modalities (row). In MM Instructions columns, the robot's current camera observation is fed directly into the VLMs. In Text Instructions columns, the camera observation is captioned by Gemini 1.5 Pro and the caption text is then fed into the VLMs. The text tour was captioned w/ Gemini 1.5 Pro.
[0086] Tu rning now to the Figures, FIG. 1 is a diagram illustrating an example of processing of a multimodal user instruction for robot navigation. FIG. 1 illustrates an overall system
architecture 100 for robot navigation, representing a hierarchical approach that combines high-level understanding with low-level motion control. The system processes multimodal inputs to facilitate effective navigation. In FIG. 1, the multimodal user instruction and a demonstration tour video of the environment are used by a long-context VLM (high-level policy) to identify the goal frame in the video. The low-level policy then uses the goal frame and an offline generated topological map (from the tour video using structure-from-motion) to compute a robot action at every timestep
[0087] FIG. 1 depicts a multimodal user instruction 101 and a demonstration tour video 102 as primary inputs to a high level goal finding with long-context VLM module 120. The high level goal finding with long-context VLM module 120 utilizes a VLM to process the multimodal user instruction 101 and the demonstration tour video 102. For example, the module 120 can generate a prompt that includes the multimodal user instruction 101 and the demonstration tour video 102 and process the prompt using the VLM. A navigation goal 103 is reflected in VLM output generated from such processing. The navigation goal 103 can correspond to a specific frame or location within the demonstration tour.
[0088] The architecture 100 further includes an offline component for generating a topological graph 104 from the demonstration tour video 102 through a structure from motion module 110. The topological graph 104 provides a simplified representation of the environment's connectivity and can be utilized for efficient low-level navigation. During operation, a localization module 124 utilizes observation data 106 (e.g., current images and/or other state data captured by sensor(s) of robot 115) from the robot 115 and the topological graph 104 to determine the robot's current pose within the environment. The pose is provided to path finding module 122
[0089] The navigation goal 103, identified by the module 120, is fed into a path finding module 122. The path finding module 122, utilizing the topological graph 104 and real-time robot pose information from localization module 124, computes a waypoint action 105. The waypoint action 105 represents the specific movement commands (e.g., change in x, y, and theta coordinates) required for the robot 115 to navigate towards the designated goal. The waypoint action 105 is then executed by the robot 115. This hierarchical design allows the
system to leverage the powerful reasoning capabilities of VLMs for goal interpretation while relying on a robust, graph-based approach for precise and efficient movement.
[0090] FIG. 2 is a flowchart illustrating an example method 200 of processing a request and prior tour image frames, using a vision language model (VLM), to determine a goal image frame that is responsive to the request and causing a robot to navigate to a goal waypoint that corresponds to the goal image frame. For example, consider a scenario where a user wants a robot to retrieve an object from a specific location in a house.
[0091] In block 252, a request is received that is directed to at least one robot in an environment. This request can include natural language, as shown in sub-block 252A, and/or one or more images, as shown in sub-block 252B. For example, the user might say, "Go to the kitchen and grab the red apple," which is a natural language input, and/or show the robot a picture of a red apple, which is an image input.
[0092] In block 254, in response to receiving the request, a vision language model (VLM) is used to process the request and prior tour image frames for the environment to generate VLM output. These prior tour image frames are captured throughout at least a portion of the environment before the request is received. For example, if the robot previously toured the house and captured images of every room, including the kitchen, these prior tour images would be processed by the VLM along with the user's request. The processing considers that the images are from the environment and the request is directed to the robot operating within that environment.
[0093] In block 256, based on the VLM output, a goal image frame is determined from the prior tour image frames that is responsive to the request. For example, the VLM, having processed the request "Go to the kitchen and grab the red apple" and the tour images, identifies a specific image frame of the kitchen showing a red apple as the goal image frame. [0094] In block 258, in response to determining the goal image frame, the at least one robot is caused to navigate, in the environment, to a goal waypoint that corresponds to the goal image frame. As one example shown in sub-block 258A, a system may determine the goal waypoint based on the goal image frame and then transmit that goal waypoint to the robot. For instance, if the goal image is of the kitchen counter where the apple is, the system
identifies the precise GPS coordinates or internal map location (waypoint) of that counter and sends those coordinates to the robot. As another example shown in sub-block 258B, the system may transmit the goal image frame directly to the robot, and the robot itself determines the corresponding goal waypoint based on its internal mapping capabilities and the received image. The robot then proceeds to navigate to that determined goal waypoint, thereby completing the user's request to retrieve the red apple.
[0095] FIG. 3 is a flowchart illustrating an example method 300 of causing a request and prior tour image frames to be processed, utilizing a vision language model (VLM), to determine a goal image frame, receiving, in response to the causing, an indication of a goal waypoint that corresponds to the goal image frame, and navigating a robot to the goal waypoint. For example, consider a scenario where a robot is located in a factory, and a human worker needs assistance finding a specific tool.
[0096] In block 352, a user request is captured via one or more input devices of the robot in the factory environment. This request can include natural language, as shown in sub-block 352A, and/or one or more images, as shown in sub-block 352B. For instance, the worker might say, "Go to the wrench rack" (natural language), and/or show the robot a picture of a wrench (image). The robot's onboard microphone or camera captures this request.
[0097] In block 354, in response to capturing the user request, the request and prior tour image frames for the environment are caused to be processed, using a VLM, to determine a goal image frame from the prior tour image frames that is responsive to the request. These prior tour image frames were captured beforehand, perhaps during an initial mapping tour of the factory, covering various areas including the tool storage section. The robot may transmit the captured request to a central server, which then utilizes the VLM to process the request against these stored tour images. The VLM determines, for example, that a particular image frame of the wrench rack best matches the worker's request.
[0098] In block 356, in response to causing this processing, the robot receives an indication of a goal waypoint in the environment that corresponds to the determined goal image frame. For instance, the central server sends back to the robot the precise coordinates of the wrench rack, which were previously associated with the identified goal image frame.
[0099] In block 358, in response to receiving the indication of the goal waypoint, the robot navigates to that goal waypoint. As shown in sub-block 358A, this navigation can involve identifying a "goal vertex" in a topological graph based on the received goal waypoint indication. This topological graph contains a network of interconnected waypoints, each linked to a prior tour image frame or derived image features. The robot then utilizes this goal vertex, along with its internal localization and navigation systems, to calculate a path and move autonomously to the wrench rack.
[00100] FIG. 4 is a flowchart illustrating an example method 400 for a robot to respond to a user request, which may involve rendering a response instead of navigating to a location. For example, consider a scenario where a robot is in a museum, and a user asks about a specific exhibit.
[00101] In block 452, a request is received that is directed to a robot in an environment, such as the museum. This request can include natural language, as shown in block 452A, and/or one or more images, as shown in block 452B. For instance, a user might point the robot's camera at a painting and verbally ask, "What can you tell me about this?" (natural language and image input).
[00102] In block 454, in response to receiving the request, a vision language model (VLM) is used to process the request and prior tour image frames for the environment to generate VLM output. These prior tour image frames were captured during an earlier detailed tour of the museum, covering all exhibits. The VLM processes the user's question and the image of the painting alongside these stored tour images.
[00103] In block 456, based on the VLM output, a goal image frame is determined from the prior tour image frames that is responsive to the request. In this example, the VLM identifies a specific prior tour image frame of the painting the user is pointing to as the goal image frame.
[00104] In block 458, based on processing the identified goal image frame and the user's request, a response to the request is determined. For example, the VLM further processes the goal image (the painting) and the question ("What can you tell me about this?") to generate a descriptive text about the painting's history, artist, and significance.
[00105] In block 460, the robot is caused to render this determined response responsive to the user's request. As shown in block 460A, this could involve determining to cause the robot to render the response, such as audibly reciting the painting's information through its speakers, in lieu of causing the robot to navigate to a goal waypoint that corresponds to the goal image frame. This illustrates that the system can decide the most appropriate response type, which may not always be navigation.
[00106] FIG. 5 schematically depicts an example architecture of a robot 520. The robot 520 includes a robot control system 560, one or more operational components 540a- 540n, and one or more sensors 542a-542m. The sensors 542a-542m may include, for example, vision sensors, light sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, and so forth. While sensors 542a-m are depicted as being integral with robot 520, this is not meant to be limiting. In some implementations, sensors 542a-m may be located external to robot 520, e.g., as standalone units.
[00107] Operational components 540a-540n may include, for example, one or more end effectors and/or one or more servo motors or other actuators to effectuate movement of one or more components of the robot. For example, the robot 520 may have multiple degrees of freedom and each of the actuators may control the actuation of the robot 520 within one or more of the degrees of freedom responsive to the control commands. As used herein, the term actuator encompasses a mechanical or electrical device that creates motion (e.g., a motor), in addition to any driver(s) that may be associated with the actuator and that translate received control commands into one or more signals for driving the actuator. Accordingly, providing a control command to an actuator may comprise providing the control command to a driver that translates the control command into appropriate signals for driving an electrical or mechanical device to create desired motion.
[00108] The robot control system 560 may be implemented in one or more processors, such as a CPU, GPU, and/or other controller(s) of the robot 520. In some implementations, the robot 520 may comprise a "brain box" that may include all or aspects of the control system 560. For example, the brain box may provide real time bursts of data
to the operational components 540a-n, with each of the real time bursts comprising a set of one or more control commands that dictate, inter alia, the parameters of motion (if any) for each of one or more of the operational components 540a-n. In some implementations, the robot control system 560 may perform one or more aspects of method(s) described herein. [00109] As described herein, in some implementations all or aspects of the control commands generated by control system 560, in controlling a robot during performance of a robotic task, can be generated based on considering different poses and/or different time stamp image(s) generated according to techniques described herein. Although control system 560 is illustrated in FIG. 5 as an integral part of the robot 520, in some implementations, all or aspects of the control system 560 may be implemented in a component that is separate from, but in communication with, robot 520. For example, all or aspects of control system 560 may be implemented on one or more computing devices that are in wired and/or wireless communication with the robot 520, such as computing device 610.
[00110] FIG. 6 is a block diagram of an example computer system 610. Computer system 610 typically includes at least one processor 614 which communicates with a number of peripheral devices via bus subsystem 612. These peripheral devices may include a storage subsystem 624, including, for example, a memory subsystem 625 and a file storage subsystem 626, user interface output devices 620, user interface input devices 622, and a network interface subsystem 616. The input and output devices allow user interaction with computer system 610. Network interface subsystem 616 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
[00111] User interface input devices 622 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and/or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computer system 610 or onto a communication network.
[00112] User interface output devices 620 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from computer system 610 to the user or to another machine or computer system.
[00113] Storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include the logic to perform selected aspects of method 200, method 300, method 400 and/or to implement one or more aspects of robot 500. Memory 625 used in the storage subsystem 624 can include a number of memories including a main randomaccess memory (RAM) 630 for storage of instructions and data during program execution and a read only memory (ROM) 632 in which fixed instructions are stored. A file storage subsystem 626 can provide persistent storage for program and data files, and may include a hard disk drive, a CD-ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 626 in the storage subsystem 624, or in other machines accessible by the processor(s) 614.
[00114] Bus subsystem 612 provides a mechanism for letting the various components and subsystems of computer system 610 communicate with each other as intended. Although bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[00115] Computer system 610 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 610 depicted in FIG. 6 is intended only as a specific example for purposes of
illustrating some implementations. Many other configurations of computer system 610 are possible having more or fewer components than the computer system depicted in FIG. 6.
[00116] In various implementations, a VLM utilized herein can be a sequence-to- sequence based machine learning models capable of processing vision data and textual data and generating generative vision data, generative audio data, generative textual data, and/or other forms of generative data. Some non-limiting examples of these sequence-to- sequence based machine learning models capable that are capable of generating one or more forms of the generative data noted above include transformer-based machine learning models (e.g., encoder-decoder transformer models, encoder-only transformer models, decoder-only transformer models, etc. that optionally employ an attention mechanism or some other form of memory), diffusion-based machine learning models, recurrent neural network-based machine learning models, generative adversarial networkbased machine learning models, etc. Some particular non-limiting examples of these sequence-to-sequence based machine learning models include the Gemini family of models (e.g., Gemini Pro 1.5).
[00117] In some implementations, a method implemented by processor(s) is provided and includes receiving a request that is directed to at least one robot in an environment. The request includes natural language and/or one or more images. The method further includes, in response to receiving the request, processing, using a vision language model (VLM), the request and prior tour image frames for the environment, to generate VLM output. The prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment. Processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment. The method further includes determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request. The method further includes, in response to determining the goal image frame, causing the at least one robot to navigate,
in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
[00118] These and other implementations of the technology disclosed herein can include one or more of the following features.
[00119] In some implementations, causing the at least one robot to navigate to the goal waypoint includes determining, based on the goal image frame, the goal waypoint that corresponds to the goal image frame and transmitting, to the at least one robot, the goal waypoint.
[00120] In some implementations, causing the at least one robot to navigate to the goal waypoint includes transmitting, to the at least one robot, the goal image frame. In some of those implementations the robot determines, based on the transmitted goal image frame, the goal waypoint that corresponds to the goal image frame.
[00121] In some implementations, the goal waypoint corresponds to the goal image frame based on a prior determination that the goal image frame was captured from the goal waypoint. In some of those implementations, the prior determination that the goal image frame was captured from the goal waypoint can be based on performing structure from motion based on the prior tour image frames. In some versions of those implementations, the prior tour image frames are captured by one or more cameras of the at least one robot. [00122] In some implementations, the at least one robot, in navigating to the goal waypoint, utilizes a topological graph, such as a topological graph that includes vertexes that are each for a corresponding waypoint and that are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames. In some of those implementations, the at least one robot, in navigating to the goal waypoint, utilizes structure-from-motion based on the topological graph. In some additional or alternative implementations, the topological graph is created utilizing structure-from-motion and based on the tour image frames.
[00123] In some implementations, the at least one robot, in navigating to the goal waypoint, utilizes a topological graph that includes vertexes that are each for a
corresponding waypoint and that are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames. In some of those implementations, causing the at least one robot to navigate to the goal waypoint includes transmitting, to the at least one robot, data that specifies a vertex, of the vertexes, that corresponds to the goal image frame. For example, the data could specify an index value for the vertex.
[00124] In some implementations, the method further includes processing, using the VLM and along with the request and the prior tour image frames, a corresponding natural language representation of each of the prior tour image frames. In some versions of those implementations, determining, based on the VLM output, the goal image frame that is responsive to the request includes determining, based on the VLM output, a given one of the corresponding natural language representations and determining the goal image frame based on the given one of the corresponding natural language representations being for the goal image frame. In some of those or other versions, the method further includes processing, using the VLM and along with the request, the prior tour image frames, and the corresponding natural language representations, instructions to output the corresponding natural language representation for one of the prior tour image frames that matches the request.
[00125] In some implementations, the method further includes processing, using the VLM and along with the request and the prior tour image frames, a corresponding narrative language description for one or more of the prior tour image frames. The corresponding narrative language descriptions can be human provided, are captured prior to receiving the request, and are each associated with a corresponding one of the prior tour image frames. For example, a subset of the prior tour image frames can each be associated with a respective corresponding narrative language description. For instance, a given prior tour image frame can have an associated narrative language description based on a human providing the narrative language description at or near a time that the given prior tour image frame was captured, or based on a human later explicitly annotating (e.g., via a user interface) the narrative language description for the given prior tour image frame.
[00126] In some implementations, the prior tour image frames are captured by one or more cameras of the at least one robot.
[00127] In some implementations, the request is a user request. In some of the user request implementations, the request includes the natural language and the natural language is based on spoken or typed input provided by a user. For example, the user request can include audio data that captures a spoken utterance of the user and the audio data can be processed utilizing the VLM and/or a transcription, generated based on the audio data, can be generated and processed utilizing the VLM. In some of the user request implementations, the request includes the one or more images. For example, the one or more images can capture a user, can capture a drawing made by a user, or can capture an automatically generated image generated by an image generation model based on input(s) of a user (e.g., based on natural language input of a user). In some of the user request implementations, the user request is captured via one or more input devices of the at least one robot, such as one or more microphones and/or one or more cameras.
[00128] In some implementations, the request is transmitted by the at least one robot in the environment and is received via one or more networks. In some of those implementations, one or more of the processors are in one or more servers that are not in the environment.
[00129] In some implementations, a method implemented by processor(s) is provided and includes capturing a user request via one or more input devices of a robot in an environment. The user request includes natural language provided by a human and/or one or more images, such as current image(s) captured in the environment. The method further includes, in response to capturing the user request, causing the user request and prior tour image frames for the environment to be processed, using a vision language model (VLM), in determining a goal image frame, of the prior tour image frames, that is responsive to the request. The prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment. The method further includes receiving, in response to the causing, an indication of a goal waypoint, in the environment, that corresponds to the determined goal image frame. The method further includes, in
response to receiving the indication of the goal waypoint, navigating the robot to the goal waypoint.
[00130] These and other implementations of the technology disclosed herein can include one or more of the following features.
[00131] In some implementations, navigating the robot to the goal waypoint includes identifying, based on the indication of the goal waypoint and in a topological graph that includes a plurality of vertexes, a goal vertex that corresponds to the goal waypoint. For example, the indication of the goal waypoint can directly or indirectly identify the goal vertex. In some versions of those implementations, each of the vertexes of the topological graph are for a corresponding waypoint and are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames. In some additional or alternative versions of those implementations, navigating the robot to the goal waypoint includes using structure-from-motion that is based on the topological graph and that utilizes the goal vertex as a goal. In some further additional or alternative versions of those implementations, navigating the robot to the goal waypoint includes identifying a sequence of the vertexes that are connected in the topological graph and that define a path between a current waypoint of the robot and the goal waypoint, and utilizing the sequence of the vertexes in navigating to the goal waypoint. In some yet further additional or alternative versions of those implementations, the goal waypoint corresponds to the goal image frame based on a prior determination that the goal image frame was captured from the goal waypoint.
[00132] In some implementations, a corresponding natural language representation of each of the prior tour image frames is processed, using the VLM and along with the user request and the prior tour image frames, in determining the goal image. In some of those implementations, instructions, to output the corresponding natural language representation for one of the prior tour image frames that matches the user request, are processed, using the VLM and along with the user request, the prior tour image frames, and the corresponding natural language representations, in determining the goal image.
[00133] In some implementations, a corresponding narrative language description for one or more of the prior tour image frames is processed, using the VLM and along with the user request and the prior tour image frames, in determining the goal image. The corresponding narrative language descriptions are human provided, are captured prior to receiving the request, and are each associated with a corresponding one of the prior tour image frames.
[00134] In some implementations, the prior tour image frames are captured by one or more cameras of the robot.
[00135] In some implementations, the user request includes the natural language and the natural language is based on spoken or typed input provided by the human.
[00136] In some implementations, the user request includes the one or more images.
[00137] In some implementations, the one or more input devices include one or more microphones and/or one or more cameras.
[00138] In some implementations, causing the user request and prior tour image frames for the environment to be processed, using the VLM, in determining the goal image that is responsive to the request includes transmitting the user request from the robot to one or more remote servers. For example, the request can be transmitted over one or more networks to one or more edge or cloud servers and optionally utilizing an application programming interface (API). In some of those implementations, causing the user request and prior tour image frames for the environment to be processed, using the VLM, in determining the goal image that is responsive to the request includes transmitting, to the one or more servers and with the user request, an indication of the robot and/or of the environment. The one or more servers utilize the indication of the robot and/or the environment in identifying the prior tour image frames. For example, a first request from a first robot in a first environment can be provided a first identifier that is used by the server(s) to identify first prior tour image frames, from a prior tour of the first environment, for processing along with the first request using the VLM. A second request from a second robot in a second environment can be provided with a second identifier that is used by the
server(s) to identify second prior tour image frames, for a prior tour of the second environment, for processing along with the second request using the VLM.
[00139] In some implementations, one or more of the processors are integrated as part of the robot.
[00140] In some implementations, a method implemented by processor(s) is provided and includes receiving a request that is directed to a robot in an environment and that includes natural language and/or one or more images. The method further includes, in response to receiving the request, processing, using a vision language model (VLM), the request and prior tour image frames for the environment, to generate VLM output. The prior tour image frames for the environment are captured, prior to the request, throughout at least a portion of the environment. Processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the robot in the environment. The method further includes determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request. The method further includes, in response to determining the goal image frame, navigating the robot, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
[00141] In some implementations, a method implemented by processor(s) is provided and includes receiving a user request that is directed to a robot in an environment and that includes natural language and/or one or more images. The method further includes, in response to receiving the user request, processing, using a vision language model (VLM), the user request and prior tour image frames for the environment, to generate VLM output. The prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment. Processing the prior tour image frames for the environment, using the VLM and along with the user request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment. The method further includes determining, based on the VLM output, a goal image frame, of the prior
tour image frames, that is responsive to the user request. The method further includes determining, based on processing the goal image frame and the user request, a response to the user request. The method further includes causing the robot to render the response responsive to the user request.
[00142] These and other implementations of the technology disclosed herein can include one or more of the following features.
[00143] In some implementations, the user request includes the natural language and/or the response to the user request is a natural language response that is audibly rendered via one or more speakers of the robot.
[00144] In some implementations, determining, based on processing the goal image frame and the user request, the response includes processing the goal image frame and the user request using the VLM or an additional VLM.
[00145] In some implementations, determining, based on processing the goal image frame and the user request, the response includes determining to cause the robot to render the response responsive to the user request in lieu of causing, responsive to the user request, the robot to navigate to a goal waypoint that corresponds to the goal image frame. [00146] Some implementations includes a system having memory storing instructions and one or more processors (e.g., graphic processing unit(s), tensor processing unit(s), and/or central processing unit(s)) operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
[00147] Some implementations include at least one transitory or non-transitory computer-readable medium including instructions that, in response to execution by one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
[00148] Some implementations includes a robot having actuators, vision component(s) (e.g., RGB camera, RGB-D camera, LIDAR, and/or other vision component(s)) memory storing instructions, and one or more processors operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
[00149] Some implementations includes a system having one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
[00150] Some implementations includes a system having memory storing instructions and one or more processors operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
[00151] Some implementations include at least one transitory or non-transitory computer-readable medium including instructions that, in response to execution by one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
Claims
1. A method implemented by one or more processors, the method comprising: receiving a request that is directed to at least one robot in an environment, the request including natural language and/or one or more images; in response to receiving the request: processing, using a vision language model (VLM), the request and prior tour image frames for the environment, to generate VLM output, wherein the prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment, and wherein processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment; determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request; and in response to determining the goal image frame: causing the at least one robot to navigate, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
2. The method of claim 1, wherein causing the at least one robot to navigate to the goal waypoint comprises: determining, based on the goal image frame, the goal waypoint that corresponds to the goal image frame; and transmitting, to the at least one robot, the goal waypoint.
3. The method of claim 1, wherein causing the at least one robot to navigate to the goal waypoint comprises: transmitting, to the at least one robot, the goal image frame; wherein the robot determines, based on the transmitted goal image frame, the goal waypoint that corresponds to the goal image frame.
4. The method of any preceding claim, wherein the goal waypoint corresponds to the goal image frame based on a prior determination that the goal image frame was captured from the goal waypoint.
5. The method of any preceding claim, wherein the at least one robot, in navigating to the goal waypoint: utilizes a topological graph that includes vertexes that are each for a corresponding waypoint and that are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames.
6. The method of claim 5, wherein the at least one robot, in navigating to the goal waypoint, utilizes structure-from-motion based on the topological graph.
7. The method of claim 1, wherein the at least one robot, in navigating to the goal waypoint: utilizes a topological graph that includes vertexes that are each for a corresponding waypoint and that are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames; and wherein causing the at least one robot to navigate to the goal waypoint comprises: transmitting, to the at least one robot, data that specifies a vertex, of the vertexes, that corresponds to the goal image frame.
8. The method of any preceding claim further comprising processing, using the VLM and along with the request and the prior tour image frames, a corresponding natural language representation of each of the prior tour image frames.
9. The method of claim 8, wherein determining, based on the VLM output, the goal image frame that is responsive to the request comprises: determining, based on the VLM output, a given one of the corresponding natural language representations; and determining the goal image frame based on the given one of the corresponding natural language representations being for the goal image frame.
10. The method of claim 8 or claim 9, further comprising processing, using the VLM and along with the request, the prior tour image frames, and the corresponding natural language representations, instructions to output the corresponding natural language representation for one of the prior tour image frames that matches the request.
11. The method of any preceding claim further comprising processing, using the VLM and along with the request and the prior tour image frames, a corresponding narrative language description for one or more of the prior tour image frames, wherein the corresponding narrative language descriptions are human provided, are captured prior to receiving the request, and are each associated with a corresponding one of the prior tour image frames.
12. The method of any preceding claim, wherein the prior tour image frames are captured by one or more cameras of the at least one robot.
13. The method of any preceding claim, wherein the request is a user request.
14. The method of claim 13, wherein the request includes the natural language and the natural language is based on spoken or typed input provided by a user.
15. The method of claim 14, wherein the request includes the one or more images and wherein the one or more images capture a user.
16. The method of any one of claims 13-15, wherein the user request is captured via one or more input devices of the at least one robot.
17. The method of claim 16, wherein the one or more input devices include one or more microphones and/or one or more cameras.
18. The method of any preceding claim, wherein the request is transmitted by the at least one robot in the environment and is received via one or more networks.
19. The method of claim 18, wherein one or more of the processors are in one or more servers that are not in the environment.
20. A method implemented by one or more processors, the method comprising: capturing a user request via one or more input devices of a robot in an environment, wherein the user request includes natural language provided by a human and/or one or more images of the human; in response to capturing the user request: causing the user request and prior tour image frames for the environment to be processed, using a vision language model (VLM), in determining a goal image frame, of the prior tour image frames, that is responsive to the request, wherein the prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment; receiving, in response to the causing, an indication of a goal waypoint, in the environment, that corresponds to the determined goal image frame; and in response to receiving the indication of the goal waypoint: navigating the robot to the goal waypoint.
21. The method of claim 20, wherein navigating the robot to the goal waypoint comprises:
identifying, based on the indication of the goal waypoint and in a topological graph that includes a plurality of vertexes, a goal vertex that corresponds to the goal waypoint, wherein each of the vertexes are for a corresponding waypoint and are each mapped to a corresponding one of the prior tour image frames and/or to corresponding image features derived from a corresponding one of the prior tour image frames.
22. The method of claim 21, wherein navigating the robot to the goal waypoint comprises using structure-from-motion that is based on the topological graph and that utilizes the goal vertex as a goal.
23. The method of claim 21 or claim 22, wherein navigating the robot to the goal waypoint comprises identifying a sequence of the vertexes that are connected in the topological graph and that define a path between a current waypoint of the robot and the goal waypoint, and utilizing the sequence of the vertexes in navigating to the goal waypoint.
24. The method of any one of claims 21 to 23, wherein the goal waypoint corresponds to the goal image frame based on a prior determination that the goal image frame was captured from the goal waypoint.
25. The method of any one of claims 20 to 24, wherein a corresponding natural language representation of each of the prior tour image frames is processed, using the VLM and along with the user request and the prior tour image frames, in determining the goal image.
26. The method of claim 25, wherein instructions, to output the corresponding natural language representation for one of the prior tour image frames that matches the user request, are processed, using the VLM and along with the user request, the prior tour image frames, and the corresponding natural language representations, in determining the goal image.
27. The method of any one of claims 20 to 26, wherein a corresponding narrative language description for one or more of the prior tour image frames is processed, using the VLM and along with the user request and the prior tour image frames, in determining the
goal image, wherein the corresponding narrative language descriptions are human provided, are captured prior to receiving the request, and are each associated with a corresponding one of the prior tour image frames.
28. The method of any one of claims 20 to 27, wherein the prior tour image frames are captured by one or more cameras of the robot.
29. The method of any one of claims 20 to 28, wherein the user request includes the natural language and the natural language is based on spoken or typed input provided by the human.
30. The method of any one of claims 20 to 29, wherein the user request includes the one or more images.
31. The method of any one of claims 20 to 30, wherein the one or more input devices include one or more microphones and/or one or more cameras.
32. The method of any one of claims 20 to 31, wherein causing the user request and prior tour image frames for the environment to be processed, using the VLM, in determining the goal image that is responsive to the request comprises transmitting the user request from the robot to one or more remote servers.
33. The method of claim 32, wherein causing the user request and prior tour image frames for the environment to be processed, using the VLM, in determining the goal image that is responsive to the request comprises transmitting, to the one or more servers and with the user request, an indication of the robot and/or of the environment, wherein the one or more servers utilize the indication of the robot and/or the environment in identifying the prior tour image frames.
34. The method of any one of claims 20 to 33, wherein one or more of the processors are integrated as part of the robot.
35. A method implemented by one or more processors, the method comprising:
receiving a request that is directed to a robot in an environment, the request including natural language and/or one or more images; in response to receiving the request: processing, using a vision language model (VLM), the request and prior tour image frames for the environment, to generate VLM output, wherein the prior tour image frames for the environment are captured, prior to the request, throughout at least a portion of the environment, and wherein processing the prior tour image frames for the environment, using the VLM and along with the request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the robot in the environment; determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the request; and in response to determining the goal image frame: navigating the robot, in the environment, to a goal waypoint, in the environment, that corresponds to the goal image frame.
36. A method implemented by one or more processors, the method comprising: receiving a user request that is directed to a robot in an environment, the user request including natural language and/or one or more images; in response to receiving the user request: processing, using a vision language model (VLM), the user request and prior tour image frames for the environment, to generate VLM output,
wherein the prior tour image frames for the environment are captured, prior to receiving the request, throughout at least a portion of the environment, and wherein processing the prior tour image frames for the environment, using the VLM and along with the user request, is based on the prior tour image frames being captured throughout at least the portion of the environment and the request being directed to the at least one robot in the environment; determining, based on the VLM output, a goal image frame, of the prior tour image frames, that is responsive to the user request; determining, based on processing the goal image frame and the user request, a response to the user request; and causing the robot to render the response responsive to the user request.
37. The method of claim 36, wherein the user request includes the natural language.
38. The method of claim 36 or 37, wherein the response to the user request is a natural language response that is audibly rendered via one or more speakers of the robot.
39. The method of any one of claims 36 to 38, wherein determining, based on processing the goal image frame and the user request, the response comprises processing the goal image frame and the user request using the VLM or an additional VLM.
40. The method of any one of claims 36 to 39, wherein determining, based on processing the goal image frame and the user request, the response comprises determining to cause the robot to render the response responsive to the user request in lieu of causing, responsive to the user request, the robot to navigate to a goal waypoint that corresponds to the goal image frame.
41. A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 40.
42. At least one non-transitory computer-readable medium comprising instructions that, in response to execution by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 40.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463657088P | 2024-06-06 | 2024-06-06 | |
| US63/657,088 | 2024-06-06 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025255236A1 true WO2025255236A1 (en) | 2025-12-11 |
Family
ID=96141208
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/032275 Pending WO2025255236A1 (en) | 2024-06-06 | 2025-06-04 | Robot navigation utilizing visual language model |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025255236A1 (en) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220343531A1 (en) * | 2021-04-27 | 2022-10-27 | Here Global B.V. | Systems and methods for synchronizing an image sensor |
| WO2024059179A1 (en) * | 2022-09-15 | 2024-03-21 | Google Llc | Robot control based on natural language instructions and on descriptors of objects that are present in the environment of the robot |
-
2025
- 2025-06-04 WO PCT/US2025/032275 patent/WO2025255236A1/en active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220343531A1 (en) * | 2021-04-27 | 2022-10-27 | Here Global B.V. | Systems and methods for synchronizing an image sensor |
| WO2024059179A1 (en) * | 2022-09-15 | 2024-03-21 | Google Llc | Robot control based on natural language instructions and on descriptors of objects that are present in the environment of the robot |
Non-Patent Citations (1)
| Title |
|---|
| SOURAV GARG ET AL: "RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 9 May 2024 (2024-05-09), XP091752195 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Chiang et al. | Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs | |
| US11397462B2 (en) | Real-time human-machine collaboration using big data driven augmented reality technologies | |
| KR102255273B1 (en) | Apparatus and method for generating map data of cleaning space | |
| Brooks | The intelligent room project | |
| Randelli et al. | Knowledge acquisition through human–robot multimodal interaction | |
| Xu et al. | Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs | |
| US20190317594A1 (en) | System and method for detecting human gaze and gesture in unconstrained environments | |
| US9744668B1 (en) | Spatiotemporal robot reservation systems and method | |
| Pavón-Pulido et al. | Cybi: A smart companion robot for elderly people: Improving teleoperation and telepresence skills by combining cloud computing technologies and fuzzy logic | |
| CN119414833A (en) | A robot path planning perception method and system based on large language model | |
| Tan et al. | Embodied scene description | |
| Roy et al. | doscenes: An autonomous driving dataset with natural language instruction for human interaction and vision-language navigation | |
| US11577396B1 (en) | Visual annotations in robot control interfaces | |
| Carroll et al. | Human-computer synergies in prosthetic interactions | |
| Rahman et al. | OpenNav: efficient open vocabulary 3D object detection for smart wheelchair navigation | |
| Banerjee et al. | Teledrive: an embodied AI based telepresence system | |
| WO2025255236A1 (en) | Robot navigation utilizing visual language model | |
| Marchionni et al. | Reem service robot: how may i help you? | |
| WO2025126663A1 (en) | Information processing device, information processing method, and program | |
| Fearn et al. | Wheelchair navigation: Automatically adapting to evolving environments | |
| Zhou et al. | MultiMap3D: A Multi-Level Semantic Perceptual Map Construction Based on SLAM and Point Cloud Detection | |
| WO2023195982A1 (en) | Keyframe downsampling for memory usage reduction in slam | |
| Kumar et al. | Sharing cognition: Human gesture and natural language grounding based planning and navigation for indoor robots | |
| Flor-Rodríguez et al. | SEMNAV: A Semantic Segmentation-Driven Approach to Visual Semantic Navigation | |
| CN108536830A (en) | Picture dynamic searching method, device, equipment, server and storage medium |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25733303 Country of ref document: EP Kind code of ref document: A1 |