EP4676696A1 - Visual prompting for few shot control - Google Patents

Visual prompting for few shot control

Info

Publication number
EP4676696A1
EP4676696A1 EP25708274.3A EP25708274A EP4676696A1 EP 4676696 A1 EP4676696 A1 EP 4676696A1 EP 25708274 A EP25708274 A EP 25708274A EP 4676696 A1 EP4676696 A1 EP 4676696A1
Authority
EP
European Patent Office
Prior art keywords
vlm
robot
candidate
distribution
visual representation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP25708274.3A
Other languages
German (de)
French (fr)
Inventor
Brian ICHTER
Ted Xiao
Wenhao YU
Fei XIA
Soroush Nasiriany
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
GDM Holding LLC
Original Assignee
GDM Holding LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by GDM Holding LLC filed Critical GDM Holding LLC
Publication of EP4676696A1 publication Critical patent/EP4676696A1/en
Pending legal-status Critical Current

Links

Classifications

    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • B25J9/1628Program controls characterised by the control loop
    • B25J9/163Program controls characterised by the control loop learning, adaptive, model based, rule based expert control
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • B25J9/1656Program controls characterised by programming, planning systems for manipulators

Definitions

  • VLMs vision-language models
  • the VLM input may include vision data from one or more vision sensors and natural language textual input identifying a task to be completed by a robot.
  • VLM output in the robotic context may include, for instance, natural language statements that describe steps for completing some task, robot control data that is usable to control a robot to complete the task, etc.
  • VLMs often require large amounts of training data to adequately support robot performance, and even still may not be sufficiently generalizable.
  • Implementations are described herein for framing tasks that involve performing actions within a two-dimensional (2D) or three-dimensional (3D) space as visual problems that can be posed to a VLM, e.g., iteratively or otherwise. More particularly, but not exclusively, techniques described herein relate to visually prompting a VLM or another multimodal generative model for few-shot and/or zero-shot control of actions performed in a 2D or 3D space. In implementations where these actions are performable in furtherance of completing a higher level task, the VLM may be prompted iteratively, with candidate sample candidate actions iteratively closing in on completion of the higher level task.
  • GUI graphical user interface
  • CAD computer-aided design
  • FIG. 1 schematically depicts an example environment in which disclosed techniques may be employed, in accordance with various implementations.
  • FIG. 2 depicts an example robot, in accordance with various implementations.
  • FIG. 3 A and Fig. 3B depict how techniques described herein may be applied for two iterations, in accordance with various implementations.
  • FIG. 4 schematically depicts an example of how components of Fig. 1 may cooperate to carry out selected aspects of the present disclosure, in accordance with various implementations.
  • Fig. 5 A and Fig. 5B depict an example of how techniques described herein may be implemented.
  • FIG. 6 depicts an example method for practicing selected aspects of the present disclosure.
  • FIG. 7 schematically depicts an example architecture of a computer system.
  • Implementations are described herein for framing tasks that involve performing actions within a two-dimensional (2D) or three-dimensional (3D) space as visual problems that can be posed to a VLM, e.g., iteratively or otherwise. More particularly, but not exclusively, techniques described herein relate to visually prompting a VLM or another multimodal generative model for few-shot and/or zero-shot control of actions performed in a 2D or 3D space. In implementations where these actions are performable in furtherance of completing a higher level task, the VLM may be prompted iteratively, with candidate sample candidate actions iteratively closing in on completion of the higher level task.
  • GUI graphical user interface
  • objects around a 2D or 3D physical space e.g., loading a truck, adding objects to or removing objects from an inventory in a warehouse, navigating an autonomous vehicle, etc.
  • a E A can, for example, include continuous coordinates, 3D spatial locations, robot control actions, or trajectories.
  • A is the set of robot actions, this may amount to finding a policy TT(-
  • VLMs visual question answering
  • the class of VLMs used in various implementations take as input an image I and a textual prefix wp from which they generate a distribution P VLM ' ⁇ TM P > I) of textual completions. Utilizing this interface to drive a policy raises the challenge of how an action from a (continuous) space A can be represented as a textual completion.
  • various implementations described herein lift low-level actions into the visual language of a VLM, i.e., a combination of images and text, such that it is closer to the training distribution of general vision-language tasks.
  • the following visual prompt mapping may be used: that transforms an image observation I and a set of candidate actions a M , a.j E A into an annotation image I and their corresponding textual labels w 1.M where Wj refers to the annotation representing a.j in the image space.
  • the camera matrices can be used to project a 3D location into the image space, and draw a visual marker at this projected location.
  • Labeling this marker with a textual reference e.g., a number, enables the VLM to not only be queried in its natural input space, namely images and text, but also to refer to spatial concepts in its natural output space by producing text that references the marker labels.
  • Techniques described herein may then be used to prompt a VLM with iterative visual optimization. Representing (continuous) robot actions and spatial concepts in image space with their associated textual labels the VLM VLM to be queried to judge if an action would be promising in solving the task. Therefore, it is possible to obtain a policy n by solving the following optimization problem:
  • a goal may be to find an action a for which the VLM would choose the corresponding label w after applying the mapping .
  • an iterative technique referred to herein as Prompting with Iterative Visual Optimization may be used.
  • the technique first samples a set of candidate actions a r L M from a distribution P A ⁇ . These candidate actions may then be mapped onto the image I producing the annotated image /W and the associated action labels w ⁇ M .
  • the VLM may then be queried on a multiple choice-style question on the labels w ⁇ M to choose which of the candidate actions are most promising.
  • actions may be represented in some implementations as arrows emanating from the robot or the center of the image, e.g., if the embodiment/robot is not visible.
  • the colors or other visual aspects of the arrows may be selected to indicate forward and backwards movement.
  • These actions may be labeled in some cases with a number label circled at the end of the arrow.
  • the VLM may be prompted to use chain-of-thought to reason through the problem and then summarize the top few labels.
  • the distributions PA in may be approximated as Gaussians.
  • parallel calls may be implemented, e.g., via ensembling, to provide a more robust experience.
  • VLMs can make mistakes, which may result in actions being selected from sub-optimal regions.
  • a parallel call strategy may be used in which E parallel instances of the VLM are called to obtain E candidate actions. The selected candidates actions may then be aggregated to identify the final action output.
  • two different approaches may be taken: 1) a new action distribution from the E action candidates may be fitted and returned, 2) the VLM may be queries again to select the single best action from the E actions.
  • Adopting parallel calls may improve the robustness of techniques described herein and mitigate local minima in the optimization process.
  • a digital visual representation of an environment in which actions are to be performed may be provided.
  • this digital visual representation may include, for instance, a digital image captured by one or more sensors onboard the robot or in the robot’s environment, a 3D point cloud captured by one or more light detection and ranging (LIDAR) sensors, a CAD drawing of a robot’s environment, a 2D or 3D snapshot of a virtual environment in which a simulated robot operates, a GUI screenshot, etc.
  • LIDAR light detection and ranging
  • a distribution of candidate actions that can be performed in a space or environment may be sampled.
  • a set of candidate actions that are performable by the robot given its current pose and/or a state of the environment may be sampled from an action space, e.g., randomly, using techniques such as cross entropy and/or particle filter optimization, and/or subject to various constraints such as a state of the environment, a natural language command from a user, an object detected in the digital visual representation, etc.
  • the action space from which the candidate actions are sampled may be constrained based on the digital visual representation of the environment.
  • the action space may be constrained to those actions that are performable by the robot within the portion of the environment that is depicted in the digital image.
  • at least an initial distribution of candidate actions may be sampled so they are evenly spaced across the entire visual representation. Iterations of candidate action distributions that are sampled subsequently may gradually approach a desired objective (e.g., an object to be acted upon by the robot), e.g., by the search spaces those subsequent distributions are sampling from gradually contracting.
  • the distribution of candidate actions may be selected to cluster about that detected object, and/or to avoid the downstream visual annotations from obscuring the object.
  • visual annotations that represent the candidate actions may be incorporated into the digital visual representation. If the digital visual representation is a digital image captured by a digital camera, then the visual annotations may include, for instance, various graphics primitives, shapes, lines, text, labels, numbers, etc. that are overlaid on top of or otherwise infused into pixels of the digital image. These visual annotations may or may not be visible to a person viewing the digital image using a digital image viewing/editing application.
  • the visual annotations may be at least partially transparent or translucent, e.g., so that objects in the environment remain at least partially visible in the digital image.
  • two copies of the digital image may be maintained: the original raw image (sans annotations) and an annotated image in which the annotations may or may not obscure underlying visual details.
  • Visual annotations may take various forms depending on a variety of factors, such as the types of candidate actions that are being explored, the capabilities of the actor (e.g., the robot, a user operating a GUI), features of objects depicted in the digital visual representation, and so forth.
  • visual annotations may take the form of, for instance, vector arrows that may or may not be tipped with shapes such as arrow heads, circles, squares, triangles, etc.
  • the visual annotations that represent the end effector trajectories may take the form of vector arrows extending from a starting point in the digital visual representation (e.g., coinciding with the robot’s current location/pose) to candidate end effector positions in the digital visual representation.
  • the visual annotations may be rendered to indicate where in the 3D space they would occur, in spite of the limitations of the 2D digital visual representation.
  • a first shape represents a first candidate action that would occur a first distance from a perspective of, for instance, a vision sensor that captured the digital visual representation and/or the robot itself.
  • a second shape represents a second candidate action that would occur a second distance from the same perspective, e.g., farther away from the vision sensor than the first candidate action.
  • the first shape may be sized differently from the second shape to convey this difference between the first and second distances.
  • the first shape may be rendered larger than the second shape to indicate that the first candidate action would occur closer to the vision sensor that captured the 2D digital visual representation than the second candidate action.
  • the first and second shapes may be colored, shaded, shaped, dashed, or otherwise visually annotated in distinctive ways to convey the difference between their respective distances from the vision sensor.
  • a VLM input prompt may be assembled that includes the annotated digital visual representation (and/or embedding(s) generated therefrom) and a request to select one or more of the candidate actions to perform in furtherance of completing the higher level task.
  • This request may take various forms depending on various features of the visual annotations themselves and/or the candidate actions they represent, as well as features of the higher level task to be completed.
  • each visual annotation may include an identifier (e.g., a number, name, character, etc.) that is unique amongst the visual annotations incorporated into the digital visual representation. These identifiers may overlay various points in the digital visual representation.
  • the higher level task is “pick up the helix-shaped chew toy,” with the objective being the “helix-shaped chew toy.”
  • the request may be something like “which identifiers are closest to and/or overlay the helix-shaped chew toy?”
  • the VLM input prompt may include other data as well.
  • This other data may include, for instance, a desired type of VLM output.
  • the VLM generate VLM output that includes tokens representing natural language describing steps that can be performed, e.g., by the robot, to complete a higher level task.
  • the VLM output contain robot control data that is used to control a robot more directly.
  • Other data that may be included in the VLM input prompt may additionally or alternatively include a raw version of the digital visual representation (e.g., where the visual annotations obscure underlying visual features of the digital visual representation), one or more indications of what the visual annotations represent (e.g., a legend), criteria and/or constraints that should be satisfied when selecting candidate actions (e.g., safety, efficiency, compliance with regulations, etc.), a reiteration of the higher level task, and so on.
  • a legend may indicate, for instance, that vector arrows represent candidate end effector trajectories, another type of shape represents a gripper closing, yet another type of shape represents rotating the end effector or other portion(s) of the robot, a line thickness representing a velocity, and so forth.
  • a visual annotation representing a candidate robot pose may take the form of a depiction of all or part of the robot (e.g., only the end effector, a portion of a robot arm, etc.) in the candidate pose.
  • the VLM input prompt may be processed using a VLM to generate VLM output.
  • the VLM output may include tokens indicative of, for instance, natural language statement(s) or other indications identifying which of the candidate actions would be most beneficial/efficient/safe for the robot to perform in furtherance of completing the higher level task, robot control data, and so forth.
  • the tokens may be indicative of, for instance, where to move and/or insert a component into an existing CAD drawing.
  • the VLM output may be used for various purposes. In some implementations, depending on the nature of the VLM output, it may be used directly (e.g., when in the form of robot control data) or indirectly (e.g., when in the form of natural language) to control a robot and/or to sample new candidate actions from progressively contracting action spaces. Additionally or alternatively, in some implementations, the VLM output may be used to train and/or further train (e.g., fine-tune) the VLM itself. This training may be based on, for instance, feedback (e.g., from a human or from a trained reward model) about the VLM output itself and/or feedback about performance of the selected candidate action(s), e.g., by the robot or by another actor.
  • feedback e.g., from a human or from a trained reward model
  • the process may repeat to leverage the VLM to iteratively predict and/or perform action(s) that collectively result in completion of a higher level task.
  • new candidate actions that correspond to the new context, e.g., of the robot and/or its environment, may be sampled to form a next distribution.
  • this next distribution of candidate actions may be selected from a new action space that is constrained, e.g., relative to prior action space, based on candidate action(s) performed up until now.
  • the new distribution of candidate actions may be fitted to the previously selected/performed candidate action(s).
  • candidate actions for a new distribution may be sampled from an action space that is constrained to an area immediately adjacent the objective, e.g., in the real world or in the digital visual representation.
  • new visual annotations representing the new distribution of candidate actions may be incorporated once again into the same digital visual representation or into an updated digital visual representation, e.g., a new digital image captured of the robot’s environment after the prior action(s) are selected/performed.
  • This new annotated digital visual representation (and/or embedding(s) generated therefrom) may once again be assembled into a VLM prompt, e.g., along with a request that is similar to or different from the previous request depending on the context (e.g., if a robot gripper is now adjacent an objective, the candidate actions and corresponding visual annotations may be considerably different than before).
  • this updated prompt may also include other elements, such as a new legend, one or more new constraints or criteria to be satisfied, etc.
  • the VLM input prompt may once again be processed using the VLM to generate a new iteration of VLM output.
  • the new iteration of VLM output may be used for purposes such as operating a robot to perform a next step in a sequence of steps that collectively result in completion of a higher level task, for training the VLM, etc.
  • a sequence of multiple actions may be iteratively sampled without the robot performing the actions during each iteration, and the robot may subsequently perform all or part of the entire sequence as a batch. For example, after a candidate action is selected from a given distribution of candidate actions, a next distribution of new candidate actions may be sampled based on the selected candidate action, without the robot performing the selected candidate action first. Similar to before, this next distribution of candidate actions may be selected from a new action space that is constrained relative to the prior action space based on candidate action(s) predicted up until now. Put another way, the new distribution of candidate actions may be fitted to the previously-predicted candidate action(s).
  • candidate actions for a new distribution may be sampled from an action space that is constrained to an area immediately adjacent the particular position in the digital visual representation.
  • techniques described herein may be used to train and/or further train (e.g., fine-tune) a pretrained VLM, e.g., using techniques such as reinforcement learning with human feedback (RLHF), to better predict which visual annotation(s) represent the best action(s).
  • RLHF reinforcement learning with human feedback
  • a real or simulated robot could be teleoperated by a human.
  • the human may be provided with a view (on a display) that renders a feed of the robot’s own onboard vision sensor(s)/cameras and/or a separate camera with a view of the robot and/or its environment.
  • the human may provide an indication (e.g., using natural language) of a task they are going to perform by teleoperating the robot, or a task may be assigned to the user to perform by teleoperating the robot.
  • techniques described herein may be used to sample actions from action space and render annotations on the user’s display representing projections of these actions.
  • the same rendition presented to the user (image plus annotations) in one or more video frames may be processed as described herein using the VLM to determine which annotation(s) represent the best actions to take currently.
  • the selected annotation(s) may be suggested/recommended to the user, e.g., using visual emphasis such as highlighting, animation, etc.
  • the selected annotations may not be suggested/recommended to the user, e.g., so that the actions the user takes are not influenced by the annotations.
  • the user may teleoperate the robot to perform the task, consciously following/ignoring the recommended annotations or without seeing the annotations at all.
  • a recommended annotation that may be positive feedback.
  • a recommended annotation that may be negative feedback.
  • the user who is not presented with annotations nonetheless teleoperates the robot in a manner that is consistent with recommended annotation(s), that may be positive feedback.
  • the user who is not presented with annotations teleoperates the robot in a manner that diverges from the recommended annotations, that may be negative feedback.
  • This human feedback may be used in some implementations to train a reward model, so that the reward model learns to predict humans’ preferences and assigns rewards to the VLM’s predictions of which annotations represent the best actions based on these predictions.
  • the VLM may then be optimized/fine-tuned using the reward model, e.g., with the goal to maximize the rewards predicted by the reward model, which in turn reflects the human evaluators' preferences.
  • Various techniques may be used to optimize the VLM, including but not limited to Proximal Policy Optimization (PPO), Trust Region Policy Optimization (TRPO), gradient descent, etc.
  • PPO Proximal Policy Optimization
  • TRPO Trust Region Policy Optimization
  • gradient descent etc.
  • conventional robot control policies may be well-suited for the specific tasks for which they were trained, but they are not generally scalable to unknown tasks, particularly with relation to objects that may not be commonly acted upon by robots.
  • Tasking a robot to find a particular toy in a child’s messy room may be challenging, not the least because the conventional robot control policy may not be trained to recognize the particular toy.
  • VLMs by contrast, are much more adept and recognizing less-common objects, including heretofore unencountered objects.
  • VLMs to perform vision-based robotic planning as described herein enables robots to be operated to perform “new” tasks, without needing to train or fine-tune robot control policies based on those new tasks (although finetuning based on those new tasks can make the VLM even more useful for robotic planning moving forward).
  • Fig. l is a schematic diagram of components that can cooperate to carry out selected aspects of the present disclosure, in accordance with various implementations.
  • the various components depicted in Fig. 1, particularly those components forming a vision language system 130 and a proprioception system 140, may be implemented using any combination of hardware and software.
  • the components of Fig. 1 are depicted as being communicatively coupled with each other via one or more networks 199, which may include one or more personal area networks, local area networks, and/or wide area networks (e.g., the Internet). However, this is not meant to be limiting.
  • Various aspects of the present disclosure that are described as being performed by and/or stored on systems 130 and/or 140 can alternatively be performed by and/or stored on a single system, such as vision language system 130, or on any combinations of systems 130 and 140.
  • a robot 100 may be in communication with systems 130 and/or 140.
  • systems 130 and/or 140 may be implemented onboard robot 100.
  • Other types of machines or apparatus that are not depicted in Fig. 1 may also be controlled using selected aspects of the present disclosure, such as autonomous vehicles, industrial equipment, climate control systems, medical systems and/or devices, video games, and so forth.
  • Robot 100 may take various forms, including but not limited to a telepresence robot (e.g., which may be as simple as a wheeled vehicle equipped with a display and a camera), a robot arm, a multi-pedal robot such as a “robot dog,” an aquatic robot, a wheeled device, a submersible vehicle, an unmanned aerial vehicle (“UAV”), and so forth.
  • a mobile robot arm is depicted in Fig. 2.
  • robot 100 may include logic 102.
  • Logic 102 may take various forms, such as a real time controller, one or more processors, one or more field-programmable gate arrays (“FPGA”), one or more applicationspecific integrated circuits (“ASIC”), and so forth.
  • logic 102 may be operably coupled with memory 103.
  • Memory 103 may take various forms, such as randomaccess memory (“RAM”), dynamic RAM (“DRAM”), read-only memory (“ROM”), Magnetoresistive RAM (“MRAM”), resistive RAM (“RRAM”), NAND flash memory, and so forth.
  • a robot controller may include, for instance, logic 102 and memory 103 of robot 100.
  • logic 102 may be operably coupled with one or more joints 104-1 to 104-N, one or more end effectors 106, and/or one or more sensors 108-1 to 108- M, e.g., via one or more buses 109.
  • joints 104 of a robot may broadly refer to actuators, motors (e.g., servo motors), shafts, gear trains, pumps (e.g., air or liquid), pistons, drives, propellers, flaps, rotors, or other components that may create and/or undergo propulsion, rotation, and/or motion.
  • Some joints 104 may be independently controllable, although this is not required.
  • end effector 106 may refer to a variety of tools that may be operated by robot 100 in order to accomplish various tasks.
  • some robots may be equipped with an end effector 106 that takes the form of a claw with two opposing “fingers” or “digits.”
  • Such a claw is one type of “gripper” known as an “impactive” gripper.
  • grippers may include but are not limited to “ingressive” (e.g., physically penetrating an object using pins, needles, etc.), “astrictive” (e.g., using suction or vacuum to pick up an object), or “contigutive” (e.g., using surface tension, freezing or adhesive to pick up object).
  • other types of end effectors may include but are not limited to drills, brushes, force-torque sensors, cutting tools, deburring tools, welding torches, containers, trays, and so forth.
  • end effector 106 may be removable, and various types of modular end effectors may be installed onto robot 100, depending on the circumstances.
  • Some robots such as some telepresence robots, may not be equipped with end effectors. Instead, some telepresence robots may include displays to render visual representations of the users controlling the telepresence robots, as well as speakers and/or microphones that facilitate the telepresence robot “acting” like the user.
  • Sensors 108-1 to 108-M may take various forms, including but not limited to 3D laser scanners (e.g., light detection and ranging, or “LIDAR”) or other 3D vision sensors (e.g., stereographic cameras used to perform stereo visual odometry) configured to provide depth measurements, two-dimensional cameras (e.g., RGB, infrared), light sensors (e.g., passive infrared), force sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors (also referred to as “distance sensors”), depth sensors, torque sensors, barcode readers, radio frequency identification (“RFID”) readers, radars, range finders, accelerometers, gyroscopes, compasses, position coordinate sensors (e.g., global positioning system, or “GPS”), speedometers, edge detectors, Geiger counters, and so forth. While sensors 108-1 to 108-M are depicted as being integral with robot 100, this is not meant to be limiting.
  • 3D laser scanners
  • vision language system 130 and/or proprioception system 140 may include one or more computing devices cooperating to perform selected aspects of the present disclosure.
  • An example of such a computing device is depicted schematically in Fig. 7.
  • one or more of systems 130 and/or 140 may include one or more servers forming part of what is often referred to as a “cloud” infrastructure, or simply “the cloud.”
  • one or more components of systems 130 and/or 140 may be operated by logic 102 of robot 100.
  • Machine learning model(s) described herein may take various forms, including, but not limited to, generative language model(s) (sometimes referred to as “large language models,” or “LLMs”) such as PaLM, PaLM-E, Gemini, BERT, LaMDA, Meena, and/or any other generative language model, such as any other generative model that is encoder-only based, decoder-only based, sequence-to-sequence based and that optionally includes an attention mechanism or other memory.
  • LLMs large language models
  • machine learning model(s) may have hundreds of millions, or even hundreds of billions of parameters.
  • machine learning model(s) may include a multi-modal model such as a VLM and/or a visual question answering (VQA) model, which can have any of the aforementioned architectures, and which can be used to process multiple modalities of data, particularly images and text, and/or images and audio for example, to generate one or more modalities of output.
  • VLM visual question answering
  • VQA visual question answering
  • Vision language system 130 may include a sampling engine 132, VLM engine 134, one or more VLMs 135, a visual annotation engine 136, and a feedback engine 138. Any of engines 134, 136, and/or 138 may be implemented using any combination of hardware and software. Moreover, any of engines 134, 136, and/or 138 may be combined with other(s) of engines 134, 136, and/or 138.
  • VLM(s) 135 that may be applied as described herein include Gemini (e.g., Nano, Pro, and/or Ultra) and/or Flamingo, to name a few.
  • sampling engine 132 may be configured to sample distribution(s) of candidate actions for robot 100 to perform in furtherance of completing task(s) in environment(s) in which robot 100 operates.
  • Sampling engine 132 may sample these candidate actions from actions spaces that may or may not be constrained based on factors such as a location of robot 100 relative to a destination robot 100 is meant to travel and/or object(s) to be acted upon by robot 100.
  • the action spaces may be constrained based on a digital visual representation of an environment in which robot 100 operates.
  • candidate actions that would occur outside of the field of view of whichever vision sensor e.g., on the robot or elsewhere in its environment) captured the digital visual representation may be excluded from or otherwise not available in the action space from which sampling engine 132 samples.
  • the action space(s) from which sampling engine 132 samples actions may be constrained based on previous candidate actions selected from previous distributions of candidate actions sampled by sampling engine 132.
  • Proprioception system 140 may be present in some implementations where robot 100 is being controlled using techniques described herein. Proprioception system 140 may be omitted in other circumstances. Proprioception system 140 may include a proprioception prediction process 142 and one or more proprioception machine learning models 144. Examples of a proprioception machine learning model that may be used include PaLM, PaLM-E, model(s) described in “RT-1 : Robotics Transformer for Real-World Control at Scale” (arXiv:2212.06817), and/or model(s) described in “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control” (arXiv:2307.15818).
  • RT-1 Robotics Transformer for Real-World Control at Scale
  • RT-2 Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
  • proprioception prediction process 142 may process input tokens indicative of a current (or past) proprioception values of robot 100, e.g., along with other data such as data indicative of a task or action to be performed (e.g., an action sampled and selected as described herein), state data of the robot’s environment, etc., to generate robot control data and/or predict future proprioception values of robot 100. These robot control data and/or future proprioception values may be used to operate robot 100.
  • Robot control data may include, for instance, low-level actuator commands (also referred to as “joint commands,” and may include torque commands) that directly control the actuators/joints 104-1 to 104-N of the robot, cartesian commands that specify directi on(s) for an end effector 106, a target robot pose, code that specifies reward functions that a motion controller can optimize (e.g., using techniques such as receding horizon optimization) to find optimal low-level actuator commands, selected predefined robot primitives, and so forth.
  • robot logic 102 may be configured to convert between joint commands and Cartesian commands, e.g., using forward and/or inverse kinematics.
  • a user 150 may control robot 100 using a client device 152. While depicted as a tablet computer or smart phone in Fig. 1, client device 152 may take other forms, such as a desktop or laptop computer, in-vehicle computing device, augmented reality (AR) and/or virtual reality (VR) headset or glasses, standalone “smart” speakers that host automated assistants that can be interacted with the control robot 100, etc.
  • client device 152 may issue one or more natural language commands, e.g., by typing the commands or uttering the commands aloud and having those spoken utterances transcribed using speech-to-text (STT) processing.
  • STT speech-to-text
  • These natural language commands may specify a task to be completed by robot 100 in an environment in which robot 100 operates. For example, user 150 may ask robot 100 to “pick up the helix-shaped dog chew toy,” “close the windows,” “take the dishes from the table to the sink,” etc.
  • Fig. 2 depicts a non-limiting example of a robot 200 in the form of a robot arm.
  • An end effector 206 in the form of a gripper claw is removably attached to a sixth joint 204-6 of robot 200.
  • six joints 204-1 to 204-6 are indicated.
  • robots may have any number of joints.
  • robot 200 may be mobile, e.g., by virtue of a wheeled base 255 or other locomotive mechanism.
  • Robot 200 is depicted in Fig. 2 in a particular selected configuration or “pose.”
  • a mobile manipulator performing navigation tasks may use the image from a fixed head camera and annotate the image with arrows originating from the bottom center of the image to represent the 2D action space.
  • an on-board depth camera of the robot may be used to map the candidate action to a 3D target location and command the robot to move toward the target (e.g., with a maximum distance of 1.0m).
  • a mobile manipulator performing manipulation tasks may use the image from a fixed head camera and annotate the image with arrows originating from the end-effector in the camera frame.
  • Each arrow may represent a 3D relative Cartesian end-effector position (x; y; z).
  • two settings may be considered: where height is represented using color grading (e.g., a red to blue spectrum) and where the arm only uses fixed-height actions.
  • Gripper closing actions may or may not be shown as visual annotations but instead may be expressed through text prompts.
  • manipulator robots such as robots having cameras mounted on their “wrists,” may have those images annotated with arrows that originate from the center of the camera frame.
  • Each arrow may represent a 3D relative Cartesian endeffector position (x, , z, where the z dimension may be captured with a color spectrum, e.g., from red to blue).
  • techniques described herein may be used in simulation as well, e.g., for pick and place manipulation. Image(s) from an overhead may be used, and those images may be annotated with pick and place locations.
  • Figs. 3 A and 3B depict how techniques described herein may be applied for two iterations when a user (e.g., 150) issues the robot command, “pick up the helix dog chew toy.”
  • a first visual representation 365 A e.g., in the form of a digital image, depicts three objects: a bow-shaped chew toy 360A, a rolling pin-shaped chew toy 360B, and a helix-shaped chew toy 360C, which is the object the user wants a robot 300 to act upon.
  • the visual representation captures the environment in which robot 300 operates from a third person perspective, rather than from the perspective of a vision sensor carried by robot 300. However, this is not meant to be limiting. Techniques described herein are also applicable to visual representations captured from other perspectives, such as from the perspective of robot 300 itself.
  • a first distribution of candidate actions for robot 300 to perform have been sampled, e.g., by sampling engine 132, from a first action space d 1 .
  • This first action space d 1 may be relatively unconstrained.
  • the first action space A 1 may only be constrained to the space that robot 300 can reach within some amount of time.
  • the first action space d 1 may be constrained to the physical space that is within the field-of-view of the vision sensor that captured first visual representation 365A.
  • each annotation takes the form of a circle representing a candidate end effector location and an arrow representing a trajectory to that location.
  • visual annotations A-H are contemplated, such as other shapes, arrows that represent non-linear (e.g., curvy) trajectories, etc.
  • each of visual annotations A-H is a circle having the same size.
  • the visual annotations may be heterogeneous in shape and/or size.
  • larger visual annotations may be rendered to represent locations that are closer to the perspective from which the visual annotation was captured, and smaller visual annotations may be rendered to represent locations that are farther away from the perspective from which the visual annotation was captured.
  • Different colored annotations may be rendered as well, e.g., to represent locations that are closer to or farther from the perspective of the vision sensor, or to represent costs or other attributes of the underlying candidate actions.
  • the candidate actions may be selected from the first action space d 1 more-or- less randomly. Consequently, the visual annotations A-H are spread relatively randomly and equally across visual representation. While some (e.g., F, G) are closer to the object (360C) to be acted upon by robot 300, others are farther away, e.g., closer to other objects and/or in random locations.
  • VLM engine 134 or another component may assemble a VLM prompt that includes data indicative of visual representation 365A (e.g., embeddings generated therefrom, pixels, extracted features, etc.), and data indicative of a request such as “which location is closest to the helix-shaped dog chew toy?”
  • VLM engine 134 may process this VLM prompt using VLM 135 to generate VLM output.
  • the VLM output may include tokens indicative of which visual annotation is closest to the object (helix-shaped chew toy 360C) to be acted upon by robot 300.
  • the VLM output tokens may be decoded to reveal location F to be the closest to helix-shaped dog chew toy 360C.
  • a new distribution of candidate actions has been sampled, e.g., by sampling engine 132, from a new action space A.
  • Second action space A may be more constrained than first action spaced 1 and the latest sampled candidate actions may be fitted to this new more constrained action space A .
  • second action space A may be constrained based on the first selected candidate action F.
  • Second visual representation 365B may or may not be a copy of first visual representation 365A (prior to the addition of visual annotations A-H) or may be captured separately, e.g., after robot 300 has moved its end effector to a location that corresponds with the visual annotation F in first visual representation 365A. As shown in Fig. 3B, the visual annotations A’-H’ are distributed more tightly around the location where visual annotation F was in first visual representation 365A.
  • VLM output may identify either visual annotation E’ or visual annotation F’.
  • both visual annotations E’/F’ are proximate enough to helix-shaped dog chew toy 360C that during the next iteration, the request that is submitted to VLM engine 134 may be different, e.g., “Is E’ or F’ better for picking up the helix-shaped dog chew toy?”
  • Fig. 4 depicts an example of how data may be exchanged between various components depicted in Fig. 1 to carry out selected aspects of the present disclosure.
  • a client device 452 may be operated by a user (not depicted) to receive a robot task.
  • the robot task may be provided as part of a natural language statement that is typed at client device 452 or recorded and transcribed at client device 452 or elsewhere.
  • the robot task may be provided to sampling engine 132.
  • Sampling engine 132 may sample an zth distribution 472 of candidate actions that may be performed by a robot 400.
  • the action space from which sampling engine 132 may sample candidate actions may be relatively unconstrained, but may converge to more constrained action spaces during subsequent iterations.
  • the zth distribution of sampled candidate actions 472 and image(s) 470 captured by a camera onboard or proximate robot 400 may be provided to or obtained by visual annotation engine 136.
  • Visual annotation engine 136 may project the zth distribution of sampled candidate actions 472 onto the image(s) 470 to generate annotated image(s) 474.
  • Annotated image(s) 474 may then be processed by VLM engine 134 using VLM 135 to generate VLM output indicative of which visual annotation represents the “best” candidate action for robot 400 to perform, or the visual annotation that is closest to an end goal of a task assigned to robot 400.
  • Feedback engine 138 may evaluate the VLM output to determine whether additional iteration(s) are needed, or whether the selected visual annotation is “close enough” or “good enough” for robot 400 to perform it. If the selected candidate action is suitable — e.g., it would place an end effector of robot 400 in a good position to grasp whatever object robot 400 is supposed to grasp — then data indicative of the selected robot action may be provided to proprioception prediction process 142, which may generate robot control data that can then be used by (e.g., transmitted to) robot 400 to cause robot 400 to perform the selected robot action. [0065] If feedback engine 138 determines that the selected candidate action is not “good enough,” on the other hand, then the process may pass back to sampling engine 132.
  • Sampling engine may now sample a new distribution z+1 of candidate actions that may be performed by a robot 400.
  • the cyclic process depicted in Fig. 4 may then repeat as necessary until the sampled candidate actions (and the visual representations that represent them) converge onto the best candidate action(s) for robot 400 to perform to carry out the task assigned to it.
  • Fig. 5 A depicts an example of how data may be processed to carry out selected aspects of the present disclosure.
  • a plurality of candidate actions are sampled, e.g., by sampling engine 132, from action space A (1) 560.
  • Each of these sampled candidate actions is represented in Fig. 5A with an arrow that originates at a robot 500.
  • Robot 500 also includes an onboard camera 558.
  • an image 562 captured by onboard camera 558 may be annotated, e.g., by visual annotation engine 136, with projections of the candidate actions sampled previously.
  • seven arrows are annotated 1-7 to represent the seven sampled candidate actions, but this is not meant to be limiting. More or less annotations/candidate actions may be sampled/incorporated into image 562.
  • the annotated image 562 is used, e.g., by VLM engine 134, to query a VLM to select one or more of the sampled candidate actions.
  • the prompt reads, “Which arrows should the robot follow to pick up the block,” but different prompts may be provided if different annotations are used, if different tasks are being performed, etc.
  • annotations 5 and 6 are the closest to the block that is to be acted upon by robot 500.
  • a new set of candidate actions may be sampled from a now reduced action space 560’.
  • This iterative process of sampling from increasingly narrower action spaces is demonstrated in Fig. 5B, in which the action space 560A gradually decreases in size to that of 560B, and eventually, to 560C.
  • Figs. 5A-B depicts a robot 500 in the form of a robot arm, and an action space 560 that includes actions performable by the robot arm.
  • this is not meant to be limiting.
  • similar techniques may be applied in other scenarios, such as to assist in a self-driving vehicle (e.g., robot, car, truck, agricultural rover, robotic dog, etc.) navigate around and/or through obstacles.
  • a self-driving vehicle e.g., robot, car, truck, agricultural rover, robotic dog, etc.
  • Techniques described herein may be equally applicable to navigation of unmanned aerial vehicles (UAVs).
  • UAVs unmanned aerial vehicles
  • FIG. 6 an example method 600 of practicing selected aspects of the present disclosure is described.
  • This system may include various components of various computer systems, including those depicted in Fig. 1.
  • operations of method 600 are shown in a particular order, this is not meant to be limiting.
  • One or more operations may be reordered, omitted or added.
  • the task may be derived from sources such as a robot operator’s natural language input, which may be typed or spoken and then processed using STT processing to generate textual output.
  • the system may incorporate, into a digital visual representation (e.g, one or more digital images 470) of the environment in which the robot operates, visual annotations (e.g., A-H and A’-H’ in Figs. 3 A-B) that represent the first distribution of candidate actions.
  • this incorporation may include overlaying and/or projecting the first distribution of candidate actions onto the visual representation.
  • the visual representation may be registered with the robot’s environment, e.g., by matching one or more locations (e.g., pixels) within the visual representation with one or more corresponding locations (e.g., points) in the robot’s environment.
  • the visual annotations may be transparent and/or translucent, e.g., so that underlying visual features of the visual representation are not entirely concealed.
  • the sampling of block 602 may be constrained to avoid objects or other visual features of potential interest in the visual representation.
  • the system may assemble a VLM input prompt that includes the digital visual representation and a request to select one or more of the candidate actions, e.g., for the robot to perform in furtherance of completing the task, and/or to facilitate sampling of a new, fitted distribution.
  • the system may process the VLM input prompt using a VLM 135 to generate VLM output.
  • This VLM output may include token(s) that are indicative of one or more selected candidate actions.
  • the system e.g., by way of VLM engine 134 or feedback engine 138, may select one of the candidate actions of the first distribution as a first selected candidate action. For instance, the VLM output may identify the visual annotation/candidate action that will advance the robot farthest towards the goal of completing the robot task.
  • this may be the visual annotation/candidate action that is closest to an object to be acted upon by the robot.
  • visual annotation F was selected because it was closest to helix-shaped dog toy 360C.
  • visual annotation E’ or F’ may be selected because they are the closest to helix-shaped dog toy 360C.
  • the system may determine whether one or more criteria are satisfied by the VLM output generated at block 608 and/or the selected candidate action of block 610. These criteria may include, for instance, whether performance of the selected candidate action will advance the robot towards completing the task assigned to it.
  • the VLM output generated at block 608 may be associated with one or more probability distributions. These probability distributions may indicate, for example, a confidence that performance of a given candidate action will result in a positive outcome, whether that outcome be advancement towards completion of the task or even final completion of the task.
  • method 600 may proceed back to block 602, and a next (e.g., zth) distribution of candidate actions for the robot to perform in furtherance of completing the task may be sampled, e.g., by sampling engine 132. In various implementations, this subsequent sampling may be constrained by the visual annotation/candidate action that was selected at block 610. In Fig. 3B, for instance, the updated distribution of visual annotations A’- G’ was constrained by the location of the previously selected visual annotation F in Fig. 3 A. Method 600 may then proceed through blocks 604-610 as described previously. However, if the criteria are/is satisfied at block 612, then in some implementations, method 600 may proceed to block 614.
  • a next distribution of candidate actions for the robot to perform in furtherance of completing the task may be sampled, e.g., by sampling engine 132. In various implementations, this subsequent sampling may be constrained by the visual annotation/candidate action that was selected at block 610. In Fig. 3B, for instance, the updated distribution of
  • VLM 135 may be trained and/or fine-tuned based on various data generated during performance of various aspects of method 600.
  • the system e.g., by way of proprioception system 140, may cause robot 100/200 to perform the most-recently selected candidate action.
  • the system e.g., by way of feedback engine 138 and/or VLM engine 134, may fine-tune the VLM 135 based on, for example, an outcome of the robot performing the selected candidate action.
  • the outcome may be derived from human feedback about the robot operation, or from feedback provided by the robot itself.
  • the system may train VLM 135 without requiring robot operation.
  • VLM 135 may be trained based on, for example, VLM output and/or based on the selected candidate robot action (which need not necessarily be performed by the robot).
  • VLM output may be evaluated by humans and/or based on a trained reward function.
  • the selected candidate robot action may be evaluated by humans and/or based on a trained reward function.
  • a human teleoperating a real or simulated robot may have a display that renders a video feed from a camera onboard the robot or elsewhere.
  • this video feed may be annotated, e.g., in real time, every //th frame, every time the visual context changes, etc., with annotations representing sampled candidate actions (and in some cases, which candidate actions are most likely to make progress towards completing a task).
  • the human may choose to teleoperate the robot in a manner consistent with the suggested annotations or not, in turn generating positive feedback or negative feedback, respectively.
  • Fig. 7 is a block diagram of an example computer system 710.
  • Computer system 710 typically includes at least one processor 714 which communicates with a number of peripheral devices via bus subsystem 712.
  • peripheral devices may include a storage subsystem 724, including, for example, a memory subsystem 725 and a file storage subsystem 726, user interface output devices 720, user interface input devices 722, and a network interface subsystem 716.
  • the input and output devices allow user interaction with computer system 710.
  • Network interface subsystem 716 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
  • User interface input devices 722 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and/or other types of input devices.
  • pointing devices such as a mouse, trackball, touchpad, or graphics tablet
  • audio input devices such as voice recognition systems, microphones, and/or other types of input devices.
  • use of the term "input device” is intended to include all possible types of devices and ways to input information into computer system 710 or onto a communication network.
  • User interface output devices 720 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices.
  • the display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image.
  • the display subsystem may also provide non-visual display such as via audio output devices.
  • output device is intended to include all possible types of devices and ways to output information from computer system 710 to the user or to another machine or computer system.
  • Storage subsystem 724 stores programming and data constructs that provide the functionality of some or all of the modules described herein.
  • the storage subsystem 724 may include the logic to perform selected aspects of method 600, and/or to implement one or more aspects of robot 100 or the various systems depicted in Fig. 1.
  • Memory 725 used in the storage subsystem 724 can include a number of memories including a main random-access memory (RAM) 730 for storage of instructions and data during program execution and a read only memory (ROM) 732 in which fixed instructions are stored.
  • a file storage subsystem 726 can provide persistent storage for program and data files, and may include a hard disk drive, a CD-ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 726 in the storage subsystem 724, or in other machines accessible by the processor(s) 714.
  • Bus subsystem 712 provides a mechanism for letting the various components and subsystems of computer system 710 communicate with each other as intended. Although bus subsystem 712 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
  • Computer system 710 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the everchanging nature of computers and networks, the description of computer system 710 depicted in Fig. 7 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer system 710 are possible having more or fewer components than the computer system depicted in Fig. 7.
  • a computer implemented method includes: sampling a first distribution of candidate actions for a robot to perform in furtherance of completing a task in an environment in which the robot operates; incorporating, into a digital visual representation of the environment in which the robot operates, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions for the robot to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; based on the VLM output, selecting one of the candidate actions of the first distribution as a first selected candidate action; and causing the robot to perform the first selected candidate action.
  • VLM vision-language model
  • the first distribution of candidate actions may be selected from an action space that is constrained based on the digital visual representation of the environment.
  • the method may include, based on the first selected candidate action, sampling a second distribution of candidate actions for the robot to perform in furtherance of completing the task.
  • the second distribution of candidate actions may be selected from an action space that is constrained based on the first selected candidate action.
  • the method may include: incorporating, into the digital visual representation, second visual annotations that represent the second distribution of candidate actions; assembling a second VLM input prompt that includes the digital visual representation and a request to select one or more of the candidate actions of the second distribution for the robot to perform in furtherance of completing the task; processing the second VLM input prompt using the VLM to generate second VLM output; based on the second VLM output, selecting one of the candidate actions of the second distribution as a second selected candidate action; and causing the robot to perform the second selected candidate action of the second distribution.
  • the robot may perform both the first and second selected candidate actions subsequent to selection of the second selected candidate action.
  • the digital visual representation may be generated based on sensor data generated by one or more sensors in the environment. In various implementations, the digital visual representation may be generated based on sensor data generated by one or more sensors carried by the robot. In various implementations, the sampling may include sampling a first plurality of points within the digital visual representation, wherein each candidate action of the first distribution corresponds to a respective point of the plurality of points.
  • the candidate actions of the first distribution may include one or more end effector trajectories.
  • the one or more visual annotations that represent the one or more end effector trajectories may include vector arrows extending from a starting point in the digital visual representation to candidate end effector positions in the digital visual representation.
  • the digital visual representation may include a three- dimensional (3D) representation of the environment.
  • the 3D representation of the environment may be a point cloud generated using a light detection and ranging (LIDAR) sensor.
  • LIDAR light detection and ranging
  • the digital visual representation may be a two-dimensional (2D) representation of the environment.
  • the visual annotations may include a first shape representing a first candidate action that would occur a first distance from a vision sensor that captured the digital visual representation, and a second shape representing a second candidate action that would occur a second distance from the vision sensor.
  • the first shape may be sized differently from the second shape to convey a difference between the first and second distances.
  • the first shape may be colored differently from the second shape to convey a difference between the first and second distances.
  • the visual annotations may overlay content depicted in the digital visual representation.
  • the visual annotations may be at least partially transparent so that underlying content remains visible in the digital visual representation.
  • the VLM input prompt may be assembled to further include a raw version of the digital visual representation without the visual annotations.
  • the robot may be simulated in a virtual environment or may be a physical robot operated in a physical environment.
  • a method may be implemented using one or more processors and may include: sampling a first distribution of candidate actions for a robot to perform in furtherance of completing a task in the environment; incorporating, into a digital visual representation of an environment in which the robot operates, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions for the robot to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; and fine-tuning the VLM directly or indirectly based on the VLM output.
  • VLM vision-language model
  • the method may include: selecting one of the candidate actions of the first distribution; and causing the robot to perform the selected candidate action; [0099] wherein the VLM is trained based at least in part on an outcome of the robot performing the selected candidate action.
  • the VLM may be trained based on human feedback provided based on the VLM output.
  • the VLM may be trained based on feedback provided based on the VLM output using a trained reward function.
  • a method may be implemented using one or more processors and may include: sampling a first distribution of candidate actions to be performed in furtherance of completing a task in a two-dimensional (2D) or three-dimensional (3D) environment; incorporating, into a digital visual representation of the environment, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions to be performed in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; based on the VLM output, selecting one of the candidate actions of the first distribution; and causing the selected candidate action to be performed in the environment.
  • VLM vision-language model
  • the environment may take the form of a canvas rendered as part of a graphical user interface (GUI).
  • GUI graphical user interface
  • the environment may be a physical space in which a robot operates, and the candidate actions comprise candidate actions to be performed by the robot in the physical space in furtherance of completing the task.
  • a method may be implemented using one or more processors and may include: sampling a first distribution of candidate actions to perform in furtherance of completing a task in an environment; incorporating, into a digital visual representation of the environment, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; and based on the VLM output, selecting one of the candidate actions of the first distribution as a first selected candidate action.
  • VLM vision-language model
  • implementations may include a non-transitory computer readable storage medium storing instructions executable by a processor to perform a method such as one or more of the methods described above.
  • implementations may include a control system including memory and one or more processors operable to execute instructions, stored in the memory, to implement one or more modules or engines that, alone or collectively, perform a method such as one or more of the methods described above.

Landscapes

  • Engineering & Computer Science (AREA)
  • Robotics (AREA)
  • Mechanical Engineering (AREA)
  • Manipulator (AREA)

Abstract

Implementations relate visually prompting a VLM or another multimodal generative model for few-shot and/or zero-shot control of actions performed in a 2D or 3D space. In various implementations, a first distribution of candidate actions to perform in furtherance of completing a task in an environment may be sampled. Visual annotations that represent the first distribution of candidate actions may be incorporated into a digital visual representation of the environment. A vision-language model (VLM) input prompt may be assembled that includes the digital visual representation and a request to select one or more of the candidate actions to perform in furtherance of completing the task. The VLM input prompt may be processed using a VLM to generate VLM output. Based on the VLM output, one of the candidate actions of the first distribution may be selected as a first selected candidate action.

Description

VISUAL PROMPTING FOR FEW SHOT CONTROL
ATTORNEY REFERENCE: DEEP-0006-WO-01 Background
[0001] Large multimodal machine learning models such as vision-language models (VLMs) may be used to process multiple modalities of data, such as text and images, to generate various types of VLM output. In the robotic context, the VLM input may include vision data from one or more vision sensors and natural language textual input identifying a task to be completed by a robot. VLM output in the robotic context may include, for instance, natural language statements that describe steps for completing some task, robot control data that is usable to control a robot to complete the task, etc. However, VLMs often require large amounts of training data to adequately support robot performance, and even still may not be sufficiently generalizable.
Summary
[0002] Implementations are described herein for framing tasks that involve performing actions within a two-dimensional (2D) or three-dimensional (3D) space as visual problems that can be posed to a VLM, e.g., iteratively or otherwise. More particularly, but not exclusively, techniques described herein relate to visually prompting a VLM or another multimodal generative model for few-shot and/or zero-shot control of actions performed in a 2D or 3D space. In implementations where these actions are performable in furtherance of completing a higher level task, the VLM may be prompted iteratively, with candidate sample candidate actions iteratively closing in on completion of the higher level task. While examples described herein relate primarily to control a real or simulated robot, this is not meant to be limiting. Techniques described herein may be applicable in any context in which actions are performed in a 2D or 3D space, such as moving graphical icons around a 2D or 3D graphical user interface (GUI, e.g., associated with computer-aided design, or “CAD”), moving objects around a 2D or 3D physical space (e.g., loading a truck, adding objects to or removing objects from an inventory in a warehouse, etc. , and so forth. Brief Description of the Drawings
[0003] Fig. 1 schematically depicts an example environment in which disclosed techniques may be employed, in accordance with various implementations.
[0004] Fig. 2 depicts an example robot, in accordance with various implementations.
[0005] Fig. 3 A and Fig. 3B depict how techniques described herein may be applied for two iterations, in accordance with various implementations.
[0006] Fig. 4 schematically depicts an example of how components of Fig. 1 may cooperate to carry out selected aspects of the present disclosure, in accordance with various implementations. [0007] Fig. 5 A and Fig. 5B depict an example of how techniques described herein may be implemented.
[0008] Fig. 6 depicts an example method for practicing selected aspects of the present disclosure.
[0009] Fig. 7 schematically depicts an example architecture of a computer system.
Detailed Description
[0010] Implementations are described herein for framing tasks that involve performing actions within a two-dimensional (2D) or three-dimensional (3D) space as visual problems that can be posed to a VLM, e.g., iteratively or otherwise. More particularly, but not exclusively, techniques described herein relate to visually prompting a VLM or another multimodal generative model for few-shot and/or zero-shot control of actions performed in a 2D or 3D space. In implementations where these actions are performable in furtherance of completing a higher level task, the VLM may be prompted iteratively, with candidate sample candidate actions iteratively closing in on completion of the higher level task. While examples described herein relate primarily to controlling a real or simulated robot, this is not meant to be limiting. Techniques described herein may be applicable in any context in which actions are performed in a 2D or 3D space, such as moving graphical icons around a 2D or 3D graphical user interface (GUI, e.g., associated with computer-aided design, or “CAD”), moving objects around a 2D or 3D physical space (e.g., loading a truck, adding objects to or removing objects from an inventory in a warehouse, navigating an autonomous vehicle, etc.), and so forth. [0011] Several implementations described herein relate to solving tasks by producing a value a E A from a set A given a task description in natural language I E L and an image observation I E RHXWX3 This set A can, for example, include continuous coordinates, 3D spatial locations, robot control actions, or trajectories. When A is the set of robot actions, this may amount to finding a policy TT(- |Z, /) that emits an action a E A. While described herein primarily in relation to a control policy for robot actions, this is not meant to be limiting. Techniques described herein may be applicable in other domains in which generating (continuous) outputs from a VLM can be useful.
[0012] Techniques described herein may be used to ground VLMs to robot actions through image annotations. In various implementations, the problem of creating the policy TI may be framed as a visual question answering (VQA) problem. The class of VLMs used in various implementations take as input an image I and a textual prefix wp from which they generate a distribution PVLM ' \™P> I) of textual completions. Utilizing this interface to drive a policy raises the challenge of how an action from a (continuous) space A can be represented as a textual completion.
[0013] Accordingly, various implementations described herein lift low-level actions into the visual language of a VLM, i.e., a combination of images and text, such that it is closer to the training distribution of general vision-language tasks. To achieve this, the following visual prompt mapping may be used: that transforms an image observation I and a set of candidate actions a M, a.j E A into an annotation image I and their corresponding textual labels w1.M where Wj refers to the annotation representing a.j in the image space. For example, and as visualized in Figs. 3 A-B and 5A-B, the camera matrices can be used to project a 3D location into the image space, and draw a visual marker at this projected location. Labeling this marker with a textual reference, e.g., a number, enables the VLM to not only be queried in its natural input space, namely images and text, but also to refer to spatial concepts in its natural output space by producing text that references the marker labels.
[0014] Techniques described herein may then be used to prompt a VLM with iterative visual optimization. Representing (continuous) robot actions and spatial concepts in image space with their associated textual labels the VLM VLM to be queried to judge if an action would be promising in solving the task. Therefore, it is possible to obtain a policy n by solving the following optimization problem:
[0015] Intuitively, a goal may be to find an action a for which the VLM would choose the corresponding label w after applying the mapping . In order to solve the optimization problem above, an iterative technique referred to herein as Prompting with Iterative Visual Optimization may be used. In each iteration i the technique first samples a set of candidate actions ar L M from a distribution PA^ . These candidate actions may then be mapped onto the image I producing the annotated image /W and the associated action labels w^M. The VLM may then be queried on a multiple choice-style question on the labels w^M to choose which of the candidate actions are most promising. This leads to set of best actions to which a new distribution PA +i) may be fit. The process may be repeated until convergence or a maximum number of steps N is reached. [0016] In some implementations, an algorithm such as the following may be used to implement visual iterative prompting as described herein:
1 : Given: image I, instruction £, action space c/Z, max iterations N, number of samples M 2: Initialize:
3 : while i <
4: Sample actions a1;M from P i
5: Project actions into image space and textual labels (/, w1;M) = 0(1, a1;M)
6: Query VLM PVLM(w\I, ) to determine the most promising actions
7: Fit distribution P^ct+i) to best actions
8: Increment iterations i <- i + 1
9: end while
10: Return: an action from the VLM best actions
[0017] These techniques may be used to query a VLM for any type of answer as long as multiple answers can be simultaneously visualized on the image. As visualized in the Figures, for the visual prompting mapping Q, actions may be represented in some implementations as arrows emanating from the robot or the center of the image, e.g., if the embodiment/robot is not visible. In some implementations, for three-dimensional (3D) problems, the colors or other visual aspects of the arrows may be selected to indicate forward and backwards movement. These actions may be labeled in some cases with a number label circled at the end of the arrow. For creating the text prompt wp, the VLM may be prompted to use chain-of-thought to reason through the problem and then summarize the top few labels. The distributions PA in may be approximated as Gaussians.
[0018] In some implementations, parallel calls may be implemented, e.g., via ensembling, to provide a more robust experience. VLMs can make mistakes, which may result in actions being selected from sub-optimal regions. To improve the robustness of techniques described herein, a parallel call strategy may be used in which E parallel instances of the VLM are called to obtain E candidate actions. The selected candidates actions may then be aggregated to identify the final action output. To aggregate the candidate actions from different instances, two different approaches may be taken: 1) a new action distribution from the E action candidates may be fitted and returned, 2) the VLM may be queries again to select the single best action from the E actions. Adopting parallel calls may improve the robustness of techniques described herein and mitigate local minima in the optimization process.
[0019] In various implementations, techniques described herein may operate as follows. A digital visual representation of an environment in which actions are to be performed may be provided. In the robotic context, this digital visual representation may include, for instance, a digital image captured by one or more sensors onboard the robot or in the robot’s environment, a 3D point cloud captured by one or more light detection and ranging (LIDAR) sensors, a CAD drawing of a robot’s environment, a 2D or 3D snapshot of a virtual environment in which a simulated robot operates, a GUI screenshot, etc.
[0020] In various implementations, a distribution of candidate actions that can be performed in a space or environment may be sampled. In the robotic context, for instance, a set of candidate actions that are performable by the robot given its current pose and/or a state of the environment may be sampled from an action space, e.g., randomly, using techniques such as cross entropy and/or particle filter optimization, and/or subject to various constraints such as a state of the environment, a natural language command from a user, an object detected in the digital visual representation, etc. In some implementations, the action space from which the candidate actions are sampled may be constrained based on the digital visual representation of the environment. For instance, if the digital visual representation of the environment is a digital image captured by a digital camera (onboard the robot or in the robot’s environment), the action space may be constrained to those actions that are performable by the robot within the portion of the environment that is depicted in the digital image. In some such implementations, at least an initial distribution of candidate actions may be sampled so they are evenly spaced across the entire visual representation. Iterations of candidate action distributions that are sampled subsequently may gradually approach a desired objective (e.g., an object to be acted upon by the robot), e.g., by the search spaces those subsequent distributions are sampling from gradually contracting. In some cases, if an object to be acted upon is detected in the digital visual representation, the distribution of candidate actions may be selected to cluster about that detected object, and/or to avoid the downstream visual annotations from obscuring the object. [0021] Once the first distribution of candidate actions is sampled, visual annotations that represent the candidate actions may be incorporated into the digital visual representation. If the digital visual representation is a digital image captured by a digital camera, then the visual annotations may include, for instance, various graphics primitives, shapes, lines, text, labels, numbers, etc. that are overlaid on top of or otherwise infused into pixels of the digital image. These visual annotations may or may not be visible to a person viewing the digital image using a digital image viewing/editing application. In some implementations, the visual annotations may be at least partially transparent or translucent, e.g., so that objects in the environment remain at least partially visible in the digital image. In other implementations, two copies of the digital image may be maintained: the original raw image (sans annotations) and an annotated image in which the annotations may or may not obscure underlying visual details.
[0022] Visual annotations may take various forms depending on a variety of factors, such as the types of candidate actions that are being explored, the capabilities of the actor (e.g., the robot, a user operating a GUI), features of objects depicted in the digital visual representation, and so forth. In the robotic context, visual annotations may take the form of, for instance, vector arrows that may or may not be tipped with shapes such as arrow heads, circles, squares, triangles, etc. In some implementations in which the candidate actions of the distribution include candidate end effector trajectories, the visual annotations that represent the end effector trajectories may take the form of vector arrows extending from a starting point in the digital visual representation (e.g., coinciding with the robot’s current location/pose) to candidate end effector positions in the digital visual representation.
[0023] In some implementations in which the digital visual representation is a 2D representation of a 3D environment (e.g., a projection, digital image, etc.), the visual annotations may be rendered to indicate where in the 3D space they would occur, in spite of the limitations of the 2D digital visual representation. Suppose a first shape represents a first candidate action that would occur a first distance from a perspective of, for instance, a vision sensor that captured the digital visual representation and/or the robot itself. Suppose further that a second shape represents a second candidate action that would occur a second distance from the same perspective, e.g., farther away from the vision sensor than the first candidate action. In various implementations, the first shape may be sized differently from the second shape to convey this difference between the first and second distances. For example, the first shape may be rendered larger than the second shape to indicate that the first candidate action would occur closer to the vision sensor that captured the 2D digital visual representation than the second candidate action. Additionally or alternatively, the first and second shapes may be colored, shaded, shaped, dashed, or otherwise visually annotated in distinctive ways to convey the difference between their respective distances from the vision sensor.
[0024] Once the visual annotations are incorporated into the digital visual representation, a VLM input prompt may be assembled that includes the annotated digital visual representation (and/or embedding(s) generated therefrom) and a request to select one or more of the candidate actions to perform in furtherance of completing the higher level task. This request may take various forms depending on various features of the visual annotations themselves and/or the candidate actions they represent, as well as features of the higher level task to be completed. For example, in some implementations, each visual annotation may include an identifier (e.g., a number, name, character, etc.) that is unique amongst the visual annotations incorporated into the digital visual representation. These identifiers may overlay various points in the digital visual representation. Suppose the higher level task is “pick up the helix-shaped chew toy,” with the objective being the “helix-shaped chew toy.” In such a scenario, and assuming the robot end effector has not yet approached the helix-shaped chew toy, the request may be something like “which identifiers are closest to and/or overlay the helix-shaped chew toy?”
[0025] In various implementations, the VLM input prompt may include other data as well. This other data may include, for instance, a desired type of VLM output. For instance, it may be desired that the VLM generate VLM output that includes tokens representing natural language describing steps that can be performed, e.g., by the robot, to complete a higher level task. Additionally or alternatively, it may be desired that the VLM output contain robot control data that is used to control a robot more directly. [0026] Other data that may be included in the VLM input prompt may additionally or alternatively include a raw version of the digital visual representation (e.g., where the visual annotations obscure underlying visual features of the digital visual representation), one or more indications of what the visual annotations represent (e.g., a legend), criteria and/or constraints that should be satisfied when selecting candidate actions (e.g., safety, efficiency, compliance with regulations, etc.), a reiteration of the higher level task, and so on. A legend may indicate, for instance, that vector arrows represent candidate end effector trajectories, another type of shape represents a gripper closing, yet another type of shape represents rotating the end effector or other portion(s) of the robot, a line thickness representing a velocity, and so forth. In some implementations, a visual annotation representing a candidate robot pose (that is different from its current pose) may take the form of a depiction of all or part of the robot (e.g., only the end effector, a portion of a robot arm, etc.) in the candidate pose.
[0027] Once assembled, the VLM input prompt may be processed using a VLM to generate VLM output. Depending on how the VLM was prompted and/or trained, the VLM output may include tokens indicative of, for instance, natural language statement(s) or other indications identifying which of the candidate actions would be most beneficial/efficient/safe for the robot to perform in furtherance of completing the higher level task, robot control data, and so forth. In a non-robotic, CAD-based example, the tokens may be indicative of, for instance, where to move and/or insert a component into an existing CAD drawing.
[0028] The VLM output may be used for various purposes. In some implementations, depending on the nature of the VLM output, it may be used directly (e.g., when in the form of robot control data) or indirectly (e.g., when in the form of natural language) to control a robot and/or to sample new candidate actions from progressively contracting action spaces. Additionally or alternatively, in some implementations, the VLM output may be used to train and/or further train (e.g., fine-tune) the VLM itself. This training may be based on, for instance, feedback (e.g., from a human or from a trained reward model) about the VLM output itself and/or feedback about performance of the selected candidate action(s), e.g., by the robot or by another actor.
[0029] In some implementations, the process may repeat to leverage the VLM to iteratively predict and/or perform action(s) that collectively result in completion of a higher level task. In some implementations, after a candidate action is selected from a given distribution of candidate actions and performed by a robot as described above, new candidate actions that correspond to the new context, e.g., of the robot and/or its environment, may be sampled to form a next distribution. In some implementations, this next distribution of candidate actions may be selected from a new action space that is constrained, e.g., relative to prior action space, based on candidate action(s) performed up until now. Put another way, the new distribution of candidate actions may be fitted to the previously selected/performed candidate action(s). For example, if the candidate action caused a robot end effector to move to a position adjacent an objective (e.g., the aforementioned helix-shaped chew toy), then candidate actions for a new distribution may be sampled from an action space that is constrained to an area immediately adjacent the objective, e.g., in the real world or in the digital visual representation.
[0030] Once the new distribution of candidate actions is sampled, new visual annotations representing the new distribution of candidate actions may be incorporated once again into the same digital visual representation or into an updated digital visual representation, e.g., a new digital image captured of the robot’s environment after the prior action(s) are selected/performed. This new annotated digital visual representation (and/or embedding(s) generated therefrom) may once again be assembled into a VLM prompt, e.g., along with a request that is similar to or different from the previous request depending on the context (e.g., if a robot gripper is now adjacent an objective, the candidate actions and corresponding visual annotations may be considerably different than before). As before, this updated prompt may also include other elements, such as a new legend, one or more new constraints or criteria to be satisfied, etc. The VLM input prompt may once again be processed using the VLM to generate a new iteration of VLM output. The new iteration of VLM output may be used for purposes such as operating a robot to perform a next step in a sequence of steps that collectively result in completion of a higher level task, for training the VLM, etc.
[0031] It is not necessary that the real or simulated robot perform each step before the next step is predicted. In some implementations, a sequence of multiple actions may be iteratively sampled without the robot performing the actions during each iteration, and the robot may subsequently perform all or part of the entire sequence as a batch. For example, after a candidate action is selected from a given distribution of candidate actions, a next distribution of new candidate actions may be sampled based on the selected candidate action, without the robot performing the selected candidate action first. Similar to before, this next distribution of candidate actions may be selected from a new action space that is constrained relative to the prior action space based on candidate action(s) predicted up until now. Put another way, the new distribution of candidate actions may be fitted to the previously-predicted candidate action(s). For example, if the previously predicted candidate action was represented by a visual annotation at a particular position (e.g., pixel coordinate) in the visual representation, then candidate actions for a new distribution may be sampled from an action space that is constrained to an area immediately adjacent the particular position in the digital visual representation.
[0032] In various implementations, techniques described herein may be used to train and/or further train (e.g., fine-tune) a pretrained VLM, e.g., using techniques such as reinforcement learning with human feedback (RLHF), to better predict which visual annotation(s) represent the best action(s). For example, a real or simulated robot could be teleoperated by a human. The human may be provided with a view (on a display) that renders a feed of the robot’s own onboard vision sensor(s)/cameras and/or a separate camera with a view of the robot and/or its environment. The human may provide an indication (e.g., using natural language) of a task they are going to perform by teleoperating the robot, or a task may be assigned to the user to perform by teleoperating the robot.
[0033] In either case, techniques described herein may be used to sample actions from action space and render annotations on the user’s display representing projections of these actions. In some implementations, the same rendition presented to the user (image plus annotations) in one or more video frames may be processed as described herein using the VLM to determine which annotation(s) represent the best actions to take currently. In some cases, the selected annotation(s) may be suggested/recommended to the user, e.g., using visual emphasis such as highlighting, animation, etc. In other cases, the selected annotations may not be suggested/recommended to the user, e.g., so that the actions the user takes are not influenced by the annotations.
[0034] In either case, the user may teleoperate the robot to perform the task, consciously following/ignoring the recommended annotations or without seeing the annotations at all. When the user follows a recommended annotation, that may be positive feedback. When the user ignores a recommended annotation, that may be negative feedback. When the user who is not presented with annotations nonetheless teleoperates the robot in a manner that is consistent with recommended annotation(s), that may be positive feedback. Likewise, if the user who is not presented with annotations teleoperates the robot in a manner that diverges from the recommended annotations, that may be negative feedback.
[0035] This human feedback may be used in some implementations to train a reward model, so that the reward model learns to predict humans’ preferences and assigns rewards to the VLM’s predictions of which annotations represent the best actions based on these predictions. The VLM may then be optimized/fine-tuned using the reward model, e.g., with the goal to maximize the rewards predicted by the reward model, which in turn reflects the human evaluators' preferences. Various techniques may be used to optimize the VLM, including but not limited to Proximal Policy Optimization (PPO), Trust Region Policy Optimization (TRPO), gradient descent, etc. [0036] Techniques described herein may give rise to various technical advantages. For example, conventional robot control policies may be well-suited for the specific tasks for which they were trained, but they are not generally scalable to unknown tasks, particularly with relation to objects that may not be commonly acted upon by robots. Tasking a robot to find a particular toy in a child’s messy room may be challenging, not the least because the conventional robot control policy may not be trained to recognize the particular toy. VLMs, by contrast, are much more adept and recognizing less-common objects, including heretofore unencountered objects. Accordingly, leveraging VLMs to perform vision-based robotic planning as described herein enables robots to be operated to perform “new” tasks, without needing to train or fine-tune robot control policies based on those new tasks (although finetuning based on those new tasks can make the VLM even more useful for robotic planning moving forward).
[0037] Fig. l is a schematic diagram of components that can cooperate to carry out selected aspects of the present disclosure, in accordance with various implementations. The various components depicted in Fig. 1, particularly those components forming a vision language system 130 and a proprioception system 140, may be implemented using any combination of hardware and software. The components of Fig. 1 are depicted as being communicatively coupled with each other via one or more networks 199, which may include one or more personal area networks, local area networks, and/or wide area networks (e.g., the Internet). However, this is not meant to be limiting. Various aspects of the present disclosure that are described as being performed by and/or stored on systems 130 and/or 140 can alternatively be performed by and/or stored on a single system, such as vision language system 130, or on any combinations of systems 130 and 140.
[0038] In some implementations, techniques described herein may be used to control various types of machines or apparatus. For example, in some implementations, a robot 100 may be in communication with systems 130 and/or 140. In various implementations, and/or all or parts of systems 130 and/or 140 may be implemented onboard robot 100. Other types of machines or apparatus that are not depicted in Fig. 1 may also be controlled using selected aspects of the present disclosure, such as autonomous vehicles, industrial equipment, climate control systems, medical systems and/or devices, video games, and so forth.
[0039] Robot 100 may take various forms, including but not limited to a telepresence robot (e.g., which may be as simple as a wheeled vehicle equipped with a display and a camera), a robot arm, a multi-pedal robot such as a “robot dog,” an aquatic robot, a wheeled device, a submersible vehicle, an unmanned aerial vehicle (“UAV”), and so forth. One non-limiting example of a mobile robot arm is depicted in Fig. 2. In various implementations, robot 100 may include logic 102. Logic 102 may take various forms, such as a real time controller, one or more processors, one or more field-programmable gate arrays (“FPGA”), one or more applicationspecific integrated circuits (“ASIC”), and so forth. In some implementations, logic 102 may be operably coupled with memory 103. Memory 103 may take various forms, such as randomaccess memory (“RAM”), dynamic RAM (“DRAM”), read-only memory (“ROM”), Magnetoresistive RAM (“MRAM”), resistive RAM (“RRAM”), NAND flash memory, and so forth. In some implementations, a robot controller may include, for instance, logic 102 and memory 103 of robot 100.
[0040] In some implementations, logic 102 may be operably coupled with one or more joints 104-1 to 104-N, one or more end effectors 106, and/or one or more sensors 108-1 to 108- M, e.g., via one or more buses 109. As used herein, “joint” 104 of a robot may broadly refer to actuators, motors (e.g., servo motors), shafts, gear trains, pumps (e.g., air or liquid), pistons, drives, propellers, flaps, rotors, or other components that may create and/or undergo propulsion, rotation, and/or motion. Some joints 104 may be independently controllable, although this is not required. In some instances, the more joints robot 100 has, the more degrees of freedom of movement it may have. [0041] As used herein, “end effector” 106 may refer to a variety of tools that may be operated by robot 100 in order to accomplish various tasks. For example, some robots may be equipped with an end effector 106 that takes the form of a claw with two opposing “fingers” or “digits.” Such a claw is one type of “gripper” known as an “impactive” gripper. Other types of grippers may include but are not limited to “ingressive” (e.g., physically penetrating an object using pins, needles, etc.), “astrictive” (e.g., using suction or vacuum to pick up an object), or “contigutive” (e.g., using surface tension, freezing or adhesive to pick up object). More generally, other types of end effectors may include but are not limited to drills, brushes, force-torque sensors, cutting tools, deburring tools, welding torches, containers, trays, and so forth. In some implementations, end effector 106 may be removable, and various types of modular end effectors may be installed onto robot 100, depending on the circumstances. Some robots, such as some telepresence robots, may not be equipped with end effectors. Instead, some telepresence robots may include displays to render visual representations of the users controlling the telepresence robots, as well as speakers and/or microphones that facilitate the telepresence robot “acting” like the user.
[0042] Sensors 108-1 to 108-M may take various forms, including but not limited to 3D laser scanners (e.g., light detection and ranging, or “LIDAR”) or other 3D vision sensors (e.g., stereographic cameras used to perform stereo visual odometry) configured to provide depth measurements, two-dimensional cameras (e.g., RGB, infrared), light sensors (e.g., passive infrared), force sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors (also referred to as “distance sensors”), depth sensors, torque sensors, barcode readers, radio frequency identification (“RFID”) readers, radars, range finders, accelerometers, gyroscopes, compasses, position coordinate sensors (e.g., global positioning system, or “GPS”), speedometers, edge detectors, Geiger counters, and so forth. While sensors 108-1 to 108-M are depicted as being integral with robot 100, this is not meant to be limiting.
[0043] In some implementations, vision language system 130 and/or proprioception system 140 may include one or more computing devices cooperating to perform selected aspects of the present disclosure. An example of such a computing device is depicted schematically in Fig. 7. In some implementations, one or more of systems 130 and/or 140 may include one or more servers forming part of what is often referred to as a “cloud” infrastructure, or simply “the cloud.” Alternatively, one or more components of systems 130 and/or 140 may be operated by logic 102 of robot 100.
[0044] Machine learning model(s) described herein may take various forms, including, but not limited to, generative language model(s) (sometimes referred to as “large language models,” or “LLMs”) such as PaLM, PaLM-E, Gemini, BERT, LaMDA, Meena, and/or any other generative language model, such as any other generative model that is encoder-only based, decoder-only based, sequence-to-sequence based and that optionally includes an attention mechanism or other memory. In generative model form, machine learning model(s) may have hundreds of millions, or even hundreds of billions of parameters. In some implementations, machine learning model(s) may include a multi-modal model such as a VLM and/or a visual question answering (VQA) model, which can have any of the aforementioned architectures, and which can be used to process multiple modalities of data, particularly images and text, and/or images and audio for example, to generate one or more modalities of output.
[0045] Vision language system 130 may include a sampling engine 132, VLM engine 134, one or more VLMs 135, a visual annotation engine 136, and a feedback engine 138. Any of engines 134, 136, and/or 138 may be implemented using any combination of hardware and software. Moreover, any of engines 134, 136, and/or 138 may be combined with other(s) of engines 134, 136, and/or 138. Non-limiting examples of VLM(s) 135 that may be applied as described herein include Gemini (e.g., Nano, Pro, and/or Ultra) and/or Flamingo, to name a few.
[0046] In various implementations, sampling engine 132 may be configured to sample distribution(s) of candidate actions for robot 100 to perform in furtherance of completing task(s) in environment(s) in which robot 100 operates. Sampling engine 132 may sample these candidate actions from actions spaces that may or may not be constrained based on factors such as a location of robot 100 relative to a destination robot 100 is meant to travel and/or object(s) to be acted upon by robot 100. Additionally, in some implementations, the action spaces may be constrained based on a digital visual representation of an environment in which robot 100 operates. For example, candidate actions that would occur outside of the field of view of whichever vision sensor e.g., on the robot or elsewhere in its environment) captured the digital visual representation may be excluded from or otherwise not available in the action space from which sampling engine 132 samples. Additionally, in some implementations, the action space(s) from which sampling engine 132 samples actions may be constrained based on previous candidate actions selected from previous distributions of candidate actions sampled by sampling engine 132.
[0047] Proprioception system 140 may be present in some implementations where robot 100 is being controlled using techniques described herein. Proprioception system 140 may be omitted in other circumstances. Proprioception system 140 may include a proprioception prediction process 142 and one or more proprioception machine learning models 144. Examples of a proprioception machine learning model that may be used include PaLM, PaLM-E, model(s) described in “RT-1 : Robotics Transformer for Real-World Control at Scale” (arXiv:2212.06817), and/or model(s) described in “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control” (arXiv:2307.15818).
[0048] In various implementations, proprioception prediction process 142 may process input tokens indicative of a current (or past) proprioception values of robot 100, e.g., along with other data such as data indicative of a task or action to be performed (e.g., an action sampled and selected as described herein), state data of the robot’s environment, etc., to generate robot control data and/or predict future proprioception values of robot 100. These robot control data and/or future proprioception values may be used to operate robot 100. “Robot control data” may include, for instance, low-level actuator commands (also referred to as “joint commands,” and may include torque commands) that directly control the actuators/joints 104-1 to 104-N of the robot, cartesian commands that specify directi on(s) for an end effector 106, a target robot pose, code that specifies reward functions that a motion controller can optimize (e.g., using techniques such as receding horizon optimization) to find optimal low-level actuator commands, selected predefined robot primitives, and so forth. In some cases, robot logic 102 may be configured to convert between joint commands and Cartesian commands, e.g., using forward and/or inverse kinematics.
[0049] In various implementations, a user 150 may control robot 100 using a client device 152. While depicted as a tablet computer or smart phone in Fig. 1, client device 152 may take other forms, such as a desktop or laptop computer, in-vehicle computing device, augmented reality (AR) and/or virtual reality (VR) headset or glasses, standalone “smart” speakers that host automated assistants that can be interacted with the control robot 100, etc. In various implementations, user 150 may issue one or more natural language commands, e.g., by typing the commands or uttering the commands aloud and having those spoken utterances transcribed using speech-to-text (STT) processing. These natural language commands may specify a task to be completed by robot 100 in an environment in which robot 100 operates. For example, user 150 may ask robot 100 to “pick up the helix-shaped dog chew toy,” “close the windows,” “take the dishes from the table to the sink,” etc.
[0050] Fig. 2 depicts a non-limiting example of a robot 200 in the form of a robot arm. An end effector 206 in the form of a gripper claw is removably attached to a sixth joint 204-6 of robot 200. In this example, six joints 204-1 to 204-6 are indicated. However, this is not meant to be limiting, and robots may have any number of joints. In some implementations, robot 200 may be mobile, e.g., by virtue of a wheeled base 255 or other locomotive mechanism. Robot 200 is depicted in Fig. 2 in a particular selected configuration or “pose.”
[0051] The embodiment depicted in Fig. 2 is not meant to be limiting. Techniques described herein may be implemented on a variety of different types of robots. In some implementations, a mobile manipulator performing navigation tasks may use the image from a fixed head camera and annotate the image with arrows originating from the bottom center of the image to represent the 2D action space. After the candidate action is identified in the pixel space, an on-board depth camera of the robot may be used to map the candidate action to a 3D target location and command the robot to move toward the target (e.g., with a maximum distance of 1.0m).
[0052] Additionally, in some implementations, a mobile manipulator performing manipulation tasks may use the image from a fixed head camera and annotate the image with arrows originating from the end-effector in the camera frame. Each arrow may represent a 3D relative Cartesian end-effector position (x; y; z). To handle the z-dimension height, two settings may be considered: where height is represented using color grading (e.g., a red to blue spectrum) and where the arm only uses fixed-height actions. Gripper closing actions may or may not be shown as visual annotations but instead may be expressed through text prompts.
[0053] In some implementations, other types of manipulator robots, such as robots having cameras mounted on their “wrists,” may have those images annotated with arrows that originate from the center of the camera frame. Each arrow may represent a 3D relative Cartesian endeffector position (x, , z, where the z dimension may be captured with a color spectrum, e.g., from red to blue). [0054] In some implementations, techniques described herein may be used in simulation as well, e.g., for pick and place manipulation. Image(s) from an overhead may be used, and those images may be annotated with pick and place locations.
[0055] Figs. 3 A and 3B depict how techniques described herein may be applied for two iterations when a user (e.g., 150) issues the robot command, “pick up the helix dog chew toy.” Fig. 3 A, a first visual representation 365 A, e.g., in the form of a digital image, depicts three objects: a bow-shaped chew toy 360A, a rolling pin-shaped chew toy 360B, and a helix-shaped chew toy 360C, which is the object the user wants a robot 300 to act upon. In Figs. 3A-B, the visual representation captures the environment in which robot 300 operates from a third person perspective, rather than from the perspective of a vision sensor carried by robot 300. However, this is not meant to be limiting. Techniques described herein are also applicable to visual representations captured from other perspectives, such as from the perspective of robot 300 itself.
[0056] In Fig. 3A, a first distribution of candidate actions for robot 300 to perform have been sampled, e.g., by sampling engine 132, from a first action space d1. This first action space d1 may be relatively unconstrained. For example, in some implementations, the first action space A1 may only be constrained to the space that robot 300 can reach within some amount of time. Additionally or alternatively, the first action space d1 may be constrained to the physical space that is within the field-of-view of the vision sensor that captured first visual representation 365A.
[0057] These sampled candidate actions have been projected, overlaid, and/or otherwise incorporated onto first visual representation 365A to form visual annotations A-H. In this example, each annotation takes the form of a circle representing a candidate end effector location and an arrow representing a trajectory to that location. However, other types of visual annotations are contemplated, such as other shapes, arrows that represent non-linear (e.g., curvy) trajectories, etc. Moreover, each of visual annotations A-H is a circle having the same size. In other implementations, the visual annotations may be heterogeneous in shape and/or size. For example, larger visual annotations may be rendered to represent locations that are closer to the perspective from which the visual annotation was captured, and smaller visual annotations may be rendered to represent locations that are farther away from the perspective from which the visual annotation was captured. Different colored annotations may be rendered as well, e.g., to represent locations that are closer to or farther from the perspective of the vision sensor, or to represent costs or other attributes of the underlying candidate actions.
[0058] In Fig. 3 A, the candidate actions may be selected from the first action space d1 more-or- less randomly. Consequently, the visual annotations A-H are spread relatively randomly and equally across visual representation. While some (e.g., F, G) are closer to the object (360C) to be acted upon by robot 300, others are farther away, e.g., closer to other objects and/or in random locations. In various implementations, VLM engine 134 or another component may assemble a VLM prompt that includes data indicative of visual representation 365A (e.g., embeddings generated therefrom, pixels, extracted features, etc.), and data indicative of a request such as “which location is closest to the helix-shaped dog chew toy?” VLM engine 134 may process this VLM prompt using VLM 135 to generate VLM output. The VLM output may include tokens indicative of which visual annotation is closest to the object (helix-shaped chew toy 360C) to be acted upon by robot 300. In this example, the VLM output tokens may be decoded to reveal location F to be the closest to helix-shaped dog chew toy 360C.
[0059] In Fig. 3B, a new distribution of candidate actions has been sampled, e.g., by sampling engine 132, from a new action space A. Second action space A may be more constrained than first action spaced1 and the latest sampled candidate actions may be fitted to this new more constrained action space A . For instance, second action space A may be constrained based on the first selected candidate action F. These newly sampled candidate actions may then be projected, overlaid, or otherwise incorporated into a second visual representation 365B. Second visual representation 365B may or may not be a copy of first visual representation 365A (prior to the addition of visual annotations A-H) or may be captured separately, e.g., after robot 300 has moved its end effector to a location that corresponds with the visual annotation F in first visual representation 365A. As shown in Fig. 3B, the visual annotations A’-H’ are distributed more tightly around the location where visual annotation F was in first visual representation 365A.
[0060] When second visual representation 365B is processed by VLM engine 134 using VLM 135, e.g., with a similar query as before (“which location is closest to the helix-shaped dog chew toy?”), the VLM output may identify either visual annotation E’ or visual annotation F’. At this point, both visual annotations E’/F’ are proximate enough to helix-shaped dog chew toy 360C that during the next iteration, the request that is submitted to VLM engine 134 may be different, e.g., “Is E’ or F’ better for picking up the helix-shaped dog chew toy?”
[0061] Fig. 4 depicts an example of how data may be exchanged between various components depicted in Fig. 1 to carry out selected aspects of the present disclosure. Starting at top right, a client device 452 may be operated by a user (not depicted) to receive a robot task. In various implementations, the robot task may be provided as part of a natural language statement that is typed at client device 452 or recorded and transcribed at client device 452 or elsewhere.
[0062] The robot task may be provided to sampling engine 132. Sampling engine 132 may sample an zth distribution 472 of candidate actions that may be performed by a robot 400. As noted previously, during a first (z=l) iteration, the action space from which sampling engine 132 may sample candidate actions may be relatively unconstrained, but may converge to more constrained action spaces during subsequent iterations.
[0063] In various implementations, the zth distribution of sampled candidate actions 472 and image(s) 470 captured by a camera onboard or proximate robot 400 may be provided to or obtained by visual annotation engine 136. Visual annotation engine 136 may project the zth distribution of sampled candidate actions 472 onto the image(s) 470 to generate annotated image(s) 474. Annotated image(s) 474 may then be processed by VLM engine 134 using VLM 135 to generate VLM output indicative of which visual annotation represents the “best” candidate action for robot 400 to perform, or the visual annotation that is closest to an end goal of a task assigned to robot 400.
[0064] Feedback engine 138 may evaluate the VLM output to determine whether additional iteration(s) are needed, or whether the selected visual annotation is “close enough” or “good enough” for robot 400 to perform it. If the selected candidate action is suitable — e.g., it would place an end effector of robot 400 in a good position to grasp whatever object robot 400 is supposed to grasp — then data indicative of the selected robot action may be provided to proprioception prediction process 142, which may generate robot control data that can then be used by (e.g., transmitted to) robot 400 to cause robot 400 to perform the selected robot action. [0065] If feedback engine 138 determines that the selected candidate action is not “good enough,” on the other hand, then the process may pass back to sampling engine 132. Sampling engine may now sample a new distribution z+1 of candidate actions that may be performed by a robot 400. The cyclic process depicted in Fig. 4 may then repeat as necessary until the sampled candidate actions (and the visual representations that represent them) converge onto the best candidate action(s) for robot 400 to perform to carry out the task assigned to it.
[0066] Fig. 5 A depicts an example of how data may be processed to carry out selected aspects of the present disclosure. Starting at left, a plurality of candidate actions are sampled, e.g., by sampling engine 132, from action space A(1) 560. Each of these sampled candidate actions is represented in Fig. 5A with an arrow that originates at a robot 500. Robot 500 also includes an onboard camera 558.
[0067] Following the circle arrow to the top, an image 562 captured by onboard camera 558 may be annotated, e.g., by visual annotation engine 136, with projections of the candidate actions sampled previously. In this example, seven arrows are annotated 1-7 to represent the seven sampled candidate actions, but this is not meant to be limiting. More or less annotations/candidate actions may be sampled/incorporated into image 562.
[0068] Following the circle arrow to the right side of Fig. 5 A, the annotated image 562 is used, e.g., by VLM engine 134, to query a VLM to select one or more of the sampled candidate actions. In this example, the prompt reads, “Which arrows should the robot follow to pick up the block,” but different prompts may be provided if different annotations are used, if different tasks are being performed, etc. In this example, annotations 5 and 6 are the closest to the block that is to be acted upon by robot 500.
[0069] Moving to the bottom of Fig. 5 A, a new set of candidate actions may be sampled from a now reduced action space 560’. This iterative process of sampling from increasingly narrower action spaces is demonstrated in Fig. 5B, in which the action space 560A gradually decreases in size to that of 560B, and eventually, to 560C.
[0070] The example of Figs. 5A-B depicts a robot 500 in the form of a robot arm, and an action space 560 that includes actions performable by the robot arm. However, this is not meant to be limiting. In various implementations, similar techniques may be applied in other scenarios, such as to assist in a self-driving vehicle (e.g., robot, car, truck, agricultural rover, robotic dog, etc.) navigate around and/or through obstacles. Techniques described herein may be equally applicable to navigation of unmanned aerial vehicles (UAVs).
[0071] Referring now to Fig. 6, an example method 600 of practicing selected aspects of the present disclosure is described. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, including those depicted in Fig. 1. Moreover, while operations of method 600 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted or added.
[0072] At block 602, the system, e.g., by way of sampling engine 132, may sample a first (e.g., 7=1) distribution of candidate actions for a real or simulated robot (e.g., 100, 200) to perform in furtherance of completing a task in a real or simulated environment in which the robot operates. In various implementations, the task may be derived from sources such as a robot operator’s natural language input, which may be typed or spoken and then processed using STT processing to generate textual output. The first distribution of candidate actions may be sampled from an action space A(i) (where 7=1 in this first iteration) that may be relatively unconstrained compared to action spaces from which candidate actions will be sampled subsequently.
[0073] At block 604, the system, e.g, by way of visual annotation engine 136, may incorporate, into a digital visual representation (e.g, one or more digital images 470) of the environment in which the robot operates, visual annotations (e.g., A-H and A’-H’ in Figs. 3 A-B) that represent the first distribution of candidate actions. In some implementations, this incorporation may include overlaying and/or projecting the first distribution of candidate actions onto the visual representation. For example, the visual representation may be registered with the robot’s environment, e.g., by matching one or more locations (e.g., pixels) within the visual representation with one or more corresponding locations (e.g., points) in the robot’s environment. In some implementations, the visual annotations may be transparent and/or translucent, e.g., so that underlying visual features of the visual representation are not entirely concealed. In some other implementations where the visual annotations are opaque, the sampling of block 602 may be constrained to avoid objects or other visual features of potential interest in the visual representation.
[0074] At block 606, the system, e.g., by way of feedback engine 138, may assemble a VLM input prompt that includes the digital visual representation and a request to select one or more of the candidate actions, e.g., for the robot to perform in furtherance of completing the task, and/or to facilitate sampling of a new, fitted distribution.
[0075] At block 608, the system, e.g., by way of VLM engine 134, may process the VLM input prompt using a VLM 135 to generate VLM output. This VLM output may include token(s) that are indicative of one or more selected candidate actions. [0076] Based on the VLM output generated at block 608, at block 610, the system, e.g., by way of VLM engine 134 or feedback engine 138, may select one of the candidate actions of the first distribution as a first selected candidate action. For instance, the VLM output may identify the visual annotation/candidate action that will advance the robot farthest towards the goal of completing the robot task. In some cases, this may be the visual annotation/candidate action that is closest to an object to be acted upon by the robot. In Fig. 3 A, for instance, visual annotation F was selected because it was closest to helix-shaped dog toy 360C. In Fig. 3B, visual annotation E’ or F’ may be selected because they are the closest to helix-shaped dog toy 360C.
[0077] At block 612, the system, e.g., by way of feedback engine 138, may determine whether one or more criteria are satisfied by the VLM output generated at block 608 and/or the selected candidate action of block 610. These criteria may include, for instance, whether performance of the selected candidate action will advance the robot towards completing the task assigned to it. For example, the VLM output generated at block 608 may be associated with one or more probability distributions. These probability distributions may indicate, for example, a confidence that performance of a given candidate action will result in a positive outcome, whether that outcome be advancement towards completion of the task or even final completion of the task.
[0078] If the answer at block 612 is no, then method 600 may proceed back to block 602, and a next (e.g., zth) distribution of candidate actions for the robot to perform in furtherance of completing the task may be sampled, e.g., by sampling engine 132. In various implementations, this subsequent sampling may be constrained by the visual annotation/candidate action that was selected at block 610. In Fig. 3B, for instance, the updated distribution of visual annotations A’- G’ was constrained by the location of the previously selected visual annotation F in Fig. 3 A. Method 600 may then proceed through blocks 604-610 as described previously. However, if the criteria are/is satisfied at block 612, then in some implementations, method 600 may proceed to block 614.
[0079] As shown by the dashed arrows, in some implementations, VLM 135 may be trained and/or fine-tuned based on various data generated during performance of various aspects of method 600. For example, at block 614, the system, e.g., by way of proprioception system 140, may cause robot 100/200 to perform the most-recently selected candidate action. At block 616, the system, e.g., by way of feedback engine 138 and/or VLM engine 134, may fine-tune the VLM 135 based on, for example, an outcome of the robot performing the selected candidate action. In some implementations, the outcome may be derived from human feedback about the robot operation, or from feedback provided by the robot itself.
[0080] Additionally or alternatively, in some implementations, the system may train VLM 135 without requiring robot operation. For instance, as shown by the dashed arrows leading from blocks 608 and 610 to block 616, in some implementations, VLM 135 may be trained based on, for example, VLM output and/or based on the selected candidate robot action (which need not necessarily be performed by the robot). VLM output may be evaluated by humans and/or based on a trained reward function. Likewise, the selected candidate robot action may be evaluated by humans and/or based on a trained reward function.
[0081] For example, and as noted previously, in some implementations, a human teleoperating a real or simulated robot may have a display that renders a video feed from a camera onboard the robot or elsewhere. In some implementations, this video feed may be annotated, e.g., in real time, every //th frame, every time the visual context changes, etc., with annotations representing sampled candidate actions (and in some cases, which candidate actions are most likely to make progress towards completing a task). The human may choose to teleoperate the robot in a manner consistent with the suggested annotations or not, in turn generating positive feedback or negative feedback, respectively. This feedback may be used, e.g., via techniques such as RLHF, to train a separate reward model to learn human preferences and assign rewards based on those preferences. The reward model may then be used to train (e.g., optimize/fine-tune) the VLM. [0082] Fig. 7 is a block diagram of an example computer system 710. Computer system 710 typically includes at least one processor 714 which communicates with a number of peripheral devices via bus subsystem 712. These peripheral devices may include a storage subsystem 724, including, for example, a memory subsystem 725 and a file storage subsystem 726, user interface output devices 720, user interface input devices 722, and a network interface subsystem 716. The input and output devices allow user interaction with computer system 710. Network interface subsystem 716 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
[0083] User interface input devices 722 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and/or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computer system 710 or onto a communication network.
[0084] User interface output devices 720 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from computer system 710 to the user or to another machine or computer system. [0085] Storage subsystem 724 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 724 may include the logic to perform selected aspects of method 600, and/or to implement one or more aspects of robot 100 or the various systems depicted in Fig. 1. Memory 725 used in the storage subsystem 724 can include a number of memories including a main random-access memory (RAM) 730 for storage of instructions and data during program execution and a read only memory (ROM) 732 in which fixed instructions are stored. A file storage subsystem 726 can provide persistent storage for program and data files, and may include a hard disk drive, a CD-ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 726 in the storage subsystem 724, or in other machines accessible by the processor(s) 714.
[0086] Bus subsystem 712 provides a mechanism for letting the various components and subsystems of computer system 710 communicate with each other as intended. Although bus subsystem 712 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0087] Computer system 710 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the everchanging nature of computers and networks, the description of computer system 710 depicted in Fig. 7 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer system 710 are possible having more or fewer components than the computer system depicted in Fig. 7.
[0088] In some implementations, a computer implemented method may be provided that includes: sampling a first distribution of candidate actions for a robot to perform in furtherance of completing a task in an environment in which the robot operates; incorporating, into a digital visual representation of the environment in which the robot operates, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions for the robot to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; based on the VLM output, selecting one of the candidate actions of the first distribution as a first selected candidate action; and causing the robot to perform the first selected candidate action.
[0089] In various implementations, the first distribution of candidate actions may be selected from an action space that is constrained based on the digital visual representation of the environment. In various implementations, the method may include, based on the first selected candidate action, sampling a second distribution of candidate actions for the robot to perform in furtherance of completing the task. In various implementations, the second distribution of candidate actions may be selected from an action space that is constrained based on the first selected candidate action.
[0090] In various implementations, the method may include: incorporating, into the digital visual representation, second visual annotations that represent the second distribution of candidate actions; assembling a second VLM input prompt that includes the digital visual representation and a request to select one or more of the candidate actions of the second distribution for the robot to perform in furtherance of completing the task; processing the second VLM input prompt using the VLM to generate second VLM output; based on the second VLM output, selecting one of the candidate actions of the second distribution as a second selected candidate action; and causing the robot to perform the second selected candidate action of the second distribution. In various implementations, the robot may perform both the first and second selected candidate actions subsequent to selection of the second selected candidate action. [0091] In various implementations, the digital visual representation may be generated based on sensor data generated by one or more sensors in the environment. In various implementations, the digital visual representation may be generated based on sensor data generated by one or more sensors carried by the robot. In various implementations, the sampling may include sampling a first plurality of points within the digital visual representation, wherein each candidate action of the first distribution corresponds to a respective point of the plurality of points.
[0092] In various implementations, the candidate actions of the first distribution may include one or more end effector trajectories. In various implementations, the one or more visual annotations that represent the one or more end effector trajectories may include vector arrows extending from a starting point in the digital visual representation to candidate end effector positions in the digital visual representation.
[0093] In various implementations, the digital visual representation may include a three- dimensional (3D) representation of the environment. In various implementations, the 3D representation of the environment may be a point cloud generated using a light detection and ranging (LIDAR) sensor.
[0094] In various implementations, the digital visual representation may be a two-dimensional (2D) representation of the environment. In various implementations, the visual annotations may include a first shape representing a first candidate action that would occur a first distance from a vision sensor that captured the digital visual representation, and a second shape representing a second candidate action that would occur a second distance from the vision sensor. In various implementations, the first shape may be sized differently from the second shape to convey a difference between the first and second distances. In various implementations, the first shape may be colored differently from the second shape to convey a difference between the first and second distances.
[0095] In various implementations, the visual annotations may overlay content depicted in the digital visual representation. In various implementations, the visual annotations may be at least partially transparent so that underlying content remains visible in the digital visual representation. In various implementations, the VLM input prompt may be assembled to further include a raw version of the digital visual representation without the visual annotations. [0096] In various implementations, the robot may be simulated in a virtual environment or may be a physical robot operated in a physical environment.
[0097] In another aspect, a method may be implemented using one or more processors and may include: sampling a first distribution of candidate actions for a robot to perform in furtherance of completing a task in the environment; incorporating, into a digital visual representation of an environment in which the robot operates, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions for the robot to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; and fine-tuning the VLM directly or indirectly based on the VLM output.
[0098] In various implementations, the method may include: selecting one of the candidate actions of the first distribution; and causing the robot to perform the selected candidate action; [0099] wherein the VLM is trained based at least in part on an outcome of the robot performing the selected candidate action. In various implementations, the VLM may be trained based on human feedback provided based on the VLM output. In various implementations, the VLM may be trained based on feedback provided based on the VLM output using a trained reward function.
[0100] In another aspect, a method may be implemented using one or more processors and may include: sampling a first distribution of candidate actions to be performed in furtherance of completing a task in a two-dimensional (2D) or three-dimensional (3D) environment; incorporating, into a digital visual representation of the environment, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions to be performed in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; based on the VLM output, selecting one of the candidate actions of the first distribution; and causing the selected candidate action to be performed in the environment. In various implementations, the environment may take the form of a canvas rendered as part of a graphical user interface (GUI). In various implementations, the environment may be a physical space in which a robot operates, and the candidate actions comprise candidate actions to be performed by the robot in the physical space in furtherance of completing the task.
[0101] In another aspect, a method may be implemented using one or more processors and may include: sampling a first distribution of candidate actions to perform in furtherance of completing a task in an environment; incorporating, into a digital visual representation of the environment, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; and based on the VLM output, selecting one of the candidate actions of the first distribution as a first selected candidate action.
[0102] Other implementations may include a non-transitory computer readable storage medium storing instructions executable by a processor to perform a method such as one or more of the methods described above. Yet another implementation may include a control system including memory and one or more processors operable to execute instructions, stored in the memory, to implement one or more modules or engines that, alone or collectively, perform a method such as one or more of the methods described above.
[0103] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in greater detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein.
[0104] While several implementations have been described and illustrated herein, a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein may be utilized, and each of such variations and/or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and/or configurations will depend upon the specific application or applications for which the teachings is/are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and/or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and/or methods, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.

Claims

CLAIMS What is claimed is:
1. A method implemented using one or more processors, comprising sampling a first distribution of candidate actions for a robot to perform in furtherance of completing a task in an environment in which the robot operates; incorporating, into a digital visual representation of the environment in which the robot operates, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions for the robot to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; based on the VLM output, selecting one of the candidate actions of the first distribution as a first selected candidate action; and causing the robot to perform the first selected candidate action.
2. The method of claim 1, wherein the first distribution of candidate actions are selected from an action space that is constrained based on the digital visual representation of the environment.
3. The method of claim 1 or 2, further comprising, based on the first selected candidate action, sampling a second distribution of candidate actions for the robot to perform in furtherance of completing the task.
4. The method of claim 3, wherein the second distribution of candidate actions are selected from an action space that is constrained based on the first selected candidate action.
5. The method of claim 4, further comprising: incorporating, into the digital visual representation, second visual annotations that represent the second distribution of candidate actions; assembling a second VLM input prompt that includes the digital visual representation and a request to select one or more of the candidate actions of the second distribution for the robot to perform in furtherance of completing the task; processing the second VLM input prompt using the VLM to generate second VLM output; based on the second VLM output, selecting one of the candidate actions of the second distribution as a second selected candidate action; and causing the robot to perform the second selected candidate action of the second distribution.
6. The method of claim 5, wherein the robot performs both the first and second selected candidate actions subsequent to selection of the second selected candidate action.
7. The method of any of the preceding claims, wherein the digital visual representation is generated based on sensor data generated by one or more sensors in the environment.
8. The method of any of the preceding claims, wherein the digital visual representation is generated based on sensor data generated by one or more sensors carried by the robot.
9. The method of any of the preceding claims, wherein the sampling comprises sampling a first plurality of points within the digital visual representation, wherein each candidate action of the first distribution corresponds to a respective point of the plurality of points.
10. The method of any of the preceding claims, wherein the candidate actions of the first distribution include one or more end effector trajectories.
11. The method of claim 10, wherein the one or more visual annotations that represent the one or more end effector trajectories comprise vector arrows extending from a starting point in the digital visual representation to candidate end effector positions in the digital visual representation.
12. The method of any of the preceding claims, wherein the digital visual representation comprises a three-dimensional (3D) representation of the environment.
13. The method of claim 12, wherein the three-dimensional representation of the environment comprises a point cloud generated using a light detection and ranging (LIDAR) sensor.
14. The method of any of the preceding claims, wherein the digital visual representation comprises a two-dimensional (2D) representation of the environment.
15. The method of claim 14, wherein the visual annotations include a first shape representing a first candidate action that would occur a first distance from a vision sensor that captured the digital visual representation, and a second shape representing a second candidate action that would occur a second distance from the vision sensor.
16. The method of claim 15, wherein the first shape is sized differently from the second shape to convey a difference between the first and second distances.
17. The method of claim 15 or 16, wherein the first shape is colored differently from the second shape to convey a difference between the first and second distances.
18. The method of any of the preceding claims, wherein the visual annotations overlay content depicted in the digital visual representation.
19. The method of claim 18, wherein the visual annotations are at least partially transparent so that underlying content remains visible in the digital visual representation.
20. The method of claim 18 or 19, wherein the VLM input prompt is assembled to further include a raw version of the digital visual representation without the visual annotations.
21. The method of any of the preceding claims, wherein the robot is simulated in a virtual environment.
22. The method of any of the preceding claims, wherein the robot is a physical robot operated in a physical environment.
23. A method implemented using one or more processors, comprising: sampling a first distribution of candidate actions for a robot to perform in furtherance of completing a task in the environment; incorporating, into a digital visual representation of an environment in which the robot operates, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions for the robot to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; and fine-tuning the VLM directly or indirectly based on the VLM output.
24. The method of claim 23, further comprising: selecting one of the candidate actions of the first distribution; and causing the robot to perform the selected candidate action; wherein the VLM is trained based at least in part on an outcome of the robot performing the selected candidate action.
25. The method of claim 23 or 24, wherein the VLM is trained based on human feedback provided based on the VLM output.
26. The method of any of claims 23-25, wherein the VLM is trained based on feedback provided based on the VLM output using a trained reward function.
27. A method implemented using one or more processors, comprising: sampling a first distribution of candidate actions to be performed in furtherance of completing a task in a two-dimensional (2D) or three-dimensional (3D) environment; incorporating, into a digital visual representation of the environment, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions to be performed in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; based on the VLM output, selecting one of the candidate actions of the first distribution; and causing the selected candidate action to be performed in the environment.
28. The method of claim 27, wherein the environment comprises a canvas rendered as part of a graphical user interface (GUI).
29. The method of claim 27 or 28, wherein the environment comprises a physical space in which a robot operates, and the candidate actions comprise candidate actions to be performed by the robot in the physical space in furtherance of completing the task.
30. A method implemented using one or more processors, comprising: sampling a first distribution of candidate actions to perform in furtherance of completing a task in an environment; incorporating, into a digital visual representation of the environment, visual annotations that represent the first distribution of candidate actions; assembling a vision-language model (VLM) input prompt that includes the digital visual representation and a request to select one or more of the candidate actions to perform in furtherance of completing the task; processing the VLM input prompt using a VLM to generate VLM output; and based on the VLM output, selecting one of the candidate actions of the first distribution as a first selected candidate action.
31. A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform any of the methods of claims 1-30.
32. At least one transitory or non-transitory computer-readable medium comprising instructions that, in response to execution by one or more processors, cause the one or more processors to perform any of the methods of claims 1-30.
EP25708274.3A 2024-02-01 2025-01-30 Visual prompting for few shot control Pending EP4676696A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202463548751P 2024-02-01 2024-02-01
PCT/US2025/013723 WO2025165948A1 (en) 2024-02-01 2025-01-30 Visual prompting for few shot control

Publications (1)

Publication Number Publication Date
EP4676696A1 true EP4676696A1 (en) 2026-01-14

Family

ID=94771936

Family Applications (1)

Application Number Title Priority Date Filing Date
EP25708274.3A Pending EP4676696A1 (en) 2024-02-01 2025-01-30 Visual prompting for few shot control

Country Status (2)

Country Link
EP (1) EP4676696A1 (en)
WO (1) WO2025165948A1 (en)

Also Published As

Publication number Publication date
WO2025165948A1 (en) 2025-08-07

Similar Documents

Publication Publication Date Title
US11554483B2 (en) Robotic grasping prediction using neural networks and geometry aware object representation
US10853646B1 (en) Generating and utilizing spatial affordances for an object in robotics applications
US12226920B2 (en) System(s) and method(s) of using imitation learning in training and refining robotic control policies
US12112494B2 (en) Robotic manipulation using domain-invariant 3D representations predicted from 2.5D vision data
JP2023164459A (en) Efficient robot control based on inputs from remote client devices
US11887363B2 (en) Training a deep neural network model to generate rich object-centric embeddings of robotic vision data
EP3769263A1 (en) Controlling a robot based on free-form natural language input
US10864633B2 (en) Automated personalized feedback for interactive learning applications
JP2019508273A (en) Deep-layer machine learning method and apparatus for grasping a robot
US20250178615A1 (en) Robot navigation using a high-level policy model and a trained low-level policy model
US12168296B1 (en) Re-simulation of recorded episodes
WO2025165948A1 (en) Visual prompting for few shot control
US20260077507A1 (en) Using affordance plans for robot control
US20250353169A1 (en) Semi-supervised learning of robot control policies
US20250312914A1 (en) Transformer diffusion for robotic task learning
US20260094437A1 (en) Systems and methods for task progress estimation using a generative model with shuffled video inputs
US12377536B1 (en) Imitation robot control stack models
US11654550B1 (en) Single iteration, multiple permutation robot simulation
US20260077489A1 (en) Robot learning through retrieval and self improvement
Gwozdz et al. Enabling semi-autonomous manipulation on iRobot’s Packbot
WO2026075823A1 (en) Adapting video generation models using generative model feedback

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251008

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR