WO2019079598A1 - Probabilistic object models for robust, repeatable pick-and-place - Google Patents
Probabilistic object models for robust, repeatable pick-and-place Download PDFInfo
- Publication number
- WO2019079598A1 WO2019079598A1 PCT/US2018/056514 US2018056514W WO2019079598A1 WO 2019079598 A1 WO2019079598 A1 WO 2019079598A1 US 2018056514 W US2018056514 W US 2018056514W WO 2019079598 A1 WO2019079598 A1 WO 2019079598A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- probabilistic
- objects
- observed
- creating
- object model
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/60—Type of objects
- G06V20/64—Three-dimensional [3D] objects
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B25—HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
- B25J—MANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
- B25J19/00—Accessories fitted to manipulators, e.g. for monitoring, for viewing; Safety devices combined with or specially adapted for use in connection with manipulators
- B25J19/02—Sensing devices
- B25J19/021—Optical sensing devices
- B25J19/023—Optical sensing devices including video camera means
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B25—HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
- B25J—MANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
- B25J9/00—Program-controlled manipulators
- B25J9/02—Program-controlled manipulators characterised by movement of the arms, e.g. cartesian coordinate type
- B25J9/04—Program-controlled manipulators characterised by movement of the arms, e.g. cartesian coordinate type by rotating at least one arm, excluding the head movement itself, e.g. cylindrical coordinate type or polar coordinate type
- B25J9/041—Cylindrical coordinate type
- B25J9/042—Cylindrical coordinate type comprising an articulated arm
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B25—HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
- B25J—MANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
- B25J9/00—Program-controlled manipulators
- B25J9/16—Program controls
- B25J9/1612—Program controls characterised by the hand, wrist, grip control
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B25—HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
- B25J—MANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
- B25J9/00—Program-controlled manipulators
- B25J9/16—Program controls
- B25J9/1628—Program controls characterised by the control loop
- B25J9/163—Program controls characterised by the control loop learning, adaptive, model based, rule based expert control
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/29—Graphical models, e.g. Bayesian networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N7/00—Computing arrangements based on specific mathematical models
- G06N7/01—Probabilistic graphical models, e.g. probabilistic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/70—Determining position or orientation of objects or cameras
- G06T7/73—Determining position or orientation of objects or cameras using feature-based methods
- G06T7/74—Determining position or orientation of objects or cameras using feature-based methods involving reference images or patches
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/772—Determining representative reference patterns, e.g. averaging or distorting patterns; Generating dictionaries
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/84—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using probabilistic graphical models from image or video features, e.g. Markov models or Bayesian networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20076—Probabilistic image processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30108—Industrial image inspection
- G06T2207/30164—Workpiece; Machine component
Definitions
- the present invention relates generally to object recognition and manipulation, and more particularly to probabilistic object models for robust, repeatable pick-and-place.
- the invention features a method including, as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
- the invention features a method including using light fields to generate a probabilistic generative model for objects, enabling a robot to use all information from a camera to achieve precision.
- the invention features a system including a robot arm, multiple cameras locations, and a process for creating a probabilistic object model to identify, localize, and manipulate an object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across the multiple camera locations.
- FIG. 1 illustrates how an exemplary model enables a robot to learn to robustly detect, localize and manipulate objects.
- FIG. 2 illustrates an exemplary probabilistic object map model for reasoning about objects.
- FIG. 3 illustrates exemplary objects.
- FIG. 4 illustrates an exemplary observed view for a scene.
- FIG. 5 illustrates automatically generated thumbnails and observed views for two configurations detected for a yellow square duplo.
- Robust object perception is a core capability for a manipulator robot.
- Current perception techniques do not reach levels of precision necessary for a home robot, which might be requested to pick hundreds of times each day and where picking errors might be very costly, resulting in broken objects and ultimately, in the person's rejection of the robot.
- the present invention enables the robot to learn a model of an object using light fields, achieving very high levels of robustness when detecting, localizing, and manipulating objects.
- Existing approaches use features, which discard information, because inference in a full generative model is intractable.
- our method uses light fields to enable efficient inference for object detection and localization, while incorporating information from every pixel observed from across multiple camera locations.
- the robot can segment objects, identify object instances, localize them, and extract three dimensional (3D) structure.
- our method enables a Baxter robot to pick an object hundreds of times in a row without failures, and to
- the robot can identify, localize and grasp objects with high accuracy using our framework; modeling one side of a novel object takes a Baxter robot approximately 30 seconds and enables detection and localization to within 2 millimeters. Furthermore, the robot can merge object models to create metrically grounded models for object categories, in order to improve accuracy on previously unencountered objects.
- a robot can automatically acquire a POM by exploiting its ability to change the environment in service of its perceptual goals. This approach enables the robot to obtain extremely high accuracy and repeatability at localizing and picking objects. Additionally, a robot can create metrically grounded models for object categories by merging POMs.
- the present invention demonstrates that a Baxter robot can autonomously create models for objects. These models enable the robot to detect, localize, and manipulate the objects with very high reliability and repeatability, localizing to within 2 mm for hundreds of successful picks in a row. Our system and method enable a Baxter robot to autonomously map objects for many hours using its wrist camera.
- a goal is for a robot to estimate objects in a scene, given observations of camera images, and associated camera poses,
- Each object instance, o N consists of a pose, x ⁇ , along with an index, tf 1 , which identifies the object type.
- Each object type, O k is a set of appearance modes. We rewrite the distribution using Bayes' rule:
- This model is known as inverse graphics because it requires assigning a probability to images given a model of the scene.
- Equation 6 corresponds to light field rendering.
- the second term corresponds to the light field distribution in a synthetic photograph given a model of the scene as objects.
- each object has only a few configurations, c ⁇ C, at which it may lie at rest. This assumption is not true in general; for example, a ball has an infinite number of such configurations. However many objects with continuously varying families of stable poses can be approximated using a few sampled configurations. Furthermore this assumption leads to straightforward representations for POMs.
- the appearance model, A c for an object is a model of the light expected to be observed from the object by a camera.
- Algorithm 1 Render the observed view from images and poses at a particular plane, z.
- Algorithm 2 Render the predicted view given a scene with objects and appearance models.
- Algorithm 4 Infer object.
- Equation 6 Using Equation 6 to find the objects that maximize a scene is still challenging because it requires integrating over all values for the variances of the map cells and searching over a full 3D configuration of all objects. Instead we approximate it by performing inference in the space of light field maps, m and integral finding the map m and scene (objects) that maximizes the
- This m * is the observed view.
- Algorithm 1 gives a formal description of how to compute it from calibrated camera images. This view can be computed by iterating through each image one time, linear in the number of pixels or light rays. For a given scene (i.e., object configuration and appearance models), we can compute an m that maximizes the second term in Equation 6.
- This ⁇ m is the predicted view; an example predicted view is shown in FIG. 4. This computation corresponds to a rendering process.
- Our model enables us to render compositionally over object maps and match in the 2D configuration space with three degrees of freedom instead of the 3D space with six.
- the system computes the observed map and then incrementally adds objects to the scene until the discrepancy is less than a predefined threshold. At this point, all of the discrepancy has been accounted for, and the robot can use this information to ground natural language referring expressions such as "next to the bowl,” to find empty space in the scene where objects can be placed, and to infer object pose for grasping.
- Modeling objects requires estimating O k for each object k, including the appearance models for each stable configuration.
- the robot creates a model for the input pile and mapping space.
- the robot picks an object from the input pile, moves it to the mapping space, maps the object, then moves it to the output pile.
- the robot To pick from the input pile before an object map has been acquired, the robot tries to pick repeatedly with generic grasp detectors. To propose grasps, the robot looks for patterns of discrepancy between the background map and the observed map. Once successful, it places the object in the mapping workspace and creates a model for the mapping workspace with the object. Regions in this model which are discrepant with the background are used to create an appearance model A for the object. After modeling has been completed, the robot clears the mapping workspace and moves on to the next object. Note that there is a trade off between object throughput and information acquired about each object; to increase throughput we can reduce the number of pick attempts during mapping. In contrast, to truly master an object, the robot might try 100 or more picks as well as other exploratory actions before moving on.
- the robot If the robot loses the object during mapping (for example because it rolls out of the workspace), the robot returns to the input pile to take the next object. If the robot is unable to grasp the current object to move it out of the mapping workspace, it uses nudging behaviors to push the object out. If its nudging behaviors fail to clear the mapping workspace, it simply updates its background model and then continues mapping objects from the input pile.
- the object model, O k is known, but the configuration, c is unknown.
- the robot needs to decide when it has encountered a new object configuration given the collection of maps it has already made. We use an approximation of a maximum likelihood estimate with a geometric prior to decide when to make a new configuration for the object.
- POMs Once POMs have been acquired, they can be merged to form models for object categories. Many standard computer vision approaches can be applied to light field views rather than images; for example deformable parts models for classification in light field space. These techniques may perform better on the light field view because they have access to metric information, variance in appearance models, as well as information about different configurations and transition information.
- Algorithm 4 we created models for several object categories. First our robot mapped a number of object instances. Then for each group of instances, we successively applied the optimization in Algorithm 4, with a bias to force the localization to overlap with the model as much as possible. This process results in a synthetic photograph for each category that is strongly biased by object shape and contains noisy information about color and internal structure. Examples of learned object categories appear in FIG. 3. We used our learned models to successfully detect novel instances of each category from a scene with distractor objects. We aim to train on larger data sets and apply more sophisticated modeling techniques to learn object categories and parts from large datasets of POMs.
- the grid cell size is 0:25 cm, and the total size of the synthetic photograph was approximately 30 cm x 30 cm.
- our approach enables a robot to adapt to the specific objects it finds and robustly pick objects many times in a row without failures.
- This robustness and repeatability outperforms existing approaches in terms of its precision by trading off recall, enabling the robot to robustly pick objects many times in a row.
- Our approach uses light fields to create a probabilistic generative model for objects, enabling the robots to use all information from the camera to achieve this high precision.
- learned models can be combined to form models for object category that are metrically grounded and can be used to perform localization and grasp prediction.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Mechanical Engineering (AREA)
- Robotics (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Computing Systems (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Probability & Statistics with Applications (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- Orthopedic Medicine & Surgery (AREA)
- Pure & Applied Mathematics (AREA)
- Algebra (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Mathematical Physics (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Image Analysis (AREA)
- Image Processing (AREA)
- Manipulator (AREA)
Abstract
A method includes, as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
Description
Probabilistic Object Models for
Robust, Repeatable Pick-and-Place
STATEMENT REGARDING GOVERNMENT INTEREST
[001] None.
CROSS REFERENCE TO RELATED APPLICATIONS
[002] This application claims benefit from U.S. Provisional Patent Application Serial No. 62/573,890, filed October 18, 2017, which is incorporated by reference in its entirety.
BACKGROUND OF THE INVENTION
[003] The present invention relates generally to object recognition and manipulation, and more particularly to probabilistic object models for robust, repeatable pick-and-place.
[004] In general, most robots cannot pick up most objects most of the time. Yet for effective human-robot collaboration, a robot must be able to detect, localize, and manipulate the specific objects that a person cares about. For example, a household robot should be able to effectively respond to a person's commands such as "Get me a glass of water in my favorite mug," or "Clean up the workshop." For these tasks, extremely high reliability is needed; if a robot can pick at 95% accuracy, it will still fail one in twenty times. Considering it might be doing hundreds of picks each day, this level of performance will not be
acceptable for an end-to-end system, because the robot will miss objects each day, potentially breaking a person's things.
SUMMARY OF THE INVENTION
[005] The following presents a simplified summary of the innovation in order to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview of the invention. It is intended to neither identify key or critical elements of the invention nor delineate the scope of the invention. Its sole purpose is to present some concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later.
[006] In an aspect, the invention features a method including, as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
[007] In another aspect, the invention features a method including using light fields to generate a probabilistic generative model for objects, enabling a robot to use all information from a camera to achieve precision.
[008] In still another aspect, the invention features a system including a robot arm, multiple cameras locations, and a process for creating a probabilistic object model to identify, localize, and manipulate an object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across the multiple camera locations.
[009] These and other features and advantages will be apparent from a reading of the following detailed description and a review of the associated drawings. It is to be understood that both the foregoing general description and the following detailed description are explanatory only and are not restrictive of aspects as claimed.
BRIEF DESCRIPTION OF THE DRAWINGS
[0010] These and other features, aspects, and advantages of the present invention will become better understood with reference to the following description, appended claims, and accompanying drawings where:
[0011] FIG. 1 illustrates how an exemplary model enables a robot to learn to robustly detect, localize and manipulate objects.
[0012] FIG. 2 illustrates an exemplary probabilistic object map model for reasoning about objects.
[0013] FIG. 3 illustrates exemplary objects.
[0014] FIG. 4 illustrates an exemplary observed view for a scene.
[0015] FIG. 5 illustrates automatically generated thumbnails and observed views for two configurations detected for a yellow square duplo.
DETAILED DESCRIPTION
[0016] The subject innovation is now described with reference to the drawings, wherein like reference numerals are used to refer to like elements throughout. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It may be evident, however, that the present invention may be practiced without these specific details. In other instances, well- known structures and devices are shown in block diagram form in order to facilitate describing the present invention.
[0017] Robust object perception is a core capability for a manipulator robot. Current perception techniques do not reach levels of precision necessary for a home robot, which might be requested to pick hundreds of times each day and where picking errors might be very costly, resulting in broken objects and ultimately, in the person's rejection of the robot. To address this problem, the present invention enables the robot to learn a model of an object using light fields, achieving very high levels of robustness when detecting, localizing, and manipulating objects. We present a generative model for object appearance and pose given observed camera images. Existing approaches use features, which discard information, because inference in a full generative model is intractable. We present a method that uses light fields to enable efficient inference for object detection and localization, while
incorporating information from every pixel observed from across multiple camera locations. Using the learned model, the robot can segment objects, identify object instances, localize them, and extract three dimensional (3D) structure. For example, our method enables a Baxter robot to pick an object hundreds of times in a row without failures, and to
autonomously create models of objects for many hours at a time. The robot can identify, localize and grasp objects with high accuracy using our framework; modeling one side of a novel object takes a Baxter robot approximately 30 seconds and enables detection and localization to within 2 millimeters. Furthermore, the robot can merge object models to create metrically grounded models for object categories, in order to improve accuracy on previously unencountered objects.
[0018] Existing systems for generic object detection and manipulation have lower accuracy than instance-based methods and cannot adapt on the fly to objects that do not work the first time. Instance based approaches can have high accuracy but require object models to be provided in advance; acquiring models is a time consuming process that is difficult for even an expert to perform. For example, the winning team from the Amazon Picking
Challenge in 2015 used a system that required specific instance-based data to do pose estimation and still found that perception was a major source of system errors. Autonomously learning 3D object models requires expensive ICP-based methods to localize objects, which are expensive to globally optimize, making them impractical for classification.
[0019] To achieve extremely high accuracy picking, we reframe the problem to one of adaptation and learning. When the robot encounters a novel object, its goal is to create a model that contains the information needed to identify, localize, and manipulate that object with extremely high accuracy and robustness. We term this information a probabilistic object model (POM). After learning a POM, the robot is able to robustly interact with the object, as shown in FIG. 1. We represent POMs using a pixel-based inverse graphics approach based around light fields. Unlike feature-based methods, pixel based inverse graphics methods use all information obtained from a camera. Inference in inverse-graphics methods is
computationally expensive because the program must search a very large space of possible scenes and marginalize over all possible images. Using the light field approach, a robot can automatically acquire a POM by exploiting its ability to change the environment in service of its perceptual goals. This approach enables the robot to obtain extremely high accuracy and
repeatability at localizing and picking objects. Additionally, a robot can create metrically grounded models for object categories by merging POMs.
[0020] The present invention demonstrates that a Baxter robot can autonomously create models for objects. These models enable the robot to detect, localize, and manipulate the objects with very high reliability and repeatability, localizing to within 2 mm for hundreds of successful picks in a row. Our system and method enable a Baxter robot to autonomously map objects for many hours using its wrist camera.
[0021] We present a probabilistic graphical model for object detection, localization, and modeling. We first describe the model, then inference in the model. The graphical model is depicted in FIG. 2, while a table of variables appears in Table I.
[0022] Probabilistic Obj ect Model
[0023] A goal is for a robot to estimate objects in a scene,
given observations of camera images,
and associated camera poses,
[0024] Each object instance, oN, consists of a pose, x^, along with an index, tf1, which identifies the object type. Each object type, Ok, is a set of appearance modes. We rewrite the distribution using Bayes' rule:
[0025] We assume a uniform prior over object locations and appearances, so that only the first term matters in the optimization.
[0026] A generative, inverse graphics approach would assume each image is independent given the object locations, since if we know the true locations and appearances of the objects, we can predict the contents of the images using graphics:
[0027] This model is known as inverse graphics because it requires assigning a probability to images given a model of the scene. However, inference under these
assumptions is intractable because of the very large space of possible images. Existing approaches turn to features such as SIFT or learned features using neural networks, but these approaches throw away information in order to achieve more generality. Instead we aim to use all information to achieve the most precise localization and manipulation possible.
[0028] To solve this problem, we introduce the light field grid, m as a latent variable. We define a distribution over the light rays emitted from a scene. Using this generative model, we can then perform inferences about the objects in the scene, conditioned on observed light rays, R, where each ray corresponds to a pixel in one of the images, Zh, combined with a camera calibration function that maps from pixel space to a light ray given camera pose. We define a synthetic photograph, m, as an L x W array of cells in a plane in space. Each cell (/, w) e m has a height z and scatters light at its (x, y, z) location. We assume each observed light ray arose from a particular cell (/, w), so that the parameters associated with each cell include its height z and a model of the intensity of light emitted from that cell.
[0029] We integrate over m as a latent variable:
[0030] Then we factor this distribution assuming images are independent given the light field model m:
[0031] Here is the bundle of rays that arose from map cell (l,w), which can be
determined finding all rays that intersect the cell using the calibration function. This factorization assumes that each bundle of rays is conditionally independent given the cell parameters, an assumption valid for cells that do not actively emit light. We can render m as an image by showing the values for
as the pixel color; however variance information μ /w is also stored. FIG. 4 shows example scenes rendered using this model. The first term in Equation 6 corresponds to light field rendering. The second term corresponds to the light field distribution in a synthetic photograph given a model of the scene as objects.
[0032] We assume each object has only a few configurations, c ε C, at which it may lie at rest. This assumption is not true in general; for example, a ball has an infinite number of such configurations. However many objects with continuously varying families of stable poses can be approximated using a few sampled configurations. Furthermore this assumption leads to straightforward representations for POMs. In particular, the appearance model, Ac, for an object is a model of the light expected to be observed from the object by a camera. At both modeling time and inference time, we render a synthetic photograph of the object from a canonical view (e.g., top down or head-on), enabling efficient inference in the a lower- dimensional subspace instead of being required to do full 3D inference as in ICP. This lower- dimensional subspace enables much deeper and finer-grained search so that we can use all information from pixels to perform very accurate pose estimation.
[0033] Inference
[0034] Algorithm 1 Render the observed view from images and poses at a particular plane, z.
[0035] Algorithm 2 Render the predicted view given a scene with objects and appearance models.
[0036] Algorithm 3 Infer scene.
[0037] Algorithm 4 Infer object.
[0038] Using Equation 6 to find the objects that maximize a scene is still challenging because it requires integrating over all values for the variances of the map cells and searching over a full 3D configuration of all objects. Instead we approximate it by performing inference in the space of light field maps, m and integral finding the map m and scene (objects) that maximizes the
integral.
[0039] First, we find the value for m that maximizes the likelihood of the observed images in the first term:
[0040] This m* is the observed view. We compute the maximum likelihood estimate from image data analytically by finding the sample mean and variance for pixel values observed at each cell using the calibration function. An example observed view is shown in FIG. 4. We use the approximation that each ray is assigned to the cell it intersects without reasoning about obstructions. Algorithm 1 gives a formal description of how to compute it from calibrated camera images. This view can be computed by iterating through each image one time, linear in the number of pixels or light rays. For a given scene (i.e., object configuration and appearance models), we can compute an m that maximizes the second term in Equation 6.
[0041] This Λm is the predicted view; an example predicted view is shown in FIG. 4. This computation corresponds to a rendering process. Our model enables us to render compositionally over object maps and match in the 2D configuration space with three degrees of freedom instead of the 3D space with six.
[0042] To maximize the product over scenes o°...oN, we can compute the discrepancy between the observed view and predicted view, shown in FIG. 4. Finding the configuration of objects that minimizes this discrepancy corresponds to maximizing the log-likelihood of the scene under the observed images. Additionally, the robot can use this discrepancy as a mask for object segmentation by mapping a region, adding an object to the scene, and then remapping the region; discrepant regions correspond to cells associated with the new object. Algorithm 2 gives a description of how to compute it given a scene defined as object poses and appearance models.
[0043] To infer a scene, the system computes the observed map and then incrementally adds objects to the scene until the discrepancy is less than a predefined threshold. At this point, all of the discrepancy has been accounted for, and the robot can use this information to ground natural language referring expressions such as "next to the bowl," to find empty space in the scene where objects can be placed, and to infer object pose for grasping.
[0044] Learning Probabilistic Object Models
[0045] Modeling objects requires estimating Ok for each object k, including the appearance models for each stable configuration. We have created a system that enables a Baxter robot to autonomously map objects for hours at a time. We first divide the workspace for one arm of the robot into three regions: an input pile, a mapping space, and an output pile. The robot creates a model for the input pile and mapping space. Then a person adds objects to the input pile, and autonomous mapping begins. The robot picks an object from the input pile, moves it to the mapping space, maps the object, then moves it to the output pile.
[0046] To pick from the input pile before an object map has been acquired, the robot tries to pick repeatedly with generic grasp detectors. To propose grasps, the robot looks for patterns of discrepancy between the background map and the observed map. Once successful,
it places the object in the mapping workspace and creates a model for the mapping workspace with the object. Regions in this model which are discrepant with the background are used to create an appearance model A for the object. After modeling has been completed, the robot clears the mapping workspace and moves on to the next object. Note that there is a trade off between object throughput and information acquired about each object; to increase throughput we can reduce the number of pick attempts during mapping. In contrast, to truly master an object, the robot might try 100 or more picks as well as other exploratory actions before moving on.
[0047] If the robot loses the object during mapping (for example because it rolls out of the workspace), the robot returns to the input pile to take the next object. If the robot is unable to grasp the current object to move it out of the mapping workspace, it uses nudging behaviors to push the object out. If its nudging behaviors fail to clear the mapping workspace, it simply updates its background model and then continues mapping objects from the input pile.
[0048] Inferring a New Configuration
[0049] After a new object has been placed in the workspace, the object model, Ok is known, but the configuration, c is unknown. The robot needs to decide when it has encountered a new object configuration given the collection of maps it has already made. We use an approximation of a maximum likelihood estimate with a geometric prior to decide when to make a new configuration for the object.
[0050] Learning Object Categories
[0051] Once POMs have been acquired, they can be merged to form models for object categories. Many standard computer vision approaches can be applied to light field views rather than images; for example deformable parts models for classification in light field space. These techniques may perform better on the light field view because they have access to metric information, variance in appearance models, as well as information about different configurations and transition information. As a proof of concept, we created models for several object categories. First our robot mapped a number of object instances. Then for each group of instances, we successively applied the optimization in Algorithm 4, with a bias to force the localization to overlap with the model as much as possible. This process results in a synthetic photograph for each category that is strongly biased by object shape and contains
noisy information about color and internal structure. Examples of learned object categories appear in FIG. 3. We used our learned models to successfully detect novel instances of each category from a scene with distractor objects. We aim to train on larger data sets and apply more sophisticated modeling techniques to learn object categories and parts from large datasets of POMs.
[0052] Evaluation
[0053] We evaluate our model's ability to detect, localize, and manipulate objects using the Baxter robot. We selected a subset of YCB objects that were rigid and pickable by Baxter with the grippers in the 6cm position as well as a standard ICRA duckie. In our
implementation the grid cell size is 0:25 cm, and the total size of the synthetic photograph was approximately 30 cm x 30 cm. We initialize the background variance to a higher value to account for changes in lighting and shadows.
[0054] Localization
[0055] To evaluate our model's ability to localize objects, we find the error of its position estimates by servoing repeatedly to the same location. For each trial, we moved the arm directly above the object, then moved to a random position and orientation within 10 cm of the true location. Next we estimated the object's position by serving: first we created a light field model at the arm's current location; then we used Algorithm 3 to estimate the object's position; then we moved the arm to the estimated position and repeated. We performed five trials in each location, then moved the object to a new location, for a total of 25 trials per object. We take the mean location estimated over the five trials as the object's true location, and report the mean distance from this location as well as 95% confidence intervals. This test records the repeatability of the servoing and pose estimation; if we are performing accurate pose estimation, then the system should find the object at the same place each time. Results appear in Table II. Our results show that using POMs, the system can localize objects to within 2 mm. We observe more error on taller objects such as the mustard, and the taller duplo structure, due to our assumption that all cells are at table height. Note that even on these objects, localization is accurate to within a centimeter, enough to pick reliably with many grippers; similarly detection accuracy is also quite high. To assess the effect of correcting for z, we computed new models for the tall yellow square duplo using the two different approaches. We found that the error reduced to 0.0013m+-2.0xl0"05 using the
maximum likelihood estimate and to 0.0019m+-1.9xl0"05 using the marginal estimate. Both methods demonstrate a significant improvement. The maximum likelihood estimate performs slightly better, perhaps because the sharper edges lead to more consistent performance.
Computing z corrections takes significant time, so we do not use it for the rest of the evaluation.
[0056] Autonomous Classification and Grasp Model Acquisition
[0057] After using our autonomous process to map our test objects, we evaluated object classification and picking performance. Due to the processing time to infer z, we used z = table for this evaluation. The robot had to identify the object type, localize the object, and then grasp it. After each grasp, it placed the object in a random position and orientation. We report accuracy at labeling the object with the correct type along with its pick success rate over the ten trials in Table II. The robot discovered 1 configuration for most objects, but for the yellow square duplo discovered a second configuration, shown in FIG. 5. We explore more deliberate elicitation of new object configurations, by rotating the hand before dropping the object or by employing bimanual manipulation. We report detection and pick accuracy for the 10 objects.
[0058] Our results show 98% accuracy at detection performance for these objects. The duplo yellow square was confused with the standard ICRA duckie which is similarly colored and sized. Other errors were due to taller objects. The robot mapped the tall duplo structure in a lying down position. It was unable to pick it when it was standing up to move it to its scanning workspace because of the error introduced by the height. After a few attempts it knocked it down; the lower height introduced less error, and it was able to pick and localize
perfectly. The padlock is challenging because it is both heavy and reflective. Also its smallest dimension just barely fits into the robot's gripper, meaning that very small amounts of position error can cause the grasp to fail. Overall our system is able to pick this data set 84% of the time.
[0059] Our automatic process successfully mapped all objects except for the mustard. The mustard is a particularly challenging object due to its height and weight; therefore we manually created an model and annotated a grasp. Our initial experiments with this model resulted in 1=10 picks due to noise from its height and its weight; however we were still able to use it for localization and detection. Next we created a new model using marginal z corrections and also performed z corrections at inference time. Additionally we changed to a different gripper configuration more appropriate to this very large object. After these changes, we were able to pick the mustard 10/10 times.
[0060] Picking Robustly
[0061] We can pick a spoon 100 times in a row.
[0062] In summary, our approach enables a robot to adapt to the specific objects it finds and robustly pick objects many times in a row without failures. This robustness and repeatability outperforms existing approaches in terms of its precision by trading off recall, enabling the robot to robustly pick objects many times in a row. Our approach uses light fields to create a probabilistic generative model for objects, enabling the robots to use all information from the camera to achieve this high precision. Additionally, learned models can be combined to form models for object category that are metrically grounded and can be used to perform localization and grasp prediction.
[0063] We can scale up this system so that the robot can come equipped with a large database of object models. This system enables the robot to automatically detect and localize novel objects. If the robot cannot pick the first time, it will automatically add a new model to its database, enabling it to increase its precision and robustness. Additionally this new model will augment the database, improving performance on novel objects.
[0064] Extending the model to full 3D perception enables it to fuse different stable configurations. We can track objects continuously over time, create a Bayes' filtering approach to object tracking. This model takes into account object affordances and actions on objects, creating a full object-oriented MDP. Ultimately objects like doors, elevators, and
drawers can be modeled in an instance-based way, and then generalized to novel instances. This model is ideal for connecting to language because it factors the world into objects, just as people do when they talk about them.
[0065] It would be appreciated by those skilled in the art that various changes and modifications can be made to the illustrated embodiments without departing from the spirit of the present invention. All such modifications and changes are intended to be within the scope of the present invention except as limited by the scope of the appended claims.
Claims
1. A method comprising:
as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
2. The method of claim 1 wherein creating the probabilistic object model comprises rendering an observed view from images and poses at a particular plane.
3. The method of claim 2 wherein rendering the observed view from images and poses at the particular plane comprises computing a maximum likelihood estimate from image data by finding a sample mean and variance for pixel values observed at each cell using a calibration function.
4. The method of claim 3 wherein creating the probabilistic object model further comprises rendering a predicted view given a scene with objects and appearance models.
5. The method of claim 4 wherein creating the probabilistic object model further comprises inferring a scene.
6. The method of claim 5 wherein inferring the scene comprises computing an observed map and incrementally adding objects to the scene until the discrepancy is less than a predefined threshold.
7. The method of claim 6 wherein creating the probabilistic object model further comprises inferring the object.
8. A method comprising:
using light fields to generate a probabilistic generative model for objects, enabling a robot to use all information from a camera to achieve precision.
9. The method of claim 8 further comprising:
combining learned models form models for object category that are metrically grounded and used to perform localization and grasp prediction.
10. A system comprising:
a robot arm;
multiple cameras locations; and
a process for creating a probabilistic object model to identify, localize, and manipulate an object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across the multiple camera locations.
11. The system of claim 10 wherein creating the probabilistic object model comprises rendering an observed view from images and poses at a particular plane.
12. The system of claim 11 wherein rendering the observed view from images and poses at the particular plane comprises computing a maximum likelihood estimate from image data by finding a sample mean and variance for pixel values observed at each cell using a calibration function.
13. The system of claim 12 wherein creating the probabilistic object model further comprises rendering a predicted view given a scene with objects and appearance models.
14. The system of claim 13 wherein creating the probabilistic object model further comprises inferring a scene.
15. The system of claim 14 wherein inferring the scene comprises computing an observed map and incrementally adding objects to the scene until the discrepancy is less than a predefined threshold.
16. The system of claim 15 wherein creating the probabilistic object model further comprises inferring the object.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/757,332 US11847841B2 (en) | 2017-10-18 | 2018-10-18 | Probabilistic object models for robust, repeatable pick-and-place |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201762573890P | 2017-10-18 | 2017-10-18 | |
| US62/573,890 | 2017-10-18 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019079598A1 true WO2019079598A1 (en) | 2019-04-25 |
Family
ID=66173485
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2018/056514 Ceased WO2019079598A1 (en) | 2017-10-18 | 2018-10-18 | Probabilistic object models for robust, repeatable pick-and-place |
Country Status (2)
| Country | Link |
|---|---|
| US (1) | US11847841B2 (en) |
| WO (1) | WO2019079598A1 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110349214A (en) * | 2019-07-01 | 2019-10-18 | 深圳前海达闼云端智能科技有限公司 | A kind of localization method of object, terminal and readable storage medium storing program for executing |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7059968B2 (en) * | 2019-03-01 | 2022-04-26 | オムロン株式会社 | Control device and alignment device |
| US11720112B2 (en) * | 2019-09-17 | 2023-08-08 | JAR Scientific LLC | Autonomous object relocation robot |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2012089928A1 (en) * | 2010-12-30 | 2012-07-05 | Zenrobotics Oy | Method, computer program and apparatus for determining a gripping location |
| RU2528140C1 (en) * | 2013-03-12 | 2014-09-10 | Открытое акционерное общество "Научно-производственное объединение "Карат" (ОАО "НПО КАРАТ") | Method for automatic recognition of objects on image |
| US20150269436A1 (en) * | 2014-03-18 | 2015-09-24 | Qualcomm Incorporated | Line segment tracking in computer vision applications |
| US9333649B1 (en) * | 2013-03-15 | 2016-05-10 | Industrial Perception, Inc. | Object pickup strategies for a robotic device |
Family Cites Families (30)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5040056A (en) | 1990-01-29 | 1991-08-13 | Technistar Corporation | Automated system for locating and transferring objects on a conveyor belt |
| US7336803B2 (en) | 2002-10-17 | 2008-02-26 | Siemens Corporate Research, Inc. | Method for scene modeling and change detection |
| US7854108B2 (en) | 2003-12-12 | 2010-12-21 | Vision Robotics Corporation | Agricultural robot system and method |
| US7933451B2 (en) | 2005-11-23 | 2011-04-26 | Leica Geosystems Ag | Feature extraction using pixel-level and object-level analysis |
| EP1956528B1 (en) | 2007-02-08 | 2018-10-03 | Samsung Electronics Co., Ltd. | Apparatus and method for expressing behavior of software robot |
| US8374388B2 (en) | 2007-12-28 | 2013-02-12 | Rustam Stolkin | Real-time tracking of non-rigid objects in image sequences for which the background may be changing |
| GB2458927B (en) * | 2008-04-02 | 2012-11-14 | Eykona Technologies Ltd | 3D Imaging system |
| GB0818561D0 (en) | 2008-10-09 | 2008-11-19 | Isis Innovation | Visual tracking of objects in images, and segmentation of images |
| US8638985B2 (en) | 2009-05-01 | 2014-01-28 | Microsoft Corporation | Human body pose estimation |
| US8244402B2 (en) | 2009-09-22 | 2012-08-14 | GM Global Technology Operations LLC | Visual perception system and method for a humanoid robot |
| US8472698B2 (en) | 2009-11-24 | 2013-06-25 | Mitsubishi Electric Research Laboratories, Inc. | System and method for determining poses of objects |
| US8918209B2 (en) | 2010-05-20 | 2014-12-23 | Irobot Corporation | Mobile human interface robot |
| US8527445B2 (en) | 2010-12-02 | 2013-09-03 | Pukoa Scientific, Llc | Apparatus, system, and method for object detection and identification |
| US9129277B2 (en) | 2011-08-30 | 2015-09-08 | Digimarc Corporation | Methods and arrangements for identifying objects |
| US8693731B2 (en) | 2012-01-17 | 2014-04-08 | Leap Motion, Inc. | Enhanced contrast for object detection and characterization by optical imaging |
| US9092698B2 (en) | 2012-06-21 | 2015-07-28 | Rethink Robotics, Inc. | Vision-guided robots and methods of training them |
| CN103051893B (en) | 2012-10-18 | 2015-05-13 | 北京航空航天大学 | Dynamic background video object extraction based on pentagonal search and five-frame background alignment |
| US9832452B1 (en) | 2013-08-12 | 2017-11-28 | Amazon Technologies, Inc. | Robust user detection and tracking |
| US10089740B2 (en) | 2014-03-07 | 2018-10-02 | Fotonation Limited | System and methods for depth regularization and semiautomatic interactive matting using RGB-D images |
| US9626566B2 (en) * | 2014-03-19 | 2017-04-18 | Neurala, Inc. | Methods and apparatus for autonomous robotic control |
| US10042047B2 (en) * | 2014-09-19 | 2018-08-07 | GM Global Technology Operations LLC | Doppler-based segmentation and optical flow in radar images |
| US9733646B1 (en) | 2014-11-10 | 2017-08-15 | X Development Llc | Heterogeneous fleet of robots for collaborative object processing |
| US9717387B1 (en) | 2015-02-26 | 2017-08-01 | Brain Corporation | Apparatus and methods for programming and training of robotic household appliances |
| US9889566B2 (en) | 2015-05-01 | 2018-02-13 | General Electric Company | Systems and methods for control of robotic manipulation |
| JP6117853B2 (en) | 2015-05-13 | 2017-04-19 | ファナック株式会社 | Article removal system and method for removing loosely stacked items |
| EP4235539A3 (en) | 2015-09-11 | 2023-09-27 | Berkshire Grey Operating Company, Inc. | Robotic systems and methods for identifying and processing a variety of objects |
| US10372968B2 (en) | 2016-01-22 | 2019-08-06 | Qualcomm Incorporated | Object-focused active three-dimensional reconstruction |
| FI127100B (en) | 2016-08-04 | 2017-11-15 | Zenrobotics Oy | A method and apparatus for separating at least one object from the multiplicity of objects |
| US10766145B2 (en) | 2017-04-14 | 2020-09-08 | Brown University | Eye in-hand robot |
| US11365068B2 (en) | 2017-09-05 | 2022-06-21 | Abb Schweiz Ag | Robotic systems and methods for operating a robot |
-
2018
- 2018-10-18 WO PCT/US2018/056514 patent/WO2019079598A1/en not_active Ceased
- 2018-10-18 US US16/757,332 patent/US11847841B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2012089928A1 (en) * | 2010-12-30 | 2012-07-05 | Zenrobotics Oy | Method, computer program and apparatus for determining a gripping location |
| RU2528140C1 (en) * | 2013-03-12 | 2014-09-10 | Открытое акционерное общество "Научно-производственное объединение "Карат" (ОАО "НПО КАРАТ") | Method for automatic recognition of objects on image |
| US9333649B1 (en) * | 2013-03-15 | 2016-05-10 | Industrial Perception, Inc. | Object pickup strategies for a robotic device |
| US20150269436A1 (en) * | 2014-03-18 | 2015-09-24 | Qualcomm Incorporated | Line segment tracking in computer vision applications |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110349214A (en) * | 2019-07-01 | 2019-10-18 | 深圳前海达闼云端智能科技有限公司 | A kind of localization method of object, terminal and readable storage medium storing program for executing |
| CN110349214B (en) * | 2019-07-01 | 2022-09-16 | 达闼机器人股份有限公司 | Object positioning method, terminal and readable storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| US20200368899A1 (en) | 2020-11-26 |
| US11847841B2 (en) | 2023-12-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112476434B (en) | Visual 3D pick-and-place method and system based on cooperative robot | |
| EP3825903B1 (en) | Method, apparatus and storage medium for detecting small obstacles | |
| CN114952809B (en) | Workpiece recognition and pose detection method, system, and grasping control method of a robotic arm | |
| Novkovic et al. | Object finding in cluttered scenes using interactive perception | |
| Mu et al. | Slam with objects using a nonparametric pose graph | |
| Popović et al. | A strategy for grasping unknown objects based on co-planarity and colour information | |
| Rusu | Semantic 3D object maps for everyday robot manipulation | |
| Rusu et al. | Laser-based perception for door and handle identification | |
| Breyer et al. | Closed-loop next-best-view planning for target-driven grasping | |
| Einhorn et al. | Finding the adequate resolution for grid mapping-cell sizes locally adapting on-the-fly | |
| Aleotti et al. | Perception and grasping of object parts from active robot exploration | |
| Kasaei et al. | Perceiving, learning, and recognizing 3d objects: An approach to cognitive service robots | |
| CN112288809B (en) | Robot grabbing detection method for multi-object complex scene | |
| US11847841B2 (en) | Probabilistic object models for robust, repeatable pick-and-place | |
| CN117095060A (en) | Accurate cleaning method and device based on garbage detection technology and cleaning robot | |
| Albrecht et al. | Seeing the unseen: Simple reconstruction of transparent objects from point cloud data | |
| Goodrich et al. | Depth by poking: Learning to estimate depth from self-supervised grasping | |
| CN118809616A (en) | A robot motion planning method, device, robot and storage medium | |
| Effendi et al. | Robot manipulation grasping of recognized objects for assistive technology support using stereo vision | |
| Bagoren et al. | Pugs: Perceptual uncertainty for grasp selection in underwater environments | |
| CN121223808B (en) | Double-mechanical-arm collaborative goods taking method and system based on multi-source visual cross-view fusion | |
| Baleia et al. | Self-supervised learning of depth-based navigation affordances from haptic cues | |
| Jafari et al. | Robotic eye-to-hand coordination: Implementing visual perception to object manipulation | |
| Useche-Murillo et al. | Obstacle Evasion Algorithm Using Convolutional Neural Networks and Kinect-V1 | |
| Wada | Robotic manipulation in clutter with object-level semantic mapping |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18868630 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18868630 Country of ref document: EP Kind code of ref document: A1 |











