WO2019079598A1 - Probabilistic object models for robust, repeatable pick-and-place - Google Patents

Probabilistic object models for robust, repeatable pick-and-place Download PDF

Info

Publication number
WO2019079598A1
WO2019079598A1 PCT/US2018/056514 US2018056514W WO2019079598A1 WO 2019079598 A1 WO2019079598 A1 WO 2019079598A1 US 2018056514 W US2018056514 W US 2018056514W WO 2019079598 A1 WO2019079598 A1 WO 2019079598A1
Authority
WO
WIPO (PCT)
Prior art keywords
probabilistic
objects
observed
creating
object model
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2018/056514
Other languages
French (fr)
Inventor
Stefanie Tellex
John Oberlin
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Brown University
Original Assignee
Brown University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Brown University filed Critical Brown University
Priority to US16/757,332 priority Critical patent/US11847841B2/en
Publication of WO2019079598A1 publication Critical patent/WO2019079598A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/60Type of objects
    • G06V20/64Three-dimensional [3D] objects
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J19/00Accessories fitted to manipulators, e.g. for monitoring, for viewing; Safety devices combined with or specially adapted for use in connection with manipulators
    • B25J19/02Sensing devices
    • B25J19/021Optical sensing devices
    • B25J19/023Optical sensing devices including video camera means
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/02Program-controlled manipulators characterised by movement of the arms, e.g. cartesian coordinate type
    • B25J9/04Program-controlled manipulators characterised by movement of the arms, e.g. cartesian coordinate type by rotating at least one arm, excluding the head movement itself, e.g. cylindrical coordinate type or polar coordinate type
    • B25J9/041Cylindrical coordinate type
    • B25J9/042Cylindrical coordinate type comprising an articulated arm
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • B25J9/1612Program controls characterised by the hand, wrist, grip control
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • B25J9/1628Program controls characterised by the control loop
    • B25J9/163Program controls characterised by the control loop learning, adaptive, model based, rule based expert control
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/29Graphical models, e.g. Bayesian networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N7/00Computing arrangements based on specific mathematical models
    • G06N7/01Probabilistic graphical models, e.g. probabilistic networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/70Determining position or orientation of objects or cameras
    • G06T7/73Determining position or orientation of objects or cameras using feature-based methods
    • G06T7/74Determining position or orientation of objects or cameras using feature-based methods involving reference images or patches
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/772Determining representative reference patterns, e.g. averaging or distorting patterns; Generating dictionaries
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/84Arrangements for image or video recognition or understanding using pattern recognition or machine learning using probabilistic graphical models from image or video features, e.g. Markov models or Bayesian networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20076Probabilistic image processing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30108Industrial image inspection
    • G06T2207/30164Workpiece; Machine component

Definitions

  • the present invention relates generally to object recognition and manipulation, and more particularly to probabilistic object models for robust, repeatable pick-and-place.
  • the invention features a method including, as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
  • the invention features a method including using light fields to generate a probabilistic generative model for objects, enabling a robot to use all information from a camera to achieve precision.
  • the invention features a system including a robot arm, multiple cameras locations, and a process for creating a probabilistic object model to identify, localize, and manipulate an object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across the multiple camera locations.
  • FIG. 1 illustrates how an exemplary model enables a robot to learn to robustly detect, localize and manipulate objects.
  • FIG. 2 illustrates an exemplary probabilistic object map model for reasoning about objects.
  • FIG. 3 illustrates exemplary objects.
  • FIG. 4 illustrates an exemplary observed view for a scene.
  • FIG. 5 illustrates automatically generated thumbnails and observed views for two configurations detected for a yellow square duplo.
  • Robust object perception is a core capability for a manipulator robot.
  • Current perception techniques do not reach levels of precision necessary for a home robot, which might be requested to pick hundreds of times each day and where picking errors might be very costly, resulting in broken objects and ultimately, in the person's rejection of the robot.
  • the present invention enables the robot to learn a model of an object using light fields, achieving very high levels of robustness when detecting, localizing, and manipulating objects.
  • Existing approaches use features, which discard information, because inference in a full generative model is intractable.
  • our method uses light fields to enable efficient inference for object detection and localization, while incorporating information from every pixel observed from across multiple camera locations.
  • the robot can segment objects, identify object instances, localize them, and extract three dimensional (3D) structure.
  • our method enables a Baxter robot to pick an object hundreds of times in a row without failures, and to
  • the robot can identify, localize and grasp objects with high accuracy using our framework; modeling one side of a novel object takes a Baxter robot approximately 30 seconds and enables detection and localization to within 2 millimeters. Furthermore, the robot can merge object models to create metrically grounded models for object categories, in order to improve accuracy on previously unencountered objects.
  • a robot can automatically acquire a POM by exploiting its ability to change the environment in service of its perceptual goals. This approach enables the robot to obtain extremely high accuracy and repeatability at localizing and picking objects. Additionally, a robot can create metrically grounded models for object categories by merging POMs.
  • the present invention demonstrates that a Baxter robot can autonomously create models for objects. These models enable the robot to detect, localize, and manipulate the objects with very high reliability and repeatability, localizing to within 2 mm for hundreds of successful picks in a row. Our system and method enable a Baxter robot to autonomously map objects for many hours using its wrist camera.
  • a goal is for a robot to estimate objects in a scene, given observations of camera images, and associated camera poses,
  • Each object instance, o N consists of a pose, x ⁇ , along with an index, tf 1 , which identifies the object type.
  • Each object type, O k is a set of appearance modes. We rewrite the distribution using Bayes' rule:
  • This model is known as inverse graphics because it requires assigning a probability to images given a model of the scene.
  • Equation 6 corresponds to light field rendering.
  • the second term corresponds to the light field distribution in a synthetic photograph given a model of the scene as objects.
  • each object has only a few configurations, c ⁇ C, at which it may lie at rest. This assumption is not true in general; for example, a ball has an infinite number of such configurations. However many objects with continuously varying families of stable poses can be approximated using a few sampled configurations. Furthermore this assumption leads to straightforward representations for POMs.
  • the appearance model, A c for an object is a model of the light expected to be observed from the object by a camera.
  • Algorithm 1 Render the observed view from images and poses at a particular plane, z.
  • Algorithm 2 Render the predicted view given a scene with objects and appearance models.
  • Algorithm 4 Infer object.
  • Equation 6 Using Equation 6 to find the objects that maximize a scene is still challenging because it requires integrating over all values for the variances of the map cells and searching over a full 3D configuration of all objects. Instead we approximate it by performing inference in the space of light field maps, m and integral finding the map m and scene (objects) that maximizes the
  • This m * is the observed view.
  • Algorithm 1 gives a formal description of how to compute it from calibrated camera images. This view can be computed by iterating through each image one time, linear in the number of pixels or light rays. For a given scene (i.e., object configuration and appearance models), we can compute an m that maximizes the second term in Equation 6.
  • This ⁇ m is the predicted view; an example predicted view is shown in FIG. 4. This computation corresponds to a rendering process.
  • Our model enables us to render compositionally over object maps and match in the 2D configuration space with three degrees of freedom instead of the 3D space with six.
  • the system computes the observed map and then incrementally adds objects to the scene until the discrepancy is less than a predefined threshold. At this point, all of the discrepancy has been accounted for, and the robot can use this information to ground natural language referring expressions such as "next to the bowl,” to find empty space in the scene where objects can be placed, and to infer object pose for grasping.
  • Modeling objects requires estimating O k for each object k, including the appearance models for each stable configuration.
  • the robot creates a model for the input pile and mapping space.
  • the robot picks an object from the input pile, moves it to the mapping space, maps the object, then moves it to the output pile.
  • the robot To pick from the input pile before an object map has been acquired, the robot tries to pick repeatedly with generic grasp detectors. To propose grasps, the robot looks for patterns of discrepancy between the background map and the observed map. Once successful, it places the object in the mapping workspace and creates a model for the mapping workspace with the object. Regions in this model which are discrepant with the background are used to create an appearance model A for the object. After modeling has been completed, the robot clears the mapping workspace and moves on to the next object. Note that there is a trade off between object throughput and information acquired about each object; to increase throughput we can reduce the number of pick attempts during mapping. In contrast, to truly master an object, the robot might try 100 or more picks as well as other exploratory actions before moving on.
  • the robot If the robot loses the object during mapping (for example because it rolls out of the workspace), the robot returns to the input pile to take the next object. If the robot is unable to grasp the current object to move it out of the mapping workspace, it uses nudging behaviors to push the object out. If its nudging behaviors fail to clear the mapping workspace, it simply updates its background model and then continues mapping objects from the input pile.
  • the object model, O k is known, but the configuration, c is unknown.
  • the robot needs to decide when it has encountered a new object configuration given the collection of maps it has already made. We use an approximation of a maximum likelihood estimate with a geometric prior to decide when to make a new configuration for the object.
  • POMs Once POMs have been acquired, they can be merged to form models for object categories. Many standard computer vision approaches can be applied to light field views rather than images; for example deformable parts models for classification in light field space. These techniques may perform better on the light field view because they have access to metric information, variance in appearance models, as well as information about different configurations and transition information.
  • Algorithm 4 we created models for several object categories. First our robot mapped a number of object instances. Then for each group of instances, we successively applied the optimization in Algorithm 4, with a bias to force the localization to overlap with the model as much as possible. This process results in a synthetic photograph for each category that is strongly biased by object shape and contains noisy information about color and internal structure. Examples of learned object categories appear in FIG. 3. We used our learned models to successfully detect novel instances of each category from a scene with distractor objects. We aim to train on larger data sets and apply more sophisticated modeling techniques to learn object categories and parts from large datasets of POMs.
  • the grid cell size is 0:25 cm, and the total size of the synthetic photograph was approximately 30 cm x 30 cm.
  • our approach enables a robot to adapt to the specific objects it finds and robustly pick objects many times in a row without failures.
  • This robustness and repeatability outperforms existing approaches in terms of its precision by trading off recall, enabling the robot to robustly pick objects many times in a row.
  • Our approach uses light fields to create a probabilistic generative model for objects, enabling the robots to use all information from the camera to achieve this high precision.
  • learned models can be combined to form models for object category that are metrically grounded and can be used to perform localization and grasp prediction.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Software Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Mechanical Engineering (AREA)
  • Robotics (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Computing Systems (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • Probability & Statistics with Applications (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Orthopedic Medicine & Surgery (AREA)
  • Pure & Applied Mathematics (AREA)
  • Algebra (AREA)
  • Computational Mathematics (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Mathematical Physics (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)
  • Manipulator (AREA)

Abstract

A method includes, as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.

Description

Probabilistic Object Models for
Robust, Repeatable Pick-and-Place
STATEMENT REGARDING GOVERNMENT INTEREST
[001] None.
CROSS REFERENCE TO RELATED APPLICATIONS
[002] This application claims benefit from U.S. Provisional Patent Application Serial No. 62/573,890, filed October 18, 2017, which is incorporated by reference in its entirety.
BACKGROUND OF THE INVENTION
[003] The present invention relates generally to object recognition and manipulation, and more particularly to probabilistic object models for robust, repeatable pick-and-place.
[004] In general, most robots cannot pick up most objects most of the time. Yet for effective human-robot collaboration, a robot must be able to detect, localize, and manipulate the specific objects that a person cares about. For example, a household robot should be able to effectively respond to a person's commands such as "Get me a glass of water in my favorite mug," or "Clean up the workshop." For these tasks, extremely high reliability is needed; if a robot can pick at 95% accuracy, it will still fail one in twenty times. Considering it might be doing hundreds of picks each day, this level of performance will not be acceptable for an end-to-end system, because the robot will miss objects each day, potentially breaking a person's things.
SUMMARY OF THE INVENTION
[005] The following presents a simplified summary of the innovation in order to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview of the invention. It is intended to neither identify key or critical elements of the invention nor delineate the scope of the invention. Its sole purpose is to present some concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later.
[006] In an aspect, the invention features a method including, as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
[007] In another aspect, the invention features a method including using light fields to generate a probabilistic generative model for objects, enabling a robot to use all information from a camera to achieve precision.
[008] In still another aspect, the invention features a system including a robot arm, multiple cameras locations, and a process for creating a probabilistic object model to identify, localize, and manipulate an object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across the multiple camera locations.
[009] These and other features and advantages will be apparent from a reading of the following detailed description and a review of the associated drawings. It is to be understood that both the foregoing general description and the following detailed description are explanatory only and are not restrictive of aspects as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] These and other features, aspects, and advantages of the present invention will become better understood with reference to the following description, appended claims, and accompanying drawings where:
[0011] FIG. 1 illustrates how an exemplary model enables a robot to learn to robustly detect, localize and manipulate objects.
[0012] FIG. 2 illustrates an exemplary probabilistic object map model for reasoning about objects.
[0013] FIG. 3 illustrates exemplary objects.
[0014] FIG. 4 illustrates an exemplary observed view for a scene.
[0015] FIG. 5 illustrates automatically generated thumbnails and observed views for two configurations detected for a yellow square duplo.
DETAILED DESCRIPTION
[0016] The subject innovation is now described with reference to the drawings, wherein like reference numerals are used to refer to like elements throughout. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It may be evident, however, that the present invention may be practiced without these specific details. In other instances, well- known structures and devices are shown in block diagram form in order to facilitate describing the present invention.
[0017] Robust object perception is a core capability for a manipulator robot. Current perception techniques do not reach levels of precision necessary for a home robot, which might be requested to pick hundreds of times each day and where picking errors might be very costly, resulting in broken objects and ultimately, in the person's rejection of the robot. To address this problem, the present invention enables the robot to learn a model of an object using light fields, achieving very high levels of robustness when detecting, localizing, and manipulating objects. We present a generative model for object appearance and pose given observed camera images. Existing approaches use features, which discard information, because inference in a full generative model is intractable. We present a method that uses light fields to enable efficient inference for object detection and localization, while incorporating information from every pixel observed from across multiple camera locations. Using the learned model, the robot can segment objects, identify object instances, localize them, and extract three dimensional (3D) structure. For example, our method enables a Baxter robot to pick an object hundreds of times in a row without failures, and to
autonomously create models of objects for many hours at a time. The robot can identify, localize and grasp objects with high accuracy using our framework; modeling one side of a novel object takes a Baxter robot approximately 30 seconds and enables detection and localization to within 2 millimeters. Furthermore, the robot can merge object models to create metrically grounded models for object categories, in order to improve accuracy on previously unencountered objects.
[0018] Existing systems for generic object detection and manipulation have lower accuracy than instance-based methods and cannot adapt on the fly to objects that do not work the first time. Instance based approaches can have high accuracy but require object models to be provided in advance; acquiring models is a time consuming process that is difficult for even an expert to perform. For example, the winning team from the Amazon Picking
Challenge in 2015 used a system that required specific instance-based data to do pose estimation and still found that perception was a major source of system errors. Autonomously learning 3D object models requires expensive ICP-based methods to localize objects, which are expensive to globally optimize, making them impractical for classification.
[0019] To achieve extremely high accuracy picking, we reframe the problem to one of adaptation and learning. When the robot encounters a novel object, its goal is to create a model that contains the information needed to identify, localize, and manipulate that object with extremely high accuracy and robustness. We term this information a probabilistic object model (POM). After learning a POM, the robot is able to robustly interact with the object, as shown in FIG. 1. We represent POMs using a pixel-based inverse graphics approach based around light fields. Unlike feature-based methods, pixel based inverse graphics methods use all information obtained from a camera. Inference in inverse-graphics methods is
computationally expensive because the program must search a very large space of possible scenes and marginalize over all possible images. Using the light field approach, a robot can automatically acquire a POM by exploiting its ability to change the environment in service of its perceptual goals. This approach enables the robot to obtain extremely high accuracy and repeatability at localizing and picking objects. Additionally, a robot can create metrically grounded models for object categories by merging POMs.
[0020] The present invention demonstrates that a Baxter robot can autonomously create models for objects. These models enable the robot to detect, localize, and manipulate the objects with very high reliability and repeatability, localizing to within 2 mm for hundreds of successful picks in a row. Our system and method enable a Baxter robot to autonomously map objects for many hours using its wrist camera.
[0021] We present a probabilistic graphical model for object detection, localization, and modeling. We first describe the model, then inference in the model. The graphical model is depicted in FIG. 2, while a table of variables appears in Table I.
Figure imgf000006_0001
[0022] Probabilistic Obj ect Model
[0023] A goal is for a robot to estimate objects in a scene,
Figure imgf000006_0004
given observations of camera images,
Figure imgf000006_0006
and associated camera poses,
Figure imgf000006_0005
Figure imgf000006_0003
[0024] Each object instance, oN, consists of a pose, x^, along with an index, tf1, which identifies the object type. Each object type, Ok, is a set of appearance modes. We rewrite the distribution using Bayes' rule:
Figure imgf000006_0002
[0025] We assume a uniform prior over object locations and appearances, so that only the first term matters in the optimization.
[0026] A generative, inverse graphics approach would assume each image is independent given the object locations, since if we know the true locations and appearances of the objects, we can predict the contents of the images using graphics:
Figure imgf000007_0001
[0027] This model is known as inverse graphics because it requires assigning a probability to images given a model of the scene. However, inference under these
assumptions is intractable because of the very large space of possible images. Existing approaches turn to features such as SIFT or learned features using neural networks, but these approaches throw away information in order to achieve more generality. Instead we aim to use all information to achieve the most precise localization and manipulation possible.
[0028] To solve this problem, we introduce the light field grid, m as a latent variable. We define a distribution over the light rays emitted from a scene. Using this generative model, we can then perform inferences about the objects in the scene, conditioned on observed light rays, R, where each ray corresponds to a pixel in one of the images, Zh, combined with a camera calibration function that maps from pixel space to a light ray given camera pose. We define a synthetic photograph, m, as an L x W array of cells in a plane in space. Each cell (/, w) e m has a height z and scatters light at its (x, y, z) location. We assume each observed light ray arose from a particular cell (/, w), so that the parameters associated with each cell include its height z and a model of the intensity of light emitted from that cell.
[0029] We integrate over m as a latent variable:
Figure imgf000007_0002
[0030] Then we factor this distribution assuming images are independent given the light field model m:
Figure imgf000008_0001
[0031] Here is the bundle of rays that arose from map cell (l,w), which can be
Figure imgf000008_0002
determined finding all rays that intersect the cell using the calibration function. This factorization assumes that each bundle of rays is conditionally independent given the cell parameters, an assumption valid for cells that do not actively emit light. We can render m as an image by showing the values for
Figure imgf000008_0003
as the pixel color; however variance information μ /w is also stored. FIG. 4 shows example scenes rendered using this model. The first term in Equation 6 corresponds to light field rendering. The second term corresponds to the light field distribution in a synthetic photograph given a model of the scene as objects.
[0032] We assume each object has only a few configurations, c ε C, at which it may lie at rest. This assumption is not true in general; for example, a ball has an infinite number of such configurations. However many objects with continuously varying families of stable poses can be approximated using a few sampled configurations. Furthermore this assumption leads to straightforward representations for POMs. In particular, the appearance model, Ac, for an object is a model of the light expected to be observed from the object by a camera. At both modeling time and inference time, we render a synthetic photograph of the object from a canonical view (e.g., top down or head-on), enabling efficient inference in the a lower- dimensional subspace instead of being required to do full 3D inference as in ICP. This lower- dimensional subspace enables much deeper and finer-grained search so that we can use all information from pixels to perform very accurate pose estimation.
[0033] Inference
[0034] Algorithm 1 Render the observed view from images and poses at a particular plane, z.
Figure imgf000009_0001
[0035] Algorithm 2 Render the predicted view given a scene with objects and appearance models.
Figure imgf000009_0002
[0036] Algorithm 3 Infer scene.
Figure imgf000009_0003
[0037] Algorithm 4 Infer object.
Figure imgf000010_0001
[0038] Using Equation 6 to find the objects that maximize a scene is still challenging because it requires integrating over all values for the variances of the map cells and searching over a full 3D configuration of all objects. Instead we approximate it by performing inference in the space of light field maps, m and integral finding the map m and scene (objects) that maximizes the
integral.
[0039] First, we find the value for m that maximizes the likelihood of the observed images in the first term:
Figure imgf000010_0002
[0040] This m* is the observed view. We compute the maximum likelihood estimate from image data analytically by finding the sample mean and variance for pixel values observed at each cell using the calibration function. An example observed view is shown in FIG. 4. We use the approximation that each ray is assigned to the cell it intersects without reasoning about obstructions. Algorithm 1 gives a formal description of how to compute it from calibrated camera images. This view can be computed by iterating through each image one time, linear in the number of pixels or light rays. For a given scene (i.e., object configuration and appearance models), we can compute an m that maximizes the second term in Equation 6.
Figure imgf000011_0001
[0041] This Λm is the predicted view; an example predicted view is shown in FIG. 4. This computation corresponds to a rendering process. Our model enables us to render compositionally over object maps and match in the 2D configuration space with three degrees of freedom instead of the 3D space with six.
[0042] To maximize the product over scenes o°...oN, we can compute the discrepancy between the observed view and predicted view, shown in FIG. 4. Finding the configuration of objects that minimizes this discrepancy corresponds to maximizing the log-likelihood of the scene under the observed images. Additionally, the robot can use this discrepancy as a mask for object segmentation by mapping a region, adding an object to the scene, and then remapping the region; discrepant regions correspond to cells associated with the new object. Algorithm 2 gives a description of how to compute it given a scene defined as object poses and appearance models.
[0043] To infer a scene, the system computes the observed map and then incrementally adds objects to the scene until the discrepancy is less than a predefined threshold. At this point, all of the discrepancy has been accounted for, and the robot can use this information to ground natural language referring expressions such as "next to the bowl," to find empty space in the scene where objects can be placed, and to infer object pose for grasping.
[0044] Learning Probabilistic Object Models
[0045] Modeling objects requires estimating Ok for each object k, including the appearance models for each stable configuration. We have created a system that enables a Baxter robot to autonomously map objects for hours at a time. We first divide the workspace for one arm of the robot into three regions: an input pile, a mapping space, and an output pile. The robot creates a model for the input pile and mapping space. Then a person adds objects to the input pile, and autonomous mapping begins. The robot picks an object from the input pile, moves it to the mapping space, maps the object, then moves it to the output pile.
[0046] To pick from the input pile before an object map has been acquired, the robot tries to pick repeatedly with generic grasp detectors. To propose grasps, the robot looks for patterns of discrepancy between the background map and the observed map. Once successful, it places the object in the mapping workspace and creates a model for the mapping workspace with the object. Regions in this model which are discrepant with the background are used to create an appearance model A for the object. After modeling has been completed, the robot clears the mapping workspace and moves on to the next object. Note that there is a trade off between object throughput and information acquired about each object; to increase throughput we can reduce the number of pick attempts during mapping. In contrast, to truly master an object, the robot might try 100 or more picks as well as other exploratory actions before moving on.
[0047] If the robot loses the object during mapping (for example because it rolls out of the workspace), the robot returns to the input pile to take the next object. If the robot is unable to grasp the current object to move it out of the mapping workspace, it uses nudging behaviors to push the object out. If its nudging behaviors fail to clear the mapping workspace, it simply updates its background model and then continues mapping objects from the input pile.
[0048] Inferring a New Configuration
[0049] After a new object has been placed in the workspace, the object model, Ok is known, but the configuration, c is unknown. The robot needs to decide when it has encountered a new object configuration given the collection of maps it has already made. We use an approximation of a maximum likelihood estimate with a geometric prior to decide when to make a new configuration for the object.
[0050] Learning Object Categories
[0051] Once POMs have been acquired, they can be merged to form models for object categories. Many standard computer vision approaches can be applied to light field views rather than images; for example deformable parts models for classification in light field space. These techniques may perform better on the light field view because they have access to metric information, variance in appearance models, as well as information about different configurations and transition information. As a proof of concept, we created models for several object categories. First our robot mapped a number of object instances. Then for each group of instances, we successively applied the optimization in Algorithm 4, with a bias to force the localization to overlap with the model as much as possible. This process results in a synthetic photograph for each category that is strongly biased by object shape and contains noisy information about color and internal structure. Examples of learned object categories appear in FIG. 3. We used our learned models to successfully detect novel instances of each category from a scene with distractor objects. We aim to train on larger data sets and apply more sophisticated modeling techniques to learn object categories and parts from large datasets of POMs.
[0052] Evaluation
[0053] We evaluate our model's ability to detect, localize, and manipulate objects using the Baxter robot. We selected a subset of YCB objects that were rigid and pickable by Baxter with the grippers in the 6cm position as well as a standard ICRA duckie. In our
implementation the grid cell size is 0:25 cm, and the total size of the synthetic photograph was approximately 30 cm x 30 cm. We initialize the background variance to a higher value to account for changes in lighting and shadows.
[0054] Localization
[0055] To evaluate our model's ability to localize objects, we find the error of its position estimates by servoing repeatedly to the same location. For each trial, we moved the arm directly above the object, then moved to a random position and orientation within 10 cm of the true location. Next we estimated the object's position by serving: first we created a light field model at the arm's current location; then we used Algorithm 3 to estimate the object's position; then we moved the arm to the estimated position and repeated. We performed five trials in each location, then moved the object to a new location, for a total of 25 trials per object. We take the mean location estimated over the five trials as the object's true location, and report the mean distance from this location as well as 95% confidence intervals. This test records the repeatability of the servoing and pose estimation; if we are performing accurate pose estimation, then the system should find the object at the same place each time. Results appear in Table II. Our results show that using POMs, the system can localize objects to within 2 mm. We observe more error on taller objects such as the mustard, and the taller duplo structure, due to our assumption that all cells are at table height. Note that even on these objects, localization is accurate to within a centimeter, enough to pick reliably with many grippers; similarly detection accuracy is also quite high. To assess the effect of correcting for z, we computed new models for the tall yellow square duplo using the two different approaches. We found that the error reduced to 0.0013m+-2.0xl0"05 using the maximum likelihood estimate and to 0.0019m+-1.9xl0"05 using the marginal estimate. Both methods demonstrate a significant improvement. The maximum likelihood estimate performs slightly better, perhaps because the sharper edges lead to more consistent performance.
Computing z corrections takes significant time, so we do not use it for the rest of the evaluation.
[0056] Autonomous Classification and Grasp Model Acquisition
[0057] After using our autonomous process to map our test objects, we evaluated object classification and picking performance. Due to the processing time to infer z, we used z = table for this evaluation. The robot had to identify the object type, localize the object, and then grasp it. After each grasp, it placed the object in a random position and orientation. We report accuracy at labeling the object with the correct type along with its pick success rate over the ten trials in Table II. The robot discovered 1 configuration for most objects, but for the yellow square duplo discovered a second configuration, shown in FIG. 5. We explore more deliberate elicitation of new object configurations, by rotating the hand before dropping the object or by employing bimanual manipulation. We report detection and pick accuracy for the 10 objects.
Figure imgf000014_0001
[0058] Our results show 98% accuracy at detection performance for these objects. The duplo yellow square was confused with the standard ICRA duckie which is similarly colored and sized. Other errors were due to taller objects. The robot mapped the tall duplo structure in a lying down position. It was unable to pick it when it was standing up to move it to its scanning workspace because of the error introduced by the height. After a few attempts it knocked it down; the lower height introduced less error, and it was able to pick and localize perfectly. The padlock is challenging because it is both heavy and reflective. Also its smallest dimension just barely fits into the robot's gripper, meaning that very small amounts of position error can cause the grasp to fail. Overall our system is able to pick this data set 84% of the time.
[0059] Our automatic process successfully mapped all objects except for the mustard. The mustard is a particularly challenging object due to its height and weight; therefore we manually created an model and annotated a grasp. Our initial experiments with this model resulted in 1=10 picks due to noise from its height and its weight; however we were still able to use it for localization and detection. Next we created a new model using marginal z corrections and also performed z corrections at inference time. Additionally we changed to a different gripper configuration more appropriate to this very large object. After these changes, we were able to pick the mustard 10/10 times.
[0060] Picking Robustly
[0061] We can pick a spoon 100 times in a row.
[0062] In summary, our approach enables a robot to adapt to the specific objects it finds and robustly pick objects many times in a row without failures. This robustness and repeatability outperforms existing approaches in terms of its precision by trading off recall, enabling the robot to robustly pick objects many times in a row. Our approach uses light fields to create a probabilistic generative model for objects, enabling the robots to use all information from the camera to achieve this high precision. Additionally, learned models can be combined to form models for object category that are metrically grounded and can be used to perform localization and grasp prediction.
[0063] We can scale up this system so that the robot can come equipped with a large database of object models. This system enables the robot to automatically detect and localize novel objects. If the robot cannot pick the first time, it will automatically add a new model to its database, enabling it to increase its precision and robustness. Additionally this new model will augment the database, improving performance on novel objects.
[0064] Extending the model to full 3D perception enables it to fuse different stable configurations. We can track objects continuously over time, create a Bayes' filtering approach to object tracking. This model takes into account object affordances and actions on objects, creating a full object-oriented MDP. Ultimately objects like doors, elevators, and drawers can be modeled in an instance-based way, and then generalized to novel instances. This model is ideal for connecting to language because it factors the world into objects, just as people do when they talk about them.
[0065] It would be appreciated by those skilled in the art that various changes and modifications can be made to the illustrated embodiments without departing from the spirit of the present invention. All such modifications and changes are intended to be within the scope of the present invention except as limited by the scope of the appended claims.

Claims

WHAT IS CLAIMED IS:
1. A method comprising:
as a robot encounters an object, creating a probabilistic object model to identify, localize, and manipulate the object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across multiple camera locations.
2. The method of claim 1 wherein creating the probabilistic object model comprises rendering an observed view from images and poses at a particular plane.
3. The method of claim 2 wherein rendering the observed view from images and poses at the particular plane comprises computing a maximum likelihood estimate from image data by finding a sample mean and variance for pixel values observed at each cell using a calibration function.
4. The method of claim 3 wherein creating the probabilistic object model further comprises rendering a predicted view given a scene with objects and appearance models.
5. The method of claim 4 wherein creating the probabilistic object model further comprises inferring a scene.
6. The method of claim 5 wherein inferring the scene comprises computing an observed map and incrementally adding objects to the scene until the discrepancy is less than a predefined threshold.
7. The method of claim 6 wherein creating the probabilistic object model further comprises inferring the object.
8. A method comprising: using light fields to generate a probabilistic generative model for objects, enabling a robot to use all information from a camera to achieve precision.
9. The method of claim 8 further comprising:
combining learned models form models for object category that are metrically grounded and used to perform localization and grasp prediction.
10. A system comprising:
a robot arm;
multiple cameras locations; and
a process for creating a probabilistic object model to identify, localize, and manipulate an object, the probabilistic object model using light fields to enable efficient inference for object detection and localization while incorporating information from every pixel observed from across the multiple camera locations.
11. The system of claim 10 wherein creating the probabilistic object model comprises rendering an observed view from images and poses at a particular plane.
12. The system of claim 11 wherein rendering the observed view from images and poses at the particular plane comprises computing a maximum likelihood estimate from image data by finding a sample mean and variance for pixel values observed at each cell using a calibration function.
13. The system of claim 12 wherein creating the probabilistic object model further comprises rendering a predicted view given a scene with objects and appearance models.
14. The system of claim 13 wherein creating the probabilistic object model further comprises inferring a scene.
15. The system of claim 14 wherein inferring the scene comprises computing an observed map and incrementally adding objects to the scene until the discrepancy is less than a predefined threshold.
16. The system of claim 15 wherein creating the probabilistic object model further comprises inferring the object.
PCT/US2018/056514 2017-10-18 2018-10-18 Probabilistic object models for robust, repeatable pick-and-place Ceased WO2019079598A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US16/757,332 US11847841B2 (en) 2017-10-18 2018-10-18 Probabilistic object models for robust, repeatable pick-and-place

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201762573890P 2017-10-18 2017-10-18
US62/573,890 2017-10-18

Publications (1)

Publication Number Publication Date
WO2019079598A1 true WO2019079598A1 (en) 2019-04-25

Family

ID=66173485

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2018/056514 Ceased WO2019079598A1 (en) 2017-10-18 2018-10-18 Probabilistic object models for robust, repeatable pick-and-place

Country Status (2)

Country Link
US (1) US11847841B2 (en)
WO (1) WO2019079598A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110349214A (en) * 2019-07-01 2019-10-18 深圳前海达闼云端智能科技有限公司 A kind of localization method of object, terminal and readable storage medium storing program for executing

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7059968B2 (en) * 2019-03-01 2022-04-26 オムロン株式会社 Control device and alignment device
US11720112B2 (en) * 2019-09-17 2023-08-08 JAR Scientific LLC Autonomous object relocation robot

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2012089928A1 (en) * 2010-12-30 2012-07-05 Zenrobotics Oy Method, computer program and apparatus for determining a gripping location
RU2528140C1 (en) * 2013-03-12 2014-09-10 Открытое акционерное общество "Научно-производственное объединение "Карат" (ОАО "НПО КАРАТ") Method for automatic recognition of objects on image
US20150269436A1 (en) * 2014-03-18 2015-09-24 Qualcomm Incorporated Line segment tracking in computer vision applications
US9333649B1 (en) * 2013-03-15 2016-05-10 Industrial Perception, Inc. Object pickup strategies for a robotic device

Family Cites Families (30)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5040056A (en) 1990-01-29 1991-08-13 Technistar Corporation Automated system for locating and transferring objects on a conveyor belt
US7336803B2 (en) 2002-10-17 2008-02-26 Siemens Corporate Research, Inc. Method for scene modeling and change detection
US7854108B2 (en) 2003-12-12 2010-12-21 Vision Robotics Corporation Agricultural robot system and method
US7933451B2 (en) 2005-11-23 2011-04-26 Leica Geosystems Ag Feature extraction using pixel-level and object-level analysis
EP1956528B1 (en) 2007-02-08 2018-10-03 Samsung Electronics Co., Ltd. Apparatus and method for expressing behavior of software robot
US8374388B2 (en) 2007-12-28 2013-02-12 Rustam Stolkin Real-time tracking of non-rigid objects in image sequences for which the background may be changing
GB2458927B (en) * 2008-04-02 2012-11-14 Eykona Technologies Ltd 3D Imaging system
GB0818561D0 (en) 2008-10-09 2008-11-19 Isis Innovation Visual tracking of objects in images, and segmentation of images
US8638985B2 (en) 2009-05-01 2014-01-28 Microsoft Corporation Human body pose estimation
US8244402B2 (en) 2009-09-22 2012-08-14 GM Global Technology Operations LLC Visual perception system and method for a humanoid robot
US8472698B2 (en) 2009-11-24 2013-06-25 Mitsubishi Electric Research Laboratories, Inc. System and method for determining poses of objects
US8918209B2 (en) 2010-05-20 2014-12-23 Irobot Corporation Mobile human interface robot
US8527445B2 (en) 2010-12-02 2013-09-03 Pukoa Scientific, Llc Apparatus, system, and method for object detection and identification
US9129277B2 (en) 2011-08-30 2015-09-08 Digimarc Corporation Methods and arrangements for identifying objects
US8693731B2 (en) 2012-01-17 2014-04-08 Leap Motion, Inc. Enhanced contrast for object detection and characterization by optical imaging
US9092698B2 (en) 2012-06-21 2015-07-28 Rethink Robotics, Inc. Vision-guided robots and methods of training them
CN103051893B (en) 2012-10-18 2015-05-13 北京航空航天大学 Dynamic background video object extraction based on pentagonal search and five-frame background alignment
US9832452B1 (en) 2013-08-12 2017-11-28 Amazon Technologies, Inc. Robust user detection and tracking
US10089740B2 (en) 2014-03-07 2018-10-02 Fotonation Limited System and methods for depth regularization and semiautomatic interactive matting using RGB-D images
US9626566B2 (en) * 2014-03-19 2017-04-18 Neurala, Inc. Methods and apparatus for autonomous robotic control
US10042047B2 (en) * 2014-09-19 2018-08-07 GM Global Technology Operations LLC Doppler-based segmentation and optical flow in radar images
US9733646B1 (en) 2014-11-10 2017-08-15 X Development Llc Heterogeneous fleet of robots for collaborative object processing
US9717387B1 (en) 2015-02-26 2017-08-01 Brain Corporation Apparatus and methods for programming and training of robotic household appliances
US9889566B2 (en) 2015-05-01 2018-02-13 General Electric Company Systems and methods for control of robotic manipulation
JP6117853B2 (en) 2015-05-13 2017-04-19 ファナック株式会社 Article removal system and method for removing loosely stacked items
EP4235539A3 (en) 2015-09-11 2023-09-27 Berkshire Grey Operating Company, Inc. Robotic systems and methods for identifying and processing a variety of objects
US10372968B2 (en) 2016-01-22 2019-08-06 Qualcomm Incorporated Object-focused active three-dimensional reconstruction
FI127100B (en) 2016-08-04 2017-11-15 Zenrobotics Oy A method and apparatus for separating at least one object from the multiplicity of objects
US10766145B2 (en) 2017-04-14 2020-09-08 Brown University Eye in-hand robot
US11365068B2 (en) 2017-09-05 2022-06-21 Abb Schweiz Ag Robotic systems and methods for operating a robot

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2012089928A1 (en) * 2010-12-30 2012-07-05 Zenrobotics Oy Method, computer program and apparatus for determining a gripping location
RU2528140C1 (en) * 2013-03-12 2014-09-10 Открытое акционерное общество "Научно-производственное объединение "Карат" (ОАО "НПО КАРАТ") Method for automatic recognition of objects on image
US9333649B1 (en) * 2013-03-15 2016-05-10 Industrial Perception, Inc. Object pickup strategies for a robotic device
US20150269436A1 (en) * 2014-03-18 2015-09-24 Qualcomm Incorporated Line segment tracking in computer vision applications

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110349214A (en) * 2019-07-01 2019-10-18 深圳前海达闼云端智能科技有限公司 A kind of localization method of object, terminal and readable storage medium storing program for executing
CN110349214B (en) * 2019-07-01 2022-09-16 达闼机器人股份有限公司 Object positioning method, terminal and readable storage medium

Also Published As

Publication number Publication date
US20200368899A1 (en) 2020-11-26
US11847841B2 (en) 2023-12-19

Similar Documents

Publication Publication Date Title
CN112476434B (en) Visual 3D pick-and-place method and system based on cooperative robot
EP3825903B1 (en) Method, apparatus and storage medium for detecting small obstacles
CN114952809B (en) Workpiece recognition and pose detection method, system, and grasping control method of a robotic arm
Novkovic et al. Object finding in cluttered scenes using interactive perception
Mu et al. Slam with objects using a nonparametric pose graph
Popović et al. A strategy for grasping unknown objects based on co-planarity and colour information
Rusu Semantic 3D object maps for everyday robot manipulation
Rusu et al. Laser-based perception for door and handle identification
Breyer et al. Closed-loop next-best-view planning for target-driven grasping
Einhorn et al. Finding the adequate resolution for grid mapping-cell sizes locally adapting on-the-fly
Aleotti et al. Perception and grasping of object parts from active robot exploration
Kasaei et al. Perceiving, learning, and recognizing 3d objects: An approach to cognitive service robots
CN112288809B (en) Robot grabbing detection method for multi-object complex scene
US11847841B2 (en) Probabilistic object models for robust, repeatable pick-and-place
CN117095060A (en) Accurate cleaning method and device based on garbage detection technology and cleaning robot
Albrecht et al. Seeing the unseen: Simple reconstruction of transparent objects from point cloud data
Goodrich et al. Depth by poking: Learning to estimate depth from self-supervised grasping
CN118809616A (en) A robot motion planning method, device, robot and storage medium
Effendi et al. Robot manipulation grasping of recognized objects for assistive technology support using stereo vision
Bagoren et al. Pugs: Perceptual uncertainty for grasp selection in underwater environments
CN121223808B (en) Double-mechanical-arm collaborative goods taking method and system based on multi-source visual cross-view fusion
Baleia et al. Self-supervised learning of depth-based navigation affordances from haptic cues
Jafari et al. Robotic eye-to-hand coordination: Implementing visual perception to object manipulation
Useche-Murillo et al. Obstacle Evasion Algorithm Using Convolutional Neural Networks and Kinect-V1
Wada Robotic manipulation in clutter with object-level semantic mapping

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18868630

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 18868630

Country of ref document: EP

Kind code of ref document: A1