EP4689711A1 - 3d sensor-based obstacle detection for autonomous vehicles and mobile robots - Google Patents

3d sensor-based obstacle detection for autonomous vehicles and mobile robots

Info

Publication number
EP4689711A1
EP4689711A1 EP23729593.6A EP23729593A EP4689711A1 EP 4689711 A1 EP4689711 A1 EP 4689711A1 EP 23729593 A EP23729593 A EP 23729593A EP 4689711 A1 EP4689711 A1 EP 4689711A1
Authority
EP
European Patent Office
Prior art keywords
points
camera
laser
autonomous system
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23729593.6A
Other languages
German (de)
French (fr)
Inventor
Jose Luis SUSA RINCON
Susanne FRESE
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Siemens AG
Siemens Corp
Original Assignee
Siemens AG
Siemens Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Siemens AG, Siemens Corp filed Critical Siemens AG
Publication of EP4689711A1 publication Critical patent/EP4689711A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G01MEASURING; TESTING
    • G01SRADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S7/00Details of systems according to groups G01S13/00, G01S15/00, G01S17/00
    • G01S7/48Details of systems according to groups G01S13/00, G01S15/00, G01S17/00 of systems according to group G01S17/00
    • G01S7/4808Evaluating distance, position or velocity data
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01SRADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S17/00Systems using the reflection or reradiation of electromagnetic waves other than radio waves, e.g. lidar systems
    • G01S17/86Combinations of lidar systems with systems other than lidar, radar or sonar, e.g. with direction finders
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01SRADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S17/00Systems using the reflection or reradiation of electromagnetic waves other than radio waves, e.g. lidar systems
    • G01S17/88Lidar systems specially adapted for specific applications
    • G01S17/89Lidar systems specially adapted for specific applications for mapping or imaging
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01SRADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S17/00Systems using the reflection or reradiation of electromagnetic waves other than radio waves, e.g. lidar systems
    • G01S17/88Lidar systems specially adapted for specific applications
    • G01S17/93Lidar systems specially adapted for specific applications for anti-collision purposes
    • G01S17/931Lidar systems specially adapted for specific applications for anti-collision purposes of land vehicles
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01SRADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S17/00Systems using the reflection or reradiation of electromagnetic waves other than radio waves, e.g. lidar systems
    • G01S17/02Systems using the reflection of electromagnetic waves other than radio waves
    • G01S17/06Systems determining position data of a target
    • G01S17/42Simultaneous measurement of distance and other co-ordinates

Definitions

  • Autonomous operations such as robotic operations or autonomous vehicle operations, in unknown or dynamic environments present various technical challenges.
  • Autonomous operations in dynamic environments may be applied to mass customization (e.g., high-mix, low-volume manufacturing), on-demand flexible manufacturing processes in smart factories, warehouse automation in smart stores, automated deliveries from distribution centers in smart logistics, and the like.
  • robots for instance mobile robots or automated guided vehicles (AGVs)
  • AGVs automated guided vehicles
  • 2D lasers for navigation that can include obstacle detection and obstacle avoidance as well as free path planning.
  • Embodiments of the invention address and overcome one or more of the described- herein shortcomings by providing methods, systems, and apparatuses that detect obstacles at various heights along various trajectories of a robot, such that the robot can avoid collisions with objects in various unknown and dynamic environments.
  • three-dimensional (3D) point cloud data can be captured and processed as two-dimensional (2D) laser data, so as to generate more accurate maps, such as maps displaying a wider range of obstacles at different heights, as compared to previous approaches that rely on 2D lasers.
  • an autonomous system is configured to operate in a dynamic industrial environment so as to define a runtime.
  • the autonomous system can include a camera configured to capture a depth image of the dynamic industrial environment, so as to define a captured image that defines a plurality of rays.
  • the autonomous system can further include one or more processors, and a memory storing instructions that, when executed by the one or more processors, cause the autonomous system to perform various operations during the runtime.
  • the operations can include computing ray information associated with the plurality of rays. Based on the ray information, the system can transform the captured image so as to remove a set of rays from the plurality of rays, thereby defining laser data from the plurality of rays.
  • the system can also process the laser data so as to detect one or more obstacles within the environment.
  • the camera defines a three-dimensional (3D) sensor.
  • the autonomous system includes a mobile robot that defines a bottom end configured to face a floor of the dynamic industrial environment, and a top end opposite the bottom end along a first direction.
  • the mobile robot can further define a front end and rear end opposite the front end along a second direction that is substantially perpendicular to the first direction, and a first side and second side opposite the first side along a third direction that is substantially perpendicular to both the first and second directions.
  • the 3D camera can be disposed at the front end, rear end, first side, second side, or top end, and the camera is configured to capture 3D point cloud data from the front end along the first, second, and third directions.
  • the captured image defines a 3D point cloud comprising a plurality of 3D points.
  • the system can project the plurality of 3D points into respective one dimensional points from the camera along the second or third directions.
  • the system can determine a predetermined distance measured from the camera along the second or third directions; identify one or more points of the plurality of one dimensional points that define a distance from the camera that is greater than the predetermined distance; and filter the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data.
  • the system can determine a predetermined height measured from the camera along the first direction; identify one or more points of the plurality of one dimensional points that define a distance above or below the camera along the first direction that is greater than the predetermined height; and filter the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data.
  • the autonomous system further comprises a laser sensor disposed at, for example, the front end and/or the rear end of the robot.
  • the system can also capture laser sensor data from the laser sensor; and merge the laser sensor data from the laser sensor with the laser data from the 3D sensor.
  • FIG. 1 shows an example automated guided vehicle (AGV) or mobile robot, for instance an autonomous mobile manipulator robot (AMMR), in an example physical environment, in accordance with an example embodiment.
  • AGV automated guided vehicle
  • AMMR autonomous mobile manipulator robot
  • FIG. 2 is a flow diagram that depicts operations that can be performed by the AGV or mobile robot to process a 3D point cloud from a 3D sensor into a laser data stream.
  • FIG. 3 illustrates an example 3D point cloud that includes points that can be transformed and filtered for processing as a 2D laser.
  • FIG. 4 illustrates an example point in a 3D point cloud that can be projected as a ID vector.
  • FIG. 5 illustrates a computing environment within which embodiments of the disclosure may be implemented.
  • Robots in an industrial setting are often described herein for purposes of examples, though it will be understood that embodiments described herein are not limited to robots in an industrial setting, and all alternative autonomous devices and settings are contemplated as being within the scope of this disclosure.
  • obstacles at or below the level or height of a laser scanner such as the 2D scanner at the fixed height described above, might be undetected by a given robot.
  • a laser scanner at a fixed height might not detect a forklift having a fork of limited thickness that moves up and down so as to be disposed at varying heights.
  • a good position of a 2D laser scanner for mapping and localization is often at a higher level than what is practical (around the height of a human eye), as static objects (e.g., walls, pillars, etc.) have a higher chance of being clearly visible around this height because static objects are less obstructed by movable objects (e.g., other robots, tables, boxes, etc.) at this height as compared to lower heights at which a laser scanner is typically disposed on a moveable robot to recognized various dynamic objects.
  • movable objects e.g., other robots, tables, boxes, etc.
  • Dynamic objects objects that move
  • objects at a range of heights should be detected and processed to prevent standstills or other operational problems.
  • Three-dimensional (3D) laser scanners and depth cameras can give information about objects in the environment within a wider field of view as compared to 2D scanners, for instance a height range of several meters and detection in three dimensions.
  • Depth cameras are typically cheaper than 3D laser scanners, however, it is recognized herein that depth cameras are often difficult to use directly as an input for a navigation system.
  • the computational cost can increase as the points in space of a given point cloud increase.
  • color information can create a data bus of information that might not be processed by a single GPU/CPU device. It is further recognized herein that current mapping and navigation software typically do not use point clouds as input to generate maps and navigate an environment.
  • an obstacle detection system that processes 2D laser data is configured to also process obstacle information from a depth camera.
  • current mapping and navigation software can process the obstacle information that is generated from a depth camera without making direct changes to the software itself.
  • an example industrial or dynamic physical environment 100 can include a computerized autonomous system 102 that can include one or more mobile robot devices or automated guided vehicles (AGVs), for instance an AGV or robot device 104, configured to perform one or more industrial tasks, such as bin picking, grasping, transport, or the like.
  • the system 102 can include one or more computing processors configured to process information and control operations of the system 102, in particular the AGV 104.
  • the AGV 104 can include one or more processors, for instance a processor 108, configured to process information and/or control various operations associated with the AGV 104.
  • An autonomous system for operating an AGV within a physical environment can further include a memory for storing modules.
  • the processors can further be configured to execute the modules so as to process information and generate models based on the information. It will be understood that the illustrated environment 100 and the system 102 are simplified for purposes of example. The environment 100 and the system 102 may vary as desired, and all such systems and environments are contemplated as being within the scope of this disclosure.
  • the AGV or robot 104 is presented as an example of a device using an obstacle avoidance system, though it will be understood that the robot 104 can define additional or alternative devices, such as cars, drones, or the like, and all devices that perform 2D navigation or obstacle avoidance are contemplated as being within the scope of this disclosure.
  • the AGV 104 can further include a robotic arm or manipulator 110 and a base 112 configured to support the robotic manipulator 110.
  • the base 112 can include wheels 114 or can otherwise be configured to move within the physical environment 100.
  • the autonomous machine 104 can further include an end effector 116 attached to the robotic manipulator 110.
  • the end effector 116 can include one or more tools configured to grasp and/or move objects.
  • Example end effectors 116 include finger grippers or vacuum-based grippers.
  • the robotic manipulator 110 can be configured to move so as to change the position of the end effector 116, for example, so as to place or move objects within the physical environment 100.
  • the AGV 104 can further include one or more cameras or sensors, for instance a depth camera 118, configured to detect or record objects within the physical environment 100.
  • the camera 118 can be mounted to the body or base of the AGV 104, for instance at a front end 117 of the AGV 104, or otherwise configured to generate images of the physical environment 100 along a trajectory of the AGV 104.
  • the camera 118 can be configured as an RGB-D camera defining a color and depth channel configured to capture images of the environment 100 as the AGV 104 moves within the environment 100.
  • the AGV 104 in particular the base 112 of the AGV 104, can define a top 109 end and a bottom end 111 opposite the top end 109 along a transverse direction 120.
  • the AGV 104 can further define a first side 113 and a second side 115 opposite the first side 113 along a second or lateral direction 122 that is substantially perpendicular to the transverse direction 120.
  • the AGV 107 can further define a front end 117 and a rear end 119 opposite the front end 117 along a third or longitudinal direction 124 that is substantially perpendicular to both the transverse and lateral directions 120 and 122, respectively.
  • the illustrated AGV 104 defines a rectangular shape, it will be understood that AGVs or robots can be alternatively shaped or sized, and all such AGVs or robots are contemplated as being within the scope of this disclosure.
  • the camera 118 can define a depth camera configured to capture depth images of the workspace 100 from the front end 117. Alternatively, or additionally, the camera 118 can be mounted or attached to another end the base 112 or portion of the robot 104 (for instance the arm 110), or to another device entirely.
  • the data captured by the camera 118 can processed by the system 102 or transmitted to a navigation system and processed as laser data.
  • example operations 200 are shown that can be performed by a computing system, for instance the autonomous system 102 or AGV 104.
  • the system 102 for instance the camera 118, can capture an image of the environment 100, so as to define a captured depth image.
  • the captured image can define RGB (color) or RGB-D information (color and depth information) corresponding to the environment 100.
  • the captured image can define a 3D camera input that includes 3D-point cloud data 300 from a depth camera associated with objects or obstacles within the physical environment 100.
  • the system can compute information associated with rays in the image or frame captured at 202, for instance the number and orientation of rays. For example, at 204, the system can compute the number of rays that are needed for a specific laser input. Such a computation can be based on minimum and maximum angles (or rays) and angle resolution.
  • a given point cloud, for instance the point cloud 300, captured at 202 can be composed of n number of three-dimensional points.
  • the 3D points can have coordinate values along the transverse direction 120, lateral direction 122, and longitudinal direction 124.
  • the camera 118 that captures the 3D point cloud 300 defines the origin 302 (0, 0, 0) of the 3D point cloud 300.
  • a range (or number) of the n points can be projected as a one-dimensional (ID) vector.
  • the range of points can be defined between two rays, for instance a minimum ray 306 and a maximum ray 304 defined with respect to a zero angle or level 308 along the transverse direction 120 from the camera 118.
  • the maximum ray 304 (or angle 306) can define a ray that travels upward along the transverse direction 120 as it travels from the camera 118 along the lateral or longitudinal directions 122 and 124
  • the minimum ray 306 (or angle) can define a ray that travels downward along the transverse direction 120 as it travels from the camera 118 along the lateral or longitudinal directions 122 and 124.
  • the maximum and minimum rays (or angles) 304 and 306 can define a section 310 of the point cloud 300.
  • the rays 304 and 306 can define boundaries of the section 310 of the point cloud 300, such that the number n of points are disposed within the boundaries defined by the maximum ray or angle 304 and the minimum ray or angle 306.
  • the number n of 3D points contained in the section 310 can vary depending on the angle resolution. In various examples, the higher the angle resolution, the greater the number n of points contained in the section 310, so more points can be processed and projected as the resolution increases. In some cases, minimum and maximum angles are user-defined, or can be based in user inputs to the system. Alternatively, or additionally, the 3D sensor 118 can define limitations, in particular a limit of angle ranges or a limit of resolution, that can determine a given section of the 3D point cloud or the number n of points in a given section.
  • the system can transform the 3D image captured at 202, for instance 3D point cloud 300, so as to remove invalid rays or points from the point cloud 300.
  • filtering is performed to remove points out of range from the main point cloud 300, for instance to remove points outside the section 310 defined by the maximum ray 304 and the minimum ray 306.
  • the section 310 can be defined by a range of distances.
  • the system might filter out points that are less than 1 meter (m) from the camera 118 or robot 104, and greater than 4m from the camera 118 or robot 104, such that the points that are processed are between Im and 4m from the camera 118.
  • clusterization can be performed to remove invalid points.
  • the 3D point cloud 300 can be transformed into a specific target frame, for instance the frame of a laser input of the robot 104.
  • This target frame may correspond to a frame of a 2D laser that is processed, for instance at 212, with the 3D point cloud 300 that is converted to laser data.
  • each part or portion of the robot 104 is defined with a frame, or a set of coordinates that specifies the location in space of each respective part.
  • One part or portion of the robot defines the center or origin of the frame.
  • the center can be defined by the center of the top end 109, though it will be understood that the center (origin) can vary as desired, for instance the center of a wheel 114, the center of the front end 117, or the like, and all such origins (frame centers) are contemplated as being within the scope of this disclosure.
  • coordinate frames can be determined for other parts of the robot 104, for instance the wheels 114, the sensor 118, and the like.
  • the laser when a laser is used to detect objects, the laser has a coordinate frame, and distances defined by the objects in the environment can be determined based on the laser coordinate frame. Those distances can be transformed based on the predefined origin, so that they can be understood.
  • the system can determine the separation between the origin of the robot 104, the sensor 118, an outermost edge of the robot 104, and the measured distance of the object itself in the laser coordinate frame to estimate when the object might collide with a portion of the robot 104, for instance the front end 117 of the robot 104.
  • the coordinate frame of the second laser can be transformed to align with the base (original) laser that might be used for navigation. If the coordinate frames are not aligned, one of the lasers might generate distances that do not make sense as compared to the distances that the other laser is generating.
  • a new sensor’s frame is transformed to match the main sensor or set of sensors, so that their respective distances are aligned with each other.
  • the frame defined by a depth camera or 3D sensor of the robot 104 can be transformed into the frame defined by the laser sensor of the robot 104.
  • the 3D sensor or depth camara can be disposed anywhere on the robot 104 in any direction relative to the robot 104, because the frame transformation matches the distance points from the 3D sensor to the 2D laser data that might be generated from the rear end 119 and/or front end 117 of the robot 104.
  • the AGV 104 in particular the sensor 118, can define a 3D sensor and a laser sensor.
  • 2D laser data can be captured by the laser sensor.
  • the laser data can be merged with the point cloud data captured by the 3D sensor at 202, such that the data captured by the laser sensor and the data captured by the 3D sensor are processed together, for example, to calculate positions of obstacles in a given environment (e.g., the environment 100).
  • the system can adapt a scan range, so as to generate laser data from the captured 3D input.
  • the scan range can define a maximum distance between objects and the 3D sensor 118 at which the which the system can process the points corresponding to those objects, so as to match any other frame.
  • the system transform the frame from the 3D sensor to the base laser sensor that might be used for navigation and obstacle avoidance, so as to merge the 3D point cloud data with laser data, so that both sets of data can be processed as lasers, at 212.
  • the maximum distance can define a predetermined distance, for instance along the lateral or longitudinal directions 122 and 124, from the robot 104 at which the robot might collide with a given object.
  • data that is outside of the predetermined distance can be filtered out before further processing at 212, thereby reducing the amount of data that is processed.
  • the maximum distance can be defined by a height range, for instance a height range 312 defined between a maximum height 314 and a minimum height 316 along the transverse direction 120. Points that are outside of the defined angles and/or height range can be removed, such that the range can be adapted, for instance to a smaller scan range.
  • parameters that determine the scan range such as maximum or minimum angles, a maximum distance or range, angular resolution, and height range, are user-defined. For example, users can define the parameters based on the particular robot and environment in which the robot interacts, such that the robot detects obstacles.
  • the user can set the respective parameters for the minimum height 316 and the maximum height 314 along the transverse direction 120.
  • the light of sight of the 3D sensor corresponds to the zero level 308.
  • multiple height ranges can be selected.
  • a first height range can be selected such that a first laser data set is generated from a range of 50 cm to 150cm
  • a second height range can be selected such that a second laser data set is generated from 0 cm to 20 cm below the level defined by a 2D laser (e.g., -20 cm from zero level 308) along the transverse direction 120, though it will be understood that various distances or height ranges can vary as desired, and all such alternative or additional distances or ranges are contemplated as being within the scope of this disclosure.
  • the system can generate one or more sets of projected laser data, for instance multiple sets based on respective different height ranges, based on a single point cloud captured at 202.
  • the system might only be interested in detecting obstacles in front of the robot.
  • the system can remove points from a point cloud, for instance a point cloud 400 (see FIG. 4), which represent points or objects that are adjacent to both sides 113 and 115 of the robot.
  • an angle range can be selected such only points are selected that are in front of the robot 104.
  • a distance range can be defined so that the system only detects obstacles that are relevant to a particular user, for instance sufficiently close to a trajectory of the robot or the robot itself.
  • the distance range can define how far, for instance how many meters or the like, from the front end 117 along the lateral direction 122 or longitudinal direction 124, within which points are selected.
  • the point cloud 400 can be captured (at 202) that includes a first 3D point 404.
  • the point can be projected as a one-dimensional (ID) value, which together with other projected points can form a ID vector of distances, so as to define a second or projected point 406.
  • ID one-dimensional
  • a distance 402 from the sensor 118 along the lateral or longitudinal directions 122 and 124, respectively, can be determined.
  • the distance 402 can be compared to the predetermined distance range (e.g., 5m). If the distance 402 is greater than the distance range, the point 406 can be filtered out, at 208, such that the point 406 is not processed at 212. If the distance 402 is less than or equal to the distance range, the point 406 can be processed, at 212, with other laser data that is captured at 210. It is recognized herein that, in some cases in which an objective of the system is to detect obstacles, a high angle resolution might not be required, as one or more points can confirm the existence of an obstacle sufficiently to avoid the obstacle without determining precise shape of the obstacle.
  • the image illustrates an example of projecting a point from 3D to ID.
  • the image also illustrates, however, that the point corresponds to two dimensions because it is projected along the lateral and longitudinal directions 122 and 124 (e.g., xy plane) so as to be defined by an angle and a distance.
  • the laser data can define a unidimensional vector having points that correspond to a specific angle.
  • the laser data can define a ID vector, because each value can be processed because each angle is known.
  • the resulting vector can include [INF, 4, 8, 80, 2] assuming the laser has a range from -90 to 90.
  • the laser can generate a value for a laser beam at -90, then -45, then 0, then 45, then 90, assuming the 0 as the vertical axis. It is recognized herein that various algorithms for navigation and obstacle avoidance need to know these parameters in order to use this information accordingly, so when the system converts from a point in 3D to ID, the system can use other dimensions, such as the angles, and parameters to make sense of those data points.
  • 3D-point cloud data from a depth camera can be projected into the data format of 2D laser scanners.
  • various laser data such as, for example, the angle, distance range of detection, and the resolution can be adjusted according to the required laser.
  • the laser data can be adjusted based on specific obstacles or objects within the environment 100.
  • the baseline laser can be adjusted based on the height (e.g., height range 312) defined by obstacles within the environment 100 along the transverse direction 120.
  • a projected point cloud captured at 202 can take the form of a 2D laser sensor.
  • the 3D camera input or image captured at 202 can define points along the lateral, longitudinal, and transverse directions (x, y, z).
  • the system can transform each data point (x, y, z) into two dimensions, for instance a first dimension representative of an angle and a second dimension representative of a distance, thereby transforming the 3D camera input into a linear array of distances.
  • the system can transform the first point 404 (aP) having coordinates (Xa, Ya, Za) into the second point 406 that is the distance 402 from the sensor 118.
  • An array of such transformed data in various examples, can be sent or fed to a laser input port of a laser-based navigation system for mobile robots, at 212.
  • the system can also capture 2D laser data from a laser input, for instance the sensor 118 that can define a 3D camera and a laser.
  • the system can process the linear array of distances transformed from the point cloud data, and the 2D layer data captured from the 2D sensor data, together for conflict avoidance using a local path planner or using a simultaneous localization and mapping (SLAM) algorithm.
  • SLAM simultaneous localization and mapping
  • depth channel data from any type of 3D camera that captures 3D point clouds can be processed by during operations 200.
  • the output at 208 can define a set of points in the same data format as lidar sensors.
  • the output can be processed as lidar sensor data, even though the 3D camera that captures the image at 202 might be less expensive than a lidar sensor.
  • a robot might scan and map a warehouse with a single 2D laser scanner.
  • the system performs operations 200 and correctly detects the obstacles (e.g., shelves) and generates the obstacle on the map, thereby protecting the robot from future collisions by computing the correct navigation trajectories through the warehouse.
  • the output at 208 can be used for reliable obstacle detection (at 212) within 2D lidar sensor-based navigation systems for mobile robots, for example, because the depth camera 118 can cover a larger height range as compared to a 2D laser scanner.
  • obstacles above and below a give 2D-laser scanner plane can reliably be detected.
  • a height range can enable detection of forklift forks, human feet, shelves at warehouses and factories, and the like, which are currently difficult to detect due to their different heights.
  • the system can also perform operations 200 so as to detect various small objects on the factory floor (e.g., screws, tools, etc.) that currently cause danger to the robots or AGVs when moving over them.
  • point cloud data from 3D-cameras can be transformed into laser data types, so that such data can be processed (at 212) by laser-based navigation systems without any adaptations to existing software, thereby creating cost and processing overhead efficiencies.
  • the data of lidar-sensors and depth cameras can also be merged together (e.g., at 212) after feeding the camera data into the same input as the laser data, such that the same type of SLAM algorithm, localization, path planning, and/or obstacle avoidance algorithm can be used that are used in current 2D laser-based systems, while adding obstacles at different height levels to the map, thereby enabling various navigation systems to process the 3D point cloud data as 2D laser data.
  • additional features above or below the 2D-laser height level can be added to the map (e.g., point cloud 300) using the depth camera so as to support the precise localization of the robot.
  • the system allows the replacement of one or more laser scanners on the mobile robot without adapting the navigation software or creating a new interface.
  • depth cameras can also be the more costefficient solution compared to laser-scanners for shorter ranges.
  • FIG. 5 illustrates an example of a computing environment within which embodiments of the present disclosure may be implemented.
  • a computing environment 600 includes a computer system 610 that may include a communication mechanism such as a system bus 621 or other communication mechanism for communicating information within the computer system 610.
  • the computer system 610 further includes one or more processors 620 coupled with the system bus 621 for processing the information.
  • the autonomous system 102 in particular the robot 104, may include, or be coupled to, the one or more processors 620.
  • the processors 620 may include one or more central processing units (CPUs), graphical processing units (GPUs), or any other processor known in the art. More generally, a processor as described herein is a device for executing machine-readable instructions stored on a computer readable medium, for performing tasks and may comprise any one or combination of, hardware and firmware. A processor may also comprise memory storing machine-readable instructions executable for performing tasks. A processor acts upon information by manipulating, analyzing, modifying, converting or transmitting information for use by an executable procedure or an information device, and/or by routing the information to an output device.
  • CPUs central processing units
  • GPUs graphical processing units
  • a processor may use or comprise the capabilities of a computer, controller or microprocessor, for example, and be conditioned using executable instructions to perform special purpose functions not performed by a general-purpose computer.
  • a processor may include any type of suitable processing unit including, but not limited to, a central processing unit, a microprocessor, a Reduced Instruction Set Computer (RISC) microprocessor, a Complex Instruction Set Computer (CISC) microprocessor, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a System-on-a-Chip (SoC), a digital signal processor (DSP), and so forth.
  • RISC Reduced Instruction Set Computer
  • CISC Complex Instruction Set Computer
  • ASIC Application Specific Integrated Circuit
  • FPGA Field-Programmable Gate Array
  • SoC System-on-a-Chip
  • DSP digital signal processor
  • processor(s) 620 may have any suitable microarchitecture design that includes any number of constituent components such as, for example, registers, multiplexers, arithmetic logic units, cache controllers for controlling read/write operations to cache memory, branch predictors, or the like.
  • the microarchitecture design of the processor may be capable of supporting any of a variety of instruction sets.
  • a processor may be coupled (electrically and/or as comprising executable components) with any other processor enabling interaction and/or communication there-between.
  • a user interface processor or generator is a known element comprising electronic circuitry or software or a combination of both for generating display images or portions thereof.
  • a user interface comprises one or more display images enabling user interaction with a processor or other device.
  • the system bus 621 may include at least one of a system bus, a memory bus, an address bus, or a message bus, and may permit exchange of information (e.g., data (including computer-executable code), signaling, etc.) between various components of the computer system 610.
  • the system bus 621 may include, without limitation, a memory bus or a memory controller, a peripheral bus, an accelerated graphics port, and so forth.
  • the system bus 621 may be associated with any suitable bus architecture including, without limitation, an Industry Standard Architecture (ISA), a Micro Channel Architecture (MCA), an Enhanced ISA (EISA), a Video Electronics Standards Association (VESA) architecture, an Accelerated Graphics Port (AGP) architecture, a Peripheral Component Interconnects (PCI) architecture, a PCI -Express architecture, a Personal Computer Memory Card International Association (PCMCIA) architecture, a Universal Serial Bus (USB) architecture, and so forth.
  • ISA Industry Standard Architecture
  • MCA Micro Channel Architecture
  • EISA Enhanced ISA
  • VESA Video Electronics Standards Association
  • AGP Accelerated Graphics Port
  • PCI Peripheral Component Interconnects
  • PCMCIA Personal Computer Memory Card International Association
  • USB Universal Serial Bus
  • the computer system 610 may also include a system memory 630 coupled to the system bus 621 for storing information and instructions to be executed by processors 620.
  • the system memory 630 may include computer readable storage media in the form of volatile and/or nonvolatile memory, such as read only memory (ROM) 631 and/or random-access memory (RAM) 632.
  • the RAM 632 may include other dynamic storage device(s) (e.g., dynamic RAM, static RAM, and synchronous DRAM).
  • the ROM 631 may include other static storage device(s) (e.g., programmable ROM, erasable PROM, and electrically erasable PROM).
  • system memory 630 may be used for storing temporary variables or other intermediate information during the execution of instructions by the processors 620.
  • a basic input/output system 633 (BIOS) containing the basic routines that help to transfer information between elements within computer system 610, such as during start-up, may be stored in the ROM 631.
  • RAM 632 may contain data and/or program modules that are immediately accessible to and/or presently being operated on by the processors 620.
  • System memory 630 may additionally include, for example, operating system 634, application programs 635, and other program modules 636.
  • Application programs 635 may also include a user portal for development of the application program, allowing input parameters to be entered and modified as necessary.
  • the operating system 634 may be loaded into the memory 630 and may provide an interface between other application software executing on the computer system 610 and hardware resources of the computer system 610. More specifically, the operating system 634 may include a set of computer-executable instructions for managing hardware resources of the computer system 610 and for providing common services to other application programs (e.g., managing memory allocation among various application programs). In certain example embodiments, the operating system 634 may control execution of one or more of the program modules depicted as being stored in the data storage 640.
  • the operating system 634 may include any operating system now known or which may be developed in the future including, but not limited to, any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.
  • the computer system 610 may also include a disk/media controller 643 coupled to the system bus 621 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 641 and/or a removable media drive 642 (e.g., floppy disk drive, compact disc drive, tape drive, flash drive, and/or solid-state drive).
  • Storage devices 640 may be added to the computer system 610 using an appropriate device interface (e.g., a small computer system interface (SCSI), integrated device electronics (IDE), Universal Serial Bus (USB), or FireWire).
  • Storage devices 641 , 642 may be external to the computer system 610.
  • the computer system 610 may also include a field device interface 665 coupled to the system bus 621 to control a field device 666, such as a device used in a production line.
  • the computer system 610 may include a user input interface or GUI 661, which may comprise one or more input devices, such as a keyboard, touchscreen, tablet and/or a pointing device, for interacting with a computer user and providing information to the processors 620.
  • the computer system 610 may perform a portion or all of the processing steps of embodiments of the invention in response to the processors 620 executing one or more sequences of one or more instructions contained in a memory, such as the system memory 630. Such instructions may be read into the system memory 630 from another computer readable medium of storage 640, such as the magnetic hard disk 641 or the removable media drive 642.
  • the magnetic hard disk 641 (or solid-state drive) and/or removable media drive 642 may contain one or more data stores and data files used by embodiments of the present disclosure.
  • the data store 640 may include, but are not limited to, databases (e.g., relational, object-oriented, etc.), file systems, flat files, distributed data stores in which data is stored on more than one node of a computer network, peer-to-peer network data stores, or the like.
  • the data stores may store various types of data such as, for example, skill data, sensor data, or any other data generated in accordance with the embodiments of the disclosure.
  • Data store contents and data files may be encrypted to improve security.
  • the processors 620 may also be employed in a multi-processing arrangement to execute the one or more sequences of instructions contained in system memory 630.
  • hard-wired circuitry may be used in place of or in combination with software instructions. Thus, embodiments are not limited to any specific combination of hardware circuitry and software.
  • the computer system 610 may include at least one computer readable medium or memory for holding instructions programmed according to embodiments of the invention and for containing data structures, tables, records, or other data described herein.
  • the term “computer readable medium” as used herein refers to any medium that participates in providing instructions to the processors 620 for execution.
  • a computer readable medium may take many forms including, but not limited to, non-transitory, non-volatile media, volatile media, and transmission media.
  • Non-limiting examples of non-volatile media include optical disks, solid state drives, magnetic disks, and magneto-optical disks, such as magnetic hard disk 641 or removable media drive 642.
  • Non-limiting examples of volatile media include dynamic memory, such as system memory 630.
  • Non-limiting examples of transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up the system bus 621. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
  • Computer readable medium instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages.
  • the computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
  • the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
  • LAN local area network
  • WAN wide area network
  • Internet Service Provider for example, AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
  • electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
  • FPGA field-programmable gate arrays
  • PLA programmable logic arrays
  • the computing environment 600 may further include the computer system 610 operating in a networked environment using logical connections to one or more remote computers, such as remote computing device 680.
  • the network interface 670 may enable communication, for example, with other remote devices 680 or systems and/or the storage devices 641, 642 via the network 671.
  • Remote computing device 680 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above relative to computer system 610.
  • computer system 610 may include modem 672 for establishing communications over a network 671, such as the Internet. Modem 672 may be connected to system bus 621 via user network interface 670, or via another appropriate mechanism.
  • Network 671 may be any network or system generally known in the art, including the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a direct connection or series of connections, a cellular telephone network, or any other network or medium capable of facilitating communication between computer system 610 and other computers (e.g., remote computing device 680).
  • the network 671 may be wired, wireless or a combination thereof. Wired connections may be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection generally known in the art.
  • Wireless connections may be implemented using Wi-Fi, WiMAX, and Bluetooth, infrared, cellular networks, satellite or any other wireless connection methodology generally known in the art. Additionally, several networks may work alone or in communication with each other to facilitate communication in the network 671.
  • program modules, applications, computer-executable instructions, code, or the like depicted in FIG. 5 as being stored in the system memory 630 are merely illustrative and not exhaustive and that processing described as being supported by any particular module may alternatively be distributed across multiple modules or performed by a different module.
  • various program module(s), script(s), plug-in(s), Application Programming Interface(s) (API(s)), or any other suitable computer-executable code hosted locally on the computer system 610, the remote device 680, and/or hosted on other computing device(s) accessible via one or more of the network(s) 671 may be provided to support functionality provided by the program modules, applications, or computer-executable code depicted in FIG.
  • program modules that support the functionality described herein may form part of one or more applications executable across any number of systems or devices in accordance with any suitable computing model such as, for example, a client-server model, a peer-to-peer model, and so forth.
  • the computer system 610 may include alternate and/or additional hardware, software, or firmware components beyond those described or depicted without departing from the scope of the disclosure.
  • functionality described as being provided by a particular module may, in various embodiments, be provided at least in part by one or more other modules. Further, one or more depicted modules may not be present in certain embodiments, while in other embodiments, additional modules not depicted may be present and may support at least a portion of the described functionality and/or additional functionality. Moreover, while certain modules may be depicted and described as sub-modules of another module, in certain embodiments, such modules may be provided as independent modules or as sub-modules of other modules.
  • any operation, element, component, data, or the like described herein as being based on another operation, element, component, data, or the like can be additionally based on one or more other operations, elements, components, data, or the like. Accordingly, the phrase “based on,” or variants thereof, should be interpreted as “based at least in part on.”
  • each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s).
  • the functions noted in the block may occur out of the order noted in the Figures.
  • two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Radar, Positioning & Navigation (AREA)
  • Remote Sensing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • General Physics & Mathematics (AREA)
  • Electromagnetism (AREA)
  • Control Of Position, Course, Altitude, Or Attitude Of Moving Bodies (AREA)

Abstract

Current approaches to obstacle detection can result in some obstacles being undetected, for instance obstacles that are higher or lower than the height of a two-dimensional (2D) laser, or obstacles defining particular thicknesses, thereby resulting in a collision or a risk of collision between the robot or autonomous device and the undetected obstacle. Using a three- dimensional (3D) camera or 3D sensor input, such as a depth camera, an automated guided vehicle (AGV), drone, or robot can detect obstacles at various heights along various trajectories of the robot, such that the robot can avoid collisions with objects in various unknown and dynamic environments, or use the extra data information for navigation and path planning.

Description

3D SENSOR-BASED OBSTACLE DETECTION FOR AUTONOMOUS VEHICLES AND
MOBILE ROBOTS
BACKGROUND
[0001] Autonomous operations, such as robotic operations or autonomous vehicle operations, in unknown or dynamic environments present various technical challenges. Autonomous operations in dynamic environments may be applied to mass customization (e.g., high-mix, low-volume manufacturing), on-demand flexible manufacturing processes in smart factories, warehouse automation in smart stores, automated deliveries from distribution centers in smart logistics, and the like. In most cases, robots, for instance mobile robots or automated guided vehicles (AGVs), rely on two-dimensional (2D) lasers for navigation that can include obstacle detection and obstacle avoidance as well as free path planning. It is recognized herein, however, that current approaches to obstacle detection can result in some obstacles being undetected, for instance obstacles at heights lower or higher than the 2D laser line, or obstacles defining particular thicknesses, thereby resulting in a collision or a risk of collision between the robot and the undetected obstacle.
BRIEF SUMMARY
[0002] Embodiments of the invention address and overcome one or more of the described- herein shortcomings by providing methods, systems, and apparatuses that detect obstacles at various heights along various trajectories of a robot, such that the robot can avoid collisions with objects in various unknown and dynamic environments. For example, in accordance with various embodiments, three-dimensional (3D) point cloud data can be captured and processed as two-dimensional (2D) laser data, so as to generate more accurate maps, such as maps displaying a wider range of obstacles at different heights, as compared to previous approaches that rely on 2D lasers.
[0003] In an example aspect, an autonomous system is configured to operate in a dynamic industrial environment so as to define a runtime. The autonomous system can include a camera configured to capture a depth image of the dynamic industrial environment, so as to define a captured image that defines a plurality of rays. The autonomous system can further include one or more processors, and a memory storing instructions that, when executed by the one or more processors, cause the autonomous system to perform various operations during the runtime. The operations can include computing ray information associated with the plurality of rays. Based on the ray information, the system can transform the captured image so as to remove a set of rays from the plurality of rays, thereby defining laser data from the plurality of rays. The system can also process the laser data so as to detect one or more obstacles within the environment. In some examples, the camera defines a three-dimensional (3D) sensor. In another example aspect, the autonomous system includes a mobile robot that defines a bottom end configured to face a floor of the dynamic industrial environment, and a top end opposite the bottom end along a first direction. The mobile robot can further define a front end and rear end opposite the front end along a second direction that is substantially perpendicular to the first direction, and a first side and second side opposite the first side along a third direction that is substantially perpendicular to both the first and second directions. In various examples, the 3D camera can be disposed at the front end, rear end, first side, second side, or top end, and the camera is configured to capture 3D point cloud data from the front end along the first, second, and third directions.
[0004] In another example aspect, the captured image defines a 3D point cloud comprising a plurality of 3D points. The system can project the plurality of 3D points into respective one dimensional points from the camera along the second or third directions. The system can determine a predetermined distance measured from the camera along the second or third directions; identify one or more points of the plurality of one dimensional points that define a distance from the camera that is greater than the predetermined distance; and filter the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data. In another example, during the runtime, the system can determine a predetermined height measured from the camera along the first direction; identify one or more points of the plurality of one dimensional points that define a distance above or below the camera along the first direction that is greater than the predetermined height; and filter the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data. In various examples, the autonomous system further comprises a laser sensor disposed at, for example, the front end and/or the rear end of the robot. Thus, the system can also capture laser sensor data from the laser sensor; and merge the laser sensor data from the laser sensor with the laser data from the 3D sensor. BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0005] The foregoing and other aspects of the present invention are best understood from the following detailed description when read in connection with the accompanying drawings. For the purpose of illustrating the invention, there is shown in the drawings embodiments that are presently preferred, it being understood, however, that the invention is not limited to the specific instrumentalities disclosed. Included in the drawings are the following Figures:
[0006] FIG. 1 shows an example automated guided vehicle (AGV) or mobile robot, for instance an autonomous mobile manipulator robot (AMMR), in an example physical environment, in accordance with an example embodiment.
[0007] FIG. 2 is a flow diagram that depicts operations that can be performed by the AGV or mobile robot to process a 3D point cloud from a 3D sensor into a laser data stream.
[0008] FIG. 3 illustrates an example 3D point cloud that includes points that can be transformed and filtered for processing as a 2D laser.
[0009] FIG. 4 illustrates an example point in a 3D point cloud that can be projected as a ID vector.
[0010] FIG. 5 illustrates a computing environment within which embodiments of the disclosure may be implemented.
DETAILED DESCRIPTION
[0011] As an initial matter, it is recognized herein that current robots typically rely on a two- dimensional (2D) laser scanner input for navigation as well as obstacle detection and avoidance. Such laser scanner input is typically provided by a 2D laser at a fixed height relative to the robot, wherein the 2D layer generally defines only one laser beam for detection on a single plane, such that maps can be computed expeditiously with low-level computing capabilities. As used herein, unless otherwise specified, autonomous mobile robots (AMRs) or robots, autonomous devices or vehicles, automated guided vehicles (AGVs), and the like can be used interchangeably herein, without limitation. Robots in an industrial setting are often described herein for purposes of examples, though it will be understood that embodiments described herein are not limited to robots in an industrial setting, and all alternative autonomous devices and settings are contemplated as being within the scope of this disclosure. [0012] It is further recognized herein, however, that obstacles at or below the level or height of a laser scanner, such as the 2D scanner at the fixed height described above, might be undetected by a given robot. By way of example, a laser scanner at a fixed height might not detect a forklift having a fork of limited thickness that moves up and down so as to be disposed at varying heights. It is further recognized herein that, when mobile robots are operating in environments with humans, a laser scanner at a fixed height can result in undetected obstacles at high or low heights (e.g., feet, shelves, tripods, and the like) colliding with the robots, thereby violating safety regulations for an industrial environment and potentially creating dangerous situations. Further still, it is recognized herein that a good position of a 2D laser scanner for mapping and localization is often at a higher level than what is practical (around the height of a human eye), as static objects (e.g., walls, pillars, etc.) have a higher chance of being clearly visible around this height because static objects are less obstructed by movable objects (e.g., other robots, tables, boxes, etc.) at this height as compared to lower heights at which a laser scanner is typically disposed on a moveable robot to recognized various dynamic objects. Dynamic objects (objects that move) are often disposed at lower heights than heights at which static objects are most visible, and thus it is recognized herein that objects at a range of heights should be detected and processed to prevent standstills or other operational problems.
[0013] Three-dimensional (3D) laser scanners and depth cameras can give information about objects in the environment within a wider field of view as compared to 2D scanners, for instance a height range of several meters and detection in three dimensions. Depth cameras are typically cheaper than 3D laser scanners, however, it is recognized herein that depth cameras are often difficult to use directly as an input for a navigation system. For example, the computational cost can increase as the points in space of a given point cloud increase. Furthermore, color information can create a data bus of information that might not be processed by a single GPU/CPU device. It is further recognized herein that current mapping and navigation software typically do not use point clouds as input to generate maps and navigate an environment. Thus, in accordance with various embodiments described herein, an obstacle detection system that processes 2D laser data is configured to also process obstacle information from a depth camera. In various examples, current mapping and navigation software can process the obstacle information that is generated from a depth camera without making direct changes to the software itself. [0014] Referring to FIG. 1, an example industrial or dynamic physical environment 100 can include a computerized autonomous system 102 that can include one or more mobile robot devices or automated guided vehicles (AGVs), for instance an AGV or robot device 104, configured to perform one or more industrial tasks, such as bin picking, grasping, transport, or the like. The system 102 can include one or more computing processors configured to process information and control operations of the system 102, in particular the AGV 104. The AGV 104 can include one or more processors, for instance a processor 108, configured to process information and/or control various operations associated with the AGV 104. An autonomous system for operating an AGV within a physical environment can further include a memory for storing modules. The processors can further be configured to execute the modules so as to process information and generate models based on the information. It will be understood that the illustrated environment 100 and the system 102 are simplified for purposes of example. The environment 100 and the system 102 may vary as desired, and all such systems and environments are contemplated as being within the scope of this disclosure. For example, the AGV or robot 104 is presented as an example of a device using an obstacle avoidance system, though it will be understood that the robot 104 can define additional or alternative devices, such as cars, drones, or the like, and all devices that perform 2D navigation or obstacle avoidance are contemplated as being within the scope of this disclosure.
[0015] Still referring to FIG. 1, the AGV 104 can further include a robotic arm or manipulator 110 and a base 112 configured to support the robotic manipulator 110. The base 112 can include wheels 114 or can otherwise be configured to move within the physical environment 100. The autonomous machine 104 can further include an end effector 116 attached to the robotic manipulator 110. The end effector 116 can include one or more tools configured to grasp and/or move objects. Example end effectors 116 include finger grippers or vacuum-based grippers. The robotic manipulator 110 can be configured to move so as to change the position of the end effector 116, for example, so as to place or move objects within the physical environment 100. The AGV 104 can further include one or more cameras or sensors, for instance a depth camera 118, configured to detect or record objects within the physical environment 100. The camera 118 can be mounted to the body or base of the AGV 104, for instance at a front end 117 of the AGV 104, or otherwise configured to generate images of the physical environment 100 along a trajectory of the AGV 104. [0016] Still referring to FIG. 1, the camera 118 can be configured as an RGB-D camera defining a color and depth channel configured to capture images of the environment 100 as the AGV 104 moves within the environment 100. For example, the AGV 104, in particular the base 112 of the AGV 104, can define a top 109 end and a bottom end 111 opposite the top end 109 along a transverse direction 120. The AGV 104 can further define a first side 113 and a second side 115 opposite the first side 113 along a second or lateral direction 122 that is substantially perpendicular to the transverse direction 120. The AGV 107 can further define a front end 117 and a rear end 119 opposite the front end 117 along a third or longitudinal direction 124 that is substantially perpendicular to both the transverse and lateral directions 120 and 122, respectively. Though the illustrated AGV 104 defines a rectangular shape, it will be understood that AGVs or robots can be alternatively shaped or sized, and all such AGVs or robots are contemplated as being within the scope of this disclosure. Thus, the camera 118 can define a depth camera configured to capture depth images of the workspace 100 from the front end 117. Alternatively, or additionally, the camera 118 can be mounted or attached to another end the base 112 or portion of the robot 104 (for instance the arm 110), or to another device entirely. Thus, in some cases, the data captured by the camera 118 can processed by the system 102 or transmitted to a navigation system and processed as laser data.
[0017] Referring now to FIG. 2, example operations 200 are shown that can be performed by a computing system, for instance the autonomous system 102 or AGV 104. At 202, the system 102, for instance the camera 118, can capture an image of the environment 100, so as to define a captured depth image. The captured image can define RGB (color) or RGB-D information (color and depth information) corresponding to the environment 100. Thus, referring also to FIG. 3, the captured image can define a 3D camera input that includes 3D-point cloud data 300 from a depth camera associated with objects or obstacles within the physical environment 100. At 204, based on the captured image (e.g., 3D point cloud 300), the system can compute information associated with rays in the image or frame captured at 202, for instance the number and orientation of rays. For example, at 204, the system can compute the number of rays that are needed for a specific laser input. Such a computation can be based on minimum and maximum angles (or rays) and angle resolution. For example, a given point cloud, for instance the point cloud 300, captured at 202 can be composed of n number of three-dimensional points. Thus, the 3D points can have coordinate values along the transverse direction 120, lateral direction 122, and longitudinal direction 124. For example, the camera 118 that captures the 3D point cloud 300 defines the origin 302 (0, 0, 0) of the 3D point cloud 300. At 204, a range (or number) of the n points can be projected as a one-dimensional (ID) vector. In various examples, the range of points can be defined between two rays, for instance a minimum ray 306 and a maximum ray 304 defined with respect to a zero angle or level 308 along the transverse direction 120 from the camera 118. For example, the maximum ray 304 (or angle 306) can define a ray that travels upward along the transverse direction 120 as it travels from the camera 118 along the lateral or longitudinal directions 122 and 124, and the minimum ray 306 (or angle) can define a ray that travels downward along the transverse direction 120 as it travels from the camera 118 along the lateral or longitudinal directions 122 and 124. Thus, the maximum and minimum rays (or angles) 304 and 306 can define a section 310 of the point cloud 300. In particular, the rays 304 and 306 can define boundaries of the section 310 of the point cloud 300, such that the number n of points are disposed within the boundaries defined by the maximum ray or angle 304 and the minimum ray or angle 306.
[0018] Furthermore, the number n of 3D points contained in the section 310 can vary depending on the angle resolution. In various examples, the higher the angle resolution, the greater the number n of points contained in the section 310, so more points can be processed and projected as the resolution increases. In some cases, minimum and maximum angles are user-defined, or can be based in user inputs to the system. Alternatively, or additionally, the 3D sensor 118 can define limitations, in particular a limit of angle ranges or a limit of resolution, that can determine a given section of the 3D point cloud or the number n of points in a given section.
[0019] Still referring to FIGs. 2 and 3, at 206, the system can transform the 3D image captured at 202, for instance 3D point cloud 300, so as to remove invalid rays or points from the point cloud 300. In an example, filtering is performed to remove points out of range from the main point cloud 300, for instance to remove points outside the section 310 defined by the maximum ray 304 and the minimum ray 306. Alternatively, or additionally, the section 310 can be defined by a range of distances. By way of example, the system might filter out points that are less than 1 meter (m) from the camera 118 or robot 104, and greater than 4m from the camera 118 or robot 104, such that the points that are processed are between Im and 4m from the camera 118. In some cases, clusterization can be performed to remove invalid points. At 208, as further described herein, the 3D point cloud 300 can be transformed into a specific target frame, for instance the frame of a laser input of the robot 104. This target frame may correspond to a frame of a 2D laser that is processed, for instance at 212, with the 3D point cloud 300 that is converted to laser data.
[0020] In various examples, each part or portion of the robot 104 is defined with a frame, or a set of coordinates that specifies the location in space of each respective part. One part or portion of the robot defines the center or origin of the frame. For example, the center can be defined by the center of the top end 109, though it will be understood that the center (origin) can vary as desired, for instance the center of a wheel 114, the center of the front end 117, or the like, and all such origins (frame centers) are contemplated as being within the scope of this disclosure. Once the center (origin) is defined, coordinate frames can be determined for other parts of the robot 104, for instance the wheels 114, the sensor 118, and the like. For example, when a laser is used to detect objects, the laser has a coordinate frame, and distances defined by the objects in the environment can be determined based on the laser coordinate frame. Those distances can be transformed based on the predefined origin, so that they can be understood. By way of example, assuming the robot 104 is avoiding collisions with a given object in the environment 100, the system can determine the separation between the origin of the robot 104, the sensor 118, an outermost edge of the robot 104, and the measured distance of the object itself in the laser coordinate frame to estimate when the object might collide with a portion of the robot 104, for instance the front end 117 of the robot 104. By way of further example, if a second laser is added or attached to the robot 104, the coordinate frame of the second laser can be transformed to align with the base (original) laser that might be used for navigation. If the coordinate frames are not aligned, one of the lasers might generate distances that do not make sense as compared to the distances that the other laser is generating. Thus, in various examples, a new sensor’s frame is transformed to match the main sensor or set of sensors, so that their respective distances are aligned with each other. In particular, for example, the frame defined by a depth camera or 3D sensor of the robot 104 can be transformed into the frame defined by the laser sensor of the robot 104. Thus, the 3D sensor or depth camara can be disposed anywhere on the robot 104 in any direction relative to the robot 104, because the frame transformation matches the distance points from the 3D sensor to the 2D laser data that might be generated from the rear end 119 and/or front end 117 of the robot 104.
[0021] In various examples, the AGV 104, in particular the sensor 118, can define a 3D sensor and a laser sensor. At 210, 2D laser data can be captured by the laser sensor. At 212, the laser data can be merged with the point cloud data captured by the 3D sensor at 202, such that the data captured by the laser sensor and the data captured by the 3D sensor are processed together, for example, to calculate positions of obstacles in a given environment (e.g., the environment 100).
[0022] With continuing reference to FIG. 2, at 208, the system can adapt a scan range, so as to generate laser data from the captured 3D input. The scan range can define a maximum distance between objects and the 3D sensor 118 at which the which the system can process the points corresponding to those objects, so as to match any other frame. In particular, for example, the system transform the frame from the 3D sensor to the base laser sensor that might be used for navigation and obstacle avoidance, so as to merge the 3D point cloud data with laser data, so that both sets of data can be processed as lasers, at 212. In an example obstacle avoidance scenario, obstacles that are sufficiently close to the 3D sensor 118, and thus sufficiently close to the robot 104, are relevant because the robot 104 is not at risk to collide with objects that are not sufficiently close. Thus, the maximum distance can define a predetermined distance, for instance along the lateral or longitudinal directions 122 and 124, from the robot 104 at which the robot might collide with a given object. At 208, data that is outside of the predetermined distance can be filtered out before further processing at 212, thereby reducing the amount of data that is processed.
[0023] Alternatively, or additionally, the maximum distance can be defined by a height range, for instance a height range 312 defined between a maximum height 314 and a minimum height 316 along the transverse direction 120. Points that are outside of the defined angles and/or height range can be removed, such that the range can be adapted, for instance to a smaller scan range. In some cases, parameters that determine the scan range, such as maximum or minimum angles, a maximum distance or range, angular resolution, and height range, are user-defined. For example, users can define the parameters based on the particular robot and environment in which the robot interacts, such that the robot detects obstacles. By way of example, if a given robot 104 needs to detect obstacles between a height range of 50cm to 150cm above the 2D laser sensor line, the user can set the respective parameters for the minimum height 316 and the maximum height 314 along the transverse direction 120. In such an example, the light of sight of the 3D sensor corresponds to the zero level 308. In various examples, multiple height ranges can be selected. By way of example, a first height range can be selected such that a first laser data set is generated from a range of 50 cm to 150cm, and a second height range can be selected such that a second laser data set is generated from 0 cm to 20 cm below the level defined by a 2D laser (e.g., -20 cm from zero level 308) along the transverse direction 120, though it will be understood that various distances or height ranges can vary as desired, and all such alternative or additional distances or ranges are contemplated as being within the scope of this disclosure. Thus, in accordance with various examples, the system can generate one or more sets of projected laser data, for instance multiple sets based on respective different height ranges, based on a single point cloud captured at 202.
[0024] By way of an example in which the robot 104 defines a non-omnidirectional robot, the system might only be interested in detecting obstacles in front of the robot. Thus, for example, the system can remove points from a point cloud, for instance a point cloud 400 (see FIG. 4), which represent points or objects that are adjacent to both sides 113 and 115 of the robot. In an example, an angle range can be selected such only points are selected that are in front of the robot 104.
[0025] Additionally, or alternatively, referring to FIG. 4, a distance range can be defined so that the system only detects obstacles that are relevant to a particular user, for instance sufficiently close to a trajectory of the robot or the robot itself. In an example, the distance range can define how far, for instance how many meters or the like, from the front end 117 along the lateral direction 122 or longitudinal direction 124, within which points are selected. Thus, referring to FIG. 4, the point cloud 400 can be captured (at 202) that includes a first 3D point 404. At 204, the point can be projected as a one-dimensional (ID) value, which together with other projected points can form a ID vector of distances, so as to define a second or projected point 406. At 206, a distance 402 from the sensor 118 along the lateral or longitudinal directions 122 and 124, respectively, can be determined. The distance 402 can be compared to the predetermined distance range (e.g., 5m). If the distance 402 is greater than the distance range, the point 406 can be filtered out, at 208, such that the point 406 is not processed at 212. If the distance 402 is less than or equal to the distance range, the point 406 can be processed, at 212, with other laser data that is captured at 210. It is recognized herein that, in some cases in which an objective of the system is to detect obstacles, a high angle resolution might not be required, as one or more points can confirm the existence of an obstacle sufficiently to avoid the obstacle without determining precise shape of the obstacle.
[0026] Referring again to FIG. 4, the image illustrates an example of projecting a point from 3D to ID. The image also illustrates, however, that the point corresponds to two dimensions because it is projected along the lateral and longitudinal directions 122 and 124 (e.g., xy plane) so as to be defined by an angle and a distance. For instance, in cylindrical coordinates, a point with radius r, angle Theta, and height z can be projected to a point defined by (r, theta, z=0). By way of further example, the laser data can define a unidimensional vector having points that correspond to a specific angle. Thus, the laser data can define a ID vector, because each value can be processed because each angle is known. By way of further example, if a given laser of the robot 104 is set up with a resolution of 45 degrees, and distances are measured at time t, the resulting vector can include [INF, 4, 8, 80, 2] assuming the laser has a range from -90 to 90. In particular, the laser can generate a value for a laser beam at -90, then -45, then 0, then 45, then 90, assuming the 0 as the vertical axis. It is recognized herein that various algorithms for navigation and obstacle avoidance need to know these parameters in order to use this information accordingly, so when the system converts from a point in 3D to ID, the system can use other dimensions, such as the angles, and parameters to make sense of those data points.
[0027] Thus, in accordance with the example operations 200, 3D-point cloud data from a depth camera can be projected into the data format of 2D laser scanners. Furthermore, various laser data such as, for example, the angle, distance range of detection, and the resolution can be adjusted according to the required laser. Additionally, or alternatively, the laser data can be adjusted based on specific obstacles or objects within the environment 100. For example, the baseline laser can be adjusted based on the height (e.g., height range 312) defined by obstacles within the environment 100 along the transverse direction 120.
[0028] Within continuing reference to FIG. 2, after 208, a projected point cloud captured at 202 can take the form of a 2D laser sensor. In particular, for example, the 3D camera input or image captured at 202 can define points along the lateral, longitudinal, and transverse directions (x, y, z). Referring also to FIGs. 3 and 4, the system can transform each data point (x, y, z) into two dimensions, for instance a first dimension representative of an angle and a second dimension representative of a distance, thereby transforming the 3D camera input into a linear array of distances. In particular, for example, the system can transform the first point 404 (aP) having coordinates (Xa, Ya, Za) into the second point 406 that is the distance 402 from the sensor 118. An array of such transformed data, in various examples, can be sent or fed to a laser input port of a laser-based navigation system for mobile robots, at 212. For example, in some cases, at 210, the system can also capture 2D laser data from a laser input, for instance the sensor 118 that can define a 3D camera and a laser. At 212, the system can process the linear array of distances transformed from the point cloud data, and the 2D layer data captured from the 2D sensor data, together for conflict avoidance using a local path planner or using a simultaneous localization and mapping (SLAM) algorithm. Thus, the 3D captured data can be processed the same way data from laser scanners can be processed.
[0029] Thus, in various examples, depth channel data from any type of 3D camera that captures 3D point clouds can be processed by during operations 200. For example, the output at 208 can define a set of points in the same data format as lidar sensors. Thus, at 212, the output can be processed as lidar sensor data, even though the 3D camera that captures the image at 202 might be less expensive than a lidar sensor.
[0030] Without being bound by theory, in current approaches, a robot might scan and map a warehouse with a single 2D laser scanner. In such an example, in some cases, only the legs of obstacles (e.g., shelves ) are visible on the map, leaving the impression that there is free space to move through the legs, which would result in a collision when the robot tries to navigate through them. Alternatively, in accordance with various embodiments, the system performs operations 200 and correctly detects the obstacles (e.g., shelves) and generates the obstacle on the map, thereby protecting the robot from future collisions by computing the correct navigation trajectories through the warehouse.
[0031] In various examples, the output at 208 can be used for reliable obstacle detection (at 212) within 2D lidar sensor-based navigation systems for mobile robots, for example, because the depth camera 118 can cover a larger height range as compared to a 2D laser scanner. Thus, obstacles above and below a give 2D-laser scanner plane can reliably be detected. By way of further example, and without limitation, such a height range can enable detection of forklift forks, human feet, shelves at warehouses and factories, and the like, which are currently difficult to detect due to their different heights. The system can also perform operations 200 so as to detect various small objects on the factory floor (e.g., screws, tools, etc.) that currently cause danger to the robots or AGVs when moving over them.
[0032] Thus, as described herein, point cloud data from 3D-cameras can be transformed into laser data types, so that such data can be processed (at 212) by laser-based navigation systems without any adaptations to existing software, thereby creating cost and processing overhead efficiencies.
[0033] Furthermore, the data of lidar-sensors and depth cameras can also be merged together (e.g., at 212) after feeding the camera data into the same input as the laser data, such that the same type of SLAM algorithm, localization, path planning, and/or obstacle avoidance algorithm can be used that are used in current 2D laser-based systems, while adding obstacles at different height levels to the map, thereby enabling various navigation systems to process the 3D point cloud data as 2D laser data. For example, additional features above or below the 2D-laser height level can be added to the map (e.g., point cloud 300) using the depth camera so as to support the precise localization of the robot. This can reduce the number of localization errors of a robot, increase the level of difficulty of environments the robot can operate, and increase the precision of localization. In some case, the integration of 3D-cameras needs little to no adaptation of existing 2D laser scanners SLAM navigation systems due to the data transformation described herein, which makes the disclosed solutions more cost effective in terms of development than other existing solutions.
[0034] Additionally, in various examples, the system allows the replacement of one or more laser scanners on the mobile robot without adapting the navigation software or creating a new interface. Apart from the height-range advantage, depth cameras can also be the more costefficient solution compared to laser-scanners for shorter ranges.
[0035] FIG. 5 illustrates an example of a computing environment within which embodiments of the present disclosure may be implemented. A computing environment 600 includes a computer system 610 that may include a communication mechanism such as a system bus 621 or other communication mechanism for communicating information within the computer system 610. The computer system 610 further includes one or more processors 620 coupled with the system bus 621 for processing the information. The autonomous system 102, in particular the robot 104, may include, or be coupled to, the one or more processors 620.
[0036] The processors 620 may include one or more central processing units (CPUs), graphical processing units (GPUs), or any other processor known in the art. More generally, a processor as described herein is a device for executing machine-readable instructions stored on a computer readable medium, for performing tasks and may comprise any one or combination of, hardware and firmware. A processor may also comprise memory storing machine-readable instructions executable for performing tasks. A processor acts upon information by manipulating, analyzing, modifying, converting or transmitting information for use by an executable procedure or an information device, and/or by routing the information to an output device. A processor may use or comprise the capabilities of a computer, controller or microprocessor, for example, and be conditioned using executable instructions to perform special purpose functions not performed by a general-purpose computer. A processor may include any type of suitable processing unit including, but not limited to, a central processing unit, a microprocessor, a Reduced Instruction Set Computer (RISC) microprocessor, a Complex Instruction Set Computer (CISC) microprocessor, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a System-on-a-Chip (SoC), a digital signal processor (DSP), and so forth. Further, the processor(s) 620 may have any suitable microarchitecture design that includes any number of constituent components such as, for example, registers, multiplexers, arithmetic logic units, cache controllers for controlling read/write operations to cache memory, branch predictors, or the like. The microarchitecture design of the processor may be capable of supporting any of a variety of instruction sets. A processor may be coupled (electrically and/or as comprising executable components) with any other processor enabling interaction and/or communication there-between. A user interface processor or generator is a known element comprising electronic circuitry or software or a combination of both for generating display images or portions thereof. A user interface comprises one or more display images enabling user interaction with a processor or other device.
[0037] The system bus 621 may include at least one of a system bus, a memory bus, an address bus, or a message bus, and may permit exchange of information (e.g., data (including computer-executable code), signaling, etc.) between various components of the computer system 610. The system bus 621 may include, without limitation, a memory bus or a memory controller, a peripheral bus, an accelerated graphics port, and so forth. The system bus 621 may be associated with any suitable bus architecture including, without limitation, an Industry Standard Architecture (ISA), a Micro Channel Architecture (MCA), an Enhanced ISA (EISA), a Video Electronics Standards Association (VESA) architecture, an Accelerated Graphics Port (AGP) architecture, a Peripheral Component Interconnects (PCI) architecture, a PCI -Express architecture, a Personal Computer Memory Card International Association (PCMCIA) architecture, a Universal Serial Bus (USB) architecture, and so forth.
[0038] Continuing with reference to FIG. 5, the computer system 610 may also include a system memory 630 coupled to the system bus 621 for storing information and instructions to be executed by processors 620. The system memory 630 may include computer readable storage media in the form of volatile and/or nonvolatile memory, such as read only memory (ROM) 631 and/or random-access memory (RAM) 632. The RAM 632 may include other dynamic storage device(s) (e.g., dynamic RAM, static RAM, and synchronous DRAM). The ROM 631 may include other static storage device(s) (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). In addition, the system memory 630 may be used for storing temporary variables or other intermediate information during the execution of instructions by the processors 620. A basic input/output system 633 (BIOS) containing the basic routines that help to transfer information between elements within computer system 610, such as during start-up, may be stored in the ROM 631. RAM 632 may contain data and/or program modules that are immediately accessible to and/or presently being operated on by the processors 620. System memory 630 may additionally include, for example, operating system 634, application programs 635, and other program modules 636. Application programs 635 may also include a user portal for development of the application program, allowing input parameters to be entered and modified as necessary.
[0039] The operating system 634 may be loaded into the memory 630 and may provide an interface between other application software executing on the computer system 610 and hardware resources of the computer system 610. More specifically, the operating system 634 may include a set of computer-executable instructions for managing hardware resources of the computer system 610 and for providing common services to other application programs (e.g., managing memory allocation among various application programs). In certain example embodiments, the operating system 634 may control execution of one or more of the program modules depicted as being stored in the data storage 640. The operating system 634 may include any operating system now known or which may be developed in the future including, but not limited to, any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.
[0040] The computer system 610 may also include a disk/media controller 643 coupled to the system bus 621 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 641 and/or a removable media drive 642 (e.g., floppy disk drive, compact disc drive, tape drive, flash drive, and/or solid-state drive). Storage devices 640 may be added to the computer system 610 using an appropriate device interface (e.g., a small computer system interface (SCSI), integrated device electronics (IDE), Universal Serial Bus (USB), or FireWire). Storage devices 641 , 642 may be external to the computer system 610.
[0041] The computer system 610 may also include a field device interface 665 coupled to the system bus 621 to control a field device 666, such as a device used in a production line. The computer system 610 may include a user input interface or GUI 661, which may comprise one or more input devices, such as a keyboard, touchscreen, tablet and/or a pointing device, for interacting with a computer user and providing information to the processors 620.
[0042] The computer system 610 may perform a portion or all of the processing steps of embodiments of the invention in response to the processors 620 executing one or more sequences of one or more instructions contained in a memory, such as the system memory 630. Such instructions may be read into the system memory 630 from another computer readable medium of storage 640, such as the magnetic hard disk 641 or the removable media drive 642. The magnetic hard disk 641 (or solid-state drive) and/or removable media drive 642 may contain one or more data stores and data files used by embodiments of the present disclosure. The data store 640 may include, but are not limited to, databases (e.g., relational, object-oriented, etc.), file systems, flat files, distributed data stores in which data is stored on more than one node of a computer network, peer-to-peer network data stores, or the like. The data stores may store various types of data such as, for example, skill data, sensor data, or any other data generated in accordance with the embodiments of the disclosure. Data store contents and data files may be encrypted to improve security. The processors 620 may also be employed in a multi-processing arrangement to execute the one or more sequences of instructions contained in system memory 630. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. Thus, embodiments are not limited to any specific combination of hardware circuitry and software.
[0043] As stated above, the computer system 610 may include at least one computer readable medium or memory for holding instructions programmed according to embodiments of the invention and for containing data structures, tables, records, or other data described herein. The term “computer readable medium” as used herein refers to any medium that participates in providing instructions to the processors 620 for execution. A computer readable medium may take many forms including, but not limited to, non-transitory, non-volatile media, volatile media, and transmission media. Non-limiting examples of non-volatile media include optical disks, solid state drives, magnetic disks, and magneto-optical disks, such as magnetic hard disk 641 or removable media drive 642. Non-limiting examples of volatile media include dynamic memory, such as system memory 630. Non-limiting examples of transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up the system bus 621. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. [0044] Computer readable medium instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0045] Aspects of the present disclosure are described herein with reference to flowchart illustrations and/or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and/or block diagrams, and combinations of blocks in the flowchart illustrations and/or block diagrams, may be implemented by computer readable medium instructions.
[0046] The computing environment 600 may further include the computer system 610 operating in a networked environment using logical connections to one or more remote computers, such as remote computing device 680. The network interface 670 may enable communication, for example, with other remote devices 680 or systems and/or the storage devices 641, 642 via the network 671. Remote computing device 680 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above relative to computer system 610. When used in a networking environment, computer system 610 may include modem 672 for establishing communications over a network 671, such as the Internet. Modem 672 may be connected to system bus 621 via user network interface 670, or via another appropriate mechanism.
[0047] Network 671 may be any network or system generally known in the art, including the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a direct connection or series of connections, a cellular telephone network, or any other network or medium capable of facilitating communication between computer system 610 and other computers (e.g., remote computing device 680). The network 671 may be wired, wireless or a combination thereof. Wired connections may be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection generally known in the art. Wireless connections may be implemented using Wi-Fi, WiMAX, and Bluetooth, infrared, cellular networks, satellite or any other wireless connection methodology generally known in the art. Additionally, several networks may work alone or in communication with each other to facilitate communication in the network 671.
[0048] It should be appreciated that the program modules, applications, computer-executable instructions, code, or the like depicted in FIG. 5 as being stored in the system memory 630 are merely illustrative and not exhaustive and that processing described as being supported by any particular module may alternatively be distributed across multiple modules or performed by a different module. In addition, various program module(s), script(s), plug-in(s), Application Programming Interface(s) (API(s)), or any other suitable computer-executable code hosted locally on the computer system 610, the remote device 680, and/or hosted on other computing device(s) accessible via one or more of the network(s) 671, may be provided to support functionality provided by the program modules, applications, or computer-executable code depicted in FIG. 5 and/or additional or alternate functionality. Further, functionality may be modularized differently such that processing described as being supported collectively by the collection of program modules depicted in FIG. 5 may be performed by a fewer or greater number of modules, or functionality described as being supported by any particular module may be supported, at least in part, by another module. In addition, program modules that support the functionality described herein may form part of one or more applications executable across any number of systems or devices in accordance with any suitable computing model such as, for example, a client-server model, a peer-to-peer model, and so forth. [0049] It should further be appreciated that the computer system 610 may include alternate and/or additional hardware, software, or firmware components beyond those described or depicted without departing from the scope of the disclosure. More particularly, it should be appreciated that software, firmware, or hardware components depicted as forming part of the computer system 610 are merely illustrative and that some components may not be present or additional components may be provided in various embodiments. While various illustrative program modules have been depicted and described as software modules stored in system memory 630, it should be appreciated that functionality described as being supported by the program modules may be enabled by any combination of hardware, software, and/or firmware. It should further be appreciated that each of the above-mentioned modules may, in various embodiments, represent a logical partitioning of supported functionality. This logical partitioning is depicted for ease of explanation of the functionality and may not be representative of the structure of software, hardware, and/or firmware for implementing the functionality. Accordingly, it should be appreciated that functionality described as being provided by a particular module may, in various embodiments, be provided at least in part by one or more other modules. Further, one or more depicted modules may not be present in certain embodiments, while in other embodiments, additional modules not depicted may be present and may support at least a portion of the described functionality and/or additional functionality. Moreover, while certain modules may be depicted and described as sub-modules of another module, in certain embodiments, such modules may be provided as independent modules or as sub-modules of other modules.
[0050] Although specific embodiments of the disclosure have been described, one of ordinary skill in the art will recognize that numerous other modifications and alternative embodiments are within the scope of the disclosure. For example, any of the functionality and/or processing capabilities described with respect to a particular device or component may be performed by any other device or component. Further, while various illustrative implementations and architectures have been described in accordance with embodiments of the disclosure, one of ordinary skill in the art will appreciate that numerous other modifications to the illustrative implementations and architectures described herein are also within the scope of this disclosure. In addition, it should be appreciated that any operation, element, component, data, or the like described herein as being based on another operation, element, component, data, or the like can be additionally based on one or more other operations, elements, components, data, or the like. Accordingly, the phrase “based on,” or variants thereof, should be interpreted as “based at least in part on.”
[0051] Although embodiments have been described in language specific to structural features and/or methodological acts, it is to be understood that the disclosure is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as illustrative forms of implementing the embodiments. Conditional language, such as, among others, “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments could include, while other embodiments do not include, certain features, elements, and/or steps. Thus, such conditional language is not generally intended to imply that features, elements, and/or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, and/or steps are included or are to be performed in any particular embodiment.
[0052] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and/or flowchart illustration, and combinations of blocks in the block diagrams and/or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

Claims

CLAIMS What is claimed is:
1. An autonomous system configured to operate in a dynamic industrial environment so as to define a runtime, the autonomous system comprising: a camera configured to capture a depth image of the dynamic industrial environment, so as to define a captured image that defines a plurality of rays; one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the autonomous system to, during the runtime: compute ray information associated with the plurality of rays; based on the ray information, transform the captured image so as to remove a set of rays from the plurality of rays, thereby defining laser data from the plurality of rays; and process the laser data so as to detect one or more obstacles within the environment.
2. The autonomous system as recited in claim 1 , wherein the camera defines a three- dimensional (3D) sensor.
3. The autonomous system as recited in claim 1 , wherein the autonomous system includes a mobile robot, the mobile robot defining: a bottom end configured to face a floor of the dynamic industrial environment, and a top end opposite the bottom end along a first direction; a front end and rear end opposite the front end along a second direction that is substantially perpendicular to the first direction; and a first side and second side opposite the first side along a third direction that is substantially perpendicular to both the first and second directions, wherein the camera is disposed at the front end, rear end, first side, second side, or top end, and the camera is configured to capture 3D point cloud data from the front end along the first, second, and third directions.
4. The autonomous system as recited in claim 3, wherein the captured image defines a 3D point cloud comprising a plurality of 3D points, the memory further storing instructions that, when executed by the one or more processors, further cause the autonomous system to, during the runtime: project the plurality of 3D points into respective one dimensional points from the camera along the second or third directions.
5. The autonomous system as recited in claim 4, the memory further storing instructions that, when executed by the one or more processors, further cause the autonomous system to, during the runtime: determine a predetermined distance measured from the camera along the second or third directions; identify one or more points of the plurality of one dimensional points that define a distance from the camera that is greater than the predetermined distance; and filter the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data.
6. The autonomous system as recited in claim 4, the memory further storing instructions that, when executed by the one or more processors, further cause the autonomous system to, during the runtime: determine a predetermined height measured from the camera along the first direction; identify one or more points of the plurality of one dimensional points that define a distance above or below the camera along the first direction that is greater than the predetermined height; and filter the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data.
7. The autonomous system as recited in claim 2, wherein the autonomous system further comprises a laser sensor.
8. The autonomous system as recited in claim 7, the memory further storing instructions that, when executed by the one or more processors, further cause the autonomous system to, during the runtime: capture laser sensor data from the laser sensor; and merge the laser sensor data from the laser sensor with the laser data from the 3D sensor.
9. A method performed by an autonomous system configured to operate in a dynamic industrial environment so as to define a runtime, the method comprising: capturing, by a camera, a depth image of the dynamic industrial environment, so as to define a captured image that defines a plurality of rays; computing ray information associated with the plurality of rays; based on the ray information, transforming the captured image so as to remove a set of rays from the plurality of rays, thereby defining laser data from the plurality of rays; and processing the laser data so as to detect one or more obstacles within the environment.
10. The method as recited in claim 9, wherein the camera defines a three-dimensional (3D) sensor, and the autonomous system includes a mobile robot that defines: a bottom end configured to face a floor of the dynamic industrial environment, and a top end opposite the bottom end along a first direction; a front end and rear end opposite the front end along a second direction that is substantially perpendicular to the first direction; and a first side and second side opposite the first side along a third direction that is substantially perpendicular to both the first and second directions, wherein the method further comprises: capturing the depth image along first, second, and third directions by the camera that is disposed at the front end, rear end, first side, second side, or top end.
11. The method as recited in claim 12, wherein the captured image defines a 3D point cloud comprising a plurality of 3D points, the method further comprising: during the runtime, projecting the plurality of 3D points into respective one dimensional points from the camera along the second or third directions.
12. The method as recited in claim 11, the memory further comprising: during the runtime, determining a predetermined distance measured from the camera along the second or third directions; identifying one or more points of the plurality of one dimensional points that define a distance from the camera that is greater than the predetermined distance; and filtering the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data.
13. The method as recited in claim 11, the method further comprising: determining a predetermined height measured from the camera along the first direction; identifying one or more points of the plurality of one dimensional points that define a distance above or below the camera along the first direction that is greater than the predetermined height; and filtering the one or more points of the plurality of one dimensional points such that the one or more points are not included in the laser data.
14. The method as recited in claim 10, wherein the autonomous system further comprises a laser sensor disposed at the front end or the rear end.
15. The method as recited in claim 14, the method further comprising capturing laser sensor data from the laser sensor; and merging the laser sensor data from the laser sensor with the laser data from the 3D sensor.
EP23729593.6A 2023-05-11 2023-05-11 3d sensor-based obstacle detection for autonomous vehicles and mobile robots Pending EP4689711A1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2023/021798 WO2024232887A1 (en) 2023-05-11 2023-05-11 3d sensor-based obstacle detection for autonomous vehicles and mobile robots

Publications (1)

Publication Number Publication Date
EP4689711A1 true EP4689711A1 (en) 2026-02-11

Family

ID=86732329

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23729593.6A Pending EP4689711A1 (en) 2023-05-11 2023-05-11 3d sensor-based obstacle detection for autonomous vehicles and mobile robots

Country Status (3)

Country Link
EP (1) EP4689711A1 (en)
CN (1) CN121263712A (en)
WO (1) WO2024232887A1 (en)

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20100066587A1 (en) * 2006-07-14 2010-03-18 Brian Masao Yamauchi Method and System for Controlling a Remote Vehicle
US10572763B2 (en) * 2017-09-07 2020-02-25 Symbol Technologies, Llc Method and apparatus for support surface edge detection
WO2022087014A1 (en) * 2020-10-20 2022-04-28 Brain Corporation Systems and methods for producing occupancy maps for robotic devices
US12442896B2 (en) * 2022-02-24 2025-10-14 Realsense Ltd. Hierarchical perception monitor for vehicle safety

Also Published As

Publication number Publication date
WO2024232887A1 (en) 2024-11-14
CN121263712A (en) 2026-01-02

Similar Documents

Publication Publication Date Title
EP3347171B1 (en) Using sensor-based observations of agents in an environment to estimate the pose of an object in the environment and to estimate an uncertainty measure for the pose
US20200264625A1 (en) Systems and methods for calibration of a pose of a sensor relative to a materials handling vehicle
KR20230066323A (en) Autonomous Robot Exploration in Storage Sites
US11112780B2 (en) Collaborative determination of a load footprint of a robotic vehicle
CN112964263B (en) Automatic mapping method, device, mobile robot and readable storage medium
WO2022000197A1 (en) Flight operation method, unmanned aerial vehicle, and storage medium
Indri et al. Sensor data fusion for smart AMRs in human-shared industrial workspaces
CN118435142A (en) Method for generating a surrounding environment map for a mobile logistics robot and mobile logistics robot
US10731970B2 (en) Method, system and apparatus for support structure detection
US20240066723A1 (en) Automatic bin detection for robotic applications
US20250387902A1 (en) Bin wall collision detection for robotic bin picking
KR20240135603A (en) Thin object detection and avoidance for aerial robots
CN121325891A (en) Multi-sensor obstacle avoidance control method and system for wind turbine blade inspection robot
JP2023054958A (en) Moving body, server device, moving body control system, moving body control method, and program
EP4689711A1 (en) 3d sensor-based obstacle detection for autonomous vehicles and mobile robots
US20210156710A1 (en) Map processing method, device, and computer-readable storage medium
JP7563455B2 (en) MOBILE BODY CONTROL SYSTEM, CONTROL DEVICE, AND MOBILE BODY CONTROL METHOD
EP4406709A1 (en) Adaptive region of interest (roi) for vision guided robotic bin picking
CN115115821A (en) Automatic recharging method, device and self-moving device
Wu et al. Enhancing automated guided vehicle navigation with multi-sensor fusion and algorithmic optimization
JP7501643B2 (en) MOBILE BODY CONTROL SYSTEM, CONTROL DEVICE, AND MOBILE BODY CONTROL METHOD
Xu et al. 3D perception for autonomous robot exploration
Ribeiro et al. Indoor Benchmark of 3-D LiDAR SLAM at Iilab—Industry and Innovation Laboratory
EP4421451A1 (en) An effective method to estimate pose, velocity and attitude with uncertainty
Biernacki et al. The method of reflection-based marker detection and identification to ensure accurate AGV docking

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251106

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR