WO2024099786A1 - Image processing method and method for predicting collisions - Google Patents
Image processing method and method for predicting collisions Download PDFInfo
- Publication number
- WO2024099786A1 WO2024099786A1 PCT/EP2023/079921 EP2023079921W WO2024099786A1 WO 2024099786 A1 WO2024099786 A1 WO 2024099786A1 EP 2023079921 W EP2023079921 W EP 2023079921W WO 2024099786 A1 WO2024099786 A1 WO 2024099786A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- disparity
- image processing
- images
- unrotated
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N13/00—Stereoscopic video systems; Multi-view video systems; Details thereof
- H04N13/10—Processing, recording or transmission of stereoscopic or multi-view image signals
- H04N13/106—Processing image signals
- H04N13/128—Adjusting depth or disparity
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/50—Depth or shape recovery
- G06T7/55—Depth or shape recovery from multiple images
- G06T7/593—Depth or shape recovery from multiple images from stereo images
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/26—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/74—Image or video pattern matching; Proximity measures in feature spaces
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
- G06V20/58—Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/60—Type of objects
- G06V20/64—Three-dimensional [3D] objects
- G06V20/647—Three-dimensional [3D] objects by matching two-dimensional images to three-dimensional objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10016—Video; Image sequence
- G06T2207/10021—Stereoscopic video; Stereoscopic image sequence
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10028—Range image; Depth image; 3D point clouds
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10048—Infrared image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20076—Probabilistic image processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30248—Vehicle exterior or interior
- G06T2207/30252—Vehicle exterior; Vicinity of vehicle
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N13/00—Stereoscopic video systems; Multi-view video systems; Details thereof
- H04N2013/0074—Stereoscopic image analysis
- H04N2013/0081—Depth or disparity estimation from stereoscopic image signals
Definitions
- Various embodiments relate to methods and devices for processing stereo images, methods and devices for predicting collisions and methods for training a machine learning model.
- Stereovision techniques typically use two cameras to look at the same object.
- the two cameras may be separated by a baseline distance.
- the baseline distance is assumed to be known accurately.
- the two cameras may simultaneously capture two images, also referred to as stereo images.
- the stereo images may be analyzed to identify the differences between the two images.
- the differences between the two images may be referred to as disparity.
- the disparity between the two images may be used to determine depth of a point in the images.
- the depth information may be used to project the point in a three-dimensional model.
- Such three-dimensional (3D) models may be useful for facilitating various driver assistance or autonomous driving functions.
- the three-dimensional model may provide information on positions of various objects near to a vehicle, thereby aiding navigation or obstacle avoidance.
- the stereo cameras may be installed onto a vehicle to capture images of the surroundings of the vehicle.
- the stereo cameras need to be precisely calibrated in order for the disparity acquired from the images to be accurate.
- regular movements of the vehicle for example, over a pothole on the road, or over a road hump, may result in vibrations that displace the stereo cameras, thereby decalibrating the stereo cameras.
- the stereo images cannot be relied on to generate accurate depth information that is needed for the three-dimensional model. Consequently, the three- dimensional model would not be suitable for use as a means to detect objects present in the vehicle’s surroundings, or to prevent collisions with the objects.
- the image processing method may include inputting an image set to a machine learning model.
- the image set may include a first image and a second image.
- the image processing method may further include generating a rotated disparity map of the image set, using the trained machine learning model.
- the image processing method may further include, for each pixel in the first image - determining, based on the rotated disparity map, a rotation angle between the pixel in the first image and its corresponding position in the second image, relative to a common image center, and determining a respective unrotated disparity based on the respective rotation angle.
- the image processing method may further include generating an unrotated disparity map based on the respective unrotated disparity of each pixel in the first image.
- an image processing device that includes a processor.
- the processor may be configured to perform the abovementioned image processing method.
- a computer-implemented method for predicting collisions may include inputting a sequence of image sets to a motion detection model. Each image set of the sequence may include a first image and a second image. The method may further include determining by the motion detection model, for each image set, a respective depth map based on the first image and the second image of the stereo image set, resulting in a sequence of depth maps. The method may further include determining by the motion detection model, an optical flow based on at least one of, the first images from the sequence of image sets and the second images of the sequence of stereo image sets. The method may further include determining by the motion detection model, motion of an object in the image set based on the optical flow. The method may further include determining time- to-collision with the object based on the sequence of depth maps and the determined motion of the object.
- a device for predicting collisions may include a processor configured to perform the abovementioned method for predicting collisions.
- FIG. 1 shows a block diagram of a computer-implemented method of training a machine learning model to determine disparity of images, according to various embodiments.
- FIG. 2 shows a block diagram of a computer-implemented method of generating a 3D model using stereo images, according to various embodiments.
- FIG. 3 shows an example of an image set according to various embodiments.
- FIG. 4 shows an example of how a position of a feature point may differ in the left image and the right image.
- FIG. 5A shows an example of an image set according to various embodiments.
- FIG. 5B shows disparity images generated by a prior art machine learning model and the trained machine learning model according to various embodiments.
- FIG. 6 shows an image processing method according to various embodiments.
- FIG. 7A shows a simplified block diagram of an image processing device according to various embodiments.
- FIG. 7B illustrates an operation of the image processing device according to various embodiments.
- FIG. 8A shows a simplified functional block diagram of a device for predicting collisions according to various embodiments.
- FIG. 8B shows a simplified hardware block diagram of the device according to various embodiments.
- FIG. 9 shows a schematic diagram of the device performing a method for predicting collisions, according to various embodiments.
- FIG. 10 shows a schematic diagram of a neural network according to various embodiments.
- FIG. 11 shows a schematic diagram of the neural network according to various embodiments, receiving different inputs from those in FIG. 10.
- FIG. 12 shows a flow diagram of a computer-implemented method for predicting collisions, according to various embodiments.
- Coupled may be understood as electrically coupled or as mechanically coupled, for example attached or fixed, or just in contact without any fixation, and it will be understood that both direct coupling or indirect coupling (in other words: coupling without direct contact) may be provided.
- the device as described in this description may include a memory which is for example used in the processing carried out in the device.
- a memory used in the embodiments may be a volatile memory, for example a DRAM (Dynamic Random Access Memory) or a non-volatile memory, for example a PROM (Programmable Read Only Memory), an EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or a flash memory, e.g., a floating gate memory, a charge trapping memory, an MRAM (Magnetoresistive Random Access Memory) or a PCRAM (Phase Change Random Access Memory).
- DRAM Dynamic Random Access Memory
- PROM Programmable Read Only Memory
- EPROM Erasable PROM
- EEPROM Electrical Erasable PROM
- flash memory e.g., a floating gate memory, a charge trapping memory, an MRAM (Magnetoresistive Random Access Memory) or a PCRAM (Phase Change Random Access Memory).
- FIG. 1 shows a block diagram of a computer-implemented method 100 of training a machine learning model 102 to determine disparity of images, according to various embodiments.
- the method of training the machine learning model 102 may include providing training image data 112 to the machine learning model 102.
- the training image data 112 may include left stereo images and right stereo images from uncalibrated stereo cameras. In other words, the training image data 112 may include uncalibrated stereo images.
- Ground truth disparity data 114 that corresponds to each set pair of left and right stereo images of the training image data 112 may also be provided to the machine learning model 102 as a training signal.
- the machine learning model 102 may be configured to extract features from the training image data 112, and may learn a pattern between the extracted features and the ground truth disparity data 114.
- the machine learning model 102 learns to compute a disparity map for each pair of left and right stereo image pair.
- the resulting machine learning model is referred herein as the trained machine learning model 104.
- the trained machine learning model 104 may include a neural network 1000 that will be described further with respect to FIG. 10.
- a feature point of the left stereo image and the corresponding feature point of the right stereo image are aligned on the same x-axis, also referred to as horizontal axis. Due to vibration or mechanical set up, one of the stereo cameras may be misaligned with the other stereo camera with at least one of the following factors: roll, i.e. rotation angle (a), pitch, yaw, and translation (x t , y t ) where x t refers to movement of an image pixel along the horizontal axis and y t refer to movement of the image pixel along the vertical axis.
- the trained machine learning model 104 may be configured to determine rotated disparities for the above-mentioned misalignments, as the training image data 112 includes uncalibrated stereo images.
- the method 100 may further include rotating stereo images to artificially uncalibrate the training image data 112.
- the training image data 112 may include images obtained from a plurality of cameras that may not be stereo cameras. These images may capture a similar scene from different angles, and hence, may similarly be used to determine depth of objects in the scene, like stereo images.
- the cameras or stereo cameras that capture the training image data 112 may be coupled to a vehicle.
- the training image data 112 may capture scenes around the vehicle.
- the method 100 may further include classifying the disparity range into near range and far range while training the machine learning model 102.
- the ground truth disparity data 114 may be split into a near range ground truth disparity data and a far range ground truth disparity data.
- the machine learning model 102 may be trained to compute near range ground truth disparity using the near range ground truth disparity data as the training signal.
- the machine learning model 102 may be further trained to compute far range ground truth disparity using the far range ground truth disparity data as the training signal, in a separate process from the training using near range ground truth disparity.
- the accuracy of the trained machine learning model 104 may be improved.
- the time taken to train the machine learning model 102 may also be reduced.
- FIG. 2 shows a block diagram of a computer-implemented method 200 of generating a 3D model using stereo images, according to various embodiments.
- the method 200 may include providing image data to the trained machine learning model 104.
- the image data may include an image set 212.
- the image set 212 may include a left stereo image and a right stereo image.
- the trained machine learning model 104 may extract feature information from the left and right stereo images, and may generate a rotated disparity map 214 based on the extracted feature information.
- a rotation correction module 202 may receive the rotated disparity map 214 and correct the rotated disparity map 214 for the relative rotation between the images in the image set 212.
- the rotation correction module 202 may generate a unrotated disparity map 216 based on the rotated disparity map 214.
- the rotation correction module 202 may provide the unrotated disparity map 216 to a 3D reconstruction module 206.
- the 3D reconstruction module 206 may generate a 3D point cloud 208 based on the unrotated disparity map 216, camera parameters 204 and the image set 212.
- the rotated disparity map 214 may include the actual disparities of every pixel between the uncalibrated left and right stereo images.
- the unrotated disparity map 214 may include corrected disparities of every pixel between the uncalibrated left and right stereo images, as if the misaligned stereo image was already corrected back to its calibrated position.
- the method 200 may be able to generate 3D point clouds with depth measurements that are more accurate, and at a lower cost, as compared to LiDAR sensors.
- the method 200 may also be able to determine the depth in scenes captured on the images regardless of dynamic decalibrations of the cameras, including rotation, horizontal shift and vertical shift.
- the method 200 may also be able to generate the 3D point clouds even if the image set 212 includes distorted images and modified intrinsics, and if the camera parameters are inaccurate, for example, incorrect focal lens information.
- the image set 212 may include images obtained from a plurality of cameras that may not be stereo cameras. These images may capture a similar scene from different angles, and hence, may similarly be used to determine depth of objects in the scene, like stereo images.
- the cameras or stereo cameras that capture the image set 212 may be coupled to a vehicle.
- the image set 212 may capture scenes around the vehicle.
- FIG. 3 shows an example of an image set 212 according to various embodiments.
- the image set 212 may include a left image 302 and a right image 304.
- the image set 212 may be uncalibrated.
- the right image 304 is rotated relative to the left image 302. This can happen, when at least one of the stereo cameras is not installed properly, or is displaced from its original position. For example, when the vehicle drives over a road hump or makes a sharp movement, the stereo camera may experience vibration that results in a slight displacement.
- the rotation correction module 202 may determine the unrotated disparity map 216 based on the rotated disparity map 214, such that the unrotated disparity map 216 provides information on the displacement between each pixel in the left image 302 and its corresponding right image 304 as if the left image 302 and the right image 304 are aligned on a horizontal axis.
- the process of determining the unrotated disparity map 216 is described with respect to FIG. 4.
- FIG. 4 shows an example of how a position of a feature point may differ in the left image 302 and the right image 304.
- an image centre 402 may be considered to be (0,0).
- a feature point 404 of the left image 302 is denoted as (x 2 ,y2)-
- the corresponding position 406 of the feature point 404 in the right image 304 as rotated, is denoted as (x' 2 , y'2).
- a line connecting the corresponding position 406 and the image center 402 may define a first hypotenuse 420, indicated herein as “hi”, of a right angle triangle 410.
- the corresponding position 406 of the feature point 404 in the right image 304 may be determined based on the rotated disparity map 214.
- the rotated disparity map may include rotated disparity values dx 2 , dy 2 ) for the feature point 404.
- the corresponding position 406, and the first hypothenuse 420 may be determined based on the feature point 404 and the rotated disparity values according to the following equations (1) to (3). x' 2 — x 2 + dx 2 Equation
- the unrotated disparity values of the feature point 404 may be expressed as (dx 2 , dy 2 ).
- the unrotated corresponding position 408 of the feature point 404 in the right image 304 when the right image 304 is adjusted to have zero rotation angle relative to the left image 302, is denoted as (x 2r ,y 2r ).
- the corresponding position 406, the unrotated corresponding position 408 and the image centre 402 may define vertices of an isosceles triangle 422.
- the distance between the corresponding position 406 and the image center 402 may be at least substantially equal to the distance between the unrotated corresponding position 408 and the image centre 402, in other words, equal to the first hypotenuse 420.
- the vertex angle 426 of the isosceles triangle 422 is denoted as a.
- the vertex angle 426 may also be referred herein as rotation angle 426.
- a line connecting the corresponding position 406 and the unrotated corresponding position 408, may define a second hypotenuse 424, denoted herein as “Z12”, of another right angle triangle 412.
- the second hypotenuse 424 may be the base of the isosceles triangle 422.
- the second hypotenuse 424 may be determined according to equation (4).
- the base 428 of the other right angle triangle 412 is denoted as d.
- the length of the base 428 may be determined according to the equation (5). uation
- the unrotated disparity value d 2 may be determined based on the base 428 and the rotated disparity value dx 2 .
- the unrotated disparity value may be determined according to the following equation (6). d 2 — d + dx 2 Equation
- the rotation correction module 202 may determine the unrotated disparity map 216 based on the abovementioned computations.
- the rotation correction module 202 may determine the corresponding position 406 based on the rotated disparity values in the rotated disparity map 214.
- the rotation correction module 202 may determine the first hypothenuse 420 based on the corresponding position 406.
- the rotation correction module 202 may determine the vertex angle 426. Determining the vertex angle 426 may include determining direction of epipolar lines of the image set 212, in the image plane without using 3D space. The vertex angle 426 may then be obtained by projecting the rotated disparity measurements towards the epipolar line directions.
- the rotation correction module 202 may determine the second hypotenuse 424 based on the vertex angle 426 and the first hypotenuse 420.
- the rotation correction module 202 may determine the base 428 based on the second hypotenuse 424 and the rotated disparity values.
- the rotation correction module 202 may determine the unrotated disparity value based on the base 428 and the rotated disparity values.
- the rotation correction module 202 may be configured to determine direction of epipolar lines of the image set 212, in the image plane without using 3D space.
- the rotation correction module 202 may determine the epipolar line directions without prior knowledge about the intrinsic and extrinsic camera parameters.
- the rotation correction module 202 may determine the epipolar line directions, by processing the images in the image set 212, region by region. In other words, the images may be segmented into a plurality of regions, and the epipolar line directions may be determined in each region of the plurality of regions.
- the rotation correction module 202 may be configured to check for infinity-negative disparities in the vertical disparity.
- the real distance represented in the images may be computed by the projection of the disparities through the epipolar line direction.
- the real distances may be determined region by region, in other words, determined in each region of the plurality of regions.
- the rotation correction module 202 be further configured to determine the translation (x t , y t ) and further configured to correct the right image with respect to the left image, based on the determined translation.
- the horizontal translation x t may be computed using tracking and filtering of the input images over time.
- the vertical translation y t may be determined based on the pixel information in the centre of y-disparity in a y-disparity array.
- the 3D reconstruction module 206 may be configured to generate an unsealed 3D point cloud based on the unrotated disparity map 216.
- the trained machine learning model 104 may receive a sequence of image sets 212 over time, and thereby generating a sequence of rotated disparity maps 214. Accordingly, the rotation correction module 202 may generate a sequence of unrotated disparity maps.
- the 3D reconstruction module 206 may correct the 3D point cloud for camera factors such as scale and yaw angle, based on comparing at least two consecutive 3D point clouds.
- Points in the 3D point cloud should move at least substantially at the same velocity, as the points move relative to the vehicle that the cameras are coupled to.
- the points in the 3D point cloud may move at a velocity that is at least substantially equal to the longitudinal speed of the vehicle.
- the 3D reconstruction module 206 may correct the 3D point cloud for camera factors based on velocity deviations of points in the 3D point cloud.
- the method 200 may further include correcting for negative disparities in the image set 212.
- the correction for negative disparities may be performed before performing 3D reconstruction.
- the negative disparities may be caused by yaw angle errors between the cameras that captured the image set 212.
- a change in the yaw angle of at least one camera may result in a horizontal shift.
- the horizontal disparity decreases as distance, i.e. depth, increases.
- the horizontal disparity of objects in infinity also referred herein as infinity objects, should be at least substantially zero.
- the infinity objects may display negative horizontal disparities instead of zero disparity when there is a negative yaw angle between the cameras.
- the method 200 may further include applying a low pass filter to the unrotated disparity map 216, to obtain the maximum negative disparity values, i.e. the negative disparity values with the largest amplitude, in the unrotated disparity map 216.
- the method 200 may include using the maximum negative disparity value to correct the disparities in the entire image region. Correction of the disparities may be performed region by region, where each image may be divided into a plurality of regions.
- FIG. 5A shows an example of an image set 212 according to various embodiments.
- the image set 212 includes a first image 502 and a second image 504. The first image 502 was taken by a first camera while the second image 504 was taken by a second camera.
- Both the first camera and the second camera captured images of the same scene, from slightly offset positions.
- the second camera is positioned at a calibrated distance away from the first camera.
- the second image 504 is rotated relative to the first image 502.
- FIG. 5B shows disparity images generated by a prior art machine learning model and the trained machine learning model 104 according to various embodiments.
- the disparity images include a first disparity image 510 and a second disparity image 520, which were both generated by machine learning models based on the example image set 212 of FIG. 5 A.
- the first disparity image 510 was generated by a prior art machine learning model.
- the prior art machine learning model failed to handle the decalibration of the second image 504, and as such, objects in the image set 212 are not visible in the first disparity image 510.
- FIG. 6 shows an image processing method 600 according to various embodiments.
- the image processing method 600 may include processes 602, 604, 606, 608 and 610.
- the process 602 may include inputting an image set 212 to a trained machine learning model 104.
- the image set 212 may include a first image and a second image.
- the process 604 may include generating a rotated disparity map 214 of the image set 212, using the trained machine learning model 104.
- the process 606 may include, for each pixel in the first image, determining based on the rotated disparity map 214, a rotation angle between the pixel in the first image and its corresponding position in the second image, relative to a common image centre.
- the process 608 may include, for each pixel in the first image, determining a respective unrotated disparity based on the respective rotation angle.
- the process 610 may include generating unrotated disparity map 216 based on the respective unrotated disparities of each pixel in the first image.
- the image processing method 600 may determine an unrotated disparity map that may be used to accurately determine depth of points in the image set 212, for example, to construct an accurate 3D point cloud, without having to calibrate the cameras that capture the image set 212.
- These cameras may be mounted on a vehicle, to capture the surroundings of the vehicle. Cameras mounted on the vehicle may shift out of their initial calibrated positions due to vibrations or movements of the vehicle. Using the image processing method 600, the depth of points in the image set 212 may be determined accurately without being affected by the decalibration of the cameras.
- the image set 212 may include stereo images.
- the first image may be a left stereo image, while the second image may be a right stereo image.
- a stereo image set may capture a similar scene from slightly different positions, such that the images within the stereo image set may be compared to obtain depth information on each point in the images. This depth information may be useful for providing situational awareness to a vehicle, such as indicating the vehicle’s distances to various objects in its surroundings.
- the trained machine learning model 104 may be configured to determine both horizontal disparities and vertical disparities of the image set 212, such that the rotated disparity map 214 may include both horizontal disparities and vertical disparity values of the image set.
- the trained machine learning model 104 may be capable of handling decalibration of the cameras in both vertical and horizontal directions. As such, the image processing method 600 may result in an accurate disparity map even if at least one of the cameras is shifted out of place in both directions, or is rotated.
- the common image centre may be an arbitrary point in the image and may be close to the image centre of at least one of the first image and the second image.
- a centre of the first image may be used as the common image centre.
- the first image may be used as a reference image, and the translation of the second image relative to the first image may be estimated.
- the image processing method 600 may further include segmenting the first image into a plurality of first regions, and segmenting the second image into a plurality of second regions corresponding to the plurality of first regions.
- the image processing method 600 may further include determining an epipolar line and its direction based on a pixel in the first image and its corresponding position in the second image, for each first region of the plurality of first regions.
- the image processing method 600 may further include for each pixel of the first image, determining a respective horizontal disparity based on projection of the respective unrotated disparity through the direction of the epipolar line.
- the horizontal disparity in the image set 212 for example, caused by misalignment of at least one camera, may be thereby corrected.
- the image processing method 600 may further include generating a 3D point cloud 208 based on the image set 212 and further based on the unrotated disparity map 216.
- the 3D point cloud may provide detailed information about the surroundings of a vehicle, and may serve to aid the vehicle in various functions such as navigation, localization and avoidance of obstacles.
- the image processing method 600 may further include inputting a further image set to the trained machine learning model 104.
- the further image set may include a further first image and a further second image captured at a consecutive time frame from the first image and the second image of the image set 212.
- the image processing method 600 may further include generating a further rotated disparity map of the further image set, using the trained machine learning model 104.
- the image processing method 600 may further include, for each pixel in the further first image, identifying a corresponding pixel in the further second image based on the further rotated disparity map, determining a respective rotation angle between the pixel in the further first image and the corresponding pixel in the further second image, relative to a further common image centre, and determining a respective unrotated disparity based on the respective rotation angle.
- the image processing method 600 may further include generating a further unrotated disparity map based on the respective unrotated disparities of each pixel in the further first image.
- the image processing method 600 may further include generating a further 3D point cloud based on the further image set and further based on the further unrotated disparity map, determining a distance that each point in the 3D point cloud moves from the 3D point cloud 208 to the further 3D point cloud, and correcting at least one of the 3D point cloud 208 and the further 3D point cloud, based on the determined distances.
- the 3D reconstruction module 206 may perform the above-described correction of the 3D point cloud 208 or the further 3D point cloud. Points in the 3D point cloud should move at least substantially at the same velocity, as the points move relative to the vehicle that the cameras are coupled to.
- the 3D reconstruction module 206 may correct the 3D point cloud for camera factors such as scale and yaw angle, based on velocity deviations of points in the 3D point cloud.
- the trained machine learning model 104 may be trained using a training dataset 112 that includes a plurality of image pairs generated by rotating a pair of calibrated images to a corresponding plurality of different angles, and using a ground truth disparity map 114 generated based on the pair of calibrated images as a training signal. As the training dataset 112 is generated using calibrated images, a precise ground truth disparity map 114 may be obtained.
- the trained machine learning model 104 may be trained to determine near range disparities, and further separately trained to determine far range disparities. Doing so may optimize the training process, reducing the training time required and improving the accuracy of the trained machine learning model 104 in determining the disparities.
- FIG. 7A shows a simplified block diagram of an image processing device 700 according to various embodiments.
- the image processing device 700 may include a processor 702.
- the processor 702 may be configured to perform the image processing method 600.
- the image processing device 700 may be any one of a server, a computer, a vehicle or a robot.
- the image processing device 700 may be an autonomous vehicle, such as a self-driving car, or a drone.
- the image processing device 700 may further include at least one camera 704.
- the camera 704 may be configured to capture the image set 212.
- the image processing device 700 may further include an engine 706 configured to drive the image processing device 700.
- the image processing device 700 may further include a steering module 708 configured to steer the image processing device 700.
- FIG. 7B illustrates an operation of the image processing device 700 according to various embodiments.
- the input 720 to the image processing device 700 may include a sequence of image sets.
- the sequence of image sets may include consecutively captured images.
- the sequence of image sets may include an image set 720a captured at t- 1, and another image set 720b captured at t, where t denotes time as a variable.
- the image set 720b may be a subsequent frame to the image set 720a.
- Each image set may include a plurality of images, each captured at a respective position. These positions may be offset from one another, such that a combination of the plurality of images may provide depth information of the objects shown in the images.
- each of the image set 720a and the image set 720b may include a pair of stereo images.
- the image set 720a may include a left image 722a and a right image 724a.
- the image set 720b may include a left image 722b and a right image 724b.
- the sequence of image sets 720a, 720b may be provided to the trained machine learning model 104.
- the trained machine learning model 104 may be trained to compute a respective rotated disparity map 214 for each image set.
- the trained machine learning model 104 may be trained to identify features in images and further configured to determine spatial offset, also referred herein as disparity data, of the features between the images.
- the spatial offset between images of the same image set may provide depth information of the features.
- a rotation correction module 202 may compute a respective unrotated disparity map 216 for each rotated disparity map 214, for example, like described with respect to FIG. 4.
- the device 700 may further include a point cloud generator 726.
- the point cloud generator 726 may include the 3D reconstruction module 206.
- the point cloud generator 726 may be configured to generate a 3D point cloud 208 based on the input 720 and the unrotated disparity map 216.
- the generated 3D point cloud 208 may be a 3D reconstruction of the environment that a vehicle is travelling in.
- the device 100 may provide the vehicle with environmental data that is not previously available, for example, previously unmapped terrain. Further, the 3D point cloud 208 may indicate to the vehicle, the presence of dynamic objects such as other traffic participants.
- the 3D point cloud 208 may be generated further based on camera intrinsic parameters such as focal length along x, camera principal point offset and axis skew.
- the point cloud generator 726 may also refine the 3D point cloud 208 to remove outliers and incorrect predictions, so that the 3D point cloud 208 may serve as an accurate dense 3D map.
- the point cloud generator 726 may remove the outliers and incorrect predictions based on probability of those data points.
- TTC time-to-collision
- a computer-implemented method 1200 for predicting collisions may estimate TTC based on stereo disparity information and optical flow.
- the stereo disparity information may be obtained using the image processing method 600.
- the method 1200 is also described with respect to FIG. 12.
- the method 1200 may include receiving a sequence of image sets.
- the sequence of image sets in the input 720 may be captured by sensors mounted on a vehicle. Examples of suitable sensors include stereo vision camera, thermal stereo camera, or dense LiDAR. LiDAR sensors may provide dense point cloud of the surrounding environment as well as the distance to each detected object in the scene.
- Each image set may include a plurality of images, and each image of the plurality of images may be captured from a different position on the vehicle, such that the plurality of images of each image set may have an offset in at least one axis, from one another.
- the plurality of images may be respectively captured by a corresponding plurality of sensors.
- the plurality of images may include a first image and a second image.
- the senor may be a pair of stereo cameras.
- the image set includes a pair of stereo images.
- the two cameras may be separated by a baseline, the distance for which is assumed to be known accurately.
- the cameras may simultaneously capture two consecutive images.
- the images may be analyzed to identify differences between the images, and to identify the corresponding pixel-positions in both images, in a stereo matching process.
- the disparity between corresponding pixel in both images may be used to estimate depth.
- the pair of stereo cameras may be mounted on the side mirror of a vehicle.
- the stereo camera may be FSC231 stereo camera with 8.3M Pixel resolution and 30° field of view.
- the method 1200 may include determining the disparity between the images in the image set, using a motion detection module that includes a machine learning model.
- the machine learning model may be, for example, the trained machine learning model 104.
- the machine learning model may include a neural network, such as graph neural network, transformers, or recurrent neural network.
- the machine learning model may include the neural network 1000 described with respect to FIG. 10.
- the distance and TTC for each given pixel or object in the scene captured by the images may be computed. The distance may be measured in meters, while the TTC may be measured in seconds.
- the features extracted for performing the stereo matching may be used to estimate TTC for each pixel in the image using deep neural network.
- the method 1200 may include recognizing moving objects, by identifying moving pixels in the sequence of image sets.
- recognizing moving objects may involve finding corresponding pixels in the sequence of image sets, over time.
- Recognition of the moving objection may be accomplished by passing two consecutive stereo image pairs through an optical flow estimation algorithm.
- the method 1200 may include combining the resulting flow information with the current stereo frame to segment moving objects in a scene, using a neural network.
- a point cloud generator may generate a 3D point cloud of the surrounding of the vehicle, based on the estimated depth, i.e. distance of each pixel.
- the method 1200 may include providing stereo camera data to a motion detection module.
- the motion detection module may include a machine learning model, also referred herein as motion detection network.
- the motion detection network may include the trained machine learning model 104.
- the motion detection network may include, for example, a Convolutional Neural Network (CNN).
- the stereo camera data may include a first stereo image pair and a second stereo image pair. The first stereo image pair and the second stereo image pair may be captured successively.
- the motion detection network may be trained using stereo camera data, ground truth disparity and ground truth TTC values.
- the stereo camera data may be images captured at night, so that the motion detection network may learn to compute the disparity values and TTC values for night images.
- the motion detection network may also be trained with images associated with other environmental conditions such as day light, rain or snow.
- the motion detection network may be trained with uncalibrated stereo images, and may generate a rotated disparity map of the stereo camera data.
- the motion detection module may include a rotation correction module 202 that converts the rotated disparity map to unrotated disparity map.
- the motion detection network may perform stereo matching, by extracting feature information of left and right stereo images, and mapping each image to a dense disparity map.
- the motion detection network may be trained to determine the rotated disparity, the optical flow, and the TTC image output of two consecutive stereo image pairs. In order to segment moving objects in a scene, the flow information obtained may be combined with a current stereo frame to train the motion detection network to predict the mask of moving objects.
- the motion detection module may determine the relative depth of the moving objects based on the unrotated disparity map, and may further determine a TTC image based on the determined relative depth.
- the TTC image may include TTC values for each pixel in the image.
- the motion detection module may further perform trajectory planning based on the TTC estimation.
- the disparity map, the input image and the flow information may be further used for estimating pose and performing simultaneous localization and mapping (SLAM) of the scene, for navigation of the vehicle.
- SLAM simultaneous localization and mapping
- FIG. 8A shows a simplified functional block diagram of a device 800 for predicting collisions according to various embodiments.
- the device 800 may be configured to receive an input 720 and may be configured to generate an output.
- the input 720 may include a sequence of image sets.
- the output may include predicted time-to-collision (TTC) 820 of a vehicle with another object or vehicle.
- TTC predicted time-to-collision
- the device 800 may be capable of determining the TTC 820 based on processing visual data, i.e. images.
- the device 800 may be configured to perform the method 1200.
- the sequence of image sets in the input 720 may be captured by sensors mounted on a vehicle.
- Each image set may include a plurality of images, and each image of the plurality of images may be captured from a different position on the vehicle, such that the plurality of images of each image set may have an offset in at least one axis, from one another.
- the plurality of images may be respectively captured by a corresponding plurality of sensors.
- the plurality of images may include a first image and a second image. In some embodiments, the first image and the second image may be a pair of stereo images.
- the device 800 may include a motion detection module 802.
- the motion detection module 802 may include a machine learning model, for example, the trained machine learning model 104.
- the motion detection module 802 may configured to generate a respective depth map 812 of each image set, based on the first image and the second image of the image set. Consequently, the motion detection module 802 may output a sequence of depth maps 812 based on the received sequence of image sets.
- Each depth map 812 may be an image or image channel that contains information relating to the distance of the surfaces of objects from a viewpoint.
- the viewpoint may be a vehicle, or more specifically, a sensor mounted on the vehicle.
- the motion detection module 802 may also be configured to determine an optical flow 814 based on at least one of, the first images from the sequence of image sets and the second images of the sequence of image sets.
- the motion detection module 802 may determine motion of an object in the image set based on the optical flow 814, to generate motion data 816.
- the machine learning model in the motion detection module 802 may be used to determine both disparity and motion.
- the machine learning model may generate the segmented mask of moving objects based on the detected motion.
- the device 800 may further include a prediction module 804.
- the prediction module 804 may be configured to receive the sequence of depth maps 812 and the motion data 816 from the motion detection module 802.
- the prediction module 804 may be configured to generate the predicted TTC 820 based on the received sequence of depth maps 812 and the motion data 816.
- the motion detection module 802 may include the device 700.
- the motion detection module 802 may determine the depth maps 812 by generating a respective unrotated disparity map 216 for each image set, for example, according to the image processing method 600.
- the device 800 may further include a point cloud generator 726 that generates a 3D point cloud 208 based on at least one image set of the sequence of image sets.
- the device 800 may be further configured to detect an object in the 3D point cloud 208.
- the point cloud generator 726 may generate a respective 3D point cloud based on each image set of the sequence of image sets, resulting in a sequence of 3D point clouds 208.
- the device 800 may compare distances moved by each point in the sequence of 3D point clouds 208, and may correct at least one 3D point cloud of the sequence of 3D point clouds 208 based on the determined distances.
- FIG. 8B shows a simplified hardware block diagram of the device 800 according to various embodiments.
- the device 800 may include at least one processor 830.
- the at least one processor 830 may be configured to carry out the functions of the machine learning model 802 and the prediction module 804.
- the device 800 may be a driver assistance system. [0093] According to various embodiments, the device 800 may be a vehicle.
- the device 800 may further include a plurality of sensors 832.
- the plurality of sensors 832 may be configured to generate the sequence of image sets 816.
- the plurality of sensors 832 may include, for example, stereo cameras, surround view cameras, infrared cameras, event cameras, or dense LiDAR.
- the device 800 may include at least one memory 834.
- the at least one memory 834 may store the machine learning model 802 and the prediction module 804.
- the at least one memory 834 may include a non-transitory computer-readable medium.
- the at least one processor 830, the plurality of sensors 832 and the at least one memory 834 may be coupled to one another, for example, mechanically or electrically, via the coupling line 840.
- FIG. 9 shows a schematic diagram of the device 800 performing a method for predicting collisions, according to various embodiments.
- the motion detection module 802 of the device 800 may include a machine learning model that may generate a sequence of depth maps 812, optical flow 814 and motion data 816, based on a received sequence of image sets.
- the sequence of image sets may include at least a first image set 720a and a second image set 720b.
- Each of the image sets may include at least a first image and a second image, for example left image 722a and right image 724a of the first image set 720a, and the left image 722b and the right image 724b of the second image set 720b.
- the motion detection module 802 may also receive a moving object mask 914, which may be generated by another machine learning model.
- the motion detection module 802 may output a segmented moving object image 910 based on the moving object mask 914 applied to the sequence of image sets.
- the device 800 may also include a tracking moving object module 906 configured to determine the motion data 816 based on the segmented moving object image 910.
- the prediction module 804 of the device 800 may include a relative depth estimation module 902.
- the relative depth estimation module 902 may be configured to receive the sequence of depth maps 812, the optical flow 814 and the motion data 816.
- the relative depth estimation module 902 may determines relative depth of pixels in the images, based on the received inputs.
- the prediction module 804 may further include a TTC determination module 904.
- the TTC determination module 904 may determine the TTC of each pixel, based on the determined relative depth and the motion data 816.
- FIG. 10 shows a schematic diagram of a neural network 1000 according to various embodiments.
- the neural network 1000 may be part of at least one of the trained machine learning model 104 and the motion detection module 802.
- the neural network 1000 may include a feature extraction network 1310 and a disparity computation network 1320.
- the feature extraction network 1310 may be configured to extract features from images in the sequence of image sets to generate feature maps.
- the disparity computation network 1320 may be configured to determine two-dimensional offsets (also referred herein as displacements or disparities) between the images based on the generated feature maps.
- the neural network 1000 may thereby determine both optical flow and depth information using a single common set of neural networks, as both the optical flow and depth information relate to two-dimensional offsets between images. As such, the neural network 1000 may be trained in a shorter time and with less resources, as compared to training two separate neural networks.
- the neural network 1000 may be trained via supervised training, using scene flow stereo images. Training the neural network 1000 may require, for example, 25,000 scene flow stereo images as the training data.
- the neural network 1000 may be fine-tuned using stereo images from the Düsseldorf Institute of Technology and Toyota Technological Institute (KITTI) dataset. As an example, about 400 stereo images from the KITTI dataset may be used.
- the feature extraction network 1310 may include a plurality of neural network branches.
- the neural network branches may include a first branch 1350 and a second branch 1360.
- Each neural network branch may include a respective convolutional stack, and a pooling module connected to the convolutional stack.
- the convolutional stack may include, for example, the CNN 1312.
- the pooling module may include, for example, the SPP module 1314.
- the plurality of neural network branches may share the same weights. By having the neural network branches share the same weights, the feature extraction network 1310 may be trained in a shorter time and with less resources, than training the neural network branches individually.
- the disparity computation network 1320 may include a 3D convolutional neural network (CNN) 1324 configured to generate three disparity maps based on the generated feature maps 1318.
- the feature extraction network 1310 may extract features at different levels.
- the 3D CNN 1324 may be configured to perform cost volume regularization on the extracted features.
- the 3D CNN 11324 may include an encoder-decoder architecture including a plurality of 3D convolution and 3D deconvolution layers with intermediate supervision.
- the 3D CNN 324 may have a stacked hourglass architecture including three hourglasses, thereby producing three disparity outputs.
- the architecture of the 3D CNN 1324 may enable it to generate accurate disparity outputs and to predict the optical flow.
- the filter size in the 3D CNN 1324 may be 3*3.
- the neural network 1000 may be configured to determine the respective depth map 812 of each image set based on the determined 2D offsets between the plurality of images of the image set.
- the depth map 812 may provide information on distance between objects in the images from the vehicle. This information may improve the localization accuracy.
- the 2D offsets may be obtained from the unrotated disparity map 216.
- the neural network 1000 may be configured to determine the optical flow 814, based on the determined 2D offsets between at least one of, the first images of adjacent image sets of and the second images of adjacent image sets.
- the optical flow 814 may provide information on the changes in position of objects in the images.
- the neural network 1000 may be configured to perform a stereo matching operation.
- the stereo matching operation may include determining estimate a pixelwise displacement map between the input images.
- the input images may include a plurality of images of the same image set. For example, when the input contains images captured by a stereo camera, the input images may include a left image 722b and a right image 724b.
- stereo images may be rectified stereo images or unrectified stereo images.
- Rectified stereo images are stereo images where the displacement of each pixel is constrained to a horizontal line.
- the displacement map may be referred herein as disparity.
- the sensors or cameras used to capture the images need to be accurately calibrated.
- Unrectified stereo images may exhibit both vertical and horizontal disparity.
- Vertical disparity may be defined as the vertical displacement between corresponding pixels in the left and right images.
- Horizontal disparity may be defined as the horizontal displacement between corresponding pixels in the left and right images.
- the neural network 1000 may be trained to be robust against rotation, shift, vibration and distortion of the input images.
- the neural network 1000 may be trained to handle both horizontal (x-disparity) and vertical (y-disparity) displacement.
- the neural network 1000 may be configured to determine both the horizontal and vertical disparity of the input images. This makes the neural network 1000 robust against vertical translation and rotation between cameras or sensors.
- the neural network 1000 may have a dual-branch neural network architecture including a first branch 1350 and a second branch 1360, such that each branch may be configured to determine disparity in a respective axis.
- Each of the feature extract network 310 and the disparity computation network 1320 may include components of the first branch 350 and the second branch 360.
- the feature extraction network 1310 may include a convolutional neural network (CNN) 1312, a spatial pyramid pooling (SPP) module 1314 and a convolution layer 1316, for each of the first branch 11350 and the second branch 1360.
- the CNN 1312 may extract feature information from the input images.
- the CNN 1312 may include three small convolution filters with kernel size (3 x 3) that are cascaded to construct a deeper network with the same receptive field.
- the CNN 1312 may include convl x, conv2_x, conv3_x, and conv4_x layers that form the basic residual blocks for learning the unitary feature extraction.
- dilated convolution may be applied to further enlarge the receptive field.
- the output feature map size may be (1/4 x 1/ 4) of the input image size.
- the SPP module 1314 may be then applied to gather context information from the output feature map.
- the SPP module 1314 may learn the relationship between objects and its sub regions to incorporate hierarchical context information.
- the SPP module 1314 may include four fixed-size average pooling blocks of size 64 x 64, 32 x 32, 16 x 16, and 8x8.
- the convolution layer 1316 may be a lxl convolution layer for reducing feature dimension.
- the feature extraction network 1310 may up-sample the feature maps to the same size as the original feature map, using bilinear interpolation.
- the size of the original feature map may be 1/4 of the input image size.
- the feature extraction network 1310 may concatenate the different levels of feature maps extracted by the various convolutional filters, as the left SPP feature map 1318 and the right SPP feature map 13
- the disparity computation network 1320 may receive the left SPP feature map 1318 and the right SPP feature map 1319.
- the disparity computation network 1320 may concatenate the left and right SPP feature maps 1318, 1319 into separate cost volumes 1322 for x and y displacements respectively.
- Each cost volume may have 4 dimensions, namely, height x width x disparity x feature size.
- the disparity computation network 1320 may include a 3D-CNN 1322 in each branch.
- the 3D-CNN 1322 may include a stack hourglass (encoderdecoder) architecture that is configured to generate three disparity maps.
- the disparity computation network 1320 may further include an upsampling module 1326 and a regression module 1328, in each branch.
- the upsampling module 1326 may upsample the three disparity maps so that their resolution matches that of the input image size.
- the regression module 1328 may apply regression to the upsampled disparity maps, to calculate an output disparity map.
- the disparity computation network 1320 may calculate the probability of each disparity based on the predicted cost via SoftMax operation.
- the predicted disparity may be calculated as the sum of each disparity weighted by its probability.
- smooth loss function may be applied between ground truth disparity and predicted disparity.
- the smooth loss function may measure how close the predictions are, to the ground truth disparity values.
- the smooth loss function may be a combination of 11 and 12 loss. It is used in deep neural network because of its robustness and low sensitivity to outliers.
- the disparity computation network 1320 then outputs the horizontal displacement 1332 at the first branch 1350, and outputs the vertical displacement 1334 at the second branch 1360.
- the machine learning model 102 may then determine a depth map 112 based on the horizontal displacement 1332 and the vertical displacement 1334, using known stereo computation methods such as semi global matching.
- FIG. 11 shows a schematic diagram of the neural network 1000 according to various embodiments, receiving different inputs from those in FIG. 10.
- the neural network 1000 may also be configured to perform optical flow computation.
- the optical flow computation may include predicting a pixelwise displacement field, such that for every pixel in a frame, the neural network 1000 may estimate its corresponding pixel in the next frame.
- the outputs of the neural network 1000 for the optical flow computation operation is the optical flow 814 that includes x-direction displacement and y-direction displacement.
- the first branch 1350 may process the left image 722b, while the second branch 1360 may process the earlier left image 722a.
- the optical flow computation operation may include extracting features using the CNN 1312 of the feature extraction network 1310.
- the SPP module 1314 may gather context information from the output feature map generated by the CNN 11312.
- the feature extraction network 1310 generate final SPP feature maps that are provided to the disparity computation network 1320.
- the 3D-CNN 1322 of each branch may generate three disparity maps based on the respective cost volume 1322.
- the upsampling module 1326 may upsample the three disparity maps.
- the regression module 1328 may apply regression to the upsampled disparity maps, to calculate an output disparity map.
- the disparity computation network 1320 may calculate the probability of each disparity based on the predicted cost via SoftMax operation. The predicted disparity may be calculated as the sum of each disparity weighted by its probability.
- suitable deep learning models for the machine learning model 102 may include, for example, PyramidStereoMatching and RAFTNet.
- training of the neural network 1000 may, for example, be based on standard backpropagation based gradient descent.
- a training dataset may be provided to the neural network 1000, and the following training processes may be carried out:
- An example of a suitable training dataset for training the neural network 1000 may be the KITTI dataset.
- the weights may be randomly initialized to numbers between 0.01 and 0.1, while the biases may be randomly initialized to numbers between 0.1 and 0.9.
- the first observations of the dataset may be loaded into the input layer of the neural network and the output value(s) is generated by forward-propagation of the input values of the input layers.
- the following loss function may be used to calculate loss with the output value(s):
- MSE Mean Square Error
- the weights and biases may subsequently be updated by an AdamOptimizer with a learning rate of 0.001.
- the steps described above may be repeated with the next set of observations until all the observations are used for training. This may represent the first training epoch, and may be repeated until 10 epochs are done.
- FIG. 12 shows a flow diagram of a computer-implemented method 1200 for predicting collisions, according to various embodiments.
- the method 1200 may include processes 1202, 1204, 1206, 1208 and 1210.
- the process 1202 may include inputting a sequence of image sets to a motion detection module, wherein each image set of the sequence comprises a first image and a second image.
- the sequence of image sets may include for example, the image sets 720a, 720b.
- the motion detection module may be for example, the motion detection module 802.
- the motion detection module may include the neural network 1000.
- the process 1204 may include determining by the motion detection module, for each image set, a respective depth map based on the first image and the second image of the image set, resulting in a sequence of depth maps.
- the image set may include a pair of stereo images, where the first image may be, for example, the left image 722a or 722b while the second image may be, for example, the right image 724a or 724b.
- the depth maps may include, for example, the depth maps 812.
- the process 1206 may include determining by the motion detection module, an optical flow based on at least one of, the first images from the sequence of image sets and the second images of the sequence of image sets.
- the optical flow may be for example, the optical flow 814.
- the process 1208 may include determining by the motion detection module, motion of an object in the image set based on the optical flow.
- the process 1210 may include determining TTC with the object based on the sequence of depth maps and the determined motion of the object.
- the method 1200 may be able to determine TTC using camera images, even when the road lighting is dim. As such, employing the method 1200 on vehicles, may result in avoidance of traffic accidents at night. Various aspects described with respect to the device 800 may be applicable to the method 1200.
- the process 1204 may include generating a respective unrotated disparity map for each image set using the image processing method 600, and determining the respective depth map based on the respective unrotated disparity map. Determining the depth maps using unrotated disparity maps generated according to the image processing method 600 allows accurate depth to be determined even when the cameras that output the images in the image sets are uncalibrated. The uncalibration of the cameras may occur as a result of vibrations or movements of the vehicle on which the cameras are mounted.
- the method 1200 may further include generating a 3D point cloud based on at least one image set of the sequence of image sets, and detecting the object in the three-dimensional point cloud.
- the 3D point cloud is a dense in data.
- the 3D point cloud provides detailed information of the surroundings of the vehicle, such that object detection in the 3D point cloud is accurate.
- the method 1200 may further include generating a respective three-dimensional point cloud based on each image set of the sequence of image sets, resulting in a sequence of three-dimensional point clouds, comparing distances moved by each point in the sequence of three-dimensional point clouds, and correcting at least one three-dimensional point cloud of the sequence of three-dimensional point clouds, based on the determined distances.
- the method 1200 may correct for distortions caused by, for example, camera intrinsics or extrinsics.
- Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A, B, and/or C, and may include multiples of A, multiples of B, or multiples of C.
- combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Artificial Intelligence (AREA)
- Databases & Information Systems (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
- Signal Processing (AREA)
- Image Analysis (AREA)
- Image Processing (AREA)
Abstract
Description
Claims
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202380073921.8A CN120092261A (en) | 2022-11-11 | 2023-10-26 | Image processing method and method for predicting collision |
| DE112023004739.1T DE112023004739T5 (en) | 2022-11-11 | 2023-10-26 | IMAGE PROCESSING METHODS AND METHODS FOR PREDICTING COLLISIONS |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| IN202241064626 | 2022-11-11 | ||
| IN202241064626 | 2022-11-11 | ||
| GB2305580.9 | 2023-04-17 | ||
| GB2305580.9A GB2624483A (en) | 2022-11-11 | 2023-04-17 | Image processing method and method for predicting collisions |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024099786A1 true WO2024099786A1 (en) | 2024-05-16 |
Family
ID=88598682
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2023/079921 Ceased WO2024099786A1 (en) | 2022-11-11 | 2023-10-26 | Image processing method and method for predicting collisions |
Country Status (3)
| Country | Link |
|---|---|
| CN (1) | CN120092261A (en) |
| DE (1) | DE112023004739T5 (en) |
| WO (1) | WO2024099786A1 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN121433264A (en) * | 2026-01-04 | 2026-01-30 | 杭州亿亿德传动设备有限公司 | A Multi-Vision-Based Intelligent Obstacle Avoidance Method and System for AMR Robots |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113077401A (en) * | 2021-04-09 | 2021-07-06 | 浙江大学 | Method for stereo correction based on viewpoint synthesis technology of novel network |
| US20220309776A1 (en) * | 2021-03-29 | 2022-09-29 | Conti Temic Microelectronic Gmbh | Method and system for determining ground level using an artificial neural network |
-
2023
- 2023-10-26 DE DE112023004739.1T patent/DE112023004739T5/en active Pending
- 2023-10-26 WO PCT/EP2023/079921 patent/WO2024099786A1/en not_active Ceased
- 2023-10-26 CN CN202380073921.8A patent/CN120092261A/en active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220309776A1 (en) * | 2021-03-29 | 2022-09-29 | Conti Temic Microelectronic Gmbh | Method and system for determining ground level using an artificial neural network |
| CN113077401A (en) * | 2021-04-09 | 2021-07-06 | 浙江大学 | Method for stereo correction based on viewpoint synthesis technology of novel network |
Non-Patent Citations (4)
| Title |
|---|
| CHANG JIA-REN ET AL: "Pyramid Stereo Matching Network", 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, IEEE, 18 June 2018 (2018-06-18), pages 5410 - 5418, XP033473454, DOI: 10.1109/CVPR.2018.00567 * |
| RATEKE THIAGO ET AL: "Road obstacles positional and dynamic features extraction combining object detection, stereo disparity maps and optical flow data", MACHINE VISION AND APPLICATIONS, SPRINGER VERLAG, DE, vol. 31, no. 7-8, 25 September 2020 (2020-09-25), XP037255193, ISSN: 0932-8092, [retrieved on 20200925], DOI: 10.1007/S00138-020-01126-W * |
| YANG WANG ET AL: "Joint Unsupervised Learning of Optical Flow and Depth by Watching Stereo Videos", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 8 October 2018 (2018-10-08), XP081058894 * |
| ZHANG XUCHONG ET AL: "End-to-end learning of self-rectification and self-supervised disparity prediction for stereo vision", ARXIV,, vol. 494, 26 April 2022 (2022-04-26), pages 308 - 319, XP087058808, DOI: 10.1016/J.NEUCOM.2022.04.095 * |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN121433264A (en) * | 2026-01-04 | 2026-01-30 | 杭州亿亿德传动设备有限公司 | A Multi-Vision-Based Intelligent Obstacle Avoidance Method and System for AMR Robots |
Also Published As
| Publication number | Publication date |
|---|---|
| CN120092261A (en) | 2025-06-03 |
| DE112023004739T5 (en) | 2025-09-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN115797454B (en) | Multi-camera fusion sensing method and device under bird's eye view angle | |
| EP3598874B1 (en) | Systems and methods for updating a high-resolution map based on binocular images | |
| US12260575B2 (en) | Scale-aware monocular localization and mapping | |
| US10909395B2 (en) | Object detection apparatus | |
| WO2020097840A1 (en) | Systems and methods for correcting a high-definition map based on detection of obstructing objects | |
| US20190387209A1 (en) | Deep Virtual Stereo Odometry | |
| CN113834463B (en) | Intelligent vehicle side pedestrian/vehicle monocular depth ranging method based on absolute size | |
| JP2017181476A (en) | Vehicle position detection device, vehicle position detection method, and computer program for vehicle position detection | |
| CN119540916A (en) | Parking visual positioning method, system and vehicle-mounted terminal based on parking space semantic information | |
| CN117576199A (en) | A driving scene visual reconstruction method, device, equipment and medium | |
| GB2624483A (en) | Image processing method and method for predicting collisions | |
| CN114972494B (en) | A method and device for constructing a map for memorizing parking scenes | |
| CN120092261A (en) | Image processing method and method for predicting collision | |
| CN111986248B (en) | Multi-vision sensing method and device and automatic driving automobile | |
| CN115690711B (en) | Target detection method and device and intelligent vehicle | |
| EP4577984A1 (en) | Method and device for generating an outsider perspective image and method of training a neural network | |
| CN117197433A (en) | Target detection method, device, electronic equipment and storage medium | |
| GB2628602A (en) | Method for detecting an object and method for trainiing a detection neural network | |
| WO2022133986A1 (en) | Accuracy estimation method and system | |
| US12549698B2 (en) | Robust stereo camera image processing method and system | |
| Rajesh et al. | Performance evaluation of low-resolution monocular vision-based velocity estimation technique for moving obstacle detection and tracking | |
| EP4623410A1 (en) | Device for localizing a vehicle and method for localizing a vehicle | |
| WO2024179658A1 (en) | Calibration of sensor devices | |
| CN120298985A (en) | A pure vision-based omnidirectional universal AEB method and system | |
| CN116071727A (en) | Target detection and early warning method, equipment, system and medium |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23798201 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202380073921.8 Country of ref document: CN |
|
| WWP | Wipo information: published in national office |
Ref document number: 202380073921.8 Country of ref document: CN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 112023004739 Country of ref document: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: 112023004739 Country of ref document: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23798201 Country of ref document: EP Kind code of ref document: A1 |